> ## Documentation Index
> Fetch the complete documentation index at: https://dify-6c0370d8-fix-template-upload-size-guidance.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# 音声をテキストに変換

> **対象アプリ**：Chatflow、Workflow、Agent、チャットボット、レガシー Agent、テキストジェネレーター。

ワークスペースのデフォルト音声認識モデルを使用して、アップロードされた音声ファイルをテキストに文字起こしします。



## OpenAPI

````yaml /ja/api-reference/openapi_service.json post /audio-to-text
openapi: 3.0.1
info:
  description: >-
    Dify アプリケーションとナレッジベースのための REST API です。アプリケーション系エンドポイントはアプリの API
    キーで、ナレッジ系エンドポイントはナレッジベースの API キーで認証します。
  title: Dify サービス API
  version: 1.0.0
servers:
  - description: Dify サービス API のベース URL です。セルフホスト環境では、独自の API ベース URL に置き換えてください。
    url: https://{api_base_url}
    variables:
      api_base_url:
        default: api.dify.ai/v1
        description: API ベース URL のホストとパス（`https://` を除く）。
security:
  - ApiKeyAuth: []
tags:
  - description: チャットメッセージとインタラクションに関連する操作です。
    name: チャットメッセージ
  - description: ファイルのアップロードとプレビューの操作です。
    name: ファイル操作
  - description: エンドユーザー情報に関連する操作です。
    name: エンドユーザー
  - description: ユーザーフィードバックの操作です。
    name: メッセージフィードバック
  - description: 会話管理に関連する操作です。
    name: 会話管理
  - description: テキスト読み上げと音声認識の操作です。
    name: 音声・テキスト変換
  - description: アプリケーション設定と情報を取得する操作です。
    name: アプリケーション設定
  - description: ダイレクト返信用のアノテーション管理に関連する操作です。
    name: アノテーション管理
  - description: 人間の入力を要する一時停止中のワークフローの再開操作です。
    name: 人間の入力
  - description: ワークフローの実行と管理のための操作です。
    name: ワークフロー実行
  - description: テキスト生成に関連する操作です。
    name: 完了メッセージ
  - description: ナレッジベースの作成、設定、取得を含むナレッジベース管理の操作です。
    name: ナレッジベース
  - description: ナレッジベース内のドキュメントの作成、更新、管理のための操作です。
    name: ドキュメント
  - description: ドキュメントチャンクと子チャンクの管理のための操作です。
    name: チャンク
  - description: ナレッジベースのメタデータフィールドとドキュメントメタデータ値の管理のための操作です。
    name: メタデータ
  - description: ナレッジベースタグとタグバインディングの管理のための操作です。
    name: タグ管理
  - description: 利用可能なモデルを取得するための操作です。
    name: モデル
  - description: データソースプラグインとパイプライン実行を含むナレッジパイプラインの管理と実行のための操作です。
    name: ナレッジパイプライン
paths:
  /audio-to-text:
    post:
      tags:
        - 音声・テキスト変換
      summary: 音声をテキストに変換
      description: |-
        **対象アプリ**：Chatflow、Workflow、Agent、チャットボット、レガシー Agent、テキストジェネレーター。

        ワークスペースのデフォルト音声認識モデルを使用して、アップロードされた音声ファイルをテキストに文字起こしします。
      operationId: basicChatAudioToTextJa
      requestBody:
        content:
          multipart/form-data:
            schema:
              $ref: '#/components/schemas/AudioToTextRequest'
        required: true
      responses:
        '200':
          content:
            application/json:
              examples:
                audioToTextSuccess:
                  summary: レスポンス例
                  value:
                    text: >-
                      Hello, I would like to know more about the iPhone 13 Pro
                      Max.
              schema:
                $ref: '#/components/schemas/AudioToTextResponse'
          description: 音声からテキストへの変換に成功しました。
        '400':
          content:
            application/json:
              examples:
                completion_request_error:
                  summary: completion_request_error
                  value:
                    code: completion_request_error
                    message: Completion request failed.
                    status: 400
                no_audio_uploaded:
                  summary: no_audio_uploaded
                  value:
                    code: no_audio_uploaded
                    message: Please upload your audio.
                    status: 400
                provider_not_initialize:
                  summary: provider_not_initialize
                  value:
                    code: provider_not_initialize
                    message: >-
                      No valid model provider credentials found. Please go to
                      Settings -> Model Provider to complete your provider
                      credentials.
                    status: 400
                provider_not_support_speech_to_text:
                  summary: provider_not_support_speech_to_text
                  value:
                    code: provider_not_support_speech_to_text
                    message: Provider not support speech to text.
                    status: 400
                speech_to_text_disabled:
                  summary: speech_to_text_disabled
                  value:
                    code: speech_to_text_disabled
                    message: Speech to text is disabled.
                    status: 400
          description: |-
            - `no_audio_uploaded` : 音声ファイルがアップロードされていません（`file` フィールドがありません）。
            - `speech_to_text_disabled` : このアプリでは音声からテキストへの変換が無効になっています。
            - `provider_not_support_speech_to_text` : モデルプロバイダーが音声認識をサポートしていません。
            - `provider_not_initialize` : 有効なモデルプロバイダーの認証情報が見つかりません。
            - `completion_request_error` : 音声認識リクエストに失敗しました。
        '413':
          content:
            application/json:
              examples:
                audio_too_large:
                  summary: audio_too_large
                  value:
                    code: audio_too_large
                    message: Audio size larger than 30 mb
                    status: 413
          description: '`audio_too_large` : 音声ファイルが `30 MB` のサイズ上限を超えています。'
        '415':
          content:
            application/json:
              examples:
                unsupported_audio_type:
                  summary: unsupported_audio_type
                  value:
                    code: unsupported_audio_type
                    message: Audio type not allowed.
                    status: 415
          description: >-
            `unsupported_audio_type` : ファイルの MIME
            タイプが、受け付ける音声タイプのいずれでもありません（`file` フィールドを参照）。
        '500':
          content:
            application/json:
              examples:
                internal_server_error:
                  summary: internal_server_error
                  value:
                    code: internal_server_error
                    message: >-
                      The server encountered an internal error and was unable to
                      complete your request. Either the server is overloaded or
                      there is an error in the application.
                    status: 500
          description: '`internal_server_error` : 内部サーバーエラー。'
components:
  schemas:
    AudioToTextRequest:
      description: 音声からテキストへの変換のリクエストボディ。
      properties:
        file:
          description: >-
            文字起こしする音声ファイルです。受け付ける MIME タイプは
            `audio/mp3`、`audio/m4a`（`audio/x-m4a`
            も受け付けます）、`audio/wav`、`audio/amr`、`audio/mpga` です。それ以外のタイプ（一般的な
            `audio/mpeg` を含む）は `unsupported_audio_type` で拒否されます。最大サイズは `30 MB`
            です。
          format: binary
          type: string
        user:
          description: >-
            エンドユーザーの識別子。アプリ側で定義し、アプリ内で一意にします。[エンドユーザーの識別](/ja/api-reference/guides/end-user-identity)
            を参照してください。
          type: string
      required:
        - file
      type: object
    AudioToTextResponse:
      properties:
        text:
          description: 音声認識からの出力テキスト。
          type: string
      type: object
  securitySchemes:
    ApiKeyAuth:
      bearerFormat: API_KEY
      description: >-
        すべてのリクエストは API キーで認証します：`Authorization: Bearer
        {API_KEY}`。アプリのエンドポイントにはアプリの API キーを、ナレッジのエンドポイントにはナレッジベースの API
        キーを使用します（[Dify API クイックスタート](/ja/api-reference/guides/get-started)）。


        キーはサーバーサイドで保管し、クライアントコードには決して埋め込まないでください。キーが欠落または無効なリクエストは HTTP
        `401`（`unauthorized`）で失敗します。
      scheme: bearer
      type: http

````