Merge branch 'develop' into u2_export

4 years ago · e3298c79ce
parent 260752aa2a 5e714ecb4a
commit e3298c79ce
353 changed files with 10603 additions and 5601 deletions
--- a/.github/ISSUE_TEMPLATE/bug-report-s2t.md
+++ b/.github/ISSUE_TEMPLATE/bug-report-s2t.md
@ -1,9 +1,9 @@
 ---
-name: Bug report
+name: "\U0001F41B S2T Bug Report"
 about: Create a report to help us improve
-title: ''
-labels: ''
-assignees: ''
+title: "[S2T]XXXX"
+labels: Bug, S2T
+assignees: zh794390558

 ---

@ -27,7 +27,7 @@ A clear and concise description of what you expected to happen.
 **Screenshots**
 If applicable, add screenshots to help explain your problem.

-** Environment (please complete the following information):**
+**Environment (please complete the following information):**
 - OS: [e.g. Ubuntu]
 - GCC/G++ Version [e.g. 8.3]
 - Python Version [e.g. 3.7]
--- a/.github/ISSUE_TEMPLATE/bug-report-tts.md
+++ b/.github/ISSUE_TEMPLATE/bug-report-tts.md
@ -0,0 +1,42 @@
+---
+name: "\U0001F41B TTS Bug Report"
+about: Create a report to help us improve
+title: "[TTS]XXXX"
+labels: Bug, T2S
+assignees: yt605155624
+
+---
+
+For support and discussions, please use our [Discourse forums](https://github.com/PaddlePaddle/DeepSpeech/discussions).
+
+If you've found a bug then please create an issue with the following information:
+
+**Describe the bug**
+A clear and concise description of what the bug is.
+
+**To Reproduce**
+Steps to reproduce the behavior:
+1. Go to '...'
+2. Click on '....'
+3. Scroll down to '....'
+4. See error
+
+**Expected behavior**
+A clear and concise description of what you expected to happen.
+
+**Screenshots**
+If applicable, add screenshots to help explain your problem.
+
+**Environment (please complete the following information):**
+ - OS: [e.g. Ubuntu]
+ - GCC/G++ Version [e.g. 8.3]
+ - Python Version [e.g. 3.7]
+ - PaddlePaddle Version [e.g. 2.0.0]
+ - Model Version [e.g. 2.0.0]
+ - GPU/DRIVER Informationo [e.g. Tesla V100-SXM2-32GB/440.64.00]
+ - CUDA/CUDNN Version [e.g. cuda-10.2]
+ - MKL Version
+- TensorRT Version
+
+**Additional context**
+Add any other context about the problem here.
--- a/.github/ISSUE_TEMPLATE/feature-request.md
+++ b/.github/ISSUE_TEMPLATE/feature-request.md
@ -0,0 +1,19 @@
+---
+name: "\U0001F680 Feature Request"
+about: As a user, I want to request a New Feature on the product.
+title: ''
+labels: feature request
+assignees: D-DanielYang, iftaken
+
+---
+
+## Feature Request
+
+**Is your feature request related to a problem? Please describe:**
+<!-- A clear and concise description of what the problem is. Ex. I'm always frustrated when [...] -->
+
+**Describe the feature you'd like:**
+<!-- A clear and concise description of what you want to happen. -->
+
+**Describe alternatives you've considered:**
+<!-- A clear and concise description of any alternative solutions or features you've considered. -->
--- a/.github/ISSUE_TEMPLATE/others.md
+++ b/.github/ISSUE_TEMPLATE/others.md
@ -0,0 +1,15 @@
+---
+name: "\U0001F9E9 Others"
+about: Report any other non-support related issues.
+title: ''
+labels: ''
+assignees: ''
+
+---
+
+## Others
+
+<!--
+你可以在这里提出任何前面几类模板不适用的问题，包括但不限于：优化性建议、框架使用体验反馈、版本兼容性问题、报错信息不清楚等。
+You can report any issues that are not applicable to the previous types of templates, including but not limited to: enhancement suggestions, feedback on the use of the framework, version compatibility issues, unclear error information, etc.
+-->
--- a/.github/ISSUE_TEMPLATE/question.md
+++ b/.github/ISSUE_TEMPLATE/question.md
@ -0,0 +1,19 @@
+---
+name: "\U0001F914 Ask a Question"
+about: I want to ask a question.
+title: ''
+labels: Question
+assignees: ''
+
+---
+
+## General Question
+
+<!--
+Before asking a question, make sure you have:
+- Baidu/Google your question.
+- Searched open and closed [GitHub issues](https://github.com/PaddlePaddle/PaddleSpeech/issues?q=is%3Aissue)
+- Read the documentation:
+  - [Readme](https://github.com/PaddlePaddle/PaddleSpeech)
+  - [Doc](https://paddlespeech.readthedocs.io/)
+-->
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@ -1,66 +0,0 @@
-# Changelog
-
-Date: 2022-3-22, Author: yt605155624.
-Add features to: CLI:
-  - Support aishell3_hifigan、vctk_hifigan
-  - PRLink: https://github.com/PaddlePaddle/PaddleSpeech/pull/1587
-
-Date: 2022-3-09, Author: yt605155624.
-Add features to: T2S:
-  - Add ljspeech hifigan egs.
-  - PRLink: https://github.com/PaddlePaddle/PaddleSpeech/pull/1549
-
-Date: 2022-3-08, Author: yt605155624.
-Add features to: T2S:
-  - Add aishell3 hifigan egs.
-  - PRLink: https://github.com/PaddlePaddle/PaddleSpeech/pull/1545
-
-Date: 2022-3-08, Author: yt605155624.
-Add features to: T2S:
-  - Add vctk hifigan egs.
-  - PRLink: https://github.com/PaddlePaddle/PaddleSpeech/pull/1544
-
-Date: 2022-1-29, Author: yt605155624.
-Add features to: T2S:
-  - Update aishell3 vc0 with new Tacotron2.
-  - PRLink: https://github.com/PaddlePaddle/PaddleSpeech/pull/1419
-
-Date: 2022-1-29, Author: yt605155624.
-Add features to: T2S:
-  - Add ljspeech Tacotron2.
-  - PRLink: https://github.com/PaddlePaddle/PaddleSpeech/pull/1416
-
-Date: 2022-1-24, Author: yt605155624.
-Add features to: T2S:
-  - Add csmsc WaveRNN.
-  - PRLink: https://github.com/PaddlePaddle/PaddleSpeech/pull/1379
-
-Date: 2022-1-19, Author: yt605155624.
-Add features to: T2S:
-  - Add csmsc Tacotron2.
-  - PRLink: https://github.com/PaddlePaddle/PaddleSpeech/pull/1314
-
-
-Date: 2022-1-10, Author: Jackwaterveg.  
-Add features to: CLI:
-  - Support English (librispeech/asr1/transformer).
-  - Support choosing `decode_method` for conformer and transformer models.  
-  - Refactor the config, using the unified config.  
-  - PRLink: https://github.com/PaddlePaddle/PaddleSpeech/pull/1297
-
-***
-
-Date: 2022-1-17, Author: Jackwaterveg.  
-Add features to: CLI:
-  - Support deepspeech2 online/offline model(aishell).
-  - PRLink: https://github.com/PaddlePaddle/PaddleSpeech/pull/1356
-
-***
-
-Date: 2022-1-24, Author: Jackwaterveg.  
-Add features to: ctc_decoders:  
-  - Support online ctc prefix-beam search decoder. 
-  - Unified ctc online decoder and ctc offline decoder.  
-  - PRLink: https://github.com/PaddlePaddle/PaddleSpeech/pull/821
-
-***
--- a/MANIFEST.in
+++ b/MANIFEST.in
@ -0,0 +1,2 @@
+include paddlespeech/t2s/exps/*.txt
+include paddlespeech/t2s/frontend/*.yaml
--- a/README.md
+++ b/README.md
@ -1,4 +1,3 @@
-
 ([简体中文](./README_cn.md)|English)
 <p align="center">
  <img src="./docs/images/PaddleSpeech_logo.png" />
@ -160,15 +159,20 @@ Via the easy-to-use, efficient, flexible and scalable implementation, our vision
  - 🧩  *Cascaded models application*: as an extension of the typical traditional audio tasks, we combine the workflows of the aforementioned tasks with other fields like Natural language processing (NLP) and Computer Vision (CV).

 ### Recent Update
- 👑 2022.05.13: Release [PP-ASR](./docs/source/asr/PPASR.md)、[PP-TTS](./docs/source/tts/PPTTS.md)、[PP-VPR](docs/source/vpr/PPVPR.md)
- 👏🏻  2022.05.06: `Streaming ASR` with `Punctuation Restoration` and `Token Timestamp`.
- 👏🏻  2022.05.06: `Server` is available for `Speaker Verification`, and `Punctuation Restoration`.
- 👏🏻  2022.04.28: `Streaming Server` is available for `Automatic Speech Recognition` and `Text-to-Speech`.
- 👏🏻  2022.03.28: `Server` is available for `Audio Classification`, `Automatic Speech Recognition` and `Text-to-Speech`.
- 👏🏻  2022.03.28: `CLI` is available for `Speaker Verification`.
+- ⚡ 2022.08.25: Release TTS [finetune](./examples/other/tts_finetune/tts3) example.
+- 🔥 2022.08.22: Add ERNIE-SAT models: [ERNIE-SAT-vctk](./examples/vctk/ernie_sat)、[ERNIE-SAT-aishell3](./examples/aishell3/ernie_sat)、[ERNIE-SAT-zh_en](./examples/aishell3_vctk/ernie_sat).
+- 🔥 2022.08.15: Add [g2pW](https://github.com/GitYCC/g2pW) into TTS Chinese Text Frontend.
+- 🔥 2022.08.09: Release [Chinese English mixed TTS](./examples/zh_en_tts/tts3).
+- ⚡ 2022.08.03: Add ONNXRuntime infer for  TTS CLI.
+- 🎉 2022.07.18: Release VITS: [VITS-csmsc](./examples/csmsc/vits)、[VITS-aishell3](./examples/aishell3/vits)、[VITS-VC](./examples/aishell3/vits-vc).
+- 🎉 2022.06.22: All TTS models support ONNX format.
+- 🍀 2022.06.17: Add [PaddleSpeech Web Demo](./demos/speech_web).
+- 👑 2022.05.13: Release [PP-ASR](./docs/source/asr/PPASR.md)、[PP-TTS](./docs/source/tts/PPTTS.md)、[PP-VPR](docs/source/vpr/PPVPR.md).
+- 👏🏻  2022.05.06: `PaddleSpeech Streaming Server` is available for `Streaming ASR` with `Punctuation Restoration` and `Token Timestamp` and `Text-to-Speech`.
+- 👏🏻  2022.05.06: `PaddleSpeech Server` is available for `Audio Classification`, `Automatic Speech Recognition` and `Text-to-Speech`, `Speaker Verification` and `Punctuation Restoration`.
+- 👏🏻  2022.03.28: `PaddleSpeech CLI` is available for `Speaker Verification`.
 - 🤗  2021.12.14: [ASR](https://huggingface.co/spaces/KPatrick/PaddleSpeechASR) and [TTS](https://huggingface.co/spaces/KPatrick/PaddleSpeechTTS) Demos on Hugging Face Spaces are available!
- 👏🏻  2021.12.10: `CLI` is available for `Audio Classification`, `Automatic Speech Recognition`, `Speech Translation (English to Chinese)` and `Text-to-Speech`.
-
+- 👏🏻  2021.12.10: `PaddleSpeech CLI` is available for `Audio Classification`, `Automatic Speech Recognition`, `Speech Translation (English to Chinese)` and `Text-to-Speech`.

 ### Community
 - Scan the QR code below with your Wechat, you can access to official technical exchange group and get the bonus ( more than 20GB learning materials, such as papers, codes and videos ) and the live link of the lessons. Look forward to your participation.
@ -180,62 +184,191 @@ Via the easy-to-use, efficient, flexible and scalable implementation, our vision
 ## Installation

 We strongly recommend our users to install PaddleSpeech in **Linux** with *python>=3.7* and *paddlepaddle>=2.3.1*.
-Up to now, **Linux** supports CLI for the all our tasks, **Mac OSX** and **Windows** only supports PaddleSpeech CLI for Audio Classification, Speech-to-Text and Text-to-Speech. To install `PaddleSpeech`, please see [installation](./docs/source/install.md).
+
+### **Dependency Introduction**
+
+ gcc >= 4.8.5
+ paddlepaddle >= 2.3.1
+ python >= 3.7
+ OS support:  Linux(recommend), Windows, Mac OSX
+
+PaddleSpeech depends on paddlepaddle. For installation, please refer to the official website of [paddlepaddle](https://www.paddlepaddle.org.cn/en) and choose according to your own machine. Here is an example of the cpu version.
+
+```bash
+pip install paddlepaddle -i https://mirror.baidu.com/pypi/simple
+```
+
+There are two quick installation methods for PaddleSpeech, one is pip installation, and the other is source code compilation (recommended).
+### pip install
+
+```shell
+pip install pytest-runner
+pip install paddlespeech
+```
+
+### source code compilation
+
+```shell
+git clone https://github.com/PaddlePaddle/PaddleSpeech.git
+cd PaddleSpeech
+pip install pytest-runner
+pip install .
+```
+
+For more installation problems, such as conda environment, librosa-dependent, gcc problems, kaldi installation, etc., you can refer to this [installation document](./docs/source/install.md). If you encounter problems during installation, you can leave a message on [#2150](https://github.com/PaddlePaddle/PaddleSpeech/issues/2150) and find related problems


 <a name="quickstart"></a>
 ## Quick Start

-Developers can have a try of our models with [PaddleSpeech Command Line](./paddlespeech/cli/README.md). Change `--input` to test your own audio/text.
+Developers can have a try of our models with [PaddleSpeech Command Line](./paddlespeech/cli/README.md) or Python. Change `--input` to test your own audio/text and support 16k wav format audio.
+
+**You can also quickly experience it in AI Studio 👉🏻 [PaddleSpeech API Demo](https://aistudio.baidu.com/aistudio/projectdetail/4353348?sUid=2470186&shared=1&ts=1660876445786)**
+
+
+Test audio sample download

-**Audio Classification**     
 ```shell
-paddlespeech cls --input input.wav
+wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav
+wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/en.wav
 ```

-**Speaker Verification**
+### Automatic Speech Recognition
+
+<details><summary>&emsp;（Click to expand）Open Source Speech Recognition</summary>
+
+**command line experience**
+
+```shell
+paddlespeech asr --lang zh --input zh.wav
 ```
-paddlespeech vector --task spk --input input_16k.wav
+
+**Python API experience**
+
+```python
+>>> from paddlespeech.cli.asr.infer import ASRExecutor
+>>> asr = ASRExecutor()
+>>> result = asr(audio_file="zh.wav")
+>>> print(result)
+我认为跑步最重要的就是给我带来了身体健康
 ```
+</details>
+
+### Text-to-Speech
+
+<details><summary>&emsp;Open Source Speech Synthesis</summary>
+
+Output 24k sample rate wav format audio
+
+
+**command line experience**

-**Automatic Speech Recognition**
 ```shell
-paddlespeech asr --lang zh --input input_16k.wav
+paddlespeech tts --input "你好，欢迎使用百度飞桨深度学习框架！" --output output.wav
 ```
- web demo for Automatic Speech Recognition is integrated to [Huggingface Spaces](https://huggingface.co/spaces) with [Gradio](https://github.com/gradio-app/gradio). See Demo: [ASR Demo](https://huggingface.co/spaces/KPatrick/PaddleSpeechASR)

-**Speech Translation** (English to Chinese)
-(not support for Mac and Windows now)
+**Python API experience**
+
+```python
+>>> from paddlespeech.cli.tts.infer import TTSExecutor
+>>> tts = TTSExecutor()
+>>> tts(text="今天天气十分不错。", output="output.wav")
+```
+- You can experience in [Huggingface Spaces](https://huggingface.co/spaces) [TTS Demo](https://huggingface.co/spaces/KPatrick/PaddleSpeechTTS)
+
+</details>
+
+### Audio Classification
+
+<details><summary>&emsp;An open-domain sound classification tool</summary>
+
+Sound classification model based on 527 categories of AudioSet dataset
+
+**command line experience**
+
 ```shell
-paddlespeech st --input input_16k.wav
+paddlespeech cls --input zh.wav
 ```

-**Text-to-Speech** 
+**Python API experience**
+
+```python
+>>> from paddlespeech.cli.cls.infer import CLSExecutor
+>>> cls = CLSExecutor()
+>>> result = cls(audio_file="zh.wav")
+>>> print(result)
+Speech 0.9027186632156372
+```
+
+</details>
+
+### Voiceprint Extraction
+
+<details><summary>&emsp;Industrial-grade voiceprint extraction tool</summary>
+
+**command line experience**
+
 ```shell
-paddlespeech tts --input "你好，欢迎使用飞桨深度学习框架！" --output output.wav
+paddlespeech vector --task spk --input zh.wav
 ```
- web demo for Text to Speech is integrated to [Huggingface Spaces](https://huggingface.co/spaces) with [Gradio](https://github.com/gradio-app/gradio). See Demo: [TTS Demo](https://huggingface.co/spaces/KPatrick/PaddleSpeechTTS)

-**Text Postprocessing** 
- Punctuation Restoration
-  ```bash
-  paddlespeech text --task punc --input 今天的天气真不错啊你下午有空吗我想约你一起去吃饭
-  ```
+**Python API experience**

-**Batch Process**
+```python
+>>> from paddlespeech.cli.vector import VectorExecutor
+>>> vec = VectorExecutor()
+>>> result = vec(audio_file="zh.wav")
+>>> print(result) # 187维向量
+[ -0.19083306   9.474295   -14.122263    -2.0916545    0.04848729
+   4.9295826    1.4780062    0.3733844   10.695862     3.2697146
+  -4.48199     -0.6617882   -9.170393   -11.1568775   -1.2358263 ...]
 ```
-echo -e "1 欢迎光临。\n2 谢谢惠顾。" | paddlespeech tts
+
+</details>
+
+### Punctuation Restoration
+
+<details><summary>&emsp;Quick recovery of text punctuation, works with ASR models</summary>
+
+**command line experience**
+
+```shell
+paddlespeech text --task punc --input 今天的天气真不错啊你下午有空吗我想约你一起去吃饭
 ```

-**Shell Pipeline**   
- ASR + Punctuation Restoration
+**Python API experience**
+
+```python
+>>> from paddlespeech.cli.text.infer import TextExecutor
+>>> text_punc = TextExecutor()
+>>> result = text_punc(text="今天的天气真不错啊你下午有空吗我想约你一起去吃饭")
+今天的天气真不错啊！你下午有空吗？我想约你一起去吃饭。
 ```
-paddlespeech asr --input ./zh.wav | paddlespeech text --task punc
+
+</details>
+
+### Speech Translation
+
+<details><summary>&emsp;End-to-end English to Chinese Speech Translation Tool</summary>
+
+Use pre-compiled kaldi related tools, only support experience in Ubuntu system
+
+**command line experience**
+
+```shell
+paddlespeech st --input en.wav
 ```

-For more command lines, please see: [demos](https://github.com/PaddlePaddle/PaddleSpeech/tree/develop/demos)
+**Python API experience**
+
+```python
+>>> from paddlespeech.cli.st.infer import STExecutor
+>>> st = STExecutor()
+>>> result = st(audio_file="en.wav")
+['我 在 这栋 建筑 的 古老 门上 敲门 。']
+```

-If you want to try more functions like training and tuning, please have a look at [Speech-to-Text Quick Start](./docs/source/asr/quick_start.md) and [Text-to-Speech Quick Start](./docs/source/tts/quick_start.md).
+</details>


 <a name="quickstartserver"></a>
@ -243,10 +376,12 @@ If you want to try more functions like training and tuning, please have a look a

 Developers can have a try of our speech server with [PaddleSpeech Server Command Line](./paddlespeech/server/README.md).

+**You can try it quickly in AI Studio (recommend): [SpeechServer](https://aistudio.baidu.com/aistudio/projectdetail/4354592?sUid=2470186&shared=1&ts=1660877827034)**
+
 **Start server**     

 ```shell
-paddlespeech_server start --config_file ./paddlespeech/server/conf/application.yaml
+paddlespeech_server start --config_file ./demos/speech_server/conf/application.yaml
 ```

 **Access Speech Recognition Services**     
@ -404,7 +539,7 @@ PaddleSpeech supports a series of most popular models. They are summarized in [r
    </td>
    </tr>
    <tr>
-      <td rowspan="4">Acoustic Model</td>
+      <td rowspan="5">Acoustic Model</td>
      <td>Tacotron2</td>
      <td>LJSpeech / CSMSC</td>
      <td>
@ -427,9 +562,16 @@ PaddleSpeech supports a series of most popular models. They are summarized in [r
    </tr>
    <tr>
      <td>FastSpeech2</td>
-      <td>LJSpeech / VCTK / CSMSC / AISHELL-3</td>
+      <td>LJSpeech / VCTK / CSMSC / AISHELL-3 / ZH_EN / finetune</td>
      <td>
-      <a href = "./examples/ljspeech/tts3">fastspeech2-ljspeech</a> / <a href = "./examples/vctk/tts3">fastspeech2-vctk</a> / <a href = "./examples/csmsc/tts3">fastspeech2-csmsc</a> / <a href = "./examples/aishell3/tts3">fastspeech2-aishell3</a>
+      <a href = "./examples/ljspeech/tts3">fastspeech2-ljspeech</a> / <a href = "./examples/vctk/tts3">fastspeech2-vctk</a> / <a href = "./examples/csmsc/tts3">fastspeech2-csmsc</a> / <a href = "./examples/aishell3/tts3">fastspeech2-aishell3</a> / <a href = "./examples/zh_en_tts/tts3">fastspeech2-zh_en</a> / <a href = "./examples/other/tts_finetune/tts3">fastspeech2-finetune</a>
+      </td>
+    </tr>
+    <tr>
+      <td>ERNIE-SAT</td>
+      <td>VCTK / AISHELL-3 / ZH_EN</td>
+      <td>
+      <a href = "./examples/vctk/ernie_sat">ERNIE-SAT-vctk</a> / <a href = "./examples/aishell3/ernie_sat">ERNIE-SAT-aishell3</a> / <a href = "./examples/aishell3_vctk/ernie_sat">ERNIE-SAT-zh_en</a>
      </td>
    </tr>
   <tr>
@ -462,47 +604,61 @@ PaddleSpeech supports a series of most popular models. They are summarized in [r
      </td>
    </tr>
    <tr>
-      <td >HiFiGAN</td>
-      <td >LJSpeech / VCTK / CSMSC / AISHELL-3</td>
+      <td>HiFiGAN</td>
+      <td>LJSpeech / VCTK / CSMSC / AISHELL-3</td>
      <td>
      <a href = "./examples/ljspeech/voc5">HiFiGAN-ljspeech</a> / <a href = "./examples/vctk/voc5">HiFiGAN-vctk</a> / <a href = "./examples/csmsc/voc5">HiFiGAN-csmsc</a> / <a href = "./examples/aishell3/voc5">HiFiGAN-aishell3</a>
      </td>
    </tr>
    <tr>
-      <td >WaveRNN</td>
-      <td >CSMSC</td>
+      <td>WaveRNN</td>
+      <td>CSMSC</td>
      <td>
      <a href = "./examples/csmsc/voc6">WaveRNN-csmsc</a>
      </td>
    </tr>
    <tr>
-      <td rowspan="3">Voice Cloning</td>
+      <td rowspan="5">Voice Cloning</td>
      <td>GE2E</td>
      <td >Librispeech, etc.</td>
      <td>
-      <a href = "./examples/other/ge2e">ge2e</a>
+      <a href = "./examples/other/ge2e">GE2E</a>
      </td>
    </tr>
    <tr>
-      <td>GE2E + Tacotron2</td>
+      <td>SV2TTS (GE2E + Tacotron2)</td>
      <td>AISHELL-3</td>
      <td>
-      <a href = "./examples/aishell3/vc0">ge2e-tacotron2-aishell3</a>
+      <a href = "./examples/aishell3/vc0">VC0</a>
      </td>
    </tr>
    <tr>
-      <td>GE2E + FastSpeech2</td>
+      <td>SV2TTS (GE2E + FastSpeech2)</td>
      <td>AISHELL-3</td>
      <td>
-      <a href = "./examples/aishell3/vc1">ge2e-fastspeech2-aishell3</a>
+      <a href = "./examples/aishell3/vc1">VC1</a>
      </td>
    </tr>
-     <tr>
+    <tr>
+      <td>SV2TTS (ECAPA-TDNN + FastSpeech2)</td>
+      <td>AISHELL-3</td>
+      <td>
+      <a href = "./examples/aishell3/vc2">VC2</a>
+      </td>
+    </tr>
+    <tr>
+      <td>GE2E + VITS</td>
+      <td>AISHELL-3</td>
+      <td>
+      <a href = "./examples/aishell3/vits-vc">VITS-VC</a>
+      </td>
+    </tr>
+    <tr>
      <td rowspan="3">End-to-End</td>
      <td>VITS</td>
-      <td >CSMSC</td>
+      <td>CSMSC / AISHELL-3</td>
      <td>
-      <a href = "./examples/csmsc/vits">VITS-csmsc</a>
+      <a href = "./examples/csmsc/vits">VITS-csmsc</a> / <a href = "./examples/aishell3/vits">VITS-aishell3</a>
      </td>
    </tr>
  </tbody>
@ -662,43 +818,79 @@ You are warmly welcome to submit questions in [discussions](https://github.com/P

 ### Contributors
 <p align="center">
-<a href="https://github.com/zh794390558"><img src="https://avatars.githubusercontent.com/u/3038472?v=4" width=75 height=75></a>
-<a href="https://github.com/Jackwaterveg"><img src="https://avatars.githubusercontent.com/u/87408988?v=4" width=75 height=75></a>
-<a href="https://github.com/yt605155624"><img src="https://avatars.githubusercontent.com/u/24568452?v=4" width=75 height=75></a>
-<a href="https://github.com/kuke"><img src="https://avatars.githubusercontent.com/u/3064195?v=4" width=75 height=75></a>
-<a href="https://github.com/xinghai-sun"><img src="https://avatars.githubusercontent.com/u/7038341?v=4" width=75 height=75></a>
-<a href="https://github.com/pkuyym"><img src="https://avatars.githubusercontent.com/u/5782283?v=4" width=75 height=75></a>
-<a href="https://github.com/KPatr1ck"><img src="https://avatars.githubusercontent.com/u/22954146?v=4" width=75 height=75></a>
-<a href="https://github.com/LittleChenCc"><img src="https://avatars.githubusercontent.com/u/10339970?v=4" width=75 height=75></a>
-<a href="https://github.com/745165806"><img src="https://avatars.githubusercontent.com/u/20623194?v=4" width=75 height=75></a>
-<a href="https://github.com/Mingxue-Xu"><img src="https://avatars.githubusercontent.com/u/92848346?v=4" width=75 height=75></a>
-<a href="https://github.com/chrisxu2016"><img src="https://avatars.githubusercontent.com/u/18379485?v=4" width=75 height=75></a>
-<a href="https://github.com/lfchener"><img src="https://avatars.githubusercontent.com/u/6771821?v=4" width=75 height=75></a>
-<a href="https://github.com/luotao1"><img src="https://avatars.githubusercontent.com/u/6836917?v=4" width=75 height=75></a>
-<a href="https://github.com/wanghaoshuang"><img src="https://avatars.githubusercontent.com/u/7534971?v=4" width=75 height=75></a>
-<a href="https://github.com/gongel"><img src="https://avatars.githubusercontent.com/u/24390500?v=4" width=75 height=75></a>
-<a href="https://github.com/mmglove"><img src="https://avatars.githubusercontent.com/u/38800877?v=4" width=75 height=75></a>
-<a href="https://github.com/iclementine"><img src="https://avatars.githubusercontent.com/u/16222986?v=4" width=75 height=75></a>
-<a href="https://github.com/ZeyuChen"><img src="https://avatars.githubusercontent.com/u/1371212?v=4" width=75 height=75></a>
-<a href="https://github.com/AK391"><img src="https://avatars.githubusercontent.com/u/81195143?v=4" width=75 height=75></a>
-<a href="https://github.com/qingqing01"><img src="https://avatars.githubusercontent.com/u/7845005?v=4" width=75 height=75></a>
-<a href="https://github.com/ericxk"><img src="https://avatars.githubusercontent.com/u/4719594?v=4" width=75 height=75></a>
-<a href="https://github.com/kvinwang"><img src="https://avatars.githubusercontent.com/u/6442159?v=4" width=75 height=75></a>
-<a href="https://github.com/jiqiren11"><img src="https://avatars.githubusercontent.com/u/82639260?v=4" width=75 height=75></a>
-<a href="https://github.com/AshishKarel"><img src="https://avatars.githubusercontent.com/u/58069375?v=4" width=75 height=75></a>
-<a href="https://github.com/chesterkuo"><img src="https://avatars.githubusercontent.com/u/6285069?v=4" width=75 height=75></a>
-<a href="https://github.com/tensor-tang"><img src="https://avatars.githubusercontent.com/u/21351065?v=4" width=75 height=75></a>
-<a href="https://github.com/hysunflower"><img src="https://avatars.githubusercontent.com/u/52739577?v=4" width=75 height=75></a>  
-<a href="https://github.com/wwhu"><img src="https://avatars.githubusercontent.com/u/6081200?v=4" width=75 height=75></a>
-<a href="https://github.com/lispc"><img src="https://avatars.githubusercontent.com/u/2833376?v=4" width=75 height=75></a>
-<a href="https://github.com/jerryuhoo"><img src="https://avatars.githubusercontent.com/u/24245709?v=4" width=75 height=75></a>
-<a href="https://github.com/harisankarh"><img src="https://avatars.githubusercontent.com/u/1307053?v=4" width=75 height=75></a>
-<a href="https://github.com/Jackiexiao"><img src="https://avatars.githubusercontent.com/u/18050469?v=4" width=75 height=75></a>
-<a href="https://github.com/limpidezza"><img src="https://avatars.githubusercontent.com/u/71760778?v=4" width=75 height=75></a>
+<a href="https://github.com/zh794390558"><img src="https://avatars.githubusercontent.com/u/3038472?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/Jackwaterveg"><img src="https://avatars.githubusercontent.com/u/87408988?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/yt605155624"><img src="https://avatars.githubusercontent.com/u/24568452?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/Honei"><img src="https://avatars.githubusercontent.com/u/11361692?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/KPatr1ck"><img src="https://avatars.githubusercontent.com/u/22954146?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/kuke"><img src="https://avatars.githubusercontent.com/u/3064195?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/lym0302"><img src="https://avatars.githubusercontent.com/u/34430015?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/SmileGoat"><img src="https://avatars.githubusercontent.com/u/56786796?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/xinghai-sun"><img src="https://avatars.githubusercontent.com/u/7038341?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/pkuyym"><img src="https://avatars.githubusercontent.com/u/5782283?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/LittleChenCc"><img src="https://avatars.githubusercontent.com/u/10339970?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/qingen"><img src="https://avatars.githubusercontent.com/u/3139179?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/D-DanielYang"><img src="https://avatars.githubusercontent.com/u/23690325?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/Mingxue-Xu"><img src="https://avatars.githubusercontent.com/u/92848346?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/745165806"><img src="https://avatars.githubusercontent.com/u/20623194?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/jerryuhoo"><img src="https://avatars.githubusercontent.com/u/24245709?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/WilliamZhang06"><img src="https://avatars.githubusercontent.com/u/97937340?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/chrisxu2016"><img src="https://avatars.githubusercontent.com/u/18379485?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/iftaken"><img src="https://avatars.githubusercontent.com/u/30135920?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/lfchener"><img src="https://avatars.githubusercontent.com/u/6771821?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/BarryKCL"><img src="https://avatars.githubusercontent.com/u/48039828?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/mmglove"><img src="https://avatars.githubusercontent.com/u/38800877?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/gongel"><img src="https://avatars.githubusercontent.com/u/24390500?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/luotao1"><img src="https://avatars.githubusercontent.com/u/6836917?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/wanghaoshuang"><img src="https://avatars.githubusercontent.com/u/7534971?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/kslz"><img src="https://avatars.githubusercontent.com/u/54951765?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/JiehangXie"><img src="https://avatars.githubusercontent.com/u/51190264?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/david-95"><img src="https://avatars.githubusercontent.com/u/15189190?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/THUzyt21"><img src="https://avatars.githubusercontent.com/u/91456992?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/buchongyu2"><img src="https://avatars.githubusercontent.com/u/29157444?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/iclementine"><img src="https://avatars.githubusercontent.com/u/16222986?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/phecda-xu"><img src="https://avatars.githubusercontent.com/u/46859427?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/freeliuzc"><img src="https://avatars.githubusercontent.com/u/23568094?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/ZeyuChen"><img src="https://avatars.githubusercontent.com/u/1371212?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/ccrrong"><img src="https://avatars.githubusercontent.com/u/101700995?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/AK391"><img src="https://avatars.githubusercontent.com/u/81195143?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/qingqing01"><img src="https://avatars.githubusercontent.com/u/7845005?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/0x45f"><img src="https://avatars.githubusercontent.com/u/23097963?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/vpegasus"><img src="https://avatars.githubusercontent.com/u/22723154?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/ericxk"><img src="https://avatars.githubusercontent.com/u/4719594?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/Betterman-qs"><img src="https://avatars.githubusercontent.com/u/61459181?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/sneaxiy"><img src="https://avatars.githubusercontent.com/u/32832641?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/Doubledongli"><img src="https://avatars.githubusercontent.com/u/20540661?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/apps/dependabot"><img src="https://avatars.githubusercontent.com/in/29110?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/kvinwang"><img src="https://avatars.githubusercontent.com/u/6442159?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/chenkui164"><img src="https://avatars.githubusercontent.com/u/34813030?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/PaddleZhang"><img src="https://avatars.githubusercontent.com/u/97284124?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/billishyahao"><img src="https://avatars.githubusercontent.com/u/96406262?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/BrightXiaoHan"><img src="https://avatars.githubusercontent.com/u/25839309?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/jiqiren11"><img src="https://avatars.githubusercontent.com/u/82639260?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/ryanrussell"><img src="https://avatars.githubusercontent.com/u/523300?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/GT-ZhangAcer"><img src="https://avatars.githubusercontent.com/u/46156734?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/tensor-tang"><img src="https://avatars.githubusercontent.com/u/21351065?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/hysunflower"><img src="https://avatars.githubusercontent.com/u/52739577?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/oyjxer"><img src="https://avatars.githubusercontent.com/u/16233945?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/JamesLim-sy"><img src="https://avatars.githubusercontent.com/u/61349199?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/limpidezza"><img src="https://avatars.githubusercontent.com/u/71760778?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/windstamp"><img src="https://avatars.githubusercontent.com/u/34057289?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/AshishKarel"><img src="https://avatars.githubusercontent.com/u/58069375?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/chesterkuo"><img src="https://avatars.githubusercontent.com/u/6285069?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/YDX-2147483647"><img src="https://avatars.githubusercontent.com/u/73375426?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/AdamBear"><img src="https://avatars.githubusercontent.com/u/2288870?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/wwhu"><img src="https://avatars.githubusercontent.com/u/6081200?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/lispc"><img src="https://avatars.githubusercontent.com/u/2833376?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/harisankarh"><img src="https://avatars.githubusercontent.com/u/1307053?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/pengzhendong"><img src="https://avatars.githubusercontent.com/u/10704539?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/Jackiexiao"><img src="https://avatars.githubusercontent.com/u/18050469?s=60&v=4" width=75 height=75></a>
 </p>

 ## Acknowledgement
-
+- Many thanks to [HighCWu](https://github.com/HighCWu) for adding [VITS-aishell3](./examples/aishell3/vits) and [VITS-VC](./examples/aishell3/vits-vc) examples.
+- Many thanks to [david-95](https://github.com/david-95) improved TTS, fixed multi-punctuation bug, and contributed to multiple program and data. 
+- Many thanks to [BarryKCL](https://github.com/BarryKCL) improved TTS Chinses frontend based on [G2PW](https://github.com/GitYCC/g2pW).
 - Many thanks to [yeyupiaoling](https://github.com/yeyupiaoling)/[PPASR](https://github.com/yeyupiaoling/PPASR)/[PaddlePaddle-DeepSpeech](https://github.com/yeyupiaoling/PaddlePaddle-DeepSpeech)/[VoiceprintRecognition-PaddlePaddle](https://github.com/yeyupiaoling/VoiceprintRecognition-PaddlePaddle)/[AudioClassification-PaddlePaddle](https://github.com/yeyupiaoling/AudioClassification-PaddlePaddle) for years of attention, constructive advice and great help.
 - Many thanks to [mymagicpower](https://github.com/mymagicpower) for the Java implementation of ASR upon [short](https://github.com/mymagicpower/AIAS/tree/main/3_audio_sdks/asr_sdk) and [long](https://github.com/mymagicpower/AIAS/tree/main/3_audio_sdks/asr_long_audio_sdk) audio files.
 - Many thanks to [JiehangXie](https://github.com/JiehangXie)/[PaddleBoBo](https://github.com/JiehangXie/PaddleBoBo) for developing Virtual Uploader(VUP)/Virtual YouTuber(VTuber) with PaddleSpeech TTS function.
--- a/README_cn.md
+++ b/README_cn.md
@ -1,4 +1,3 @@
-
 (简体中文|[English](./README.md))
 <p align="center">
  <img src="./docs/images/PaddleSpeech_logo.png" />
@ -165,13 +164,37 @@
  - 🧩 级联模型应用: 作为传统语音任务的扩展，我们结合了自然语言处理、计算机视觉等任务，实现更接近实际需求的产业级应用。


-### 近期更新
+### 近期活动
+
+ ❗️重磅❗️飞桨智慧金融行业系列直播课
+✅ 覆盖智能风控、智能运维、智能营销、智能客服四大金融主流场景
+
+📆 9月6日-9月29日每周二、四19:00
+ 智慧金融行业深入洞察
+ 8节理论+实践精品直播课
+ 10+真实产业场景范例教学及实践
+ 更有免费算力+结业证书等礼品等你来拿
+扫码报名码住直播链接，与行业精英深度交流
+
+<div align="center">
+<img src="https://user-images.githubusercontent.com/30135920/188431897-a02f028f-dd13-41e8-8ff6-749468cdc850.jpg"  width = "200"  />
+</div>

+### 近期更新
+- ⚡ 2022.08.25: 发布 TTS [finetune](./examples/other/tts_finetune/tts3) 示例。
+- 🔥 2022.08.22: 新增 ERNIE-SAT 模型: [ERNIE-SAT-vctk](./examples/vctk/ernie_sat)、[ERNIE-SAT-aishell3](./examples/aishell3/ernie_sat)、[ERNIE-SAT-zh_en](./examples/aishell3_vctk/ernie_sat)。
+- 🔥 2022.08.15: 将 [g2pW](https://github.com/GitYCC/g2pW) 引入 TTS 中文文本前端。
+- 🔥 2022.08.09: 发布[中英文混合 TTS](./examples/zh_en_tts/tts3)。
+- ⚡ 2022.08.03: TTS CLI 新增 ONNXRuntime 推理方式。
+- 🎉 2022.07.18: 发布 VITS 模型: [VITS-csmsc](./examples/csmsc/vits)、[VITS-aishell3](./examples/aishell3/vits)、[VITS-VC](./examples/aishell3/vits-vc)。
+- 🎉 2022.06.22: 所有 TTS 模型支持了 ONNX 格式。
+- 🍀 2022.06.17: 新增 [PaddleSpeech 网页应用](./demos/speech_web)。
 - 👑 2022.05.13: PaddleSpeech 发布 [PP-ASR](./docs/source/asr/PPASR_cn.md) 流式语音识别系统、[PP-TTS](./docs/source/tts/PPTTS_cn.md) 流式语音合成系统、[PP-VPR](docs/source/vpr/PPVPR_cn.md) 全链路声纹识别系统
- 👏🏻 2022.05.06: PaddleSpeech Streaming Server 上线! 覆盖了语音识别（标点恢复、时间戳），和语音合成。
- 👏🏻 2022.05.06: PaddleSpeech Server 上线! 覆盖了声音分类、语音识别、语音合成、声纹识别，标点恢复。
- 👏🏻 2022.03.28: PaddleSpeech CLI 覆盖声音分类、语音识别、语音翻译（英译中）、语音合成，声纹验证。
- 🤗 2021.12.14: PaddleSpeech [ASR](https://huggingface.co/spaces/KPatrick/PaddleSpeechASR) and [TTS](https://huggingface.co/spaces/KPatrick/PaddleSpeechTTS) Demos on Hugging Face Spaces are available!
+- 👏🏻 2022.05.06: PaddleSpeech Streaming Server 上线！覆盖了语音识别（标点恢复、时间戳）和语音合成。
+- 👏🏻 2022.05.06: PaddleSpeech Server 上线！覆盖了声音分类、语音识别、语音合成、声纹识别，标点恢复。
+- 👏🏻 2022.03.28: PaddleSpeech CLI 覆盖声音分类、语音识别、语音翻译（英译中）、语音合成和声纹验证。
+- 🤗 2021.12.14: PaddleSpeech [ASR](https://huggingface.co/spaces/KPatrick/PaddleSpeechASR) 和 [TTS](https://huggingface.co/spaces/KPatrick/PaddleSpeechTTS) 可在 Hugging Face Spaces 上体验！
+- 👏🏻 2021.12.10: PaddleSpeech CLI 支持语音分类, 语音识别, 语音翻译（英译中）和语音合成。


 ### 🔥 加入技术交流群获取入群福利
@ -196,13 +219,13 @@
 + python >= 3.7
 + linux(推荐), mac, windows

-PaddleSpeech依赖于paddlepaddle，安装可以参考[paddlepaddle官网](https://www.paddlepaddle.org.cn/)，根据自己机器的情况进行选择。这里给出cpu版本示例，其它版本大家可以根据自己机器的情况进行安装。
+PaddleSpeech 依赖于 paddlepaddle，安装可以参考[ paddlepaddle 官网](https://www.paddlepaddle.org.cn/)，根据自己机器的情况进行选择。这里给出 cpu 版本示例，其它版本大家可以根据自己机器的情况进行安装。

 ```shell
 pip install paddlepaddle -i https://mirror.baidu.com/pypi/simple
 ```

-PaddleSpeech快速安装方式有两种，一种是pip安装，一种是源码编译（推荐）。
+PaddleSpeech 快速安装方式有两种，一种是 pip 安装，一种是源码编译（推荐）。

 ### pip 安装
 ```shell
@ -222,10 +245,9 @@ pip install .

 <a name="快速开始"></a>
 ## 快速开始
+安装完成后，开发者可以通过命令行或者 Python 快速开始，命令行模式下改变 `--input` 可以尝试用自己的音频或文本测试，支持 16k wav 格式音频。

-安装完成后，开发者可以通过命令行或者Python快速开始，命令行模式下改变 `--input` 可以尝试用自己的音频或文本测试，支持16k wav格式音频。
-
-你也可以在`aistudio`中快速体验 👉🏻[PaddleSpeech API Demo ](https://aistudio.baidu.com/aistudio/projectdetail/4281335?shared=1)。
+你也可以在 `aistudio` 中快速体验 👉🏻[一键预测，快速上手 Speech 开发任务](https://aistudio.baidu.com/aistudio/projectdetail/4353348?sUid=2470186&shared=1&ts=1660878142250)。

 测试音频示例下载
 ```shell
@ -281,7 +303,7 @@ Python API 一键预测

 <details><summary>&emsp;适配多场景的开放领域声音分类工具</summary>

-基于AudioSet数据集527个类别的声音分类模型
+基于 AudioSet 数据集 527 个类别的声音分类模型

 命令行一键体验

@ -350,7 +372,7 @@ Python API 一键预测

 <details><summary>&emsp;端到端英译中语音翻译工具</summary>

-使用预编译的kaldi相关工具，只支持在Ubuntu系统中体验
+使用预编译的 kaldi 相关工具，只支持在 Ubuntu 系统中体验

 命令行一键体验

@ -370,14 +392,15 @@ python API 一键预测
 </details>


-
 <a name="快速使用服务"></a>
 ## 快速使用服务
-安装完成后，开发者可以通过命令行一键启动语音识别，语音合成，音频分类三种服务。
+安装完成后，开发者可以通过命令行一键启动语音识别，语音合成，音频分类等多种服务。
+
+你可以在 AI Studio 中快速体验：[SpeechServer 一键部署](https://aistudio.baidu.com/aistudio/projectdetail/4354592?sUid=2470186&shared=1&ts=1660878208266)

 **启动服务**     
 ```shell
-paddlespeech_server start --config_file ./paddlespeech/server/conf/application.yaml
+paddlespeech_server start --config_file ./demos/speech_server/conf/application.yaml
 ```

 **访问语音识别服务**     
@ -529,7 +552,7 @@ PaddleSpeech 的 **语音合成** 主要包含三个模块：文本前端、声
    </td>
    </tr>
    <tr>
-      <td rowspan="4">声学模型</td>
+      <td rowspan="5">声学模型</td>
      <td>Tacotron2</td>
      <td>LJSpeech / CSMSC</td>
      <td>
@ -552,9 +575,16 @@ PaddleSpeech 的 **语音合成** 主要包含三个模块：文本前端、声
    </tr>
    <tr>
      <td>FastSpeech2</td>
-      <td>LJSpeech / VCTK / CSMSC / AISHELL-3</td>
+      <td>LJSpeech / VCTK / CSMSC / AISHELL-3 / ZH_EN / finetune</td>
+      <td>
+      <a href = "./examples/ljspeech/tts3">fastspeech2-ljspeech</a> / <a href = "./examples/vctk/tts3">fastspeech2-vctk</a> / <a href = "./examples/csmsc/tts3">fastspeech2-csmsc</a> / <a href = "./examples/aishell3/tts3">fastspeech2-aishell3</a> / <a href = "./examples/zh_en_tts/tts3">fastspeech2-zh_en</a> / <a href = "./examples/other/tts_finetune/tts3">fastspeech2-finetune</a>
+      </td>
+    </tr>
+    <tr>
+      <td>ERNIE-SAT</td>
+      <td>VCTK / AISHELL-3 / ZH_EN</td>
      <td>
-      <a href = "./examples/ljspeech/tts3">fastspeech2-ljspeech</a> / <a href = "./examples/vctk/tts3">fastspeech2-vctk</a> / <a href = "./examples/csmsc/tts3">fastspeech2-csmsc</a> / <a href = "./examples/aishell3/tts3">fastspeech2-aishell3</a>
+      <a href = "./examples/vctk/ernie_sat">ERNIE-SAT-vctk</a> / <a href = "./examples/aishell3/ernie_sat">ERNIE-SAT-aishell3</a> / <a href = "./examples/aishell3_vctk/ernie_sat">ERNIE-SAT-zh_en</a>
      </td>
    </tr>
   <tr>
@ -601,34 +631,47 @@ PaddleSpeech 的 **语音合成** 主要包含三个模块：文本前端、声
      </td>
    </tr>
    <tr>
-      <td rowspan="3">声音克隆</td>
+      <td rowspan="5">声音克隆</td>
      <td>GE2E</td>
      <td >Librispeech, etc.</td>
      <td>
-      <a href = "./examples/other/ge2e">ge2e</a>
+      <a href = "./examples/other/ge2e">GE2E</a>
      </td>
    </tr>
    <tr>
-      <td>GE2E + Tacotron2</td>
+      <td>SV2TTS (GE2E + Tacotron2)</td>
      <td>AISHELL-3</td>
      <td>
-      <a href = "./examples/aishell3/vc0">ge2e-tacotron2-aishell3</a>
+      <a href = "./examples/aishell3/vc0">VC0</a>
      </td>
    </tr>
    <tr>
-      <td>GE2E + FastSpeech2</td>
+      <td>SV2TTS (GE2E + FastSpeech2)</td>
      <td>AISHELL-3</td>
      <td>
-      <a href = "./examples/aishell3/vc1">ge2e-fastspeech2-aishell3</a>
+      <a href = "./examples/aishell3/vc1">VC1</a>
      </td>
    </tr>
+    <tr>
+      <td>SV2TTS (ECAPA-TDNN + FastSpeech2)</td>
+      <td>AISHELL-3</td>
+      <td>
+      <a href = "./examples/aishell3/vc2">VC2</a>
+      </td>
+    </tr>
+    <tr>
+      <td>GE2E + VITS</td>
+      <td>AISHELL-3</td>
+      <td>
+      <a href = "./examples/aishell3/vits-vc">VITS-VC</a>
+      </td>
    </tr>
     <tr>
      <td rowspan="3">端到端</td>
      <td>VITS</td>
-      <td >CSMSC</td>
+      <td>CSMSC / AISHELL-3</td>
      <td>
-      <a href = "./examples/csmsc/vits">VITS-csmsc</a>
+      <a href = "./examples/csmsc/vits">VITS-csmsc</a> / <a href = "./examples/aishell3/vits">VITS-aishell3</a>
      </td>
    </tr>
  </tbody>
@ -796,43 +839,79 @@ PaddleSpeech 的 **语音合成** 主要包含三个模块：文本前端、声

 ### 贡献者
 <p align="center">
-<a href="https://github.com/zh794390558"><img src="https://avatars.githubusercontent.com/u/3038472?v=4" width=75 height=75></a>
-<a href="https://github.com/Jackwaterveg"><img src="https://avatars.githubusercontent.com/u/87408988?v=4" width=75 height=75></a>
-<a href="https://github.com/yt605155624"><img src="https://avatars.githubusercontent.com/u/24568452?v=4" width=75 height=75></a>
-<a href="https://github.com/kuke"><img src="https://avatars.githubusercontent.com/u/3064195?v=4" width=75 height=75></a>
-<a href="https://github.com/xinghai-sun"><img src="https://avatars.githubusercontent.com/u/7038341?v=4" width=75 height=75></a>
-<a href="https://github.com/pkuyym"><img src="https://avatars.githubusercontent.com/u/5782283?v=4" width=75 height=75></a>
-<a href="https://github.com/KPatr1ck"><img src="https://avatars.githubusercontent.com/u/22954146?v=4" width=75 height=75></a>
-<a href="https://github.com/LittleChenCc"><img src="https://avatars.githubusercontent.com/u/10339970?v=4" width=75 height=75></a>
-<a href="https://github.com/745165806"><img src="https://avatars.githubusercontent.com/u/20623194?v=4" width=75 height=75></a>
-<a href="https://github.com/Mingxue-Xu"><img src="https://avatars.githubusercontent.com/u/92848346?v=4" width=75 height=75></a>
-<a href="https://github.com/chrisxu2016"><img src="https://avatars.githubusercontent.com/u/18379485?v=4" width=75 height=75></a>
-<a href="https://github.com/lfchener"><img src="https://avatars.githubusercontent.com/u/6771821?v=4" width=75 height=75></a>
-<a href="https://github.com/luotao1"><img src="https://avatars.githubusercontent.com/u/6836917?v=4" width=75 height=75></a>
-<a href="https://github.com/wanghaoshuang"><img src="https://avatars.githubusercontent.com/u/7534971?v=4" width=75 height=75></a>
-<a href="https://github.com/gongel"><img src="https://avatars.githubusercontent.com/u/24390500?v=4" width=75 height=75></a>
-<a href="https://github.com/mmglove"><img src="https://avatars.githubusercontent.com/u/38800877?v=4" width=75 height=75></a>
-<a href="https://github.com/iclementine"><img src="https://avatars.githubusercontent.com/u/16222986?v=4" width=75 height=75></a>
-<a href="https://github.com/ZeyuChen"><img src="https://avatars.githubusercontent.com/u/1371212?v=4" width=75 height=75></a>
-<a href="https://github.com/AK391"><img src="https://avatars.githubusercontent.com/u/81195143?v=4" width=75 height=75></a>
-<a href="https://github.com/qingqing01"><img src="https://avatars.githubusercontent.com/u/7845005?v=4" width=75 height=75></a>
-<a href="https://github.com/ericxk"><img src="https://avatars.githubusercontent.com/u/4719594?v=4" width=75 height=75></a>
-<a href="https://github.com/kvinwang"><img src="https://avatars.githubusercontent.com/u/6442159?v=4" width=75 height=75></a>
-<a href="https://github.com/jiqiren11"><img src="https://avatars.githubusercontent.com/u/82639260?v=4" width=75 height=75></a>
-<a href="https://github.com/AshishKarel"><img src="https://avatars.githubusercontent.com/u/58069375?v=4" width=75 height=75></a>
-<a href="https://github.com/chesterkuo"><img src="https://avatars.githubusercontent.com/u/6285069?v=4" width=75 height=75></a>
-<a href="https://github.com/tensor-tang"><img src="https://avatars.githubusercontent.com/u/21351065?v=4" width=75 height=75></a>
-<a href="https://github.com/hysunflower"><img src="https://avatars.githubusercontent.com/u/52739577?v=4" width=75 height=75></a>  
-<a href="https://github.com/wwhu"><img src="https://avatars.githubusercontent.com/u/6081200?v=4" width=75 height=75></a>
-<a href="https://github.com/lispc"><img src="https://avatars.githubusercontent.com/u/2833376?v=4" width=75 height=75></a>
-<a href="https://github.com/jerryuhoo"><img src="https://avatars.githubusercontent.com/u/24245709?v=4" width=75 height=75></a>
-<a href="https://github.com/harisankarh"><img src="https://avatars.githubusercontent.com/u/1307053?v=4" width=75 height=75></a>
-<a href="https://github.com/Jackiexiao"><img src="https://avatars.githubusercontent.com/u/18050469?v=4" width=75 height=75></a>
-<a href="https://github.com/limpidezza"><img src="https://avatars.githubusercontent.com/u/71760778?v=4" width=75 height=75></a>
+<a href="https://github.com/zh794390558"><img src="https://avatars.githubusercontent.com/u/3038472?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/Jackwaterveg"><img src="https://avatars.githubusercontent.com/u/87408988?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/yt605155624"><img src="https://avatars.githubusercontent.com/u/24568452?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/Honei"><img src="https://avatars.githubusercontent.com/u/11361692?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/KPatr1ck"><img src="https://avatars.githubusercontent.com/u/22954146?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/kuke"><img src="https://avatars.githubusercontent.com/u/3064195?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/lym0302"><img src="https://avatars.githubusercontent.com/u/34430015?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/SmileGoat"><img src="https://avatars.githubusercontent.com/u/56786796?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/xinghai-sun"><img src="https://avatars.githubusercontent.com/u/7038341?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/pkuyym"><img src="https://avatars.githubusercontent.com/u/5782283?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/LittleChenCc"><img src="https://avatars.githubusercontent.com/u/10339970?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/qingen"><img src="https://avatars.githubusercontent.com/u/3139179?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/D-DanielYang"><img src="https://avatars.githubusercontent.com/u/23690325?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/Mingxue-Xu"><img src="https://avatars.githubusercontent.com/u/92848346?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/745165806"><img src="https://avatars.githubusercontent.com/u/20623194?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/jerryuhoo"><img src="https://avatars.githubusercontent.com/u/24245709?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/WilliamZhang06"><img src="https://avatars.githubusercontent.com/u/97937340?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/chrisxu2016"><img src="https://avatars.githubusercontent.com/u/18379485?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/iftaken"><img src="https://avatars.githubusercontent.com/u/30135920?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/lfchener"><img src="https://avatars.githubusercontent.com/u/6771821?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/BarryKCL"><img src="https://avatars.githubusercontent.com/u/48039828?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/mmglove"><img src="https://avatars.githubusercontent.com/u/38800877?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/gongel"><img src="https://avatars.githubusercontent.com/u/24390500?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/luotao1"><img src="https://avatars.githubusercontent.com/u/6836917?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/wanghaoshuang"><img src="https://avatars.githubusercontent.com/u/7534971?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/kslz"><img src="https://avatars.githubusercontent.com/u/54951765?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/JiehangXie"><img src="https://avatars.githubusercontent.com/u/51190264?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/david-95"><img src="https://avatars.githubusercontent.com/u/15189190?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/THUzyt21"><img src="https://avatars.githubusercontent.com/u/91456992?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/buchongyu2"><img src="https://avatars.githubusercontent.com/u/29157444?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/iclementine"><img src="https://avatars.githubusercontent.com/u/16222986?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/phecda-xu"><img src="https://avatars.githubusercontent.com/u/46859427?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/freeliuzc"><img src="https://avatars.githubusercontent.com/u/23568094?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/ZeyuChen"><img src="https://avatars.githubusercontent.com/u/1371212?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/ccrrong"><img src="https://avatars.githubusercontent.com/u/101700995?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/AK391"><img src="https://avatars.githubusercontent.com/u/81195143?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/qingqing01"><img src="https://avatars.githubusercontent.com/u/7845005?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/0x45f"><img src="https://avatars.githubusercontent.com/u/23097963?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/vpegasus"><img src="https://avatars.githubusercontent.com/u/22723154?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/ericxk"><img src="https://avatars.githubusercontent.com/u/4719594?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/Betterman-qs"><img src="https://avatars.githubusercontent.com/u/61459181?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/sneaxiy"><img src="https://avatars.githubusercontent.com/u/32832641?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/Doubledongli"><img src="https://avatars.githubusercontent.com/u/20540661?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/apps/dependabot"><img src="https://avatars.githubusercontent.com/in/29110?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/kvinwang"><img src="https://avatars.githubusercontent.com/u/6442159?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/chenkui164"><img src="https://avatars.githubusercontent.com/u/34813030?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/PaddleZhang"><img src="https://avatars.githubusercontent.com/u/97284124?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/billishyahao"><img src="https://avatars.githubusercontent.com/u/96406262?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/BrightXiaoHan"><img src="https://avatars.githubusercontent.com/u/25839309?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/jiqiren11"><img src="https://avatars.githubusercontent.com/u/82639260?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/ryanrussell"><img src="https://avatars.githubusercontent.com/u/523300?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/GT-ZhangAcer"><img src="https://avatars.githubusercontent.com/u/46156734?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/tensor-tang"><img src="https://avatars.githubusercontent.com/u/21351065?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/hysunflower"><img src="https://avatars.githubusercontent.com/u/52739577?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/oyjxer"><img src="https://avatars.githubusercontent.com/u/16233945?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/JamesLim-sy"><img src="https://avatars.githubusercontent.com/u/61349199?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/limpidezza"><img src="https://avatars.githubusercontent.com/u/71760778?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/windstamp"><img src="https://avatars.githubusercontent.com/u/34057289?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/AshishKarel"><img src="https://avatars.githubusercontent.com/u/58069375?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/chesterkuo"><img src="https://avatars.githubusercontent.com/u/6285069?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/YDX-2147483647"><img src="https://avatars.githubusercontent.com/u/73375426?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/AdamBear"><img src="https://avatars.githubusercontent.com/u/2288870?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/wwhu"><img src="https://avatars.githubusercontent.com/u/6081200?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/lispc"><img src="https://avatars.githubusercontent.com/u/2833376?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/harisankarh"><img src="https://avatars.githubusercontent.com/u/1307053?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/pengzhendong"><img src="https://avatars.githubusercontent.com/u/10704539?s=60&v=4" width=75 height=75></a>
+<a href="https://github.com/Jackiexiao"><img src="https://avatars.githubusercontent.com/u/18050469?s=60&v=4" width=75 height=75></a>
 </p>

 ## 致谢
-
+- 非常感谢 [HighCWu](https://github.com/HighCWu) 新增 [VITS-aishell3](./examples/aishell3/vits) 和 [VITS-VC](./examples/aishell3/vits-vc) 代码示例。
+- 非常感谢 [david-95](https://github.com/david-95) 修复句尾多标点符号出错的问题，贡献补充多条程序和数据。
+- 非常感谢 [BarryKCL](https://github.com/BarryKCL) 基于 [G2PW](https://github.com/GitYCC/g2pW) 对 TTS 中文文本前端的优化。
 - 非常感谢 [yeyupiaoling](https://github.com/yeyupiaoling)/[PPASR](https://github.com/yeyupiaoling/PPASR)/[PaddlePaddle-DeepSpeech](https://github.com/yeyupiaoling/PaddlePaddle-DeepSpeech)/[VoiceprintRecognition-PaddlePaddle](https://github.com/yeyupiaoling/VoiceprintRecognition-PaddlePaddle)/[AudioClassification-PaddlePaddle](https://github.com/yeyupiaoling/AudioClassification-PaddlePaddle) 多年来的关注和建议，以及在诸多问题上的帮助。
 - 非常感谢 [mymagicpower](https://github.com/mymagicpower) 采用PaddleSpeech 对 ASR 的[短语音](https://github.com/mymagicpower/AIAS/tree/main/3_audio_sdks/asr_sdk)及[长语音](https://github.com/mymagicpower/AIAS/tree/main/3_audio_sdks/asr_long_audio_sdk)进行 Java 实现。
 - 非常感谢 [JiehangXie](https://github.com/JiehangXie)/[PaddleBoBo](https://github.com/JiehangXie/PaddleBoBo) 采用 PaddleSpeech 语音合成功能实现 Virtual Uploader(VUP)/Virtual YouTuber(VTuber) 虚拟主播。
--- a/demos/audio_searching/README.md
+++ b/demos/audio_searching/README.md
@ -226,6 +226,12 @@ recall and elapsed time statistics are shown in the following figure：

 The retrieval framework based on Milvus takes about 2.9 milliseconds to retrieve on the premise of 90% recall rate, and it takes about 500 milliseconds for feature extraction (testing audio takes about 5 seconds), that is, a single audio test takes about 503 milliseconds in total, which can meet most application scenarios.

+* compute embeding takes 500 ms
+* retrieval with cosine takes 2.9 ms
+* total takes 503 ms
+
+> test audio is 5 sec
+
 ### 6.Pretrained Models

 Here is a list of pretrained models released by PaddleSpeech :
--- a/demos/audio_searching/src/operations/load.py
+++ b/demos/audio_searching/src/operations/load.py
@ -26,8 +26,9 @@ def get_audios(path):
    """
    supported_formats = [".wav", ".mp3", ".ogg", ".flac", ".m4a"]
    return [
-        item for sublist in [[os.path.join(dir, file) for file in files]
-                             for dir, _, files in list(os.walk(path))]
+        item
+        for sublist in [[os.path.join(dir, file) for file in files]
+                        for dir, _, files in list(os.walk(path))]
        for item in sublist if os.path.splitext(item)[1] in supported_formats
    ]

--- a/demos/metaverse/README.md
+++ b/demos/metaverse/README.md
@ -1,3 +1,5 @@
+([简体中文](./README_cn.md)|English)
+
 # Metaverse
 ## Introduction
 Metaverse is a new Internet application and social form integrating virtual reality produced by integrating a variety of new technologies. 
--- a/demos/metaverse/README_cn.md
+++ b/demos/metaverse/README_cn.md
@ -0,0 +1,27 @@
+(简体中文|[English](./README.md))
+
+# Metaverse
+
+## 简介
+
+Metaverse 是一种新的互联网应用和社交形式，融合了多种新技术，产生了虚拟现实。
+
+这个演示是一个让图片中的名人“说话”的实现。通过 `PaddleSpeech` 的 `TTS` 模块和 `PaddleGAN` 的组合，我们集成了安装和特定模块到一个 shell 脚本中。
+
+## 使用
+
+您可以使用 `PaddleSpeech` 的 `TTS` 模块和 `PaddleGAN` 让您最喜欢的人说出指定的内容，并构建您的虚拟人。
+
+运行 `run.sh` 完成所有基本程序，包括安装。
+
+```bash
+./run.sh
+```
+
+在 `run.sh`, 先会执行 `source path.sh` 来设置好环境变量。
+
+如果您想尝试您的句子，请替换 `sentences.txt` 中的句子。
+
+如果您想尝试图像，请将图像替换 shell 脚本中的 `download/Lamarr.png` 。
+
+结果已显示在我们的 [notebook](https://github.com/PaddlePaddle/PaddleSpeech/blob/develop/docs/tutorial/tts/tts_tutorial.ipynb)。
--- a/demos/speaker_verification/README.md
+++ b/demos/speaker_verification/README.md
@ -19,6 +19,7 @@ The input of this cli demo should be a WAV file(`.wav`), and the sample rate mus
 Here are sample files for this demo that can be downloaded:
 ```bash
 wget -c https://paddlespeech.bj.bcebos.com/vector/audio/85236145389.wav
+wget -c https://paddlespeech.bj.bcebos.com/vector/audio/123456789.wav
 ```

 ### 3. Usage
--- a/demos/speaker_verification/README_cn.md
+++ b/demos/speaker_verification/README_cn.md
@ -19,6 +19,7 @@
 ```bash
 # 该音频的内容是数字串 85236145389
 wget -c https://paddlespeech.bj.bcebos.com/vector/audio/85236145389.wav
+wget -c https://paddlespeech.bj.bcebos.com/vector/audio/123456789.wav
 ```
 ### 3. 使用方法
 - 命令行 (推荐使用)
--- a/demos/speech_server/README.md
+++ b/demos/speech_server/README.md
@ -26,14 +26,6 @@ At present, the speech tasks integrated by the service include: asr (speech reco
 Currently the engine type supports two forms: python and inference (Paddle Inference)
 **Note:** If the service can be started normally in the container, but the client access IP is unreachable, you can try to replace the `host` address in the configuration file with the local IP address.

-
-The input of  ASR client demo should be a WAV file(`.wav`), and the sample rate must be the same as the model.
-
-Here are sample files for thisASR client demo that can be downloaded:
-```bash
-wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespeech.bj.bcebos.com/PaddleAudio/en.wav
-```
-
 ### 3. Server Usage
 - Command Line (Recommended)

@ -86,6 +78,15 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee


 ### 4. ASR Client Usage
+
+The input of  ASR client demo should be a WAV file(`.wav`), and the sample rate must be the same as the model.
+
+Here are sample files for this ASR client demo that can be downloaded:
+```bash
+wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav
+wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/en.wav
+```
+
 **Note:** The response time will be slightly longer when using the client for the first time
 - Command Line (Recommended)

@ -110,15 +111,13 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee

  Output:
  ```text
-  [2022-02-23 18:11:22,819] [    INFO] - {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'transcription': '我认为跑步最重要的就是给我带来了身体健康'}}
-  [2022-02-23 18:11:22,820] [    INFO] - time cost 0.689145 s.
-
+  [2022-08-01 07:54:01,646] [    INFO] - ASR result: 我认为跑步最重要的就是给我带来了身体健康
+  [2022-08-01 07:54:01,646] [    INFO] - Response time 4.898965 s.
  ```

 - Python API
  ```python
  from paddlespeech.server.bin.paddlespeech_client import ASRClientExecutor
-  import json

  asrclient_executor = ASRClientExecutor()
  res = asrclient_executor(
@ -128,12 +127,11 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee
      sample_rate=16000,
      lang="zh_cn",
      audio_format="wav")
-  print(res.json())
+  print(res)
  ```
-
  Output:
  ```text
-  {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'transcription': '我认为跑步最重要的就是给我带来了身体健康'}}
+  我认为跑步最重要的就是给我带来了身体健康
  ```
 
 ### 5. TTS Client Usage
@ -162,7 +160,6 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee

    Output:
    ```text
-    [2022-02-23 15:20:37,875] [    INFO] - {'description': 'success.'}
    [2022-02-23 15:20:37,875] [    INFO] - Save synthesized audio successfully on output.wav.
    [2022-02-23 15:20:37,875] [    INFO] - Audio duration: 3.612500 s.
    [2022-02-23 15:20:37,875] [    INFO] - Response time: 0.348050 s.
@ -198,6 +195,12 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee
  ```

 ### 6. CLS Client Usage
+
+Here are sample files for this CLS Client demo that can be downloaded:
+```bash
+wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav 
+```
+
 **Note:** The response time will be slightly longer when using the client for the first time
 - Command Line (Recommended)

@ -246,6 +249,12 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee

 ### 7. Speaker Verification Client Usage

+Here are sample files for this Speaker Verification Client demo that can be downloaded:
+```bash
+wget -c https://paddlespeech.bj.bcebos.com/vector/audio/85236145389.wav
+wget -c https://paddlespeech.bj.bcebos.com/vector/audio/123456789.wav
+```
+
 #### 7.1 Extract speaker embedding
 **Note:** The response time will be slightly longer when using the client for the first time
 - Command Line (Recommended)
@ -273,18 +282,18 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee
  Output:

  ```text
-  [2022-05-25 12:25:36,165] [    INFO] - vector http client start
-  [2022-05-25 12:25:36,165] [    INFO] - the input audio: 85236145389.wav
-  [2022-05-25 12:25:36,165] [    INFO] - endpoint: http://127.0.0.1:8790/paddlespeech/vector
-  [2022-05-25 12:25:36,166] [    INFO] - http://127.0.0.1:8790/paddlespeech/vector
-  [2022-05-25 12:25:36,324] [    INFO] - The vector: {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'vec': [-1.3251205682754517, 7.860682487487793, -4.620625972747803, 0.3000721037387848, 2.2648534774780273, -1.1931440830230713, 3.064713716506958, 7.673594951629639, -6.004472732543945, -12.024259567260742, -1.9496068954467773, 3.126953601837158, 1.6188379526138306, -7.638310432434082, -1.2299772500991821, -12.33833122253418, 2.1373026371002197, -5.395712375640869, 9.717328071594238, 5.675230503082275, 3.7805123329162598, 3.0597171783447266, 3.429692029953003, 8.9760103225708, 13.174124717712402, -0.5313228368759155, 8.942471504211426, 4.465109825134277, -4.426247596740723, -9.726503372192383, 8.399328231811523, 7.223917484283447, -7.435853958129883, 2.9441683292388916, -4.343039512634277, -13.886964797973633, -1.6346734762191772, -10.902740478515625, -5.311244964599609, 3.800722122192383, 3.897603750228882, -2.123077392578125, -2.3521194458007812, 4.151031017303467, -7.404866695404053, 0.13911646604537964, 2.4626107215881348, 4.96645450592041, 0.9897574186325073, 5.483975410461426, -3.3574001789093018, 10.13400650024414, -0.6120170950889587, -10.403095245361328, 4.600754261016846, 16.009349822998047, -7.78369140625, -4.194530487060547, -6.93686056137085, 1.1789555549621582, 11.490800857543945, 4.23802375793457, 9.550930976867676, 8.375045776367188, 7.508914470672607, -0.6570729613304138, -0.3005157709121704, 2.8406054973602295, 3.0828027725219727, 0.7308170199394226, 6.1483540534973145, 0.1376611888408661, -13.424735069274902, -7.746140480041504, -2.322798252105713, -8.305252075195312, 2.98791241645813, -10.99522876739502, 0.15211068093776703, -2.3820347785949707, -1.7984174489974976, 8.49562931060791, -5.852236747741699, -3.755497932434082, 0.6989710927009583, -5.270299434661865, -2.6188621520996094, -1.8828465938568115, -4.6466498374938965, 14.078543663024902, -0.5495333075523376, 10.579157829284668, -3.216050148010254, 9.349003791809082, -4.381077766418457, -11.675816535949707, -2.863020658493042, 4.5721755027771, 2.246612071990967, -4.574341773986816, 1.8610187768936157, 2.3767874240875244, 5.625787734985352, -9.784077644348145, 0.6496725678443909, -1.457950472831726, 0.4263263940811157, -4.921126365661621, -2.4547839164733887, 3.4869801998138428, -0.4265422224998474, 8.341268539428711, 1.356552004814148, 7.096688270568848, -13.102828979492188, 8.01673412322998, -7.115934371948242, 1.8699780702590942, 0.20872099697589874, 14.699383735656738, -1.0252779722213745, -2.6107232570648193, -2.5082311630249023, 8.427192687988281, 6.913852691650391, -6.29124641418457, 0.6157366037368774, 2.489687919616699, -3.4668266773223877, 9.92176342010498, 11.200815200805664, -0.19664029777050018, 7.491600513458252, -0.6231271624565125, -0.2584814429283142, -9.947997093200684, -0.9611040949821472, 1.1649218797683716, -2.1907122135162354, -1.502848744392395, -0.5192610621452332, 15.165953636169434, 2.4649462699890137, -0.998044490814209, 7.44166374206543, -2.0768048763275146, 3.5896823406219482, -7.305543422698975, -7.562084674835205, 4.32333517074585, 0.08044180274009705, -6.564010143280029, -2.314805269241333, -1.7642345428466797, -2.470881700515747, -7.6756181716918945, -9.548877716064453, -1.017755389213562, 0.1698644608259201, 2.5877134799957275, -1.8752295970916748, -0.36614322662353516, -6.049378395080566, -2.3965611457824707, -5.945338726043701, 0.9424033164978027, -13.155974388122559, -7.45780086517334, 0.14658108353614807, -3.7427968978881836, 5.841492652893066, -1.2872905731201172, 5.569431304931641, 12.570590019226074, 1.0939218997955322, 2.2142086029052734, 1.9181575775146484, 6.991420745849609, -5.888138771057129, 3.1409823894500732, -2.0036280155181885, 2.4434285163879395, 9.973138809204102, 5.036680221557617, 2.005120277404785, 2.861560344696045, 5.860223770141602, 2.917618751525879, -1.63111412525177, 2.0292205810546875, -4.070415019989014, -6.831437110900879]}}
-  [2022-05-25 12:25:36,324] [    INFO] - Response time 0.159053 s.
+  [2022-08-01 09:01:22,151] [    INFO] - vector http client start
+  [2022-08-01 09:01:22,152] [    INFO] - the input audio: 85236145389.wav
+  [2022-08-01 09:01:22,152] [    INFO] - endpoint: http://127.0.0.1:8090/paddlespeech/vector
+  [2022-08-01 09:01:27,093] [    INFO] - {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'vec': [1.4217487573623657, 5.626248836517334, -5.342073440551758, 1.177390217781067, 3.308061122894287, 1.7565997838974, 5.1678876876831055, 10.806346893310547, -3.822679042816162, -5.614130973815918, 2.6238481998443604, -0.8072965741157532, 1.963512659072876, -7.312864780426025, 0.011034967377781868, -9.723127365112305, 0.661963164806366, -6.976816654205322, 10.213465690612793, 7.494767189025879, 2.9105641841888428, 3.894925117492676, 3.7999846935272217, 7.106173992156982, 16.905324935913086, -7.149376392364502, 8.733112335205078, 3.423002004623413, -4.831653118133545, -11.403371810913086, 11.232216835021973, 7.127464771270752, -4.282831192016602, 2.4523589611053467, -5.13075065612793, -18.17765998840332, -2.611666440963745, -11.00034236907959, -6.731431007385254, 1.6564655303955078, 0.7618184685707092, 1.1253058910369873, -2.0838277339935303, 4.725739002227783, -8.782590866088867, -3.5398736000061035, 3.8142387866973877, 5.142062664031982, 2.162053346633911, 4.09642219543457, -6.416221618652344, 12.747454643249512, 1.9429889917373657, -15.152948379516602, 6.417416572570801, 16.097013473510742, -9.716649055480957, -1.9920448064804077, -3.364956855773926, -1.8719490766525269, 11.567351341247559, 3.6978795528411865, 11.258269309997559, 7.442364692687988, 9.183405876159668, 4.528151512145996, -1.2417811155319214, 4.395910263061523, 6.672768592834473, 5.889888763427734, 7.627115249633789, -0.6692016124725342, -11.889703750610352, -9.208883285522461, -7.427401542663574, -3.777655601501465, 6.917237758636475, -9.848749160766602, -2.094479560852051, -5.1351189613342285, 0.49564215540885925, 9.317541122436523, -5.9141845703125, -1.809845209121704, -0.11738205701112747, -7.169270992279053, -1.0578246116638184, -5.721685886383057, -5.117387294769287, 16.137670516967773, -4.473618984222412, 7.66243314743042, -0.5538089871406555, 9.631582260131836, -6.470466613769531, -8.54850959777832, 4.371622085571289, -0.7970349192619324, 4.479003429412842, -2.9758646488189697, 3.2721707820892334, 2.8382749557495117, 5.1345953941345215, -9.19078254699707, -0.5657423138618469, -4.874573230743408, 2.316561460494995, -5.984307289123535, -2.1798791885375977, 0.35541653633117676, -0.3178458511829376, 9.493547439575195, 2.114448070526123, 4.358088493347168, -12.089820861816406, 8.451695442199707, -7.925461769104004, 4.624246120452881, 4.428938388824463, 18.691999435424805, -2.620460033416748, -5.149182319641113, -0.3582168221473694, 8.488557815551758, 4.98148250579834, -9.326834678649902, -2.2544236183166504, 6.64176607131958, 1.2119656801223755, 10.977132797241211, 16.55504035949707, 3.323848247528076, 9.55185317993164, -1.6677050590515137, -0.7953923940658569, -8.605660438537598, -0.4735637903213501, 2.6741855144500732, -5.359188079833984, -2.6673784255981445, 0.6660736799240112, 15.443212509155273, 4.740597724914551, -3.4725306034088135, 11.592561721801758, -2.05450701713562, 1.7361239194869995, -8.26533031463623, -9.304476737976074, 5.406835079193115, -1.5180232524871826, -7.746610641479492, -6.089605331420898, 0.07112561166286469, -0.34904858469963074, -8.649889945983887, -9.998958587646484, -2.5648481845855713, -0.5399898886680603, 2.6018145084381104, -0.31927648186683655, -1.8815231323242188, -2.0721378326416016, -3.4105639457702637, -8.299802780151367, 1.4836379289627075, -15.366002082824707, -8.288193702697754, 3.884773015975952, -3.4876506328582764, 7.362995624542236, 0.4657321572303772, 3.1326000690460205, 12.438883781433105, -1.8337029218673706, 4.532927513122559, 2.726433277130127, 10.145345687866211, -6.521956920623779, 2.8971481323242188, -3.3925881385803223, 5.079156398773193, 7.759725093841553, 4.677562236785889, 5.8457818031311035, 2.4023921489715576, 7.707108974456787, 3.9711389541625977, -6.390035152435303, 6.126871109008789, -3.776031017303467, -11.118141174316406]}}
+  [2022-08-01 09:01:27,094] [    INFO] - Response time 4.941739 s.
  ```

 * Python API

  ``` python
  from paddlespeech.server.bin.paddlespeech_client import VectorClientExecutor
+  import json

  vectorclient_executor = VectorClientExecutor()
  res = vectorclient_executor(
@ -292,13 +301,13 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee
      server_ip="127.0.0.1",
      port=8090,
      task="spk")
-  print(res)
+  print(res.json())
  ```

  Output:

  ```text
-  {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'vec': [-1.3251205682754517, 7.860682487487793, -4.620625972747803, 0.3000721037387848, 2.2648534774780273, -1.1931440830230713, 3.064713716506958, 7.673594951629639, -6.004472732543945, -12.024259567260742, -1.9496068954467773, 3.126953601837158, 1.6188379526138306, -7.638310432434082, -1.2299772500991821, -12.33833122253418, 2.1373026371002197, -5.395712375640869, 9.717328071594238, 5.675230503082275, 3.7805123329162598, 3.0597171783447266, 3.429692029953003, 8.9760103225708, 13.174124717712402, -0.5313228368759155, 8.942471504211426, 4.465109825134277, -4.426247596740723, -9.726503372192383, 8.399328231811523, 7.223917484283447, -7.435853958129883, 2.9441683292388916, -4.343039512634277, -13.886964797973633, -1.6346734762191772, -10.902740478515625, -5.311244964599609, 3.800722122192383, 3.897603750228882, -2.123077392578125, -2.3521194458007812, 4.151031017303467, -7.404866695404053, 0.13911646604537964, 2.4626107215881348, 4.96645450592041, 0.9897574186325073, 5.483975410461426, -3.3574001789093018, 10.13400650024414, -0.6120170950889587, -10.403095245361328, 4.600754261016846, 16.009349822998047, -7.78369140625, -4.194530487060547, -6.93686056137085, 1.1789555549621582, 11.490800857543945, 4.23802375793457, 9.550930976867676, 8.375045776367188, 7.508914470672607, -0.6570729613304138, -0.3005157709121704, 2.8406054973602295, 3.0828027725219727, 0.7308170199394226, 6.1483540534973145, 0.1376611888408661, -13.424735069274902, -7.746140480041504, -2.322798252105713, -8.305252075195312, 2.98791241645813, -10.99522876739502, 0.15211068093776703, -2.3820347785949707, -1.7984174489974976, 8.49562931060791, -5.852236747741699, -3.755497932434082, 0.6989710927009583, -5.270299434661865, -2.6188621520996094, -1.8828465938568115, -4.6466498374938965, 14.078543663024902, -0.5495333075523376, 10.579157829284668, -3.216050148010254, 9.349003791809082, -4.381077766418457, -11.675816535949707, -2.863020658493042, 4.5721755027771, 2.246612071990967, -4.574341773986816, 1.8610187768936157, 2.3767874240875244, 5.625787734985352, -9.784077644348145, 0.6496725678443909, -1.457950472831726, 0.4263263940811157, -4.921126365661621, -2.4547839164733887, 3.4869801998138428, -0.4265422224998474, 8.341268539428711, 1.356552004814148, 7.096688270568848, -13.102828979492188, 8.01673412322998, -7.115934371948242, 1.8699780702590942, 0.20872099697589874, 14.699383735656738, -1.0252779722213745, -2.6107232570648193, -2.5082311630249023, 8.427192687988281, 6.913852691650391, -6.29124641418457, 0.6157366037368774, 2.489687919616699, -3.4668266773223877, 9.92176342010498, 11.200815200805664, -0.19664029777050018, 7.491600513458252, -0.6231271624565125, -0.2584814429283142, -9.947997093200684, -0.9611040949821472, 1.1649218797683716, -2.1907122135162354, -1.502848744392395, -0.5192610621452332, 15.165953636169434, 2.4649462699890137, -0.998044490814209, 7.44166374206543, -2.0768048763275146, 3.5896823406219482, -7.305543422698975, -7.562084674835205, 4.32333517074585, 0.08044180274009705, -6.564010143280029, -2.314805269241333, -1.7642345428466797, -2.470881700515747, -7.6756181716918945, -9.548877716064453, -1.017755389213562, 0.1698644608259201, 2.5877134799957275, -1.8752295970916748, -0.36614322662353516, -6.049378395080566, -2.3965611457824707, -5.945338726043701, 0.9424033164978027, -13.155974388122559, -7.45780086517334, 0.14658108353614807, -3.7427968978881836, 5.841492652893066, -1.2872905731201172, 5.569431304931641, 12.570590019226074, 1.0939218997955322, 2.2142086029052734, 1.9181575775146484, 6.991420745849609, -5.888138771057129, 3.1409823894500732, -2.0036280155181885, 2.4434285163879395, 9.973138809204102, 5.036680221557617, 2.005120277404785, 2.861560344696045, 5.860223770141602, 2.917618751525879, -1.63111412525177, 2.0292205810546875, -4.070415019989014, -6.831437110900879]}}
+  {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'vec': [1.4217487573623657, 5.626248836517334, -5.342073440551758, 1.177390217781067, 3.308061122894287, 1.7565997838974, 5.1678876876831055, 10.806346893310547, -3.822679042816162, -5.614130973815918, 2.6238481998443604, -0.8072965741157532, 1.963512659072876, -7.312864780426025, 0.011034967377781868, -9.723127365112305, 0.661963164806366, -6.976816654205322, 10.213465690612793, 7.494767189025879, 2.9105641841888428, 3.894925117492676, 3.7999846935272217, 7.106173992156982, 16.905324935913086, -7.149376392364502, 8.733112335205078, 3.423002004623413, -4.831653118133545, -11.403371810913086, 11.232216835021973, 7.127464771270752, -4.282831192016602, 2.4523589611053467, -5.13075065612793, -18.17765998840332, -2.611666440963745, -11.00034236907959, -6.731431007385254, 1.6564655303955078, 0.7618184685707092, 1.1253058910369873, -2.0838277339935303, 4.725739002227783, -8.782590866088867, -3.5398736000061035, 3.8142387866973877, 5.142062664031982, 2.162053346633911, 4.09642219543457, -6.416221618652344, 12.747454643249512, 1.9429889917373657, -15.152948379516602, 6.417416572570801, 16.097013473510742, -9.716649055480957, -1.9920448064804077, -3.364956855773926, -1.8719490766525269, 11.567351341247559, 3.6978795528411865, 11.258269309997559, 7.442364692687988, 9.183405876159668, 4.528151512145996, -1.2417811155319214, 4.395910263061523, 6.672768592834473, 5.889888763427734, 7.627115249633789, -0.6692016124725342, -11.889703750610352, -9.208883285522461, -7.427401542663574, -3.777655601501465, 6.917237758636475, -9.848749160766602, -2.094479560852051, -5.1351189613342285, 0.49564215540885925, 9.317541122436523, -5.9141845703125, -1.809845209121704, -0.11738205701112747, -7.169270992279053, -1.0578246116638184, -5.721685886383057, -5.117387294769287, 16.137670516967773, -4.473618984222412, 7.66243314743042, -0.5538089871406555, 9.631582260131836, -6.470466613769531, -8.54850959777832, 4.371622085571289, -0.7970349192619324, 4.479003429412842, -2.9758646488189697, 3.2721707820892334, 2.8382749557495117, 5.1345953941345215, -9.19078254699707, -0.5657423138618469, -4.874573230743408, 2.316561460494995, -5.984307289123535, -2.1798791885375977, 0.35541653633117676, -0.3178458511829376, 9.493547439575195, 2.114448070526123, 4.358088493347168, -12.089820861816406, 8.451695442199707, -7.925461769104004, 4.624246120452881, 4.428938388824463, 18.691999435424805, -2.620460033416748, -5.149182319641113, -0.3582168221473694, 8.488557815551758, 4.98148250579834, -9.326834678649902, -2.2544236183166504, 6.64176607131958, 1.2119656801223755, 10.977132797241211, 16.55504035949707, 3.323848247528076, 9.55185317993164, -1.6677050590515137, -0.7953923940658569, -8.605660438537598, -0.4735637903213501, 2.6741855144500732, -5.359188079833984, -2.6673784255981445, 0.6660736799240112, 15.443212509155273, 4.740597724914551, -3.4725306034088135, 11.592561721801758, -2.05450701713562, 1.7361239194869995, -8.26533031463623, -9.304476737976074, 5.406835079193115, -1.5180232524871826, -7.746610641479492, -6.089605331420898, 0.07112561166286469, -0.34904858469963074, -8.649889945983887, -9.998958587646484, -2.5648481845855713, -0.5399898886680603, 2.6018145084381104, -0.31927648186683655, -1.8815231323242188, -2.0721378326416016, -3.4105639457702637, -8.299802780151367, 1.4836379289627075, -15.366002082824707, -8.288193702697754, 3.884773015975952, -3.4876506328582764, 7.362995624542236, 0.4657321572303772, 3.1326000690460205, 12.438883781433105, -1.8337029218673706, 4.532927513122559, 2.726433277130127, 10.145345687866211, -6.521956920623779, 2.8971481323242188, -3.3925881385803223, 5.079156398773193, 7.759725093841553, 4.677562236785889, 5.8457818031311035, 2.4023921489715576, 7.707108974456787, 3.9711389541625977, -6.390035152435303, 6.126871109008789, -3.776031017303467, -11.118141174316406]}}
  ```

 #### 7.2 Get the score between speaker audio embedding
@ -330,18 +339,18 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee
  Output:

  ```text
-  [2022-05-25 12:33:24,527] [    INFO] - vector score http client start
-  [2022-05-25 12:33:24,527] [    INFO] - enroll audio: 85236145389.wav, test audio: 123456789.wav
-  [2022-05-25 12:33:24,528] [    INFO] - endpoint: http://127.0.0.1:8790/paddlespeech/vector/score
-  [2022-05-25 12:33:24,695] [    INFO] - The vector score is: {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'score': 0.45332613587379456}}
-  [2022-05-25 12:33:24,696] [    INFO] - The vector: {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'score': 0.45332613587379456}}
-  [2022-05-25 12:33:24,696] [    INFO] - Response time 0.168271 s.
+  [2022-08-01 09:04:42,275] [    INFO] - vector score http client start
+  [2022-08-01 09:04:42,275] [    INFO] - enroll audio: 85236145389.wav, test audio: 123456789.wav
+  [2022-08-01 09:04:42,275] [    INFO] - endpoint: http://127.0.0.1:8090/paddlespeech/vector/score
+  [2022-08-01 09:04:44,611] [    INFO] - {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'score': 0.4292638897895813}}
+  [2022-08-01 09:04:44,611] [    INFO] - Response time 2.336258 s.
  ```

 * Python API

  ``` python 
  from paddlespeech.server.bin.paddlespeech_client import VectorClientExecutor
+  import json

  vectorclient_executor = VectorClientExecutor()
  res = vectorclient_executor(
@ -351,17 +360,13 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee
      server_ip="127.0.0.1",
      port=8090,
      task="score")
-  print(res)
+  print(res.json())
  ```

  Output:

  ```text
-  [2022-05-25 12:30:14,143] [    INFO] - vector score http client start
-  [2022-05-25 12:30:14,143] [    INFO] - enroll audio: 85236145389.wav, test audio: 123456789.wav
-  [2022-05-25 12:30:14,143] [    INFO] - endpoint: http://127.0.0.1:8790/paddlespeech/vector/score
-  [2022-05-25 12:30:14,363] [    INFO] - The vector score is: {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'score': 0.45332613587379456}}
-  {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'score': 0.45332613587379456}}
+  {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'score': 0.4292638897895813}}
  ```

 ### 8. Punctuation prediction
@ -402,7 +407,6 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee
      server_ip="127.0.0.1",
      port=8090,)
  print(res)
-
  ```

  Output:
--- a/demos/speech_server/README_cn.md
+++ b/demos/speech_server/README_cn.md
@ -29,14 +29,6 @@
 目前引擎类型支持两种形式：python 及 inference (Paddle Inference)
 **注意：** 如果在容器里可正常启动服务，但客户端访问 ip 不可达，可尝试将配置文件中 `host` 地址换成本地 ip 地址。

-
-ASR client 的输入是一个 WAV 文件（`.wav`），并且采样率必须与模型的采样率相同。
-
-可以下载此 ASR client 的示例音频：
-```bash
-wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespeech.bj.bcebos.com/PaddleAudio/en.wav
-```
-
 ### 3. 服务端使用方法
 - 命令行 (推荐使用)

@ -88,6 +80,15 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee
  ```

 ### 4. ASR 客户端使用方法
+
+ASR 客户端的输入是一个 WAV 文件（`.wav`），并且采样率必须与模型的采样率相同。
+
+可以下载 ASR 客户端的示例音频：
+```bash
+wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav
+wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/en.wav
+```
+
 **注意：** 初次使用客户端时响应时间会略长
 - 命令行 (推荐使用)

@ -95,7 +96,6 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee

  ```
  paddlespeech_client asr --server_ip 127.0.0.1 --port 8090 --input ./zh.wav
-
  ```

  使用帮助:
@ -114,14 +114,13 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee

  输出:
  ```text
-  [2022-02-23 18:11:22,819] [    INFO] - {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'transcription': '我认为跑步最重要的就是给我带来了身体健康'}}
-  [2022-02-23 18:11:22,820] [    INFO] - time cost 0.689145 s.
+  [2022-08-01 07:54:01,646] [    INFO] - ASR result: 我认为跑步最重要的就是给我带来了身体健康
+  [2022-08-01 07:54:01,646] [    INFO] - Response time 4.898965 s.
  ```

 - Python API
  ```python
  from paddlespeech.server.bin.paddlespeech_client import ASRClientExecutor
-  import json

  asrclient_executor = ASRClientExecutor()
  res = asrclient_executor(
@ -131,12 +130,12 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee
      sample_rate=16000,
      lang="zh_cn",
      audio_format="wav")
-  print(res.json())
+  print(res)
  ```

  输出:
  ```text
-  {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'transcription': '我认为跑步最重要的就是给我带来了身体健康'}}
+  我认为跑步最重要的就是给我带来了身体健康
  ```
 
 ### 5. TTS 客户端使用方法
@ -166,7 +165,6 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee

  输出:
  ```text
-  [2022-02-23 15:20:37,875] [    INFO] - {'description': 'success.'}
  [2022-02-23 15:20:37,875] [    INFO] - Save synthesized audio successfully on output.wav.
  [2022-02-23 15:20:37,875] [    INFO] - Audio duration: 3.612500 s.
  [2022-02-23 15:20:37,875] [    INFO] - Response time: 0.348050 s.
@ -203,6 +201,11 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee

 ### 6. CLS 客户端使用方法

+可以下载 CLS 客户端的示例音频：
+```bash
+wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav
+```
+
 **注意：** 初次使用客户端时响应时间会略长

 - 命令行 (推荐使用)
@ -251,8 +254,14 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee

 ### 7. 声纹客户端使用方法

+可以下载声纹客户端的示例音频：
+```bash
+wget -c https://paddlespeech.bj.bcebos.com/vector/audio/85236145389.wav
+wget -c https://paddlespeech.bj.bcebos.com/vector/audio/123456789.wav
+```
+
 #### 7.1 提取声纹特征
-注意： 初次使用客户端时响应时间会略长
+**注意：** 初次使用客户端时响应时间会略长
 * 命令行 (推荐使用)

  若 `127.0.0.1` 不能访问，则需要使用实际服务 IP 地址
@ -276,18 +285,18 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee

  输出:
  ```text
-  [2022-05-25 12:25:36,165] [    INFO] - vector http client start
-  [2022-05-25 12:25:36,165] [    INFO] - the input audio: 85236145389.wav
-  [2022-05-25 12:25:36,165] [    INFO] - endpoint: http://127.0.0.1:8790/paddlespeech/vector
-  [2022-05-25 12:25:36,166] [    INFO] - http://127.0.0.1:8790/paddlespeech/vector
-  [2022-05-25 12:25:36,324] [    INFO] - The vector: {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'vec': [-1.3251205682754517, 7.860682487487793, -4.620625972747803, 0.3000721037387848, 2.2648534774780273, -1.1931440830230713, 3.064713716506958, 7.673594951629639, -6.004472732543945, -12.024259567260742, -1.9496068954467773, 3.126953601837158, 1.6188379526138306, -7.638310432434082, -1.2299772500991821, -12.33833122253418, 2.1373026371002197, -5.395712375640869, 9.717328071594238, 5.675230503082275, 3.7805123329162598, 3.0597171783447266, 3.429692029953003, 8.9760103225708, 13.174124717712402, -0.5313228368759155, 8.942471504211426, 4.465109825134277, -4.426247596740723, -9.726503372192383, 8.399328231811523, 7.223917484283447, -7.435853958129883, 2.9441683292388916, -4.343039512634277, -13.886964797973633, -1.6346734762191772, -10.902740478515625, -5.311244964599609, 3.800722122192383, 3.897603750228882, -2.123077392578125, -2.3521194458007812, 4.151031017303467, -7.404866695404053, 0.13911646604537964, 2.4626107215881348, 4.96645450592041, 0.9897574186325073, 5.483975410461426, -3.3574001789093018, 10.13400650024414, -0.6120170950889587, -10.403095245361328, 4.600754261016846, 16.009349822998047, -7.78369140625, -4.194530487060547, -6.93686056137085, 1.1789555549621582, 11.490800857543945, 4.23802375793457, 9.550930976867676, 8.375045776367188, 7.508914470672607, -0.6570729613304138, -0.3005157709121704, 2.8406054973602295, 3.0828027725219727, 0.7308170199394226, 6.1483540534973145, 0.1376611888408661, -13.424735069274902, -7.746140480041504, -2.322798252105713, -8.305252075195312, 2.98791241645813, -10.99522876739502, 0.15211068093776703, -2.3820347785949707, -1.7984174489974976, 8.49562931060791, -5.852236747741699, -3.755497932434082, 0.6989710927009583, -5.270299434661865, -2.6188621520996094, -1.8828465938568115, -4.6466498374938965, 14.078543663024902, -0.5495333075523376, 10.579157829284668, -3.216050148010254, 9.349003791809082, -4.381077766418457, -11.675816535949707, -2.863020658493042, 4.5721755027771, 2.246612071990967, -4.574341773986816, 1.8610187768936157, 2.3767874240875244, 5.625787734985352, -9.784077644348145, 0.6496725678443909, -1.457950472831726, 0.4263263940811157, -4.921126365661621, -2.4547839164733887, 3.4869801998138428, -0.4265422224998474, 8.341268539428711, 1.356552004814148, 7.096688270568848, -13.102828979492188, 8.01673412322998, -7.115934371948242, 1.8699780702590942, 0.20872099697589874, 14.699383735656738, -1.0252779722213745, -2.6107232570648193, -2.5082311630249023, 8.427192687988281, 6.913852691650391, -6.29124641418457, 0.6157366037368774, 2.489687919616699, -3.4668266773223877, 9.92176342010498, 11.200815200805664, -0.19664029777050018, 7.491600513458252, -0.6231271624565125, -0.2584814429283142, -9.947997093200684, -0.9611040949821472, 1.1649218797683716, -2.1907122135162354, -1.502848744392395, -0.5192610621452332, 15.165953636169434, 2.4649462699890137, -0.998044490814209, 7.44166374206543, -2.0768048763275146, 3.5896823406219482, -7.305543422698975, -7.562084674835205, 4.32333517074585, 0.08044180274009705, -6.564010143280029, -2.314805269241333, -1.7642345428466797, -2.470881700515747, -7.6756181716918945, -9.548877716064453, -1.017755389213562, 0.1698644608259201, 2.5877134799957275, -1.8752295970916748, -0.36614322662353516, -6.049378395080566, -2.3965611457824707, -5.945338726043701, 0.9424033164978027, -13.155974388122559, -7.45780086517334, 0.14658108353614807, -3.7427968978881836, 5.841492652893066, -1.2872905731201172, 5.569431304931641, 12.570590019226074, 1.0939218997955322, 2.2142086029052734, 1.9181575775146484, 6.991420745849609, -5.888138771057129, 3.1409823894500732, -2.0036280155181885, 2.4434285163879395, 9.973138809204102, 5.036680221557617, 2.005120277404785, 2.861560344696045, 5.860223770141602, 2.917618751525879, -1.63111412525177, 2.0292205810546875, -4.070415019989014, -6.831437110900879]}}
-  [2022-05-25 12:25:36,324] [    INFO] - Response time 0.159053 s.
+  [2022-08-01 09:01:22,151] [    INFO] - vector http client start
+  [2022-08-01 09:01:22,152] [    INFO] - the input audio: 85236145389.wav
+  [2022-08-01 09:01:22,152] [    INFO] - endpoint: http://127.0.0.1:8090/paddlespeech/vector
+  [2022-08-01 09:01:27,093] [    INFO] - {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'vec': [1.4217487573623657, 5.626248836517334, -5.342073440551758, 1.177390217781067, 3.308061122894287, 1.7565997838974, 5.1678876876831055, 10.806346893310547, -3.822679042816162, -5.614130973815918, 2.6238481998443604, -0.8072965741157532, 1.963512659072876, -7.312864780426025, 0.011034967377781868, -9.723127365112305, 0.661963164806366, -6.976816654205322, 10.213465690612793, 7.494767189025879, 2.9105641841888428, 3.894925117492676, 3.7999846935272217, 7.106173992156982, 16.905324935913086, -7.149376392364502, 8.733112335205078, 3.423002004623413, -4.831653118133545, -11.403371810913086, 11.232216835021973, 7.127464771270752, -4.282831192016602, 2.4523589611053467, -5.13075065612793, -18.17765998840332, -2.611666440963745, -11.00034236907959, -6.731431007385254, 1.6564655303955078, 0.7618184685707092, 1.1253058910369873, -2.0838277339935303, 4.725739002227783, -8.782590866088867, -3.5398736000061035, 3.8142387866973877, 5.142062664031982, 2.162053346633911, 4.09642219543457, -6.416221618652344, 12.747454643249512, 1.9429889917373657, -15.152948379516602, 6.417416572570801, 16.097013473510742, -9.716649055480957, -1.9920448064804077, -3.364956855773926, -1.8719490766525269, 11.567351341247559, 3.6978795528411865, 11.258269309997559, 7.442364692687988, 9.183405876159668, 4.528151512145996, -1.2417811155319214, 4.395910263061523, 6.672768592834473, 5.889888763427734, 7.627115249633789, -0.6692016124725342, -11.889703750610352, -9.208883285522461, -7.427401542663574, -3.777655601501465, 6.917237758636475, -9.848749160766602, -2.094479560852051, -5.1351189613342285, 0.49564215540885925, 9.317541122436523, -5.9141845703125, -1.809845209121704, -0.11738205701112747, -7.169270992279053, -1.0578246116638184, -5.721685886383057, -5.117387294769287, 16.137670516967773, -4.473618984222412, 7.66243314743042, -0.5538089871406555, 9.631582260131836, -6.470466613769531, -8.54850959777832, 4.371622085571289, -0.7970349192619324, 4.479003429412842, -2.9758646488189697, 3.2721707820892334, 2.8382749557495117, 5.1345953941345215, -9.19078254699707, -0.5657423138618469, -4.874573230743408, 2.316561460494995, -5.984307289123535, -2.1798791885375977, 0.35541653633117676, -0.3178458511829376, 9.493547439575195, 2.114448070526123, 4.358088493347168, -12.089820861816406, 8.451695442199707, -7.925461769104004, 4.624246120452881, 4.428938388824463, 18.691999435424805, -2.620460033416748, -5.149182319641113, -0.3582168221473694, 8.488557815551758, 4.98148250579834, -9.326834678649902, -2.2544236183166504, 6.64176607131958, 1.2119656801223755, 10.977132797241211, 16.55504035949707, 3.323848247528076, 9.55185317993164, -1.6677050590515137, -0.7953923940658569, -8.605660438537598, -0.4735637903213501, 2.6741855144500732, -5.359188079833984, -2.6673784255981445, 0.6660736799240112, 15.443212509155273, 4.740597724914551, -3.4725306034088135, 11.592561721801758, -2.05450701713562, 1.7361239194869995, -8.26533031463623, -9.304476737976074, 5.406835079193115, -1.5180232524871826, -7.746610641479492, -6.089605331420898, 0.07112561166286469, -0.34904858469963074, -8.649889945983887, -9.998958587646484, -2.5648481845855713, -0.5399898886680603, 2.6018145084381104, -0.31927648186683655, -1.8815231323242188, -2.0721378326416016, -3.4105639457702637, -8.299802780151367, 1.4836379289627075, -15.366002082824707, -8.288193702697754, 3.884773015975952, -3.4876506328582764, 7.362995624542236, 0.4657321572303772, 3.1326000690460205, 12.438883781433105, -1.8337029218673706, 4.532927513122559, 2.726433277130127, 10.145345687866211, -6.521956920623779, 2.8971481323242188, -3.3925881385803223, 5.079156398773193, 7.759725093841553, 4.677562236785889, 5.8457818031311035, 2.4023921489715576, 7.707108974456787, 3.9711389541625977, -6.390035152435303, 6.126871109008789, -3.776031017303467, -11.118141174316406]}}
+  [2022-08-01 09:01:27,094] [    INFO] - Response time 4.941739 s.
  ```

 * Python API

  ``` python
  from paddlespeech.server.bin.paddlespeech_client import VectorClientExecutor
+  import json

  vectorclient_executor = VectorClientExecutor()
  res = vectorclient_executor(
@ -295,17 +304,17 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee
      server_ip="127.0.0.1",
      port=8090,
      task="spk")
-  print(res)
+  print(res.json())
  ```

  输出:
  ```text
-  {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'vec': [-1.3251205682754517, 7.860682487487793, -4.620625972747803, 0.3000721037387848, 2.2648534774780273, -1.1931440830230713, 3.064713716506958, 7.673594951629639, -6.004472732543945, -12.024259567260742, -1.9496068954467773, 3.126953601837158, 1.6188379526138306, -7.638310432434082, -1.2299772500991821, -12.33833122253418, 2.1373026371002197, -5.395712375640869, 9.717328071594238, 5.675230503082275, 3.7805123329162598, 3.0597171783447266, 3.429692029953003, 8.9760103225708, 13.174124717712402, -0.5313228368759155, 8.942471504211426, 4.465109825134277, -4.426247596740723, -9.726503372192383, 8.399328231811523, 7.223917484283447, -7.435853958129883, 2.9441683292388916, -4.343039512634277, -13.886964797973633, -1.6346734762191772, -10.902740478515625, -5.311244964599609, 3.800722122192383, 3.897603750228882, -2.123077392578125, -2.3521194458007812, 4.151031017303467, -7.404866695404053, 0.13911646604537964, 2.4626107215881348, 4.96645450592041, 0.9897574186325073, 5.483975410461426, -3.3574001789093018, 10.13400650024414, -0.6120170950889587, -10.403095245361328, 4.600754261016846, 16.009349822998047, -7.78369140625, -4.194530487060547, -6.93686056137085, 1.1789555549621582, 11.490800857543945, 4.23802375793457, 9.550930976867676, 8.375045776367188, 7.508914470672607, -0.6570729613304138, -0.3005157709121704, 2.8406054973602295, 3.0828027725219727, 0.7308170199394226, 6.1483540534973145, 0.1376611888408661, -13.424735069274902, -7.746140480041504, -2.322798252105713, -8.305252075195312, 2.98791241645813, -10.99522876739502, 0.15211068093776703, -2.3820347785949707, -1.7984174489974976, 8.49562931060791, -5.852236747741699, -3.755497932434082, 0.6989710927009583, -5.270299434661865, -2.6188621520996094, -1.8828465938568115, -4.6466498374938965, 14.078543663024902, -0.5495333075523376, 10.579157829284668, -3.216050148010254, 9.349003791809082, -4.381077766418457, -11.675816535949707, -2.863020658493042, 4.5721755027771, 2.246612071990967, -4.574341773986816, 1.8610187768936157, 2.3767874240875244, 5.625787734985352, -9.784077644348145, 0.6496725678443909, -1.457950472831726, 0.4263263940811157, -4.921126365661621, -2.4547839164733887, 3.4869801998138428, -0.4265422224998474, 8.341268539428711, 1.356552004814148, 7.096688270568848, -13.102828979492188, 8.01673412322998, -7.115934371948242, 1.8699780702590942, 0.20872099697589874, 14.699383735656738, -1.0252779722213745, -2.6107232570648193, -2.5082311630249023, 8.427192687988281, 6.913852691650391, -6.29124641418457, 0.6157366037368774, 2.489687919616699, -3.4668266773223877, 9.92176342010498, 11.200815200805664, -0.19664029777050018, 7.491600513458252, -0.6231271624565125, -0.2584814429283142, -9.947997093200684, -0.9611040949821472, 1.1649218797683716, -2.1907122135162354, -1.502848744392395, -0.5192610621452332, 15.165953636169434, 2.4649462699890137, -0.998044490814209, 7.44166374206543, -2.0768048763275146, 3.5896823406219482, -7.305543422698975, -7.562084674835205, 4.32333517074585, 0.08044180274009705, -6.564010143280029, -2.314805269241333, -1.7642345428466797, -2.470881700515747, -7.6756181716918945, -9.548877716064453, -1.017755389213562, 0.1698644608259201, 2.5877134799957275, -1.8752295970916748, -0.36614322662353516, -6.049378395080566, -2.3965611457824707, -5.945338726043701, 0.9424033164978027, -13.155974388122559, -7.45780086517334, 0.14658108353614807, -3.7427968978881836, 5.841492652893066, -1.2872905731201172, 5.569431304931641, 12.570590019226074, 1.0939218997955322, 2.2142086029052734, 1.9181575775146484, 6.991420745849609, -5.888138771057129, 3.1409823894500732, -2.0036280155181885, 2.4434285163879395, 9.973138809204102, 5.036680221557617, 2.005120277404785, 2.861560344696045, 5.860223770141602, 2.917618751525879, -1.63111412525177, 2.0292205810546875, -4.070415019989014, -6.831437110900879]}}
+  {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'vec': [1.4217487573623657, 5.626248836517334, -5.342073440551758, 1.177390217781067, 3.308061122894287, 1.7565997838974, 5.1678876876831055, 10.806346893310547, -3.822679042816162, -5.614130973815918, 2.6238481998443604, -0.8072965741157532, 1.963512659072876, -7.312864780426025, 0.011034967377781868, -9.723127365112305, 0.661963164806366, -6.976816654205322, 10.213465690612793, 7.494767189025879, 2.9105641841888428, 3.894925117492676, 3.7999846935272217, 7.106173992156982, 16.905324935913086, -7.149376392364502, 8.733112335205078, 3.423002004623413, -4.831653118133545, -11.403371810913086, 11.232216835021973, 7.127464771270752, -4.282831192016602, 2.4523589611053467, -5.13075065612793, -18.17765998840332, -2.611666440963745, -11.00034236907959, -6.731431007385254, 1.6564655303955078, 0.7618184685707092, 1.1253058910369873, -2.0838277339935303, 4.725739002227783, -8.782590866088867, -3.5398736000061035, 3.8142387866973877, 5.142062664031982, 2.162053346633911, 4.09642219543457, -6.416221618652344, 12.747454643249512, 1.9429889917373657, -15.152948379516602, 6.417416572570801, 16.097013473510742, -9.716649055480957, -1.9920448064804077, -3.364956855773926, -1.8719490766525269, 11.567351341247559, 3.6978795528411865, 11.258269309997559, 7.442364692687988, 9.183405876159668, 4.528151512145996, -1.2417811155319214, 4.395910263061523, 6.672768592834473, 5.889888763427734, 7.627115249633789, -0.6692016124725342, -11.889703750610352, -9.208883285522461, -7.427401542663574, -3.777655601501465, 6.917237758636475, -9.848749160766602, -2.094479560852051, -5.1351189613342285, 0.49564215540885925, 9.317541122436523, -5.9141845703125, -1.809845209121704, -0.11738205701112747, -7.169270992279053, -1.0578246116638184, -5.721685886383057, -5.117387294769287, 16.137670516967773, -4.473618984222412, 7.66243314743042, -0.5538089871406555, 9.631582260131836, -6.470466613769531, -8.54850959777832, 4.371622085571289, -0.7970349192619324, 4.479003429412842, -2.9758646488189697, 3.2721707820892334, 2.8382749557495117, 5.1345953941345215, -9.19078254699707, -0.5657423138618469, -4.874573230743408, 2.316561460494995, -5.984307289123535, -2.1798791885375977, 0.35541653633117676, -0.3178458511829376, 9.493547439575195, 2.114448070526123, 4.358088493347168, -12.089820861816406, 8.451695442199707, -7.925461769104004, 4.624246120452881, 4.428938388824463, 18.691999435424805, -2.620460033416748, -5.149182319641113, -0.3582168221473694, 8.488557815551758, 4.98148250579834, -9.326834678649902, -2.2544236183166504, 6.64176607131958, 1.2119656801223755, 10.977132797241211, 16.55504035949707, 3.323848247528076, 9.55185317993164, -1.6677050590515137, -0.7953923940658569, -8.605660438537598, -0.4735637903213501, 2.6741855144500732, -5.359188079833984, -2.6673784255981445, 0.6660736799240112, 15.443212509155273, 4.740597724914551, -3.4725306034088135, 11.592561721801758, -2.05450701713562, 1.7361239194869995, -8.26533031463623, -9.304476737976074, 5.406835079193115, -1.5180232524871826, -7.746610641479492, -6.089605331420898, 0.07112561166286469, -0.34904858469963074, -8.649889945983887, -9.998958587646484, -2.5648481845855713, -0.5399898886680603, 2.6018145084381104, -0.31927648186683655, -1.8815231323242188, -2.0721378326416016, -3.4105639457702637, -8.299802780151367, 1.4836379289627075, -15.366002082824707, -8.288193702697754, 3.884773015975952, -3.4876506328582764, 7.362995624542236, 0.4657321572303772, 3.1326000690460205, 12.438883781433105, -1.8337029218673706, 4.532927513122559, 2.726433277130127, 10.145345687866211, -6.521956920623779, 2.8971481323242188, -3.3925881385803223, 5.079156398773193, 7.759725093841553, 4.677562236785889, 5.8457818031311035, 2.4023921489715576, 7.707108974456787, 3.9711389541625977, -6.390035152435303, 6.126871109008789, -3.776031017303467, -11.118141174316406]}}
  ```

 #### 7.2 音频声纹打分

-注意： 初次使用客户端时响应时间会略长
+**注意：** 初次使用客户端时响应时间会略长
 * 命令行 (推荐使用)

  若 `127.0.0.1` 不能访问，则需要使用实际服务 IP 地址
@ -330,18 +339,18 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee

  输出:
  ```text
-  [2022-05-25 12:33:24,527] [    INFO] - vector score http client start
-  [2022-05-25 12:33:24,527] [    INFO] - enroll audio: 85236145389.wav, test audio: 123456789.wav
-  [2022-05-25 12:33:24,528] [    INFO] - endpoint: http://127.0.0.1:8790/paddlespeech/vector/score
-  [2022-05-25 12:33:24,695] [    INFO] - The vector score is: {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'score': 0.45332613587379456}}
-  [2022-05-25 12:33:24,696] [    INFO] - The vector: {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'score': 0.45332613587379456}}
-  [2022-05-25 12:33:24,696] [    INFO] - Response time 0.168271 s.
+  [2022-08-01 09:04:42,275] [    INFO] - vector score http client start
+  [2022-08-01 09:04:42,275] [    INFO] - enroll audio: 85236145389.wav, test audio: 123456789.wav
+  [2022-08-01 09:04:42,275] [    INFO] - endpoint: http://127.0.0.1:8090/paddlespeech/vector/score
+  [2022-08-01 09:04:44,611] [    INFO] - {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'score': 0.4292638897895813}}
+  [2022-08-01 09:04:44,611] [    INFO] - Response time 2.336258 s.
  ```

 * Python API

  ```python 
  from paddlespeech.server.bin.paddlespeech_client import VectorClientExecutor
+  import json

  vectorclient_executor = VectorClientExecutor()
  res = vectorclient_executor(
@ -351,16 +360,12 @@ wget -c https://paddlespeech.bj.bcebos.com/PaddleAudio/zh.wav https://paddlespee
      server_ip="127.0.0.1",
      port=8090,
      task="score")
-  print(res)
+  print(res.json())
  ```

  输出:
  ```text
-  [2022-05-25 12:30:14,143] [    INFO] - vector score http client start
-  [2022-05-25 12:30:14,143] [    INFO] - enroll audio: 85236145389.wav, test audio: 123456789.wav
-  [2022-05-25 12:30:14,143] [    INFO] - endpoint: http://127.0.0.1:8790/paddlespeech/vector/score
-  [2022-05-25 12:30:14,363] [    INFO] - The vector score is: {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'score': 0.45332613587379456}}
-  {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'score': 0.45332613587379456}}
+  {'success': True, 'code': 200, 'message': {'description': 'success'}, 'result': {'score': 0.4292638897895813}}
  ```

 ### 8. 标点预测
--- a/demos/speech_server/conf/application.yaml
+++ b/demos/speech_server/conf/application.yaml
@ -7,7 +7,7 @@ host: 0.0.0.0
 port: 8090

 # The task format in the engin_list is: <speech task>_<engine type>
-# task choices = ['asr_python', 'asr_inference', 'tts_python', 'tts_inference', 'cls_python', 'cls_inference']
+# task choices = ['asr_python', 'asr_inference', 'tts_python', 'tts_inference', 'cls_python', 'cls_inference', 'text_python', 'vector_python']
 protocol: 'http'
 engine_list: ['asr_python', 'tts_python', 'cls_python', 'text_python', 'vector_python']

@ -28,7 +28,6 @@ asr_python:
    force_yes: True
    device:  # set 'gpu:id' or 'cpu'

-
 ################### speech task: asr; engine_type: inference #######################
 asr_inference:
    # model_type choices=['deepspeech2offline_aishell']
@ -50,10 +49,11 @@ asr_inference:

 ################################### TTS #########################################
 ################### speech task: tts; engine_type: python #######################
-tts_python: 
-    # am (acoustic model) choices=['speedyspeech_csmsc', 'fastspeech2_csmsc', 
-    #                              'fastspeech2_ljspeech', 'fastspeech2_aishell3',
-    #                              'fastspeech2_vctk']        
+tts_python:
+    # am (acoustic model) choices=['speedyspeech_csmsc', 'fastspeech2_csmsc',
+    #                             'fastspeech2_ljspeech', 'fastspeech2_aishell3',
+    #                             'fastspeech2_vctk', 'fastspeech2_mix',
+    #                             'tacotron2_csmsc', 'tacotron2_ljspeech']
    am: 'fastspeech2_csmsc'   
    am_config: 
    am_ckpt: 
@ -61,11 +61,13 @@ tts_python:
    phones_dict: 
    tones_dict: 
    speaker_dict: 
-    spk_id: 0
+

    # voc (vocoder) choices=['pwgan_csmsc', 'pwgan_ljspeech', 'pwgan_aishell3',
-    #                        'pwgan_vctk', 'mb_melgan_csmsc']
-    voc: 'pwgan_csmsc'
+    #                        'pwgan_vctk', 'mb_melgan_csmsc', 'style_melgan_csmsc',
+    #                        'hifigan_csmsc', 'hifigan_ljspeech', 'hifigan_aishell3',
+    #                        'hifigan_vctk', 'wavernn_csmsc']
+    voc: 'mb_melgan_csmsc'
    voc_config: 
    voc_ckpt: 
    voc_stat: 
@ -85,7 +87,7 @@ tts_inference:
    phones_dict: 
    tones_dict: 
    speaker_dict: 
-    spk_id: 0
+

    am_predictor_conf:
        device:  # set 'gpu:id' or 'cpu'
@ -94,7 +96,7 @@ tts_inference:
        summary: True  # False -> do not show predictor config

    # voc (vocoder) choices=['pwgan_csmsc', 'mb_melgan_csmsc','hifigan_csmsc']
-    voc: 'pwgan_csmsc'
+    voc: 'mb_melgan_csmsc'
    voc_model: # the pdmodel file of your vocoder static model (XX.pdmodel)
    voc_params: # the pdiparams file of your vocoder static model (XX.pdipparams)
    voc_sample_rate: 24000
--- a/demos/speech_web/API.md
+++ b/demos/speech_web/API.md
@ -401,4 +401,4 @@ curl -X 'GET' \
  "code": 0,
  "result":"AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA",
  "message": "ok"
-```
+```
--- a/demos/speech_web/speech_server/main.py
+++ b/demos/speech_web/speech_server/main.py
@ -3,48 +3,48 @@
 # 2. 接收录音音频，返回识别结果
 # 3. 接收ASR识别结果，返回NLP对话结果
 # 4. 接收NLP对话结果，返回TTS音频
-
+import argparse
 import base64
-import yaml
-import os
-import json
 import datetime
+import json
+import os
+from typing import List
+
+import aiofiles
 import librosa
 import soundfile as sf
-import numpy as np
-import argparse
 import uvicorn
-import aiofiles
-from typing import Optional, List 
-from pydantic import BaseModel
-from fastapi import FastAPI, Header, File, UploadFile, Form, Cookie, WebSocket, WebSocketDisconnect
+from fastapi import FastAPI
+from fastapi import File
+from fastapi import Form
+from fastapi import UploadFile
+from fastapi import WebSocket
+from fastapi import WebSocketDisconnect
 from fastapi.responses import StreamingResponse
-from starlette.responses import FileResponse
-from starlette.middleware.cors import CORSMiddleware
-from starlette.requests import Request
-from starlette.websockets import WebSocketState as WebSocketState
-
+from pydantic import BaseModel
 from src.AudioManeger import AudioMannger
-from src.util import *
 from src.robot import Robot
-from src.WebsocketManeger import ConnectionManager
 from src.SpeechBase.vpr import VPR
+from src.util import *
+from src.WebsocketManeger import ConnectionManager
+from starlette.middleware.cors import CORSMiddleware
+from starlette.requests import Request
+from starlette.responses import FileResponse
+from starlette.websockets import WebSocketState as WebSocketState

 from paddlespeech.server.engine.asr.online.python.asr_engine import PaddleASRConnectionHanddler
 from paddlespeech.server.utils.audio_process import float2pcm

-
 # 解析配置
-parser = argparse.ArgumentParser(
-        prog='PaddleSpeechDemo', add_help=True)
+parser = argparse.ArgumentParser(prog='PaddleSpeechDemo', add_help=True)

 parser.add_argument(
-        "--port",
-        action="store",
-        type=int,
-        help="port of the app",
-        default=8010,
-        required=False)
+    "--port",
+    action="store",
+    type=int,
+    help="port of the app",
+    default=8010,
+    required=False)

 args = parser.parse_args()
 port = args.port
@ -60,39 +60,41 @@ ie_model_path = "source/model"
 UPLOAD_PATH = "source/vpr"
 WAV_PATH = "source/wav"

-
-base_sources = [
-    UPLOAD_PATH, WAV_PATH
-]
+base_sources = [UPLOAD_PATH, WAV_PATH]
 for path in base_sources:
    os.makedirs(path, exist_ok=True)

-
 # 初始化
 app = FastAPI()
-chatbot = Robot(asr_config, tts_config, asr_init_path, ie_model_path=ie_model_path)
+chatbot = Robot(
+    asr_config, tts_config, asr_init_path, ie_model_path=ie_model_path)
 manager = ConnectionManager()
 aumanager = AudioMannger(chatbot)
 aumanager.init()
-vpr = VPR(db_path, dim = 192, top_k = 5)
+vpr = VPR(db_path, dim=192, top_k=5)
+

 # 服务配置
 class NlpBase(BaseModel):
    chat: str

+
 class TtsBase(BaseModel):
-    text: str 
+    text: str
+

 class Audios:
    def __init__(self) -> None:
        self.audios = b""

+
 audios = Audios()

 ######################################################################
 ########################### ASR 服务 #################################
 #####################################################################

+
 # 接收文件，返回ASR结果
 # 上传文件
@app.post("/asr/offline")
@ -101,7 +103,8 @@ async def speech2textOffline(files: List[UploadFile]):
    asr_res = ""
    for file in files[:1]:
        # 生成时间戳
-        now_name = "asr_offline_" + datetime.datetime.strftime(datetime.datetime.now(), '%Y%m%d%H%M%S') + randName() + ".wav"
+        now_name = "asr_offline_" + datetime.datetime.strftime(
+            datetime.datetime.now(), '%Y%m%d%H%M%S') + randName() + ".wav"
        out_file_path = os.path.join(WAV_PATH, now_name)
        async with aiofiles.open(out_file_path, 'wb') as out_file:
            content = await file.read()  # async read
@ -110,10 +113,9 @@ async def speech2textOffline(files: List[UploadFile]):
        # 返回ASR识别结果
        asr_res = chatbot.speech2text(out_file_path)
        return SuccessRequest(result=asr_res)
-        # else:
-            # return ErrorRequest(message="文件不是.wav格式")
    return ErrorRequest(message="上传文件为空")

+
 # 接收文件，同时将wav强制转成16k, int16类型
@app.post("/asr/offlinefile")
 async def speech2textOfflineFile(files: List[UploadFile]):
@ -121,7 +123,8 @@ async def speech2textOfflineFile(files: List[UploadFile]):
    asr_res = ""
    for file in files[:1]:
        # 生成时间戳
-        now_name = "asr_offline_" + datetime.datetime.strftime(datetime.datetime.now(), '%Y%m%d%H%M%S') + randName() + ".wav"
+        now_name = "asr_offline_" + datetime.datetime.strftime(
+            datetime.datetime.now(), '%Y%m%d%H%M%S') + randName() + ".wav"
        out_file_path = os.path.join(WAV_PATH, now_name)
        async with aiofiles.open(out_file_path, 'wb') as out_file:
            content = await file.read()  # async read
@ -132,22 +135,18 @@ async def speech2textOfflineFile(files: List[UploadFile]):
        wav = float2pcm(wav)  # float32 to int16
        wav_bytes = wav.tobytes()  # to bytes
        wav_base64 = base64.b64encode(wav_bytes).decode('utf8')
-        
+
        # 将文件重新写入
        now_name = now_name[:-4] + "_16k" + ".wav"
        out_file_path = os.path.join(WAV_PATH, now_name)
-        sf.write(out_file_path,wav,16000)
+        sf.write(out_file_path, wav, 16000)

        # 返回ASR识别结果
        asr_res = chatbot.speech2text(out_file_path)
-        response_res = {
-            "asr_result": asr_res,
-            "wav_base64": wav_base64
-        }
+        response_res = {"asr_result": asr_res, "wav_base64": wav_base64}
        return SuccessRequest(result=response_res)
-        
-    return ErrorRequest(message="上传文件为空")

+    return ErrorRequest(message="上传文件为空")


 # 流式接收测试
@ -161,15 +160,17 @@ async def speech2textOnlineRecive(files: List[UploadFile]):
    print(f"audios长度变化: {len(audios.audios)}")
    return SuccessRequest(message="接收成功")

+
 # 采集环境噪音大小
@app.post("/asr/collectEnv")
 async def collectEnv(files: List[UploadFile]):
-     for file in files[:1]:
+    for file in files[:1]:
        content = await file.read()  # async read
        # 初始化, wav 前44字节是头部信息
        aumanager.compute_env_volume(content[44:])
        vad_ = aumanager.vad_threshold
-        return SuccessRequest(result=vad_,message="采集环境噪音成功")
+        return SuccessRequest(result=vad_, message="采集环境噪音成功")
+

 # 停止录音
@app.get("/asr/stopRecord")
@ -179,6 +180,7 @@ async def stopRecord():
    print("Online录音暂停")
    return SuccessRequest(message="停止成功")

+
 # 恢复录音
@app.get("/asr/resumeRecord")
 async def resumeRecord():
@ -187,7 +189,7 @@ async def resumeRecord():
    return SuccessRequest(message="Online录音恢复")


-# 聊天用的ASR
+# 聊天用的 ASR
@app.websocket("/ws/asr/offlineStream")
 async def websocket_endpoint(websocket: WebSocket):
    await manager.connect(websocket)
@ -210,9 +212,9 @@ async def websocket_endpoint(websocket: WebSocket):
        # print(f"用户-{user}-离开")


-# Online识别的ASR
+    # 流式识别的 ASR
@app.websocket('/ws/asr/onlineStream')
-async def websocket_endpoint(websocket: WebSocket):
+async def websocket_endpoint_online(websocket: WebSocket):
    """PaddleSpeech Online ASR Server api

    Args:
@ -298,12 +300,14 @@ async def websocket_endpoint(websocket: WebSocket):
    except WebSocketDisconnect:
        pass

+
 ######################################################################
 ########################### NLP 服务 #################################
 #####################################################################

+
@app.post("/nlp/chat")
-async def chatOffline(nlp_base:NlpBase):
+async def chatOffline(nlp_base: NlpBase):
    chat = nlp_base.chat
    if not chat:
        return ErrorRequest(message="传入文本为空")
@ -311,8 +315,9 @@ async def chatOffline(nlp_base:NlpBase):
        res = chatbot.chat(chat)
        return SuccessRequest(result=res)

+
@app.post("/nlp/ie")
-async def ieOffline(nlp_base:NlpBase):
+async def ieOffline(nlp_base: NlpBase):
    nlp_text = nlp_base.chat
    if not nlp_text:
        return ErrorRequest(message="传入文本为空")
@ -320,17 +325,20 @@ async def ieOffline(nlp_base:NlpBase):
        res = chatbot.ie(nlp_text)
        return SuccessRequest(result=res)

+
 ######################################################################
 ########################### TTS 服务 #################################
 #####################################################################

+
@app.post("/tts/offline")
-async def text2speechOffline(tts_base:TtsBase):
+async def text2speechOffline(tts_base: TtsBase):
    text = tts_base.text
    if not text:
        return ErrorRequest(message="文本为空")
    else:
-        now_name = "tts_"+ datetime.datetime.strftime(datetime.datetime.now(), '%Y%m%d%H%M%S') + randName() + ".wav"
+        now_name = "tts_" + datetime.datetime.strftime(
+            datetime.datetime.now(), '%Y%m%d%H%M%S') + randName() + ".wav"
        out_file_path = os.path.join(WAV_PATH, now_name)
        # 保存为文件，再转成base64传输
        chatbot.text2speech(text, outpath=out_file_path)
@ -339,12 +347,14 @@ async def text2speechOffline(tts_base:TtsBase):
        base_str = base64.b64encode(data_bin)
        return SuccessRequest(result=base_str)

+
 # http流式TTS
@app.post("/tts/online")
 async def stream_tts(request_body: TtsBase):
    text = request_body.text
    return StreamingResponse(chatbot.text2speechStreamBytes(text=text))

+
 # ws流式TTS
@app.websocket("/ws/tts/online")
 async def stream_ttsWS(websocket: WebSocket):
@ -356,17 +366,11 @@ async def stream_ttsWS(websocket: WebSocket):
            if text:
                for sub_wav in chatbot.text2speechStream(text=text):
                    # print("发送sub wav: ", len(sub_wav))
-                    res = {
-                        "wav": sub_wav,
-                        "done": False
-                    }
+                    res = {"wav": sub_wav, "done": False}
                    await websocket.send_json(res)
-                
+
                # 输送结束
-                res = {
-                        "wav": sub_wav,
-                        "done": True
-                    }
+                res = {"wav": sub_wav, "done": True}
                await websocket.send_json(res)
            # manager.disconnect(websocket)

@ -396,8 +400,9 @@ async def vpr_enroll(table_name: str=None,
            return {'status': False, 'msg': "spk_id can not be None"}
        # Save the upload data to server.
        content = await audio.read()
-        now_name = "vpr_enroll_" + datetime.datetime.strftime(datetime.datetime.now(), '%Y%m%d%H%M%S') + randName() + ".wav"
-        audio_path =  os.path.join(UPLOAD_PATH, now_name)
+        now_name = "vpr_enroll_" + datetime.datetime.strftime(
+            datetime.datetime.now(), '%Y%m%d%H%M%S') + randName() + ".wav"
+        audio_path = os.path.join(UPLOAD_PATH, now_name)

        with open(audio_path, "wb+") as f:
            f.write(content)
@ -413,20 +418,19 @@ async def vpr_recog(request: Request,
                    audio: UploadFile=File(...)):
    # Voice print recognition online
    # try:
-        # Save the upload data to server.
+    # Save the upload data to server.
    content = await audio.read()
-    now_name = "vpr_query_" + datetime.datetime.strftime(datetime.datetime.now(), '%Y%m%d%H%M%S') + randName() + ".wav"
-    query_audio_path =  os.path.join(UPLOAD_PATH, now_name)
+    now_name = "vpr_query_" + datetime.datetime.strftime(
+        datetime.datetime.now(), '%Y%m%d%H%M%S') + randName() + ".wav"
+    query_audio_path = os.path.join(UPLOAD_PATH, now_name)
    with open(query_audio_path, "wb+") as f:
-        f.write(content)        
+        f.write(content)
    spk_ids, paths, scores = vpr.do_search_vpr(query_audio_path)

    res = dict(zip(spk_ids, zip(paths, scores)))
    # Sort results by distance metric, closest distances first
    res = sorted(res.items(), key=lambda item: item[1][1], reverse=True)
    return res
-    # except Exception as e:
-        # return {'status': False, 'msg': e}, 400


@app.post('/vpr/del')
@ -460,17 +464,18 @@ async def vpr_database64(vprId: int):
            return {'status': False, 'msg': "vpr_id can not be None"}
        audio_path = vpr.do_get_wav(vprId)
        # 返回base64
-        
+
        # 将文件转成16k, 16bit类型的wav文件
        wav, sr = librosa.load(audio_path, sr=16000)
        wav = float2pcm(wav)  # float32 to int16
        wav_bytes = wav.tobytes()  # to bytes
        wav_base64 = base64.b64encode(wav_bytes).decode('utf8')
-        
+
        return SuccessRequest(result=wav_base64)
    except Exception as e:
        return {'status': False, 'msg': e}, 400

+
@app.get('/vpr/data')
 async def vpr_data(vprId: int):
    # Get the audio file from path by spk_id in MySQL
@ -482,11 +487,6 @@ async def vpr_data(vprId: int):
    except Exception as e:
        return {'status': False, 'msg': e}, 400

+
 if __name__ == '__main__':
    uvicorn.run(app=app, host='0.0.0.0', port=port)
-    
-
-
-
-
-
--- a/demos/speech_web/speech_server/requirements.txt
+++ b/demos/speech_web/speech_server/requirements.txt
@ -1,14 +1,13 @@
 aiofiles
+faiss-cpu
 fastapi
 librosa
 numpy
+paddlenlp
+paddlepaddle
+paddlespeech
 pydantic
-scikit_learn
+python-multipartscikit_learn
 SoundFile
 starlette
 uvicorn
-paddlepaddle
-paddlespeech
-paddlenlp
-faiss-cpu
-python-multipart
--- a/demos/speech_web/speech_server/src/AudioManeger.py
+++ b/demos/speech_web/speech_server/src/AudioManeger.py
@ -1,15 +1,19 @@
-import imp
-from queue import Queue
-import numpy as np
+import datetime
 import os
 import wave
-import random
-import datetime
+
+import numpy as np
+
 from .util import randName


 class AudioMannger:
-    def __init__(self, robot, frame_length=160, frame=10, data_width=2, vad_default = 300):
+    def __init__(self,
+                 robot,
+                 frame_length=160,
+                 frame=10,
+                 data_width=2,
+                 vad_default=300):
        # 二进制 pcm 流 
        self.audios = b''
        self.asr_result = ""
@ -20,8 +24,9 @@ class AudioMannger:
        os.makedirs(self.file_dir, exist_ok=True)
        self.vad_deafult = vad_default
        self.vad_threshold = vad_default
-        self.vad_threshold_path = os.path.join(self.file_dir, "vad_threshold.npy")
-        
+        self.vad_threshold_path = os.path.join(self.file_dir,
+                                               "vad_threshold.npy")
+
        # 10ms 一帧
        self.frame_length = frame_length
        # 10帧，检测一次 vad
@ -30,67 +35,64 @@ class AudioMannger:
        self.data_width = data_width
        # window
        self.window_length = frame_length * frame * data_width
-        
+
        # 是否开始录音
        self.on_asr = False
-        self.silence_cnt  = 0
+        self.silence_cnt = 0
        self.max_silence_cnt = 4
        self.is_pause = False  # 录音暂停与恢复
-        
-        
-    
+
    def init(self):
        if os.path.exists(self.vad_threshold_path):
            # 平均响度文件存在
            self.vad_threshold = np.load(self.vad_threshold_path)
-        
-    
+
    def clear_audio(self):
        # 清空 pcm 累积片段与 asr 识别结果
        self.audios = b''
-    
+
    def clear_asr(self):
        self.asr_result = ""
-    
-    
+
    def compute_chunk_volume(self, start_index, pcm_bins):
        # 根据帧长计算能量平均值
-        pcm_bin = pcm_bins[start_index: start_index + self.window_length]
+        pcm_bin = pcm_bins[start_index:start_index + self.window_length]
        # 转成 numpy
        pcm_np = np.frombuffer(pcm_bin, np.int16)
        # 归一化 + 计算响度
        x = pcm_np.astype(np.float32)
        x = np.abs(x)
-        return np.mean(x) 
-        
-    
+        return np.mean(x)
+
    def is_speech(self, start_index, pcm_bins):
        # 检查是否没
        if start_index > len(pcm_bins):
            return False
        # 检查从这个 start 开始是否为静音帧
-        energy = self.compute_chunk_volume(start_index=start_index, pcm_bins=pcm_bins)
+        energy = self.compute_chunk_volume(
+            start_index=start_index, pcm_bins=pcm_bins)
        # print(energy)
        if energy > self.vad_threshold:
            return True
        else:
            return False
-    
+
    def compute_env_volume(self, pcm_bins):
        max_energy = 0
        start = 0
        while start < len(pcm_bins):
-            energy = self.compute_chunk_volume(start_index=start, pcm_bins=pcm_bins)
+            energy = self.compute_chunk_volume(
+                start_index=start, pcm_bins=pcm_bins)
            if energy > max_energy:
                max_energy = energy
            start += self.window_length
        self.vad_threshold = max_energy + 100 if max_energy > self.vad_deafult else self.vad_deafult
-        
+
        # 保存成文件
        np.save(self.vad_threshold_path, self.vad_threshold)
        print(f"vad 阈值大小: {self.vad_threshold}")
        print(f"环境采样保存: {os.path.realpath(self.vad_threshold_path)}")
-    
+
    def stream_asr(self, pcm_bin):
        # 先把 pcm_bin 送进去做端点检测
        start = 0
@ -99,7 +101,7 @@ class AudioMannger:
                self.on_asr = True
                self.silence_cnt = 0
                print("录音中")
-                self.audios += pcm_bin[ start : start + self.window_length]
+                self.audios += pcm_bin[start:start + self.window_length]
            else:
                if self.on_asr:
                    self.silence_cnt += 1
@ -110,41 +112,42 @@ class AudioMannger:
                        print("录音停止")
                        # audios 保存为 wav, 送入 ASR
                        if len(self.audios) > 2 * 16000:
-                            file_path = os.path.join(self.file_dir, "asr_" + datetime.datetime.strftime(datetime.datetime.now(), '%Y%m%d%H%M%S') + randName() + ".wav")
+                            file_path = os.path.join(
+                                self.file_dir,
+                                "asr_" + datetime.datetime.strftime(
+                                    datetime.datetime.now(),
+                                    '%Y%m%d%H%M%S') + randName() + ".wav")
                            self.save_audio(file_path=file_path)
                            self.asr_result = self.robot.speech2text(file_path)
                        self.clear_audio()
-                        return self.asr_result   
+                        return self.asr_result
                    else:
                        # 正常接收
                        print("录音中 静音")
-                        self.audios += pcm_bin[ start : start + self.window_length]
+                        self.audios += pcm_bin[start:start + self.window_length]
            start += self.window_length
        return ""
-    
+
    def save_audio(self, file_path):
        print("保存音频")
-        wf = wave.open(file_path, 'wb')          # 创建一个音频文件，名字为“01.wav"
-        wf.setnchannels(1)                      # 设置声道数为2
-        wf.setsampwidth(2)                      # 设置采样深度为
-        wf.setframerate(16000)                  # 设置采样率为16000
+        wf = wave.open(file_path, 'wb')  # 创建一个音频文件，名字为“01.wav"
+        wf.setnchannels(1)  # 设置声道数为2
+        wf.setsampwidth(2)  # 设置采样深度为
+        wf.setframerate(16000)  # 设置采样率为16000
        # 将数据写入创建的音频文件
        wf.writeframes(self.audios)
        # 写完后将文件关闭
        wf.close()
-    
+
    def end(self):
        # audios 保存为 wav, 送入 ASR
        file_path = os.path.join(self.file_dir, "asr.wav")
        self.save_audio(file_path=file_path)
        return self.robot.speech2text(file_path)
-    
+
    def stop(self):
        self.is_pause = True
        self.audios = b''
-    
+
    def resume(self):
        self.is_pause = False
-    
-    
-    
--- a/demos/speech_web/speech_server/src/SpeechBase/asr.py
+++ b/demos/speech_web/speech_server/src/SpeechBase/asr.py
@ -1,13 +1,10 @@
-from re import sub
 import numpy as np
-import paddle
-import librosa
-import soundfile

 from paddlespeech.server.engine.asr.online.python.asr_engine import ASREngine
 from paddlespeech.server.engine.asr.online.python.asr_engine import PaddleASRConnectionHanddler
 from paddlespeech.server.utils.config import get_config

+
 def readWave(samples):
    x_len = len(samples)

@ -31,20 +28,23 @@ def readWave(samples):


 class ASR:
-    def __init__(self, config_path, ) -> None:
+    def __init__(
+            self,
+            config_path, ) -> None:
        self.config = get_config(config_path)['asr_online']
        self.engine = ASREngine()
        self.engine.init(self.config)
        self.connection_handler = PaddleASRConnectionHanddler(self.engine)
-    
+
    def offlineASR(self, samples, sample_rate=16000):
-        x_chunk, x_chunk_lens = self.engine.preprocess(samples=samples, sample_rate=sample_rate)
+        x_chunk, x_chunk_lens = self.engine.preprocess(
+            samples=samples, sample_rate=sample_rate)
        self.engine.run(x_chunk, x_chunk_lens)
        result = self.engine.postprocess()
        self.engine.reset()
        return result

-    def onlineASR(self, samples:bytes=None, is_finished=False):
+    def onlineASR(self, samples: bytes=None, is_finished=False):
        if not is_finished:
            # 流式开始
            self.connection_handler.extract_feat(samples)
@ -58,5 +58,3 @@ class ASR:
            asr_results = self.connection_handler.get_result()
            self.connection_handler.reset()
            return asr_results
-
-        
--- a/demos/speech_web/speech_server/src/SpeechBase/nlp.py
+++ b/demos/speech_web/speech_server/src/SpeechBase/nlp.py
@ -1,23 +1,23 @@
 from paddlenlp import Taskflow

+
 class NLP:
    def __init__(self, ie_model_path=None):
        schema = ["时间", "出发地", "目的地", "费用"]
        if ie_model_path:
-            self.ie_model = Taskflow("information_extraction",
-                                    schema=schema, task_path=ie_model_path)
+            self.ie_model = Taskflow(
+                "information_extraction",
+                schema=schema,
+                task_path=ie_model_path)
        else:
-            self.ie_model = Taskflow("information_extraction",
-                                    schema=schema)
-            
+            self.ie_model = Taskflow("information_extraction", schema=schema)
+
        self.dialogue_model = Taskflow("dialogue")
-    
+
    def chat(self, text):
        result = self.dialogue_model([text])
        return result[0]
-    
+
    def ie(self, text):
        result = self.ie_model(text)
        return result
-
-    
--- a/demos/speech_web/speech_server/src/SpeechBase/sql_helper.py
+++ b/demos/speech_web/speech_server/src/SpeechBase/sql_helper.py
@ -1,18 +1,19 @@
 import base64
-import sqlite3
 import os
+import sqlite3
+
 import numpy as np
-from pkg_resources import resource_stream


-def dict_factory(cursor, row):  
-    d = {}  
-    for idx, col in enumerate(cursor.description):  
-        d[col[0]] = row[idx]  
-    return d 
+def dict_factory(cursor, row):
+    d = {}
+    for idx, col in enumerate(cursor.description):
+        d[col[0]] = row[idx]
+    return d
+

 class DataBase(object):
-    def __init__(self, db_path:str):
+    def __init__(self, db_path: str):
        db_path = os.path.realpath(db_path)

        if os.path.exists(db_path):
@ -21,12 +22,12 @@ class DataBase(object):
            db_path_dir = os.path.dirname(db_path)
            os.makedirs(db_path_dir, exist_ok=True)
            self.db_path = db_path
-        
+
        self.conn = sqlite3.connect(self.db_path)
        self.conn.row_factory = dict_factory
        self.cursor = self.conn.cursor()
        self.init_database()
-    
+
    def init_database(self):
        """
        初始化数据库， 若表不存在则创建
@ -41,20 +42,21 @@ class DataBase(object):
        """
        self.cursor.execute(sql)
        self.conn.commit()
-    
+
    def execute_base(self, sql, data_dict):
        self.cursor.execute(sql, data_dict)
        self.conn.commit()
-    
-    def insert_one(self, username, vector_base64:str, wav_path):
+
+    def insert_one(self, username, vector_base64: str, wav_path):
        if not os.path.exists(wav_path):
            return None, "wav not exists"
        else:
-            sql = f"""
+            sql = """
            insert into 
            vprtable (username, vector, wavpath)
            values (?, ?, ?)
            """
+
            try:
                self.cursor.execute(sql, (username, vector_base64, wav_path))
                self.conn.commit()
@ -63,25 +65,27 @@ class DataBase(object):
            except Exception as e:
                print(e)
                return None, e
-            
+
    def select_all(self):
        sql = """
        SELECT * from vprtable
        """
        result = self.cursor.execute(sql).fetchall()
        return result
-    
+
    def select_by_id(self, vpr_id):
        sql = f"""
        SELECT * from vprtable WHERE `id` = {vpr_id}
        """
+
        result = self.cursor.execute(sql).fetchall()
        return result
-    
+
    def select_by_username(self, username):
        sql = f"""
        SELECT * from vprtable WHERE `username` = '{username}'
        """
+
        result = self.cursor.execute(sql).fetchall()
        return result

@ -89,28 +93,30 @@ class DataBase(object):
        sql = f"""
        DELETE from vprtable WHERE `username`='{username}'
        """
+
        self.cursor.execute(sql)
        self.conn.commit()
-    
+
    def drop_all(self):
-        sql = f"""
+        sql = """
        DELETE from vprtable
        """
+
        self.cursor.execute(sql)
        self.conn.commit()
-    
+
    def drop_table(self):
-        sql = f"""
+        sql = """
            DROP TABLE vprtable
        """
+
        self.cursor.execute(sql)
        self.conn.commit()
-    
-    def encode_vector(self, vector:np.ndarray):
+
+    def encode_vector(self, vector: np.ndarray):
        return base64.b64encode(vector).decode('utf8')
-    
+
    def decode_vector(self, vector_base64, dtype=np.float32):
        b = base64.b64decode(vector_base64)
        vc = np.frombuffer(b, dtype=dtype)
        return vc
-    
--- a/demos/speech_web/speech_server/src/SpeechBase/tts.py
+++ b/demos/speech_web/speech_server/src/SpeechBase/tts.py
@ -5,18 +5,19 @@
 # 2. 加载模型
 # 3. 端到端推理
 # 4. 流式推理
-
 import base64
-import math
 import logging
+import math
+
 import numpy as np
-from paddlespeech.server.utils.onnx_infer import get_sess
-from paddlespeech.t2s.frontend.zh_frontend import Frontend
-from paddlespeech.server.utils.util import denorm, get_chunks
+
+from paddlespeech.server.engine.tts.online.onnx.tts_engine import TTSEngine
 from paddlespeech.server.utils.audio_process import float2pcm
 from paddlespeech.server.utils.config import get_config
+from paddlespeech.server.utils.util import denorm
+from paddlespeech.server.utils.util import get_chunks
+from paddlespeech.t2s.frontend.zh_frontend import Frontend

-from paddlespeech.server.engine.tts.online.onnx.tts_engine import TTSEngine

 class TTS:
    def __init__(self, config_path):
@ -26,12 +27,12 @@ class TTS:
        self.engine.init(self.config)
        self.executor = self.engine.executor
        #self.engine.warm_up()
-        
+
        # 前端初始化
        self.frontend = Frontend(
-                phone_vocab_path=self.engine.executor.phones_dict,
-                tone_vocab_path=None)
-    
+            phone_vocab_path=self.engine.executor.phones_dict,
+            tone_vocab_path=None)
+
    def depadding(self, data, chunk_num, chunk_id, block, pad, upsample):
        """ 
        Streaming inference removes the result of pad inference
@ -48,39 +49,37 @@ class TTS:
            data = data[front_pad * upsample:(front_pad + block) * upsample]

        return data
-         
+
    def offlineTTS(self, text):
        get_tone_ids = False
        merge_sentences = False
-        
+
        input_ids = self.frontend.get_input_ids(
-                text,
-                merge_sentences=merge_sentences,
-                get_tone_ids=get_tone_ids)
+            text, merge_sentences=merge_sentences, get_tone_ids=get_tone_ids)
        phone_ids = input_ids["phone_ids"]
        wav_list = []
        for i in range(len(phone_ids)):
            orig_hs = self.engine.executor.am_encoder_infer_sess.run(
-                            None, input_feed={'text': phone_ids[i].numpy()}
-                            )
+                None, input_feed={'text': phone_ids[i].numpy()})
            hs = orig_hs[0]
            am_decoder_output = self.engine.executor.am_decoder_sess.run(
-                        None, input_feed={'xs': hs})
+                None, input_feed={'xs': hs})
            am_postnet_output = self.engine.executor.am_postnet_sess.run(
-                        None,
-                        input_feed={
-                            'xs': np.transpose(am_decoder_output[0], (0, 2, 1))
-                        })
+                None,
+                input_feed={
+                    'xs': np.transpose(am_decoder_output[0], (0, 2, 1))
+                })
            am_output_data = am_decoder_output + np.transpose(
                am_postnet_output[0], (0, 2, 1))
            normalized_mel = am_output_data[0][0]
-            mel = denorm(normalized_mel, self.engine.executor.am_mu, self.engine.executor.am_std)
+            mel = denorm(normalized_mel, self.engine.executor.am_mu,
+                         self.engine.executor.am_std)
            wav = self.engine.executor.voc_sess.run(
-                            output_names=None, input_feed={'logmel': mel})[0]
+                output_names=None, input_feed={'logmel': mel})[0]
            wav_list.append(wav)
        wavs = np.concatenate(wav_list)
        return wavs
-    
+
    def streamTTS(self, text):

        get_tone_ids = False
@ -88,9 +87,7 @@ class TTS:

        # front 
        input_ids = self.frontend.get_input_ids(
-                text,
-                merge_sentences=merge_sentences,
-                get_tone_ids=get_tone_ids)
+            text, merge_sentences=merge_sentences, get_tone_ids=get_tone_ids)
        phone_ids = input_ids["phone_ids"]

        for i in range(len(phone_ids)):
@ -105,14 +102,15 @@ class TTS:
                mel = mel[0]

                # voc streaming
-                mel_chunks = get_chunks(mel, self.config.voc_block, self.config.voc_pad, "voc")
+                mel_chunks = get_chunks(mel, self.config.voc_block,
+                                        self.config.voc_pad, "voc")
                voc_chunk_num = len(mel_chunks)
                for i, mel_chunk in enumerate(mel_chunks):
                    sub_wav = self.executor.voc_sess.run(
                        output_names=None, input_feed={'logmel': mel_chunk})
-                    sub_wav = self.depadding(sub_wav[0], voc_chunk_num, i,
-                                             self.config.voc_block, self.config.voc_pad,
-                                             self.config.voc_upsample)
+                    sub_wav = self.depadding(
+                        sub_wav[0], voc_chunk_num, i, self.config.voc_block,
+                        self.config.voc_pad, self.config.voc_upsample)

                    yield self.after_process(sub_wav)

@ -130,7 +128,8 @@ class TTS:
                end = min(self.config.voc_block + self.config.voc_pad, mel_len)

                # streaming am
-                hss = get_chunks(orig_hs, self.config.am_block, self.config.am_pad, "am")
+                hss = get_chunks(orig_hs, self.config.am_block,
+                                 self.config.am_pad, "am")
                am_chunk_num = len(hss)
                for i, hs in enumerate(hss):
                    am_decoder_output = self.executor.am_decoder_sess.run(
@ -147,7 +146,8 @@ class TTS:
                    sub_mel = denorm(normalized_mel, self.executor.am_mu,
                                     self.executor.am_std)
                    sub_mel = self.depadding(sub_mel, am_chunk_num, i,
-                                             self.config.am_block, self.config.am_pad, 1)
+                                             self.config.am_block,
+                                             self.config.am_pad, 1)

                    if i == 0:
                        mel_streaming = sub_mel
@ -165,23 +165,22 @@ class TTS:
                            output_names=None, input_feed={'logmel': voc_chunk})
                        sub_wav = self.depadding(
                            sub_wav[0], voc_chunk_num, voc_chunk_id,
-                            self.config.voc_block, self.config.voc_pad, self.config.voc_upsample)
+                            self.config.voc_block, self.config.voc_pad,
+                            self.config.voc_upsample)

                        yield self.after_process(sub_wav)

                        voc_chunk_id += 1
-                        start = max(
-                            0, voc_chunk_id * self.config.voc_block - self.config.voc_pad)
-                        end = min(
-                            (voc_chunk_id + 1) * self.config.voc_block + self.config.voc_pad,
-                            mel_len)
+                        start = max(0, voc_chunk_id * self.config.voc_block -
+                                    self.config.voc_pad)
+                        end = min((voc_chunk_id + 1) * self.config.voc_block +
+                                  self.config.voc_pad, mel_len)

            else:
                logging.error(
                    "Only support fastspeech2_csmsc or fastspeech2_cnndecoder_csmsc on streaming tts."
-                ) 
+                )

-    
    def streamTTSBytes(self, text):
        for wav in self.engine.executor.infer(
                text=text,
@ -191,19 +190,14 @@ class TTS:
            wav = float2pcm(wav)  # float32 to int16
            wav_bytes = wav.tobytes()  # to bytes
            yield wav_bytes
-        
-    
+
    def after_process(self, wav):
        # for tvm
        wav = float2pcm(wav)  # float32 to int16
        wav_bytes = wav.tobytes()  # to bytes
        wav_base64 = base64.b64encode(wav_bytes).decode('utf8')  # to base64
        return wav_base64
-    
+
    def streamTTS_TVM(self, text):
        # 用 TVM 优化
        pass
-
-    
-    
-        
--- a/demos/speech_web/speech_server/src/SpeechBase/vpr.py
+++ b/demos/speech_web/speech_server/src/SpeechBase/vpr.py
@ -1,11 +1,13 @@
 # vpr Demo 没有使用 mysql 与 muilvs, 仅用于docker演示
 import logging
+
 import faiss
-from matplotlib import use
 import numpy as np
+
 from .sql_helper import DataBase
 from .vpr_encode import get_audio_embedding

+
 class VPR:
    def __init__(self, db_path, dim, top_k) -> None:
        # 初始化
@ -14,15 +16,15 @@ class VPR:
        self.top_k = top_k
        self.dtype = np.float32
        self.vpr_idx = 0
-        
+
        # db 初始化
        self.db = DataBase(db_path)
-        
+
        # faiss 初始化
        index_ip = faiss.IndexFlatIP(dim)
        self.index_ip = faiss.IndexIDMap(index_ip)
        self.init()
-    
+
    def init(self):
        # demo 初始化，把 mysql中的向量注册到 faiss 中
        sql_dbs = self.db.select_all()
@ -34,12 +36,13 @@ class VPR:
                if len(vc.shape) == 1:
                    vc = np.expand_dims(vc, axis=0)
                # 构建数据库
-                self.index_ip.add_with_ids(vc, np.array((idx,)).astype('int64'))
+                self.index_ip.add_with_ids(vc, np.array(
+                    (idx, )).astype('int64'))
            logging.info("faiss 构建完毕")
-    
+
    def faiss_enroll(self, idx, vc):
-        self.index_ip.add_with_ids(vc, np.array((idx,)).astype('int64'))
-    
+        self.index_ip.add_with_ids(vc, np.array((idx, )).astype('int64'))
+
    def vpr_enroll(self, username, wav_path):
        # 注册声纹
        emb = get_audio_embedding(wav_path)
@ -53,21 +56,22 @@ class VPR:
        else:
            last_idx, mess = None
        return last_idx
-    
+
    def vpr_recog(self, wav_path):
        # 识别声纹
        emb_search = get_audio_embedding(wav_path)
-        
+
        if emb_search is not None:
            emb_search = np.expand_dims(emb_search, axis=0)
            D, I = self.index_ip.search(emb_search, self.top_k)
            D = D.tolist()[0]
-            I = I.tolist()[0]            
-            return [(round(D[i] * 100, 2 ), I[i]) for i in range(len(D)) if I[i] != -1]
+            I = I.tolist()[0]
+            return [(round(D[i] * 100, 2), I[i]) for i in range(len(D))
+                    if I[i] != -1]
        else:
            logging.error("识别失败")
            return None
-    
+
    def do_search_vpr(self, wav_path):
        spk_ids, paths, scores = [], [], []
        recog_result = self.vpr_recog(wav_path)
@ -78,41 +82,39 @@ class VPR:
                scores.append(score)
                paths.append("")
        return spk_ids, paths, scores
-    
+
    def vpr_del(self, username):
        # 根据用户username, 删除声纹
        # 查用户ID，删除对应向量
        res = self.db.select_by_username(username)
        for r in res:
            idx = r['id']
-            self.index_ip.remove_ids(np.array((idx,)).astype('int64'))
-        
+            self.index_ip.remove_ids(np.array((idx, )).astype('int64'))
+
        self.db.drop_by_username(username)
-    
+
    def vpr_list(self):
        # 获取数据列表
        return self.db.select_all()
-    
+
    def do_list(self):
        spk_ids, vpr_ids = [], []
        for res in self.db.select_all():
            spk_ids.append(res['username'])
            vpr_ids.append(res['id'])
-        return spk_ids, vpr_ids 
-    
+        return spk_ids, vpr_ids
+
    def do_get_wav(self, vpr_idx):
-         res = self.db.select_by_id(vpr_idx)
-         return res[0]['wavpath']
-         
-    
+        res = self.db.select_by_id(vpr_idx)
+        return res[0]['wavpath']
+
    def vpr_data(self, idx):
        # 获取对应ID的数据
        res = self.db.select_by_id(idx)
        return res
-    
+
    def vpr_droptable(self):
        # 删除表
        self.db.drop_table()
        # 清空 faiss
        self.index_ip.reset()
-        
--- a/demos/speech_web/speech_server/src/SpeechBase/vpr_encode.py
+++ b/demos/speech_web/speech_server/src/SpeechBase/vpr_encode.py
@ -1,9 +1,12 @@
-from paddlespeech.cli.vector import VectorExecutor
-import numpy as np
 import logging

+import numpy as np
+
+from paddlespeech.cli.vector import VectorExecutor
+
 vector_executor = VectorExecutor()

+
 def get_audio_embedding(path):
    """
    Use vpr_inference to generate embedding of audio
@ -16,5 +19,3 @@ def get_audio_embedding(path):
    except Exception as e:
        logging.error(f"Error with embedding:{e}")
        return None
-
-    
--- a/demos/speech_web/speech_server/src/WebsocketManeger.py
+++ b/demos/speech_web/speech_server/src/WebsocketManeger.py
@ -2,6 +2,7 @@ from typing import List

 from fastapi import WebSocket

+
 class ConnectionManager:
    def __init__(self):
        # 存放激活的ws连接对象
@ -28,4 +29,4 @@ class ConnectionManager:
            await connection.send_text(message)


-manager = ConnectionManager()
+manager = ConnectionManager()
--- a/demos/speech_web/speech_server/src/robot.py
+++ b/demos/speech_web/speech_server/src/robot.py
@ -1,60 +1,64 @@
-from paddlespeech.cli.asr.infer import ASRExecutor
-import soundfile as sf
 import os
-import librosa

+import soundfile as sf
 from src.SpeechBase.asr import ASR
-from src.SpeechBase.tts import TTS
 from src.SpeechBase.nlp import NLP
+from src.SpeechBase.tts import TTS
+
+from paddlespeech.cli.asr.infer import ASRExecutor


 class Robot:
-    def __init__(self, asr_config, tts_config,asr_init_path,
+    def __init__(self,
+                 asr_config,
+                 tts_config,
+                 asr_init_path,
                 ie_model_path=None) -> None:
        self.nlp = NLP(ie_model_path=ie_model_path)
        self.asr = ASR(config_path=asr_config)
        self.tts = TTS(config_path=tts_config)
        self.tts_sample_rate = 24000
        self.asr_sample_rate = 16000
-        
+
        # 流式识别效果不如端到端的模型，这里流式模型与端到端模型分开
        self.asr_model = ASRExecutor()
        self.asr_name = "conformer_wenetspeech"
        self.warm_up_asrmodel(asr_init_path)
-        

-    def warm_up_asrmodel(self, asr_init_path):        
+    def warm_up_asrmodel(self, asr_init_path):
        if not os.path.exists(asr_init_path):
            path_dir = os.path.dirname(asr_init_path)
            if not os.path.exists(path_dir):
                os.makedirs(path_dir, exist_ok=True)
-            
+
            # TTS生成，采样率24000
            text = "生成初始音频"
            self.text2speech(text, asr_init_path)
-            
+
        # asr model初始化
-        self.asr_model(asr_init_path, model=self.asr_name,lang='zh',
-                 sample_rate=16000, force_yes=True)
-        
-    
+        self.asr_model(
+            asr_init_path,
+            model=self.asr_name,
+            lang='zh',
+            sample_rate=16000,
+            force_yes=True)
+
    def speech2text(self, audio_file):
        self.asr_model.preprocess(self.asr_name, audio_file)
        self.asr_model.infer(self.asr_name)
        res = self.asr_model.postprocess()
        return res
-    
+
    def text2speech(self, text, outpath):
        wav = self.tts.offlineTTS(text)
-        sf.write(
-            outpath, wav, samplerate=self.tts_sample_rate)
+        sf.write(outpath, wav, samplerate=self.tts_sample_rate)
        res = wav
        return res
-    
+
    def text2speechStream(self, text):
        for sub_wav_base64 in self.tts.streamTTS(text=text):
            yield sub_wav_base64
-    
+
    def text2speechStreamBytes(self, text):
        for wav_bytes in self.tts.streamTTSBytes(text=text):
            yield wav_bytes
@ -66,5 +70,3 @@ class Robot:
    def ie(self, text):
        result = self.nlp.ie(text)
        return result
-    
-    
--- a/demos/speech_web/speech_server/src/util.py
+++ b/demos/speech_web/speech_server/src/util.py
@ -1,18 +1,13 @@
 import random

+
 def randName(n=5):
-    return "".join(random.sample('zyxwvutsrqponmlkjihgfedcba',n))
+    return "".join(random.sample('zyxwvutsrqponmlkjihgfedcba', n))
+

 def SuccessRequest(result=None, message="ok"):
-    return {
-        "code": 0,
-        "result":result,
-        "message": message
-    }
+    return {"code": 0, "result": result, "message": message}
+

 def ErrorRequest(result=None, message="error"):
-    return {
-        "code": -1,
-        "result":result,
-        "message": message
-    }
+    return {"code": -1, "result": result, "message": message}
--- a/demos/story_talker/README.md
+++ b/demos/story_talker/README.md
@ -1,3 +1,5 @@
+([简体中文](./README_cn.md)|English)
+
 # Story Talker
 ## Introduction
 Storybooks are very important children's enlightenment books, but parents usually don't have enough time to read storybooks for their children. For very young children, they may not understand the Chinese characters in storybooks. Or sometimes, children just want to "listen" but don't want to "read".
--- a/demos/story_talker/README_cn.md
+++ b/demos/story_talker/README_cn.md
@ -0,0 +1,20 @@
+
+(简体中文|[English](./README.md))
+
+# Story Talker
+
+## 简介
+
+故事书是非常重要的儿童启蒙书，但家长通常没有足够的时间为孩子读故事书。对于非常小的孩子，他们可能不理解故事书中的汉字。或有时，孩子们只是想“听”，而不想“读”。
+
+您可以使用 `PaddleOCR` 获取故事书的文本，并通过 `PaddleSpeech` 的 `TTS` 模块进行阅读。
+
+## 使用
+
+运行以下命令行开始：
+
+```
+./run.sh
+```
+
+结果已显示在 [notebook](https://github.com/PaddlePaddle/PaddleSpeech/blob/develop/docs/tutorial/tts/tts_tutorial.ipynb)。
--- a/demos/streaming_asr_server/conf/application.yaml
+++ b/demos/streaming_asr_server/conf/application.yaml
@ -28,6 +28,7 @@ asr_online:
    sample_rate: 16000
    cfg_path: 
    decode_method: 
+    num_decoding_left_chunks: -1
    force_yes: True
    device: 'cpu' # cpu or gpu:id
    decode_method: "attention_rescoring"
--- a/demos/streaming_asr_server/local/rtf_from_log.py
+++ b/demos/streaming_asr_server/local/rtf_from_log.py
@ -34,7 +34,7 @@ if __name__ == '__main__':
    n = 0
    for m in rtfs:
        # not accurate, may have duplicate log
-        n += 1  
+        n += 1
        T += m['T']
        P += m['P']

--- a/demos/streaming_tts_server/conf/tts_online_application.yaml
+++ b/demos/streaming_tts_server/conf/tts_online_application.yaml
@ -29,7 +29,7 @@ tts_online:
    phones_dict: 
    tones_dict: 
    speaker_dict: 
-    spk_id: 0
+    

    # voc (vocoder) choices=['mb_melgan_csmsc, hifigan_csmsc']
    # Both mb_melgan_csmsc and hifigan_csmsc support streaming voc inference
@ -70,7 +70,6 @@ tts_online-onnx:
    phones_dict: 
    tones_dict: 
    speaker_dict: 
-    spk_id: 0
    am_sample_rate: 24000
    am_sess_conf:
        device: "cpu" # set 'gpu:id' or 'cpu'
@ -79,7 +78,7 @@ tts_online-onnx:

    # voc (vocoder) choices=['mb_melgan_csmsc_onnx, hifigan_csmsc_onnx']
    # Both mb_melgan_csmsc_onnx and hifigan_csmsc_onnx support streaming voc inference
-    voc: 'hifigan_csmsc_onnx'
+    voc: 'mb_melgan_csmsc_onnx'
    voc_ckpt: 
    voc_sample_rate: 24000
    voc_sess_conf:
@ -100,4 +99,4 @@ tts_online-onnx:
    voc_pad: 14
    # voc_upsample should be same as n_shift on voc config.
    voc_upsample: 300
-    
+    
--- a/demos/streaming_tts_server/conf/tts_online_ws_application.yaml
+++ b/demos/streaming_tts_server/conf/tts_online_ws_application.yaml
@ -29,7 +29,7 @@ tts_online:
    phones_dict: 
    tones_dict: 
    speaker_dict: 
-    spk_id: 0
+        

    # voc (vocoder) choices=['mb_melgan_csmsc, hifigan_csmsc']
    # Both mb_melgan_csmsc and hifigan_csmsc support streaming voc inference
@ -70,7 +70,6 @@ tts_online-onnx:
    phones_dict: 
    tones_dict: 
    speaker_dict: 
-    spk_id: 0
    am_sample_rate: 24000
    am_sess_conf:
        device: "cpu" # set 'gpu:id' or 'cpu'
@ -79,7 +78,7 @@ tts_online-onnx:

    # voc (vocoder) choices=['mb_melgan_csmsc_onnx, hifigan_csmsc_onnx']
    # Both mb_melgan_csmsc_onnx and hifigan_csmsc_onnx support streaming voc inference
-    voc: 'hifigan_csmsc_onnx'
+    voc: 'mb_melgan_csmsc_onnx'
    voc_ckpt: 
    voc_sample_rate: 24000
    voc_sess_conf:
@ -100,4 +99,4 @@ tts_online-onnx:
    voc_pad: 14
    # voc_upsample should be same as n_shift on voc config.
    voc_upsample: 300
-    
+    
--- a/demos/style_fs2/README.md
+++ b/demos/style_fs2/README.md
@ -1,3 +1,5 @@
+([简体中文](./README_cn.md)|English)
+
 # Style FastSpeech2
 ## Introduction
 [FastSpeech2](https://arxiv.org/abs/2006.04558)  is a classical acoustic model for Text-to-Speech synthesis, which introduces controllable speech input, including `phoneme duration`、 `energy` and `pitch`. 
--- a/demos/style_fs2/README_cn.md
+++ b/demos/style_fs2/README_cn.md
@ -0,0 +1,33 @@
+(简体中文|[English](./README.md))
+
+# Style FastSpeech2
+
+## 简介
+
+[FastSpeech2](https://arxiv.org/abs/2006.04558)  是用于语音合成的经典声学模型，它引入了可控语音输入，包括 `phoneme duration` 、 `energy` 和 `pitch` 。
+
+在预测阶段，您可以更改这些变量以获得一些有趣的结果。
+
+例如:
+
+1.  `FastSpeech2` 中的 `duration` 可以控制音频的速度 ，并保持 `pitch` 。（在某些语音工具中，增加速度将增加音调，反之亦然。）
+2. 当我们将一个句子的 `pitch` 设置为平均值并将音素的 `tones` 设置为 `1` 时，我们将获得 `robot-style` 的音色。
+3. 当我们提高成年女性的 `pitch` （比例固定）时，我们会得到 `child-style` 的音色。
+
+句子中不同音素的 `duration` 和 `pitch` 可以具有不同的比例。您可以设置不同的音阶比例来强调或削弱某些音素的发音。
+
+## 运行
+
+运行以下命令行开始：
+
+```
+./run.sh
+```
+
+在 `run.sh`, 会首先执行 `source path.sh` 去设置好环境变量。
+
+如果您想尝试您的句子，请替换 `sentences.txt`中的句子。
+
+更多的细节，请查看 `style_syn.py`。
+
+语音样例可以在 [style-control-in-fastspeech2](https://paddlespeech.readthedocs.io/en/latest/tts/demo.html#style-control-in-fastspeech2) 查看。
--- a/demos/text_to_speech/README.md
+++ b/demos/text_to_speech/README.md
@ -16,8 +16,8 @@ You can choose one way from easy, meduim and hard to install paddlespeech.
 The input of this demo should be a text of the specific language that can be passed via argument.
 ### 3. Usage
 - Command Line (Recommended)
+    The default acoustic model is `Fastspeech2`, and the default vocoder is `HiFiGAN`, the default inference method is dygraph inference. 
    - Chinese
-        The default acoustic model is `Fastspeech2`, and the default vocoder is `Parallel WaveGAN`.
        ```bash
        paddlespeech tts --input "你好，欢迎使用百度飞桨深度学习框架！"
        ```
@ -45,7 +45,33 @@ The input of this demo should be a text of the specific language that can be pas
        You can change `spk_id` here.
        ```bash
        paddlespeech tts --am fastspeech2_vctk --voc pwgan_vctk --input "hello, boys" --lang en --spk_id 0
-        ```   
+        ```
+    - Chinese English Mixed, multi-speaker
+        You can change `spk_id` here.
+        ```bash
+        # The `am` must be `fastspeech2_mix`!
+        # The `lang` must be `mix`!
+        # The voc must be chinese datasets' voc now!
+        # spk 174 is csmcc, spk 175 is ljspeech
+        paddlespeech tts --am fastspeech2_mix --voc hifigan_csmsc --lang mix --input "热烈欢迎您在 Discussions 中提交问题，并在 Issues 中指出发现的 bug。此外，我们非常希望您参与到 Paddle Speech 的开发中！" --spk_id 174 --output mix_spk174.wav
+        paddlespeech tts --am fastspeech2_mix --voc hifigan_aishell3 --lang mix --input "热烈欢迎您在 Discussions 中提交问题，并在 Issues 中指出发现的 bug。此外，我们非常希望您参与到 Paddle Speech 的开发中！" --spk_id 174 --output mix_spk174_aishell3.wav
+        paddlespeech tts --am fastspeech2_mix --voc pwgan_csmsc --lang mix --input "我们的声学模型使用了 Fast Speech Two, 声码器使用了 Parallel Wave GAN and Hifi GAN." --spk_id 175 --output mix_spk175_pwgan.wav
+        paddlespeech tts --am fastspeech2_mix --voc hifigan_csmsc --lang mix --input "我们的声学模型使用了 Fast Speech Two, 声码器使用了 Parallel Wave GAN and Hifi GAN." --spk_id 175 --output mix_spk175.wav
+        ```
+     - Use ONNXRuntime infer：
+        ```bash
+        paddlespeech tts --input "你好，欢迎使用百度飞桨深度学习框架！" --output default.wav --use_onnx True
+        paddlespeech tts --am speedyspeech_csmsc --input "你好，欢迎使用百度飞桨深度学习框架！" --output ss.wav --use_onnx True
+        paddlespeech tts --voc mb_melgan_csmsc --input "你好，欢迎使用百度飞桨深度学习框架！" --output mb.wav --use_onnx True
+        paddlespeech tts --voc pwgan_csmsc --input "你好，欢迎使用百度飞桨深度学习框架！" --output pwgan.wav --use_onnx True
+        paddlespeech tts --am fastspeech2_aishell3 --voc pwgan_aishell3 --input "你好，欢迎使用百度飞桨深度学习框架！" --spk_id 0 --output aishell3_fs2_pwgan.wav --use_onnx True
+        paddlespeech tts --am fastspeech2_aishell3 --voc hifigan_aishell3 --input "你好，欢迎使用百度飞桨深度学习框架！" --spk_id 0 --output aishell3_fs2_hifigan.wav --use_onnx True
+        paddlespeech tts --am fastspeech2_ljspeech --voc pwgan_ljspeech --lang en --input "Life was like a box of chocolates, you never know what you're gonna get." --output lj_fs2_pwgan.wav --use_onnx True
+        paddlespeech tts --am fastspeech2_ljspeech --voc hifigan_ljspeech --lang en --input "Life was like a box of chocolates, you never know what you're gonna get." --output lj_fs2_hifigan.wav --use_onnx True
+        paddlespeech tts --am fastspeech2_vctk --voc pwgan_vctk --input "Life was like a box of chocolates, you never know what you're gonna get." --lang en --spk_id 0 --output vctk_fs2_pwgan.wav --use_onnx True
+        paddlespeech tts --am fastspeech2_vctk --voc hifigan_vctk --input "Life was like a box of chocolates, you never know what you're gonna get." --lang en --spk_id 0 --output vctk_fs2_hifigan.wav --use_onnx True
+         ```
+
  Usage:
  
  ```bash
@ -68,6 +94,8 @@ The input of this demo should be a text of the specific language that can be pas
  - `lang`: Language of tts task. Default: `zh`.
  - `device`: Choose device to execute model inference. Default: default device of paddlepaddle in current environment.
  - `output`: Output wave filepath. Default: `output.wav`.
+  - `use_onnx`: whether to usen ONNXRuntime inference.
+  - `fs`: sample rate for ONNX models when use specified model files.

  Output:
  ```bash
@ -75,54 +103,76 @@ The input of this demo should be a text of the specific language that can be pas
  ```

 - Python API
-  ```python
-  import paddle
-  from paddlespeech.cli.tts import TTSExecutor
-
-  tts_executor = TTSExecutor()
-  wav_file = tts_executor(
-      text='今天的天气不错啊',
-      output='output.wav',
-      am='fastspeech2_csmsc',
-      am_config=None,
-      am_ckpt=None,
-      am_stat=None,
-      spk_id=0,
-      phones_dict=None,
-      tones_dict=None,
-      speaker_dict=None,
-      voc='pwgan_csmsc',
-      voc_config=None,
-      voc_ckpt=None,
-      voc_stat=None,
-      lang='zh',
-      device=paddle.get_device())
-  print('Wave file has been generated: {}'.format(wav_file))
-  ```
-
+    - Dygraph infer:
+        ```python
+        import paddle
+        from paddlespeech.cli.tts import TTSExecutor
+        tts_executor = TTSExecutor()
+        wav_file = tts_executor(
+            text='今天的天气不错啊',
+            output='output.wav',
+            am='fastspeech2_csmsc',
+            am_config=None,
+            am_ckpt=None,
+            am_stat=None,
+            spk_id=0,
+            phones_dict=None,
+            tones_dict=None,
+            speaker_dict=None,
+            voc='pwgan_csmsc',
+            voc_config=None,
+            voc_ckpt=None,
+            voc_stat=None,
+            lang='zh',
+            device=paddle.get_device())
+        print('Wave file has been generated: {}'.format(wav_file))
+        ```
+    - ONNXRuntime infer:
+        ```python
+        from paddlespeech.cli.tts import TTSExecutor
+        tts_executor = TTSExecutor()
+        wav_file = tts_executor(
+            text='对数据集进行预处理',
+            output='output.wav',
+            am='fastspeech2_csmsc',
+            voc='hifigan_csmsc',
+            lang='zh',
+            use_onnx=True,
+            cpu_threads=2)
+        ```
+ 
  Output:
  ```bash
  Wave file has been generated: output.wav
  ```

 ### 4. Pretrained Models
-
 Here is a list of pretrained models released by PaddleSpeech that can be used by command and python API:

 - Acoustic model
-  | Model | Language
+  | Model | Language |
  | :--- | :---: |
-  | speedyspeech_csmsc| zh
-  | fastspeech2_csmsc| zh
-  | fastspeech2_aishell3| zh
-  | fastspeech2_ljspeech| en
-  | fastspeech2_vctk| en
+  |      speedyspeech_csmsc      |    zh    |
+  |      fastspeech2_csmsc       |    zh    |
+  |     fastspeech2_ljspeech     |    en    |
+  |     fastspeech2_aishell3     |    zh    |
+  |       fastspeech2_vctk       |    en    |
+  | fastspeech2_cnndecoder_csmsc |    zh    |
+  |       fastspeech2_mix        |   mix    |
+  |       tacotron2_csmsc        |    zh    |
+  |      tacotron2_ljspeech      |    en    |

 - Vocoder
-  | Model | Language
+  | Model | Language |
  | :--- | :---: |
-  | pwgan_csmsc| zh
-  | pwgan_aishell3| zh
-  | pwgan_ljspeech| en
-  | pwgan_vctk| en
-  | mb_melgan_csmsc| zh
+  |         pwgan_csmsc          |    zh    |
+  |        pwgan_ljspeech        |    en    |
+  |        pwgan_aishell3        |    zh    |
+  |          pwgan_vctk          |    en    |
+  |       mb_melgan_csmsc        |    zh    |
+  |      style_melgan_csmsc      |    zh    |
+  |        hifigan_csmsc         |    zh    |
+  |       hifigan_ljspeech       |    en    |
+  |       hifigan_aishell3       |    zh    |
+  |         hifigan_vctk         |    en    |
+  |        wavernn_csmsc         |    zh    |
--- a/demos/text_to_speech/README_cn.md
+++ b/demos/text_to_speech/README_cn.md
@ -1,26 +1,23 @@
 (简体中文|[English](./README.md))

 # 语音合成
-
 ## 介绍
 语音合成是一种自然语言建模过程，其将文本转换为语音以进行音频演示。

 这个 demo 是一个从给定文本生成音频的实现，它可以通过使用 `PaddleSpeech` 的单个命令或 python 中的几行代码来实现。
-
 ## 使用方法
 ### 1. 安装
 请看[安装文档](https://github.com/PaddlePaddle/PaddleSpeech/blob/develop/docs/source/install_cn.md)。

-你可以从 easy，medium，hard 三中方式中选择一种方式安装。
+你可以从 easy，medium，hard 三种方式中选择一种方式安装。

 ### 2. 准备输入

 这个 demo 的输入是通过参数传递的特定语言的文本。
 ### 3. 使用方法
 - 命令行 (推荐使用)
+     默认的声学模型是 `Fastspeech2`，默认的声码器是 `HiFiGAN`，默认推理方式是动态图推理。
    - 中文
-    
-       默认的声学模型是 `Fastspeech2`，默认的声码器是 `Parallel WaveGAN`.
        ```bash
        paddlespeech tts --input "你好，欢迎使用百度飞桨深度学习框架！"
        ```
@ -34,7 +31,7 @@
        ```
    - 中文， 多说话人
    
-        你可以改变 `spk_id` 。
+        你可以改变 `spk_id`。
        ```bash
        paddlespeech tts --am fastspeech2_aishell3 --voc pwgan_aishell3 --input "你好，欢迎使用百度飞桨深度学习框架！" --spk_id 0
        ```
@ -45,10 +42,36 @@
        ```
    - 英文，多说话人
    
-        你可以改变 `spk_id` 。
+        你可以改变 `spk_id`。
        ```bash
        paddlespeech tts --am fastspeech2_vctk --voc pwgan_vctk --input "hello, boys" --lang en --spk_id 0
        ```
+    - 中英文混合，多说话人
+        你可以改变 `spk_id`。
+        ```bash
+        # The `am` must be `fastspeech2_mix`!
+        # The `lang` must be `mix`!
+        # The voc must be chinese datasets' voc now!
+        # spk 174 is csmcc, spk 175 is ljspeech
+        paddlespeech tts --am fastspeech2_mix --voc hifigan_csmsc --lang mix --input "热烈欢迎您在 Discussions 中提交问题，并在 Issues 中指出发现的 bug。此外，我们非常希望您参与到 Paddle Speech 的开发中！" --spk_id 174 --output mix_spk174.wav
+        paddlespeech tts --am fastspeech2_mix --voc hifigan_aishell3 --lang mix --input "热烈欢迎您在 Discussions 中提交问题，并在 Issues 中指出发现的 bug。此外，我们非常希望您参与到 Paddle Speech 的开发中！" --spk_id 174 --output mix_spk174_aishell3.wav
+        paddlespeech tts --am fastspeech2_mix --voc pwgan_csmsc --lang mix --input "我们的声学模型使用了 Fast Speech Two, 声码器使用了 Parallel Wave GAN and Hifi GAN." --spk_id 175 --output mix_spk175_pwgan.wav
+        paddlespeech tts --am fastspeech2_mix --voc hifigan_csmsc --lang mix --input "我们的声学模型使用了 Fast Speech Two, 声码器使用了 Parallel Wave GAN and Hifi GAN." --spk_id 175 --output mix_spk175.wav
+        ```
+     - 使用 ONNXRuntime 推理：
+        ```bash
+        paddlespeech tts --input "你好，欢迎使用百度飞桨深度学习框架！" --output default.wav --use_onnx True
+        paddlespeech tts --am speedyspeech_csmsc --input "你好，欢迎使用百度飞桨深度学习框架！" --output ss.wav --use_onnx True
+        paddlespeech tts --voc mb_melgan_csmsc --input "你好，欢迎使用百度飞桨深度学习框架！" --output mb.wav --use_onnx True
+        paddlespeech tts --voc pwgan_csmsc --input "你好，欢迎使用百度飞桨深度学习框架！" --output pwgan.wav --use_onnx True
+        paddlespeech tts --am fastspeech2_aishell3 --voc pwgan_aishell3 --input "你好，欢迎使用百度飞桨深度学习框架！" --spk_id 0 --output aishell3_fs2_pwgan.wav --use_onnx True
+        paddlespeech tts --am fastspeech2_aishell3 --voc hifigan_aishell3 --input "你好，欢迎使用百度飞桨深度学习框架！" --spk_id 0 --output aishell3_fs2_hifigan.wav --use_onnx True
+        paddlespeech tts --am fastspeech2_ljspeech --voc pwgan_ljspeech --lang en --input "Life was like a box of chocolates, you never know what you're gonna get." --output lj_fs2_pwgan.wav --use_onnx True
+        paddlespeech tts --am fastspeech2_ljspeech --voc hifigan_ljspeech --lang en --input "Life was like a box of chocolates, you never know what you're gonna get." --output lj_fs2_hifigan.wav --use_onnx True
+        paddlespeech tts --am fastspeech2_vctk --voc pwgan_vctk --input "Life was like a box of chocolates, you never know what you're gonna get." --lang en --spk_id 0 --output vctk_fs2_pwgan.wav --use_onnx True
+        paddlespeech tts --am fastspeech2_vctk --voc hifigan_vctk --input "Life was like a box of chocolates, you never know what you're gonna get." --lang en --spk_id 0 --output vctk_fs2_hifigan.wav --use_onnx True
+         ```
+
  使用方法：
  
  ```bash
@ -71,6 +94,8 @@
  - `lang`：TTS 任务的语言， 默认值：`zh`。
  - `device`：执行预测的设备， 默认值：当前系统下 paddlepaddle 的默认 device。
  - `output`：输出音频的路径， 默认值：`output.wav`。
+  - `use_onnx`: 是否使用 ONNXRuntime 进行推理。
+  - `fs`: 使用特定 ONNX 模型时的采样率。

  输出：
  ```bash
@ -78,31 +103,44 @@
  ```

 - Python API
-  ```python
-  import paddle
-  from paddlespeech.cli.tts import TTSExecutor
-
-  tts_executor = TTSExecutor()
-  wav_file = tts_executor(
-      text='今天的天气不错啊',
-      output='output.wav',
-      am='fastspeech2_csmsc',
-      am_config=None,
-      am_ckpt=None,
-      am_stat=None,
-      spk_id=0,
-      phones_dict=None,
-      tones_dict=None,
-      speaker_dict=None,
-      voc='pwgan_csmsc',
-      voc_config=None,
-      voc_ckpt=None,
-      voc_stat=None,
-      lang='zh',
-      device=paddle.get_device())
-  print('Wave file has been generated: {}'.format(wav_file))
-  ```
-
+     - 动态图推理:
+        ```python
+        import paddle
+        from paddlespeech.cli.tts import TTSExecutor
+        tts_executor = TTSExecutor()
+        wav_file = tts_executor(
+            text='今天的天气不错啊',
+            output='output.wav',
+            am='fastspeech2_csmsc',
+            am_config=None,
+            am_ckpt=None,
+            am_stat=None,
+            spk_id=0,
+            phones_dict=None,
+            tones_dict=None,
+            speaker_dict=None,
+            voc='pwgan_csmsc',
+            voc_config=None,
+            voc_ckpt=None,
+            voc_stat=None,
+            lang='zh',
+            device=paddle.get_device())
+        print('Wave file has been generated: {}'.format(wav_file))
+        ```
+    -  ONNXRuntime 推理:
+        ```python
+        from paddlespeech.cli.tts import TTSExecutor
+        tts_executor = TTSExecutor()
+        wav_file = tts_executor(
+            text='对数据集进行预处理',
+            output='output.wav',
+            am='fastspeech2_csmsc',
+            voc='hifigan_csmsc',
+            lang='zh',
+            use_onnx=True,
+            cpu_threads=2)
+        ```
+ 
  输出：
  ```bash
  Wave file has been generated: output.wav
@ -112,19 +150,29 @@
 以下是 PaddleSpeech 提供的可以被命令行和 python API 使用的预训练模型列表：

 - 声学模型
-  | 模型 | 语言
+  | 模型 | 语言 |
  | :--- | :---: |
-  | speedyspeech_csmsc| zh
-  | fastspeech2_csmsc| zh
-  | fastspeech2_aishell3| zh
-  | fastspeech2_ljspeech| en
-  | fastspeech2_vctk| en
+  |      speedyspeech_csmsc      |    zh    |
+  |      fastspeech2_csmsc       |    zh    |
+  |     fastspeech2_ljspeech     |    en    |
+  |     fastspeech2_aishell3     |    zh    |
+  |       fastspeech2_vctk       |    en    |
+  | fastspeech2_cnndecoder_csmsc |    zh    |
+  |       fastspeech2_mix        |   mix    |
+  |       tacotron2_csmsc        |    zh    |
+  |      tacotron2_ljspeech      |    en    |

 - 声码器
-  | 模型 | 语言
+  | 模型 | 语言 |
  | :--- | :---: |
-  | pwgan_csmsc| zh
-  | pwgan_aishell3| zh
-  | pwgan_ljspeech| en
-  | pwgan_vctk| en
-  | mb_melgan_csmsc| zh
+  |         pwgan_csmsc          |    zh    |
+  |        pwgan_ljspeech        |    en    |
+  |        pwgan_aishell3        |    zh    |
+  |          pwgan_vctk          |    en    |
+  |       mb_melgan_csmsc        |    zh    |
+  |      style_melgan_csmsc      |    zh    |
+  |        hifigan_csmsc         |    zh    |
+  |       hifigan_ljspeech       |    en    |
+  |       hifigan_aishell3       |    zh    |
+  |         hifigan_vctk         |    en    |
+  |        wavernn_csmsc         |    zh    |
--- a/docs/requirements.txt
+++ b/docs/requirements.txt
@ -1,12 +1,7 @@
-myst-parser
-numpydoc
-recommonmark>=0.5.0
-sphinx
-sphinx-autobuild
-sphinx-markdown-tables
-sphinx_rtd_theme
-paddlepaddle>=2.2.2
+braceexpand
+colorlog
 editdistance
+fastapi
 g2p_en
 g2pM
 h5py
@ -14,39 +9,45 @@ inflect
 jieba
 jsonlines
 kaldiio
+keyboard
 librosa==0.8.1
 loguru
 matplotlib
+myst-parser
 nara_wpe
-onnxruntime
-pandas
+numpydoc
+onnxruntime==1.10.0
+opencc
 paddlenlp
+paddlepaddle>=2.2.2
 paddlespeech_feat
+pandas
+pathos == 0.2.8
+pattern_singleton
 Pillow>=9.0.0
 praatio==5.0.0
-pypinyin
+prettytable
+pypinyin<=0.44.0
 pypinyin-dict
 python-dateutil
 pyworld==0.2.12
+recommonmark>=0.5.0
 resampy==0.2.2
 sacrebleu
 scipy
 sentencepiece~=0.1.96
 soundfile~=0.10
+sphinx
+sphinx-autobuild
+sphinx-markdown-tables
+sphinx_rtd_theme
 textgrid
 timer
 tqdm
 typeguard
+uvicorn
 visualdl
 webrtcvad
+websockets
 yacs~=0.1.8
-prettytable
 zhon
-colorlog
-pathos == 0.2.8
-fastapi
-websockets
-keyboard
-uvicorn
-pattern_singleton
-braceexpand
--- a/docs/source/api/paddlespeech.audio.rst
+++ b/docs/source/api/paddlespeech.audio.rst
@ -20,4 +20,7 @@ Subpackages
   paddlespeech.audio.io
   paddlespeech.audio.metric
   paddlespeech.audio.sox_effects
+   paddlespeech.audio.streamdata
+   paddlespeech.audio.text
+   paddlespeech.audio.transform
   paddlespeech.audio.utils
--- a/docs/source/api/paddlespeech.audio.streamdata.autodecode.rst
+++ b/docs/source/api/paddlespeech.audio.streamdata.autodecode.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.streamdata.autodecode module
+===============================================
+
+.. automodule:: paddlespeech.audio.streamdata.autodecode
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.streamdata.cache.rst
+++ b/docs/source/api/paddlespeech.audio.streamdata.cache.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.streamdata.cache module
+==========================================
+
+.. automodule:: paddlespeech.audio.streamdata.cache
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.streamdata.compat.rst
+++ b/docs/source/api/paddlespeech.audio.streamdata.compat.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.streamdata.compat module
+===========================================
+
+.. automodule:: paddlespeech.audio.streamdata.compat
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.streamdata.extradatasets.rst
+++ b/docs/source/api/paddlespeech.audio.streamdata.extradatasets.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.streamdata.extradatasets module
+==================================================
+
+.. automodule:: paddlespeech.audio.streamdata.extradatasets
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.streamdata.filters.rst
+++ b/docs/source/api/paddlespeech.audio.streamdata.filters.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.streamdata.filters module
+============================================
+
+.. automodule:: paddlespeech.audio.streamdata.filters
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.streamdata.gopen.rst
+++ b/docs/source/api/paddlespeech.audio.streamdata.gopen.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.streamdata.gopen module
+==========================================
+
+.. automodule:: paddlespeech.audio.streamdata.gopen
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.streamdata.handlers.rst
+++ b/docs/source/api/paddlespeech.audio.streamdata.handlers.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.streamdata.handlers module
+=============================================
+
+.. automodule:: paddlespeech.audio.streamdata.handlers
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.streamdata.mix.rst
+++ b/docs/source/api/paddlespeech.audio.streamdata.mix.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.streamdata.mix module
+========================================
+
+.. automodule:: paddlespeech.audio.streamdata.mix
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.streamdata.paddle_utils.rst
+++ b/docs/source/api/paddlespeech.audio.streamdata.paddle_utils.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.streamdata.paddle\_utils module
+==================================================
+
+.. automodule:: paddlespeech.audio.streamdata.paddle_utils
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.streamdata.pipeline.rst
+++ b/docs/source/api/paddlespeech.audio.streamdata.pipeline.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.streamdata.pipeline module
+=============================================
+
+.. automodule:: paddlespeech.audio.streamdata.pipeline
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.streamdata.rst
+++ b/docs/source/api/paddlespeech.audio.streamdata.rst
@ -0,0 +1,28 @@
+paddlespeech.audio.streamdata package
+=====================================
+
+.. automodule:: paddlespeech.audio.streamdata
+   :members:
+   :undoc-members:
+   :show-inheritance:
+
+Submodules
+----------
+
+.. toctree::
+   :maxdepth: 4
+
+   paddlespeech.audio.streamdata.autodecode
+   paddlespeech.audio.streamdata.cache
+   paddlespeech.audio.streamdata.compat
+   paddlespeech.audio.streamdata.extradatasets
+   paddlespeech.audio.streamdata.filters
+   paddlespeech.audio.streamdata.gopen
+   paddlespeech.audio.streamdata.handlers
+   paddlespeech.audio.streamdata.mix
+   paddlespeech.audio.streamdata.paddle_utils
+   paddlespeech.audio.streamdata.pipeline
+   paddlespeech.audio.streamdata.shardlists
+   paddlespeech.audio.streamdata.tariterators
+   paddlespeech.audio.streamdata.utils
+   paddlespeech.audio.streamdata.writer
--- a/docs/source/api/paddlespeech.audio.streamdata.shardlists.rst
+++ b/docs/source/api/paddlespeech.audio.streamdata.shardlists.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.streamdata.shardlists module
+===============================================
+
+.. automodule:: paddlespeech.audio.streamdata.shardlists
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.streamdata.tariterators.rst
+++ b/docs/source/api/paddlespeech.audio.streamdata.tariterators.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.streamdata.tariterators module
+=================================================
+
+.. automodule:: paddlespeech.audio.streamdata.tariterators
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.streamdata.utils.rst
+++ b/docs/source/api/paddlespeech.audio.streamdata.utils.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.streamdata.utils module
+==========================================
+
+.. automodule:: paddlespeech.audio.streamdata.utils
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.streamdata.writer.rst
+++ b/docs/source/api/paddlespeech.audio.streamdata.writer.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.streamdata.writer module
+===========================================
+
+.. automodule:: paddlespeech.audio.streamdata.writer
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.text.rst
+++ b/docs/source/api/paddlespeech.audio.text.rst
@ -0,0 +1,16 @@
+paddlespeech.audio.text package
+===============================
+
+.. automodule:: paddlespeech.audio.text
+   :members:
+   :undoc-members:
+   :show-inheritance:
+
+Submodules
+----------
+
+.. toctree::
+   :maxdepth: 4
+
+   paddlespeech.audio.text.text_featurizer
+   paddlespeech.audio.text.utility
--- a/docs/source/api/paddlespeech.audio.text.text_featurizer.rst
+++ b/docs/source/api/paddlespeech.audio.text.text_featurizer.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.text.text\_featurizer module
+===============================================
+
+.. automodule:: paddlespeech.audio.text.text_featurizer
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.text.utility.rst
+++ b/docs/source/api/paddlespeech.audio.text.utility.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.text.utility module
+======================================
+
+.. automodule:: paddlespeech.audio.text.utility
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.transform.add_deltas.rst
+++ b/docs/source/api/paddlespeech.audio.transform.add_deltas.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.transform.add\_deltas module
+===============================================
+
+.. automodule:: paddlespeech.audio.transform.add_deltas
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.transform.channel_selector.rst
+++ b/docs/source/api/paddlespeech.audio.transform.channel_selector.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.transform.channel\_selector module
+=====================================================
+
+.. automodule:: paddlespeech.audio.transform.channel_selector
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.transform.cmvn.rst
+++ b/docs/source/api/paddlespeech.audio.transform.cmvn.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.transform.cmvn module
+========================================
+
+.. automodule:: paddlespeech.audio.transform.cmvn
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.transform.functional.rst
+++ b/docs/source/api/paddlespeech.audio.transform.functional.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.transform.functional module
+==============================================
+
+.. automodule:: paddlespeech.audio.transform.functional
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.transform.perturb.rst
+++ b/docs/source/api/paddlespeech.audio.transform.perturb.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.transform.perturb module
+===========================================
+
+.. automodule:: paddlespeech.audio.transform.perturb
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.transform.rst
+++ b/docs/source/api/paddlespeech.audio.transform.rst
@ -0,0 +1,24 @@
+paddlespeech.audio.transform package
+====================================
+
+.. automodule:: paddlespeech.audio.transform
+   :members:
+   :undoc-members:
+   :show-inheritance:
+
+Submodules
+----------
+
+.. toctree::
+   :maxdepth: 4
+
+   paddlespeech.audio.transform.add_deltas
+   paddlespeech.audio.transform.channel_selector
+   paddlespeech.audio.transform.cmvn
+   paddlespeech.audio.transform.functional
+   paddlespeech.audio.transform.perturb
+   paddlespeech.audio.transform.spec_augment
+   paddlespeech.audio.transform.spectrogram
+   paddlespeech.audio.transform.transform_interface
+   paddlespeech.audio.transform.transformation
+   paddlespeech.audio.transform.wpe
--- a/docs/source/api/paddlespeech.audio.transform.spec_augment.rst
+++ b/docs/source/api/paddlespeech.audio.transform.spec_augment.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.transform.spec\_augment module
+=================================================
+
+.. automodule:: paddlespeech.audio.transform.spec_augment
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.transform.spectrogram.rst
+++ b/docs/source/api/paddlespeech.audio.transform.spectrogram.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.transform.spectrogram module
+===============================================
+
+.. automodule:: paddlespeech.audio.transform.spectrogram
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.transform.transform_interface.rst
+++ b/docs/source/api/paddlespeech.audio.transform.transform_interface.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.transform.transform\_interface module
+========================================================
+
+.. automodule:: paddlespeech.audio.transform.transform_interface
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.transform.transformation.rst
+++ b/docs/source/api/paddlespeech.audio.transform.transformation.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.transform.transformation module
+==================================================
+
+.. automodule:: paddlespeech.audio.transform.transformation
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.transform.wpe.rst
+++ b/docs/source/api/paddlespeech.audio.transform.wpe.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.transform.wpe module
+=======================================
+
+.. automodule:: paddlespeech.audio.transform.wpe
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.utils.check_kwargs.rst
+++ b/docs/source/api/paddlespeech.audio.utils.check_kwargs.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.utils.check\_kwargs module
+=============================================
+
+.. automodule:: paddlespeech.audio.utils.check_kwargs
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.utils.dynamic_import.rst
+++ b/docs/source/api/paddlespeech.audio.utils.dynamic_import.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.utils.dynamic\_import module
+===============================================
+
+.. automodule:: paddlespeech.audio.utils.dynamic_import
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.audio.utils.rst
+++ b/docs/source/api/paddlespeech.audio.utils.rst
@ -12,8 +12,11 @@ Submodules
 .. toctree::
   :maxdepth: 4

+   paddlespeech.audio.utils.check_kwargs
   paddlespeech.audio.utils.download
+   paddlespeech.audio.utils.dynamic_import
   paddlespeech.audio.utils.error
   paddlespeech.audio.utils.log
   paddlespeech.audio.utils.numeric
+   paddlespeech.audio.utils.tensor_utils
   paddlespeech.audio.utils.time
--- a/docs/source/api/paddlespeech.audio.utils.tensor_utils.rst
+++ b/docs/source/api/paddlespeech.audio.utils.tensor_utils.rst
@ -0,0 +1,7 @@
+paddlespeech.audio.utils.tensor\_utils module
+=============================================
+
+.. automodule:: paddlespeech.audio.utils.tensor_utils
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.kws.exps.mdtc.collate.rst
+++ b/docs/source/api/paddlespeech.kws.exps.mdtc.collate.rst
@ -0,0 +1,7 @@
+paddlespeech.kws.exps.mdtc.collate module
+=========================================
+
+.. automodule:: paddlespeech.kws.exps.mdtc.collate
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.kws.exps.mdtc.compute_det.rst
+++ b/docs/source/api/paddlespeech.kws.exps.mdtc.compute_det.rst
@ -0,0 +1,7 @@
+paddlespeech.kws.exps.mdtc.compute\_det module
+==============================================
+
+.. automodule:: paddlespeech.kws.exps.mdtc.compute_det
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.kws.exps.mdtc.plot_det_curve.rst
+++ b/docs/source/api/paddlespeech.kws.exps.mdtc.plot_det_curve.rst
@ -0,0 +1,7 @@
+paddlespeech.kws.exps.mdtc.plot\_det\_curve module
+==================================================
+
+.. automodule:: paddlespeech.kws.exps.mdtc.plot_det_curve
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.kws.exps.mdtc.rst
+++ b/docs/source/api/paddlespeech.kws.exps.mdtc.rst
@ -0,0 +1,19 @@
+paddlespeech.kws.exps.mdtc package
+==================================
+
+.. automodule:: paddlespeech.kws.exps.mdtc
+   :members:
+   :undoc-members:
+   :show-inheritance:
+
+Submodules
+----------
+
+.. toctree::
+   :maxdepth: 4
+
+   paddlespeech.kws.exps.mdtc.collate
+   paddlespeech.kws.exps.mdtc.compute_det
+   paddlespeech.kws.exps.mdtc.plot_det_curve
+   paddlespeech.kws.exps.mdtc.score
+   paddlespeech.kws.exps.mdtc.train
--- a/docs/source/api/paddlespeech.kws.exps.mdtc.score.rst
+++ b/docs/source/api/paddlespeech.kws.exps.mdtc.score.rst
@ -0,0 +1,7 @@
+paddlespeech.kws.exps.mdtc.score module
+=======================================
+
+.. automodule:: paddlespeech.kws.exps.mdtc.score
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.kws.exps.mdtc.train.rst
+++ b/docs/source/api/paddlespeech.kws.exps.mdtc.train.rst
@ -0,0 +1,7 @@
+paddlespeech.kws.exps.mdtc.train module
+=======================================
+
+.. automodule:: paddlespeech.kws.exps.mdtc.train
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.kws.exps.rst
+++ b/docs/source/api/paddlespeech.kws.exps.rst
@ -0,0 +1,15 @@
+paddlespeech.kws.exps package
+=============================
+
+.. automodule:: paddlespeech.kws.exps
+   :members:
+   :undoc-members:
+   :show-inheritance:
+
+Subpackages
+-----------
+
+.. toctree::
+   :maxdepth: 4
+
+   paddlespeech.kws.exps.mdtc
--- a/docs/source/api/paddlespeech.kws.rst
+++ b/docs/source/api/paddlespeech.kws.rst
@ -12,4 +12,5 @@ Subpackages
 .. toctree::
   :maxdepth: 4

+   paddlespeech.kws.exps
   paddlespeech.kws.models
--- a/docs/source/api/paddlespeech.resource.model_alias.rst
+++ b/docs/source/api/paddlespeech.resource.model_alias.rst
@ -0,0 +1,7 @@
+paddlespeech.resource.model\_alias module
+=========================================
+
+.. automodule:: paddlespeech.resource.model_alias
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.resource.pretrained_models.rst
+++ b/docs/source/api/paddlespeech.resource.pretrained_models.rst
@ -0,0 +1,7 @@
+paddlespeech.resource.pretrained\_models module
+===============================================
+
+.. automodule:: paddlespeech.resource.pretrained_models
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.resource.resource.rst
+++ b/docs/source/api/paddlespeech.resource.resource.rst
@ -0,0 +1,7 @@
+paddlespeech.resource.resource module
+=====================================
+
+.. automodule:: paddlespeech.resource.resource
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.resource.rst
+++ b/docs/source/api/paddlespeech.resource.rst
@ -0,0 +1,17 @@
+paddlespeech.resource package
+=============================
+
+.. automodule:: paddlespeech.resource
+   :members:
+   :undoc-members:
+   :show-inheritance:
+
+Submodules
+----------
+
+.. toctree::
+   :maxdepth: 4
+
+   paddlespeech.resource.model_alias
+   paddlespeech.resource.pretrained_models
+   paddlespeech.resource.resource
--- a/docs/source/api/paddlespeech.rst
+++ b/docs/source/api/paddlespeech.rst
@ -16,8 +16,10 @@ Subpackages
   paddlespeech.cli
   paddlespeech.cls
   paddlespeech.kws
+   paddlespeech.resource
   paddlespeech.s2t
   paddlespeech.server
   paddlespeech.t2s
   paddlespeech.text
+   paddlespeech.utils
   paddlespeech.vector
--- a/docs/source/api/paddlespeech.s2t.rst
+++ b/docs/source/api/paddlespeech.s2t.rst
@ -19,5 +19,4 @@ Subpackages
   paddlespeech.s2t.models
   paddlespeech.s2t.modules
   paddlespeech.s2t.training
-   paddlespeech.s2t.transform
   paddlespeech.s2t.utils
--- a/docs/source/api/paddlespeech.server.utils.rst
+++ b/docs/source/api/paddlespeech.server.utils.rst
@ -18,7 +18,6 @@ Submodules
   paddlespeech.server.utils.config
   paddlespeech.server.utils.errors
   paddlespeech.server.utils.exception
-   paddlespeech.server.utils.log
   paddlespeech.server.utils.onnx_infer
   paddlespeech.server.utils.paddle_predictor
   paddlespeech.server.utils.util
--- a/docs/source/api/paddlespeech.t2s.datasets.rst
+++ b/docs/source/api/paddlespeech.t2s.datasets.rst
@ -19,4 +19,5 @@ Submodules
   paddlespeech.t2s.datasets.get_feats
   paddlespeech.t2s.datasets.ljspeech
   paddlespeech.t2s.datasets.preprocess_utils
+   paddlespeech.t2s.datasets.sampler
   paddlespeech.t2s.datasets.vocoder_batch_fn
--- a/docs/source/api/paddlespeech.t2s.datasets.sampler.rst
+++ b/docs/source/api/paddlespeech.t2s.datasets.sampler.rst
@ -0,0 +1,7 @@
+paddlespeech.t2s.datasets.sampler module
+========================================
+
+.. automodule:: paddlespeech.t2s.datasets.sampler
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.t2s.exps.ernie_sat.align.rst
+++ b/docs/source/api/paddlespeech.t2s.exps.ernie_sat.align.rst
@ -0,0 +1,7 @@
+paddlespeech.t2s.exps.ernie\_sat.align module
+=============================================
+
+.. automodule:: paddlespeech.t2s.exps.ernie_sat.align
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.t2s.exps.ernie_sat.normalize.rst
+++ b/docs/source/api/paddlespeech.t2s.exps.ernie_sat.normalize.rst
@ -0,0 +1,7 @@
+paddlespeech.t2s.exps.ernie\_sat.normalize module
+=================================================
+
+.. automodule:: paddlespeech.t2s.exps.ernie_sat.normalize
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.t2s.exps.ernie_sat.preprocess.rst
+++ b/docs/source/api/paddlespeech.t2s.exps.ernie_sat.preprocess.rst
@ -0,0 +1,7 @@
+paddlespeech.t2s.exps.ernie\_sat.preprocess module
+==================================================
+
+.. automodule:: paddlespeech.t2s.exps.ernie_sat.preprocess
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.t2s.exps.ernie_sat.rst
+++ b/docs/source/api/paddlespeech.t2s.exps.ernie_sat.rst
@ -0,0 +1,21 @@
+paddlespeech.t2s.exps.ernie\_sat package
+========================================
+
+.. automodule:: paddlespeech.t2s.exps.ernie_sat
+   :members:
+   :undoc-members:
+   :show-inheritance:
+
+Submodules
+----------
+
+.. toctree::
+   :maxdepth: 4
+
+   paddlespeech.t2s.exps.ernie_sat.align
+   paddlespeech.t2s.exps.ernie_sat.normalize
+   paddlespeech.t2s.exps.ernie_sat.preprocess
+   paddlespeech.t2s.exps.ernie_sat.synthesize
+   paddlespeech.t2s.exps.ernie_sat.synthesize_e2e
+   paddlespeech.t2s.exps.ernie_sat.train
+   paddlespeech.t2s.exps.ernie_sat.utils
--- a/docs/source/api/paddlespeech.t2s.exps.ernie_sat.synthesize.rst
+++ b/docs/source/api/paddlespeech.t2s.exps.ernie_sat.synthesize.rst
@ -0,0 +1,7 @@
+paddlespeech.t2s.exps.ernie\_sat.synthesize module
+==================================================
+
+.. automodule:: paddlespeech.t2s.exps.ernie_sat.synthesize
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.t2s.exps.ernie_sat.synthesize_e2e.rst
+++ b/docs/source/api/paddlespeech.t2s.exps.ernie_sat.synthesize_e2e.rst
@ -0,0 +1,7 @@
+paddlespeech.t2s.exps.ernie\_sat.synthesize\_e2e module
+=======================================================
+
+.. automodule:: paddlespeech.t2s.exps.ernie_sat.synthesize_e2e
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/docs/source/api/paddlespeech.t2s.exps.ernie_sat.train.rst
+++ b/docs/source/api/paddlespeech.t2s.exps.ernie_sat.train.rst
@ -0,0 +1,7 @@
+paddlespeech.t2s.exps.ernie\_sat.train module
+=============================================
+
+.. automodule:: paddlespeech.t2s.exps.ernie_sat.train
+   :members:
+   :undoc-members:
+   :show-inheritance:
--- a/Show More
+++ b/Show More