Voice evaluation scoring method based on scene demand description, storage medium and equipment
By encoded scene description data into scene vectors and fused with acoustic features, a voice evaluation model based on Transformer network is trained, which solves the problem of insufficient adaptability of speech evaluation models in the prior art in different scenarios, and realizes efficient adaptation of the model in multiple scenarios.
Patent Information
- Application Number
- CN202510583970.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing voice evaluation model has insufficient adaptability in scoring requirements in different schools and regions, resulting in frequent retraining of the model to adapt to the needs of various scenarios.
By inputting scene description data for training, the encoding obtains scene vectors and acoustic features to integrate them, and the speech evaluation model is trained to meet different scene needs. This method uses embedded coding technology to vectorize discrete and continuous scene dimension features, and uses a voice evaluation model based on the network structure of Transformer codec.
Through large amounts of data pre-training, the voice evaluation model has the ability to adapt under various scenario conditions, which reduces the work of model retraining and can effectively adapt to various scenario needs.
Smart Images

Figure CN120220664A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent education technology, and particularly relates to a voice evaluation scoring method, a storage medium and a device based on scenario requirement description. Background Art
[0002] In current various oral exams, voice evaluation models have been used for automatic scoring, which is of great significance for improving the scoring efficiency. However, the scoring of existing voice evaluation models is unstable. Especially for different schools and regions, it is difficult for a single model to meet various scoring requirements. Therefore, the voice evaluation model needs to be retrained with each adjustment of requirements. Summary of the Invention
[0003] Aiming at the deficiencies of the prior art, the present invention aims to provide a voice evaluation scoring method, a storage medium and a device based on scenario requirement description.
[0004] To achieve the above object, the present invention adopts the following technical solutions:
[0005] A voice evaluation scoring method based on scenario requirement description includes the following steps:
[0006] S1. Input scenario description data for training;
[0007] S2. Encode the scenario description input in step S1 to obtain a scenario vector: for discrete scenario dimension features, use an embedding encoding technique for vectorization; for continuous scenario dimension features, divide them into discrete features of N levels, and then use the embedding encoding technique to vectorize the obtained discrete features, where N is a positive integer;
[0008] S3. Fuse each scenario vector obtained by encoding the scenario description in step S2 with the acoustic features for training;
[0009] S4. Input the fusion result of the acoustic features and the scenario vectors obtained in step S3 into a voice evaluation model to train the voice evaluation model;
[0010] S5. When using the voice evaluation model trained in step S4 for scoring, input the scenario description data and the voice file data to be scored into the voice evaluation model, and the voice evaluation model outputs a scoring result.
[0011] Further, in step S1, the scenario description includes vocabulary level, test paper difficulty level, grade, region code or dialect code, and scoring tightness.
[0012] Further, in step S2, the discrete scenario dimension features include the difficulty level of the test paper, grade, area code or dialect code, and vocabulary level, and the continuous scenario dimension feature includes the scoring tightness.
[0013] Further, in step S4, the speech evaluation model adopts a network structure based on the Transformer encoder-decoder.
[0014] The present invention also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented.
[0015] The present invention also provides a computer device, including a processor and a memory, where the memory is used to store a computer program; when the processor executes the computer program, the above method is implemented.
[0016] The beneficial effect of the present invention is that: in the method of the present invention, through a large amount of data pre-training, including vocabulary level, test paper difficulty level, grade, area code or dialect code, and scoring tightness, etc., the obtained speech evaluation model already has the adaptation ability under various scenario conditions. Only by adjusting the specific scenario parameters can it be adapted in this scenario, so as to effectively adapt to various scenario requirements and reduce the work of model retraining. Description of the Drawings
[0017] Figure 1 It is a flowchart of the method according to the embodiment of the present invention. Detailed Embodiment
[0018] The following will further describe the present invention with reference to the drawings. It should be noted that this embodiment is based on the technical solution of the present application, and detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to this embodiment.
[0019] This embodiment provides a speech evaluation scoring method based on scenario requirement description, as Figure 1 shown, including the following steps:
[0020] S1. Input the scenario description data for training, including vocabulary level, test paper difficulty level, grade, area code or dialect code, and scoring tightness; the scoring tightness can be any decimal between 0 and 1.
[0021] S2. Encode the scenario description input in step S1 to obtain a scenario vector: For discrete scenario dimension features (including test paper difficulty level, grade, region code or dialect code, and vocabulary level), use the Embedding encoding technique for vectorization; for continuous scenario dimension features (including marking tightness), divide them into discrete features with N levels, and then use the Embedding encoding technique to vectorize the divided discrete features, where N is a positive integer.
[0022] S3. Fuse each scenario vector obtained by encoding the scenario description in step S2 with the acoustic features used for training.
[0023] In this embodiment, the acoustic features use the FBank features commonly used in speech evaluation and recognition modeling.
[0024] In this embodiment, each scenario vector and acoustic feature are fused by concatenating them in the length direction. Denote the acoustic feature as T*D and the scenario vector as M*D, then the feature after concatenating the acoustic feature and the scenario vector is (T + M)*D, where T represents the length of the acoustic feature, M represents the total number of scenario vectors obtained by encoding the scenario description, and D represents the dimension of the sound feature and the scenario vector, and the dimensions of the sound feature and the scenario vector are the same.
[0025] S4. Input the fusion result of the acoustic feature and the scenario vector obtained in step S3 into the speech evaluation model to train the speech evaluation model.
[0026] In this embodiment, the speech evaluation model uses a network structure based on the Transformer encoder-decoder. The encoder receives T acoustic features and M scenario vectors and encodes them, and then the decoder outputs T scoring results. There can be one or more decoders. When there are multiple decoders, they respectively correspond to different scoring dimensions (such as accuracy, fluency, prosody, etc.).
[0027] S5. When using the speech evaluation model trained in step S4 for scoring, input the scenario description data and the speech file data to be scored into the speech evaluation model, and the speech evaluation model outputs the scoring result.
[0028] For those skilled in the art, various corresponding changes and deformations can be given according to the above technical solutions and concepts, and all these changes and deformations should be included in the protection scope of the claims of the present invention.
Claims
1. A speech evaluation and scoring method based on scenario requirement description, characterized in that: The steps include: S1, input scene description data for training; S2. Encode the scene description input in step S1 to obtain a scene vector: for discrete scene dimension features, use embedded coding technology to vectorize; for continuous scene dimension features, divide them into N levels of discrete features, and then use embedded coding technology to vectorize the discrete features obtained by division, where N is a positive integer; S3, fusing each scene vector obtained by encoding the scene description in step S2 with the acoustic features used for training; S4, inputting the fusion result of the acoustic features and the scene vector obtained in step S3 into the speech evaluation model to train the speech evaluation model; S5. When using the speech evaluation model trained in step S4 for scoring, the scene description data and the speech file data to be scored are input into the speech evaluation model, and the speech evaluation model outputs the scoring result.
2. The method according to claim 1, characterized in that In step S1, the scenario description includes vocabulary level, test paper difficulty level, grade, region code or dialect code, and scoring tightness.
3. The method according to claim 2, characterized in that In step S2, the discrete scenario dimension features include the test paper difficulty level, grade, region code or dialect code and vocabulary level, and the continuous scenario dimension features include the tightness of scoring.
4. The method according to claim 1, characterized in that In step S4, the speech evaluation model adopts a Transformer codec network structure.
5. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
6. A computer device, characterized in that: The method comprises a processor and a memory, wherein the memory is used to store a computer program; when the processor is used to execute the computer program, the method described in any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Rapid speech cognition evaluation method and device
CN114916921A
Speech recognition method and device, electronic equipment and storage medium
CN116469390A
Depression state evaluation method and device based on voice signal, terminal and medium
CN116978409A
Automatic voice testing method and device, electronic equipment and storage medium
CN117877510A
Intelligent conversation shorthand method and system based on voice recognition and medium
CN119724245A