Court hearing voice recognition engine training method and device supporting law and law model
By identifying bias analysis and using trial duration threshold filtering, the training data of the trial speech recognition engine was optimized, solving the problem of insufficient accuracy in recognizing professional terms in trial speech recognition, and achieving efficient and accurate trial speech recognition training.
Patent Information
- Application Number
- CN202511112664.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-09
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-08-09
AI Technical Summary
Existing technologies for courtroom speech recognition suffer from insufficient accuracy in recognizing specialized terms, resulting in low training efficiency and difficulty in meeting recognition reliability requirements. This is especially true when courtroom speech contains non-specialized terms, making it difficult to adjust the training data range.
By identifying bias analysis, we can determine professional terms and similar phonetic words, set a trial duration threshold, screen effective training data, adjust differentiated training data, and combine professional terminology analysis in the field of court trials to optimize training strategies and improve recognition accuracy and training efficiency.
It has achieved efficient and accurate court hearing speech recognition training, improved the recognition accuracy of the court hearing speech recognition engine and the relevance of the training data, and avoided the low annotation efficiency caused by recognition problems.
Smart Images

Figure CN120833779A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of speech recognition, and particularly relates to a court speech recognition engine training method and device supporting a law language model. BACKGROUND
[0002] The court speech recognition is different from other speech recognitions in that a large number of professional vocabularies exist, and therefore the recognition processing effect of a general speech recognition model is often difficult to meet the requirements, so that how to conduct targeted training processing of the court speech recognition engine and improve the recognition processing accuracy becomes a technical problem to be solved.
[0003] To solve the above technical problem, in the prior art patent application CN202310646456.4 "Language field model self-adaptive classification method and device for court record automatic generation", the recognition text of the historical court speech before the current court speech is used to perform court field classification to obtain a classification result, and a language field model corresponding to the classification result is used to perform speech recognition on the current court speech, so that the recognition processing accuracy is improved. However, the above technical solution has the following technical defects: During the recognition processing of the court speech recognition engine, there are still some non-professional vocabularies in the court speech, so if all the court speeches are used as training data to improve the training quantity of the speech recognition engine, on the one hand, the training processing efficiency will be slow, and on the other hand, the training result of the speech recognition engine may be difficult to meet the requirements, so that how to adjust the training data range according to the recognition reliability of the speech recognition engine to improve the training processing and artificial labeling processing efficiency becomes a technical problem to be solved.
[0004] To solve the above technical problem, the application provides a court speech recognition engine training method and device supporting a law language model. SUMMARY
[0005] To achieve the object of the application, the application adopts the following technical solutions: Specifically, the application provides a court speech recognition engine training method supporting a law language model, which specifically includes the following steps: S1, when the recognition accuracy of a court speech recognition engine does not meet the requirements according to the recognition deviation of the court speech recognition engine under different professional vocabularies, the next step is entered; S2, based on the recognition deviation, the recognition deviation vocabulary is determined according to the composition data of the recognition deviation vocabulary in the corresponding court field of the court speech recognition engine, and the court duration threshold corresponding to the training data is determined by combining the similar speech vocabularies under different recognition deviation vocabularies, and the effective training data is acquired by using the court duration threshold. S3 utilizes the court trial speech recognition engine to perform recognition processing on the effective training data, and utilizes recognition data of professional vocabulary in different time periods in the recognition processing result, and determines training processing data in the effective training data in combination with recognition deviation conditions of different professional vocabulary. S4 according to the training processing data, after marking, the training processing of the court trial speech recognition engine is carried out, and according to the use data of the training processing data in the effective training data after different training processing times of the court trial speech recognition engine, the adjustment processing strategy of the court trial duration threshold corresponding to the training data is determined.
[0006] The beneficial effects of the present application are: The recognition data of professional vocabulary in different time periods in the recognition processing result, the recognition deviation conditions of different professional vocabulary are utilized to determine the training processing data in the effective training data, so as to realize the evaluation of the demand degree of artificial marking processing of the effective training data from the number and distribution data of the unit time period with higher recognition reliability, ensure the pertinence of artificial marking processing, and avoid the technical problems that the efficiency of artificial marking processing cannot meet the requirements due to recognition problems.
[0007] According to the use data of the training processing data in the effective training data after different training processing times of the court trial speech recognition engine, the adjustment processing strategy of the court trial duration threshold corresponding to the training data is determined, which fully considers the change of the recognition accuracy of professional vocabulary of the court trial speech recognition engine after updating the training processing, further realizes the determination of the differentiated court trial duration threshold adjustment processing strategy, not only ensures that the range of training data can be expanded in time on the basis of higher recognition accuracy, improves the efficiency of training processing, but also avoids the expansion of training data on the basis of lower recognition accuracy, which leads to that the marking and recognition processing efficiency of the effective training data cannot meet the requirements.
[0008] Further technical solutions are that the professional vocabulary is determined according to the court trial field, and the professional vocabulary in the legal documents in the court trial field is determined according to the analysis result.
[0009] It can be understood that for the civil court trial field, professional vocabulary including plaintiff, defendant, litigation third party, recusal, jurisdiction objection, counterclaim, evidence presentation and proof, proof responsibility, illegal evidence exclusion, res judicata, appeal, execution, litigation preservation, etc.
[0010] It should be noted that the court trial field is determined according to the case type, including criminal cases, civil cases, administrative litigation cases.
[0011] The further technical solution is that the identification deviation situation includes the number of identification deviations of the court speech recognition engine in history for different professional vocabularies.
[0012] The further technical solution is that the determination of the identification accuracy of the court speech recognition engine not meeting the requirement specifically includes: The number of identification deviations of the court speech recognition engine in history for different professional vocabularies is determined based on the identification deviation situation for different professional vocabularies. The professional vocabulary with the number of identification deviations is taken as a defect vocabulary. The identification accuracy of the court speech recognition engine is determined whether to meet the requirement according to the number of the defect vocabularies.
[0013] The further technical solution is that the method for determining the adjustment processing strategy of the court duration threshold corresponding to the training data is: The variation of the training processing data in the effective training data is determined based on the identification result of the court speech recognition engine in the effective training data after different training processing times. The training processing time of the newly added effective training data is determined based on the variation, and the training processing time is taken as a newly added processing time. The adjustment processing strategy of the court duration threshold corresponding to the training data is determined according to the newly added processing time and the newly added data of the effective training data in different newly added processing times.
[0014] In a second aspect, the present application provides a computer device, which comprises a memory and a processor connected in communication, and a computer program stored on the memory and capable of running on the processor, and the processor executes the computer program to perform the court speech recognition engine training method supporting the law language model.
[0015] Other features and advantages will be set forth in the following description of the application, and in part will be apparent from the description or can be learned by practice of the application, the objects and other advantages of the application will be realized and obtained by the structures particularly pointed out in the description and the accompanying drawings.
[0016] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the following preferred embodiments are specifically described, and the accompanying drawings are described in detail as follows. BRIEF DESCRIPTION OF DRAWINGS
[0017] The above and other features and advantages of the present application will become more apparent from the following detailed description of example embodiments thereof, taken in conjunction with the accompanying drawings.
[0018] Figure 1 A flowchart of a court speech recognition engine training method supporting a law language model; Figure 2is a flowchart of a process of determining that the recognition accuracy of a court session speech recognition engine does not meet requirements; Figure 3 is a flowchart of a method of determining a court session duration threshold corresponding to training data; Figure 4 is a flowchart of a method of determining a monitoring deviation flue gas parameter at a variable flue gas temperature. DETAILED DESCRIPTION
[0019] In order for those skilled in the art to better understand the technical solutions in the specification, the technical solutions in the specification will be clearly and completely described below in combination with the drawings in the specification. Obviously, the described embodiments are only part of the embodiments of the specification, not all. Based on the embodiments of the specification, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the specification.
[0020] In the present application, by adjusting the range of differentiated training data according to the reliability of the recognition processing of the court session speech recognition engine on professional vocabulary, the technical problem of low recognition reliability is avoided when determining the time period that needs to be manually annotated by using the court session speech recognition engine, and the technical problem of slow efficiency of manual annotation processing is further avoided.
[0021] Embodiment 1 As shown in Figure 1 The present application provides a court session speech recognition engine training method supporting legal language and legal model, specifically including: S1, when the recognition accuracy of the court session speech recognition engine does not meet the requirements according to the recognition deviation of the court session speech recognition engine under different professional vocabulary, proceed to the next step; Further, the professional vocabulary is determined according to the court session field, and the professional vocabulary in the legal documents in the court session field is determined according to the analysis result.
[0022] It can be understood that for the civil court session field, the professional vocabulary includes plaintiff, defendant, litigation third party, recusal, jurisdiction objection, counterclaim, evidence presentation and evidence, proof of responsibility, illegal evidence exclusion, res judicata, appeal, execution, litigation preservation, etc.
[0023] It should be noted that the court session field is determined according to the case type, including criminal cases, civil cases, administrative litigation cases.
[0024] Further, the recognition deviation includes the number of recognition deviations of the court session speech recognition engine under different professional vocabulary in history.
[0025] Specifically, as shown in Figure 2 determining that the recognition accuracy of the court speech recognition engine does not meet the requirement, specifically including: determining the number of recognition deviations of the court speech recognition engine in history for different professional vocabularies based on the recognition deviation situation of different professional vocabularies; taking the professional vocabulary with the number of recognition deviations as a defect vocabulary; determining whether the recognition accuracy of the court speech recognition engine meets the requirement according to the number of defect vocabularies.
[0026] It can be understood that the number of recognition deviations is the number of historical recognition processing of the professional vocabulary whose recognition result does not belong to the professional vocabulary.
[0027] Specifically, when the number of defect vocabularies does not meet the requirement, in a possible embodiment, when the proportion of defect vocabularies in the professional vocabulary under the court speech recognition engine is more than 0.25, it is determined that the recognition accuracy of the court speech recognition engine does not meet the requirement.
[0028] Specifically, when the recognition accuracy of the court speech recognition engine meets the requirement, there is no need to determine the court duration threshold, and all historical court speeches are taken as effective training data.
[0029] Optionally, determining that the recognition accuracy of the court speech recognition engine does not meet the requirement, specifically including: determining the number of recognition deviations of the court speech recognition engine in history for different professional vocabularies based on the recognition deviation situation of different professional vocabularies; determining the recognition deviation vocabulary in the professional vocabulary based on the number of recognition deviations; determining whether the recognition accuracy of the court speech recognition engine meets the requirement according to the number of recognition deviation vocabularies.
[0030] It can be understood that the recognition deviation vocabulary is the professional vocabulary whose number of recognition deviations does not meet the requirement, and specifically, the professional vocabulary whose number of recognition deviations is greater than a preset deviation number threshold is taken as the recognition deviation vocabulary. In a possible embodiment, the professional vocabulary whose number of recognition deviations in history of the court speech recognition engine is more than 20 is taken as the recognition deviation vocabulary.
[0031] Specifically, when the number of recognition deviation vocabularies does not meet the requirement, that is, it is greater than a preset deviation vocabulary number threshold, it is determined that the recognition accuracy of the court speech recognition engine does not meet the requirement.
[0032] S2 determines the recognition bias vocabulary based on the identified recognition bias situation, determines the training data corresponding to the court session duration threshold according to the composition data of the recognition bias vocabulary in the court session field corresponding to the court session speech recognition engine, and in combination with the similar speech vocabulary under different recognition bias vocabularies, and determines the training data corresponding to the court session duration threshold, and uses the court session duration threshold to obtain effective training data; Further, the composition data of the recognition bias vocabulary in the corresponding court session field is determined according to the proportion of the number of the recognition bias vocabulary in the corresponding professional vocabulary in the court session field.
[0033] Specifically, the similar speech vocabulary is a vocabulary similar to the recognition bias vocabulary, and specific examples include deposit and deposit, legal person and person, arbitration and final award, vegetable planting, litigation and lawsuit.
[0034] Specifically, as shown in Figure 3 The method for determining the training data corresponding to the court session duration threshold is: Determine the proportion of the number of the recognition bias vocabulary according to the composition data of the recognition bias vocabulary in the court session field corresponding to the court session speech recognition engine; Determine the recognition bias vocabulary with similar speech vocabulary according to the similar speech vocabulary under different recognition bias vocabularies; Determine the training data corresponding to the court session duration threshold according to the recognition bias vocabulary with similar speech vocabulary and the proportion of the number of the recognition bias vocabulary.
[0035] It can be understood that the training data corresponding to the court session duration threshold is determined according to the recognition bias vocabulary with similar speech vocabulary and the proportion of the number of the recognition bias vocabulary, which specifically includes: Determine the recognition risk value according to the proportion of the recognition bias vocabulary with similar speech vocabulary in the corresponding professional vocabulary in the court session field and the average of the proportion of the recognition bias vocabulary in the corresponding professional vocabulary in the court session field; Determine the court session duration threshold based on the preset court session duration threshold corresponding to the recognition risk value.
[0036] Specifically, the larger the recognition risk value is, the more difficult it is to label and process due to the fact that the court session speech recognition engine may not be able to effectively recognize the period of court session speech with professional vocabulary, so the court session duration threshold is smaller at this time. In one possible embodiment, the court session duration threshold is determined according to the one-to-one correspondence relationship in the preset table, or it can be determined according to the average of the existing court session speech duration that can be used as training data, the product of the difference between 1 and the recognition risk value.
[0037] Optionally, the method for determining the court session duration threshold corresponding to the training data is: S21 determines the number of recognition bias words and the proportion of the number of recognition bias words in the recognition bias words in the court session field corresponding to the court session speech recognition engine according to the composition data of the recognition bias words; Optionally, if the number of recognition bias words or the proportion of the number of recognition bias words is large, i.e., greater than a threshold, in the above step, a preset duration threshold is used to determine the court session duration threshold. It can be understood that the value of the preset duration threshold is 1 hour or the court session duration of the first 10% of the training data in the training data is determined.
[0038] In addition, it should be noted that if the number of recognition bias words is not large and the proportion of the number of recognition bias words is not large, the sum of the recognition bias times of different recognition bias words in the court session speech recognition engine is determined. It can be understood that when the sum of the recognition bias times of different recognition bias words in the court session speech recognition engine is greater than a preset number threshold, a preset duration threshold is used to determine the court session duration threshold.
[0039] In addition, it should be noted that if the sum of the recognition bias times of different recognition bias words in the court session speech recognition engine is not greater than the preset number threshold, the next step is entered.
[0040] S22 determines the recognition bias words with similar speech words according to the similar speech words under different recognition bias words, and determines the number of similar speech words of different recognition bias words. It should be noted that in the above step, the number of recognition bias words with similar speech words is obtained. When the number of recognition bias words with similar speech words does not meet the requirement, i.e., greater than a preset threshold, the recognition bias risk is greater, and therefore a preset duration threshold is used to determine the court session duration threshold on this basis. In addition, it can be understood that when the number of recognition bias words with similar speech words meets the requirement, the number of similar speech words of different recognition bias words also needs to be determined. When the total number of similar speech words of different recognition bias words does not meet the requirement, i.e., greater than a preset number threshold, the recognition bias risk is greater, and therefore a preset duration threshold is used to determine the court session duration threshold on this basis. Further, even if the total number of similar speech words of different recognition bias words meets the requirement, it is also necessary to determine that the number of similar speech words is more than 2 in the recognition bias words, and when the proportion of the number of similar speech words in the professional words corresponding to the 2 or more recognition bias words in the court field does not meet the requirement, in a possible embodiment, more than 0.05, then it is determined that the preset time threshold is used to determine the court time threshold, and otherwise it is transferred to the next step.
[0041] S23 determines the court time threshold corresponding to the training data according to the number of similar speech words of different recognition bias words, in combination with the constituent number proportion of the recognition bias words and the number of the recognition bias words.
[0042] In a possible embodiment, the recognition bias weight value of different recognition bias words is determined according to the number of similar speech words of different recognition bias words, the recognition risk value is determined according to the sum of the recognition bias weight values of the recognition bias times and the constituent number proportion of the recognition bias words, and the court time threshold is determined based on the preset court time threshold corresponding to the recognition risk value.
[0043] It can be understood that, in a possible embodiment, the recognition bias weight value is determined according to the number of similar speech words, specifically according to a preset proportion factor and the number of similar speech words, or determined according to an expert scoring method or a data fitting method.
[0044] In another possible embodiment, the recognition risk value is determined according to the sum of the recognition bias weight values of the recognition bias times and the constituent number proportion of the recognition bias words, and a mathematical function is constructed, in a possible embodiment, the mathematical function includes any one of the product, the mean value or the mathematical model of the analytic hierarchy process.
[0045] Further, the effective training data is the training data with a court time less than the court time threshold.
[0046] S3 uses the court speech recognition engine to perform recognition processing on the effective training data, and uses the recognition data of the professional words in different time periods in the recognition processing result, and combines the recognition bias situations of different professional words to determine the training processing data in the effective training data. Specifically, as shown in Figure 4 The method for determining the training processing data in the effective training data is: The recognition data of the professional words in different time periods is used to determine a unit time period in which there is a professional word, and the unit time period is used as a matching unit time period. determining the matching unit time period without the recognition deviation words according to the recognition deviation of the professional words in different matching unit time periods; determining whether the effective training data is the training processing data based on the matching unit time period without the recognition deviation words and the number of the matching unit time periods in the effective training data.
[0047] It can be understood that when the number of the matching unit time period without the recognition deviation words does not meet the requirement, it is determined that the effective training data is not the training processing data. It can be understood that when the number of the matching unit time period without the recognition deviation words accounts for less than 0.1 in the number of the unit time periods in the effective training data, it is determined that the number of the matching unit time period without the recognition deviation words does not meet the requirement.
[0048] In another possible embodiment, when the number of the matching unit time period without the recognition deviation words meets the requirement, it is determined that the effective training data is the training processing data when the number of the matching unit time period without the recognition deviation words and the number of the matching unit time periods accounts for more than 0.7 in the number of the unit time periods in the effective training data, and otherwise, it is not the training processing data.
[0049] Optionally, the method for determining the training processing data in the effective training data is as follows: S31 determining the unit time period with the professional words based on the recognition data of the professional words in different time periods, taking the unit time period with the professional words as the matching unit time period, and determining the matching unit time period without the recognition deviation words according to the recognition deviation of the professional words in different matching unit time periods, and taking the matching unit time period without the recognition deviation words as the reliable recognition time period; Optionally, in the above step, when the number of the matching unit time period does not meet the requirement, i.e. the number of the matching unit time period in the effective training data is small, it is determined that the effective data is not the training processing data.
[0050] Further, in the above step, when the number of the matching unit time period meets the requirement, i.e. the number of the matching unit time period is greater than the preset matching time period threshold, if the number of the matching unit time period without the recognition deviation words does not meet the requirement, it is determined that the effective training data is not the training processing data. It can be understood that when the number of the matching unit time period without the recognition deviation words accounts for less than 0.1 in the number of the unit time periods in the effective training data, it is determined that the number of the matching unit time period without the recognition deviation words does not meet the requirement.
[0051] In another possible embodiment, if the number of identified reliable time periods meets the requirement, and based on the number of time periods without a matching unit period of an identified bias vocabulary and the number of matching unit periods, and the proportion of the number of unit periods in the valid training data, when the proportion is greater than 0.7 or more, it is determined that the valid training data is training processing data, otherwise, it goes to the next step.
[0052] S32 determines the number of matching unit periods corresponding to different identified bias vocabularies according to the composition data of the matching unit periods of different identified bias vocabularies, and determines the identification matching value of different identified bias vocabularies in combination with the composition data of professional vocabularies in different matching unit periods that do not belong to the identified bias vocabulary. Optionally, in the above step, in a possible embodiment, the identification reliability value of different matching unit periods is determined according to the proportion of the number of professional vocabularies in different matching unit periods that do not belong to the identified bias vocabulary, and the identification matching value of different identified bias vocabularies is determined according to the sum of the identification reliability values of the matching unit periods corresponding to different identified bias vocabularies.
[0053] It can be understood that when there is an identified bias vocabulary with an identification matching value greater than a preset matching threshold, it indicates that the demand for labeling processing in the matching unit period in which the identified bias vocabulary exists is also high, and therefore, on this basis, the valid training data can be directly determined as training processing data.
[0054] In addition, it should be noted that when there is no identified bias vocabulary with an identification matching value greater than a preset matching threshold, it is further determined whether there is an identified bias vocabulary with an identification matching value meeting the requirement, and specifically, when there is no identified bias vocabulary with an identification matching value meeting the requirement, in a possible embodiment, when the identification matching value of different identified bias vocabularies is less than 0.15, it can be determined that the valid training data does not belong to training processing data. Further, when there is an identified bias vocabulary with an identification matching value meeting the requirement, it is further determined whether the sum of the identification matching values of different identified bias vocabularies is greater than the matching threshold, and when the sum of the identification matching values of different identified bias vocabularies is greater than the matching threshold, it can be determined that the valid training data belongs to training processing data, otherwise, it goes to the next step.
[0055] S33 determines whether the valid training data is training processing data based on the number of identified reliable time periods and matching unit periods in the valid training data, and in combination with the identification matching value of different identified bias vocabularies.
[0056] It can be understood that in a possible embodiment, when there is no number of matching unit periods of the identification bias vocabulary and the proportion of the number of the sum of the number of matching unit periods in the unit period in the valid training data as the period matching value, the data matching value of the valid training data is determined according to the average of the sum of the period matching value and the identification matching value of different identification bias vocabularies, and when the data matching value is greater than a preset matching threshold, it is determined that the valid training data is training processing data.
[0057] Specifically, based on the identification data of the professional vocabulary of the training processing data, the unit period with the professional vocabulary is determined, and after the unit period with the professional vocabulary is manually annotated, a training data set is formed, and the court trial speech recognition engine is incrementally trained based on the training data set to obtain the court trial speech recognition engine after training processing.
[0058] It can be understood that the value range of the unit period is 1 minute.
[0059] S4, according to the training processing data, the training processing of the court trial speech recognition engine after annotation, and according to the use data of the training processing data in the valid training data of the court trial speech recognition engine after different training processing times, the adjustment processing strategy of the training data corresponding to the court trial duration threshold is determined.
[0060] Specifically, the method for determining the adjustment processing strategy of the training data corresponding to the court trial duration threshold is: determine the variation of the training processing data in the valid training data according to the recognition result of the court trial speech recognition engine in the valid training data after different training processing times; determine the training processing number of the newly added valid training data based on the variation and take it as a new processing number; determine the adjustment processing strategy of the training data corresponding to the court trial duration threshold according to the new processing number and the new data of the valid training data in different new processing numbers.
[0061] Optionally, if the new processing number is large, that is, the proportion of the new processing number in the different training processing numbers of the court trial speech recognition engine is greater than a preset new proportion threshold and the new processing number is greater than a preset new processing number threshold, in a possible embodiment, greater than 0.7 and the new processing number is greater than 10, then there is no need to determine the court trial duration threshold, directly take all the training data as the basis, and perform the screening of the training processing data and the valid training data, and then perform the update training processing of the court trial speech recognition engine.
[0062] It should be noted that if the proportion of the newly added processing times in the different training processing times of the court trial speech recognition engine is not greater than a preset newly added proportion threshold or the newly added processing times are not greater than a preset newly added processing times threshold, if the proportion of the newly added processing times in the different training processing times of the court trial speech recognition engine is greater than the preset newly added proportion threshold or the newly added processing times are greater than the preset newly added processing times threshold, as long as there are newly added processing times, the recognition results of the effective training data in the training data in different time interval are used to determine whether the court trial time threshold needs to be updated, and specifically, when the proportion of the effective training data satisfies more than one time interval of the required time interval, the maximum value of the end points of the time interval of the effective training data satisfying the required time interval is taken as the new court trial time threshold, and in possible embodiments, if the time interval is less than the preset time threshold, the time interval is less than the second threshold value, and the time interval is less than the court trial time threshold, the maximum preset time threshold is taken as the new court trial time threshold, and when the proportion of the effective training data in the training data in the time interval is greater than 0.5, the maximum value of the end points corresponding to the time interval is determined as the new court trial time threshold.
[0063] If the proportion of the newly added processing times in the different training processing times of the court trial speech recognition engine is not greater than the preset newly added proportion threshold and the newly added processing times are not greater than the preset newly added processing times threshold, if the proportion of the newly added processing times is less than the second proportion threshold and the newly added processing times are less than the preset processing times threshold, that is, less than 0.2 and the newly added processing times are 2 times or less, the updating of the court trial time threshold is temporarily not performed; Further, if the proportion of the newly added processing times is not less than 0.2 or the newly added processing times are not 2 times or less, whether the court trial time threshold needs to be updated is determined according to the proportion of the newly added processing times in the preset number of times, and it can be understood that if the proportion of the newly added processing times in the preset number of times is greater than the preset newly added proportion threshold, the maximum value of the end points of the time interval of the effective training data satisfying the required time interval is taken as the new court trial time threshold, and otherwise, the updating of the court trial time threshold is temporarily not performed.
[0064] Embodiment 2 In a second aspect, the present application provides a computer device, comprising a memory and a processor connected in communication, and a computer program stored on the memory and capable of running on the processor, wherein the processor executes the computer program to perform the court trial speech recognition engine training method supporting the legal language and legal language model.
[0065] The various embodiments in this specification describe the application in progressive stages. Each stage builds upon the previous stages, and each stage can be described in terms of the differences between that stage and the previous stage. For example, the device, apparatus, and non-transitory computer storage medium embodiments are described more quickly because they are substantially similar to the method embodiments. The relevant portions of the method embodiments are referenced.
[0066] The above description describes certain embodiments of the application. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve desirable results. Additionally, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve desirable results. In certain implementations, multitasking and parallel processing can be advantageous.
[0067] The above description is of one or more embodiments of the application and is not intended to limit the application. Those skilled in the art will be able to make various changes and modifications without departing from the spirit and scope of the one or more embodiments of the application. Any further modifications, equivalents and / or alternatives come within the scope of the claims of the application.
Claims
1. A court speech recognition engine training method supporting a French-English model, characterized in that, Specifically comprising: When the recognition accuracy of the court trial speech recognition engine does not meet the requirements based on the recognition bias of different professional vocabulary, the next step is entered; Based on the recognition bias, the recognition bias vocabulary is determined according to the composition data of the recognition bias vocabulary in the corresponding court trial field, and the training data corresponding to the court trial duration threshold is determined by combining the similar speech vocabulary under different recognition bias vocabulary and the recognition data of professional vocabulary in different time periods of the recognition processing result of the court trial speech recognition engine, and the effective training data is obtained by using the court trial duration threshold; The recognition processing of the effective training data is performed by using the court trial speech recognition engine, and the recognition data of professional vocabulary in different time periods is used, and the training processing data in the effective training data is determined by combining the recognition bias of different professional vocabulary. According to the training processing data, the training processing of the court trial speech recognition engine is performed after labeling, and the adjustment processing strategy of the training data corresponding to the court trial duration threshold is determined according to the use data of the training processing data in the effective training data after different training processing times of the court trial speech recognition engine.
2. The court speech recognition engine training method supporting a language model of claim 1, wherein, The professional vocabulary is determined according to the court trial field, and the professional vocabulary in the legal documents in the court trial field is determined according to the analysis result.
3. The court speech recognition engine training method supporting a language model of claim 2, wherein, The court trial field is determined according to the case type, including criminal cases, civil cases and administrative litigation cases.
4. The court speech recognition engine training method supporting a language model of claim 1, wherein, The recognition bias situation includes the recognition bias times of the court trial speech recognition engine in history under different professional vocabulary.
5. The court speech recognition engine training method supporting a language model of claim 1, wherein, When the recognition accuracy of the court trial speech recognition engine does not meet the requirements, the recognition bias times of the court trial speech recognition engine in history under different professional vocabulary is determined based on the recognition bias under different professional vocabulary. The professional vocabulary with recognition bias times is regarded as a defect vocabulary. According to the number of defect vocabulary, it is determined whether the recognition accuracy of the court trial speech recognition engine meets the requirements. When the number of defect vocabulary does not meet the requirements, it is determined that the recognition accuracy of the court trial speech recognition engine does not meet the requirements.
6. The court speech recognition engine training method supporting a language model of claim 5, wherein, The composition data of the recognition bias vocabulary in the corresponding court trial field is determined according to the composition quantity proportion of the recognition bias vocabulary in the professional vocabulary corresponding to the court trial field.
7. The court speech recognition engine training method supporting a language model of claim 1, wherein, The method for determining the training data corresponding to the court trial duration threshold is:
8. The court speech recognition engine training method supporting a language model of claim 1, wherein, The composition quantity proportion of the recognition bias vocabulary is determined according to the composition data of the recognition bias vocabulary in the corresponding court trial field of the court trial speech recognition engine; According to the similar speech vocabulary under different recognition bias vocabulary, the recognition bias vocabulary with similar speech vocabulary is determined; According to the recognition bias vocabulary with similar speech vocabulary and the composition quantity proportion of the recognition bias vocabulary, the training data corresponding to the court trial duration threshold is determined. The method for determining the adjustment processing strategy of the training data corresponding to the court trial duration threshold is:
9. The court speech recognition engine training method supporting a language model of claim 1, wherein, The variation of the training processing data in the effective training data is determined according to the recognition result of the effective training data after different training processing times of the court trial speech recognition engine. Determine the training processing times of the newly added valid training data based on the change condition, and take it as the newly added processing times; According to the newly added processing times and the newly added data of the valid training data in different newly added processing times, determine the adjustment processing strategy of the trial duration threshold corresponding to the training data.
10. A computer apparatus comprising: A memory and a processor connected in communication, and a computer program stored on the memory and capable of running on the processor, characterized in that the processor executes the computer program to perform the court speech recognition engine training method supporting the legal language and legal text model according to any one of claims 1-9.
Citation Information
Patent Citations
Speech recognition result processing method and device
CN107945802A
Automatic error correction method for real-time court hearing speech recognition, storage medium and computing device
CN108984529A
Course trial auxiliary processing method and device, trial auxiliary processing method and device, equipment and medium
CN110704571A
Court trial process real-time inspection method and system through voice recognition
CN112686782A
Audio and video processing method, device and equipment for remote court trial of court and storage medium
CN115665438A