Courtroom speech recognition engine training method and apparatus supporting law speech language models
By adjusting the training data range and duration threshold of the court hearing speech recognition engine, the problem of inaccurate recognition of professional terms in court hearing speech was solved, improving the accuracy and training efficiency of court hearing speech recognition, and ensuring recognition reliability and annotation efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- YUNJIACLOUD COM
- Filing Date
- 2025-08-09
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies for courtroom speech recognition suffer from inaccurate recognition of specialized terms, resulting in low training efficiency and insufficient recognition accuracy, especially when courtroom speech contains non-specialized terms.
By identifying the recognition deviation of the court hearing speech recognition engine under professional vocabulary, and determining when the recognition accuracy does not meet the requirements, the training data range is adjusted. Effective training data is obtained by using the court hearing duration threshold, and the training data is determined and adjusted in combination with the recognition deviation of professional vocabulary, so as to ensure recognition reliability and training efficiency.
The recognition accuracy of the court hearing speech recognition engine was improved, the scope and efficiency of training data were optimized, and the low annotation efficiency caused by inaccurate recognition was avoided, thus achieving efficient court hearing speech recognition training.
Smart Images

Figure CN120833779B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of speech recognition technology, and in particular relates to a training method and apparatus for a court speech recognition engine that supports legal language models. Background Technology
[0002] Unlike other speech recognition methods, courtroom speech recognition involves a large number of specialized terms. As a result, the recognition and processing effects of general speech recognition models are often insufficient. Therefore, how to train and process courtroom speech recognition engines in a targeted manner to improve the accuracy of recognition and processing has become an urgent technical problem to be solved.
[0003] To address the aforementioned technical problems, the existing technical solution, as described in invention patent application CN202310646456.4, "Adaptive Classification Method and Apparatus for Language Domain Model in Automatic Generation of Court Transcripts," classifies the court hearing domain based on the recognized text of historical court hearing transcripts preceding the current transcript, obtains the classification result, and then performs speech recognition on the current transcript based on the language domain model corresponding to the classification result, thus improving the accuracy of the recognition process. However, the above technical solution has the following technical defects:
[0004] When processing court hearing speech using a speech recognition engine, there are often some non-professional words in the court hearing speech. Therefore, if all court hearing speech is used as training data to increase the training quantity of the speech recognition engine, it will slow down the training efficiency and may also make it difficult for the training results of the speech recognition engine to meet the requirements. This makes it an urgent technical problem to solve how to adjust the training data range according to the recognition reliability of the speech recognition engine, thereby improving the efficiency of training and manual annotation.
[0005] To address the aforementioned technical issues, this application provides a method and apparatus for training a courtroom speech recognition engine that supports legal language models. Summary of the Invention
[0006] To achieve the objectives of this invention, the following technical solution is adopted:
[0007] Specifically, this application provides a method for training a courtroom speech recognition engine that supports a legal language model, including:
[0008] S1 determines that the recognition accuracy of the court hearing speech recognition engine does not meet the requirements based on the recognition deviation of the engine under different professional terms, and then proceeds to the next step.
[0009] S2 determines the identification deviation words based on the identification deviation situation, and determines the court hearing duration threshold corresponding to the training data based on the composition data of the identification deviation words in the court hearing domain corresponding to the court hearing speech recognition engine, and combined with the similar speech words under different identification deviation words. The court hearing duration threshold is used to obtain effective training data.
[0010] S3 uses the court hearing speech recognition engine to perform recognition processing on the effective training data, and uses the recognition data of professional words in different time periods in the recognition processing results, and combines the recognition deviation of different professional words to determine the training processing data in the effective training data.
[0011] S4, based on the training data, performs training processing on the court hearing speech recognition engine after annotation, and determines the adjustment processing strategy for the court hearing duration threshold corresponding to the training data based on the usage data of the training data in the effective training data after different training processing times of the court hearing speech recognition engine.
[0012] The beneficial effects of this invention are as follows:
[0013] By utilizing the recognition data of professional terms in different time periods and the recognition deviation of different professional terms in the recognition processing results, the training data in the effective training data is determined. This enables the assessment of the need for manual annotation processing of effective training data based on the quantity and distribution of unit time periods with high recognition reliability. This ensures the pertinence of manual annotation processing and avoids the technical problem that the efficiency of manual annotation processing is difficult to meet the requirements due to recognition problems.
[0014] Based on the usage data of the effective training data of the court hearing speech recognition engine after different training processing times, the adjustment strategy for the court hearing duration threshold corresponding to the training data is determined. This fully considers the changes in the accuracy of the court hearing speech recognition engine in recognizing professional terms after the training processing is updated, and further realizes the determination of a differentiated adjustment strategy for the court hearing duration threshold. This not only ensures that the range of training data can be expanded in a timely manner while maintaining a high recognition accuracy, thus improving the efficiency of training processing, but also avoids the problem of expanding the training data when the recognition accuracy is low, which would result in the labeling and recognition processing efficiency of the effective training data being difficult to meet the requirements.
[0015] A further technical solution is that the specialized terms are determined based on the field of court proceedings, specifically by analyzing the specialized terms in legal documents within the court proceedings.
[0016] Understandably, in the field of civil trials, there are professional terms such as plaintiff, defendant, third party in litigation, recusal, objection to jurisdiction, counterclaim, presentation and examination of evidence, burden of proof, exclusion of illegally obtained evidence, res judicata, appeal, enforcement, and litigation preservation.
[0017] It should be noted that the court hearing areas are determined according to the type of case, including criminal cases, civil cases, and administrative litigation cases.
[0018] A further technical solution is that the recognition deviation includes the number of times the court hearing speech recognition engine has recognized different professional terms in history.
[0019] A further technical solution involves determining that the recognition accuracy of the court hearing speech recognition engine does not meet the requirements, specifically including:
[0020] Based on the recognition deviation under different professional terms, determine the number of recognition deviations of the court hearing speech recognition engine for different professional terms in history;
[0021] Specialized terms that exhibit a high number of identification biases are classified as defective terms.
[0022] Based on the number of defective words, determine whether the recognition accuracy of the court hearing speech recognition engine meets the requirements.
[0023] A further technical solution involves determining the method for adjusting the trial duration threshold corresponding to the training data as follows:
[0024] Based on the recognition results of the court hearing speech recognition engine in the effective training data after different training processing times, the changes in the training processing data in the effective training data are determined.
[0025] Based on the aforementioned changes, determine the number of training processes for which new valid training data exists, and use this number as the new processing count.
[0026] Based on the number of new processing steps and the new data from the effective training data in different number of new processing steps, determine the adjustment strategy for the court hearing duration threshold corresponding to the training data.
[0027] In a second aspect, the present invention provides a computer device comprising: a memory and a processor connected in communication, and a computer program stored in the memory and capable of running on the processor, wherein the processor executes the above-described method for training a courtroom speech recognition engine that supports a legal language model when running the computer program.
[0028] Other features and advantages will be set forth in the following description, and the objects and other advantages of the invention are realized and obtained through the structures particularly pointed out in the description and the drawings.
[0029] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0030] The above and other features and advantages of the present invention will become more apparent from a detailed description of exemplary embodiments thereof with reference to the accompanying drawings.
[0031] Figure 1 This is a flowchart of a training method for a courtroom speech recognition engine that supports legal language models;
[0032] Figure 2 This is a flowchart for determining whether the accuracy of the court hearing speech recognition engine does not meet the requirements;
[0033] Figure 3 This is a flowchart illustrating the method for determining the court hearing duration threshold corresponding to the training data;
[0034] Figure 4 This is a flowchart illustrating the method for determining training data in the effective training data set. Detailed Implementation
[0035] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0036] In this application, by adjusting the range of training data according to the reliability of the court hearing speech recognition engine in recognizing professional terms, the technical problem of low recognition reliability when using the court hearing speech recognition engine to determine the time period that requires manual annotation is avoided, which leads to the technical problem of slow efficiency of manual annotation.
[0037] Example 1
[0038] like Figure 1 As shown, this application provides a method for training a courtroom speech recognition engine that supports a legal language model, specifically including:
[0039] S1 determines that the recognition accuracy of the court hearing speech recognition engine does not meet the requirements based on the recognition deviation of the engine under different professional terms, and then proceeds to the next step.
[0040] Furthermore, the terminology is determined based on the field of court proceedings, specifically by analyzing the terminology used in legal documents within that field.
[0041] Understandably, in the field of civil trials, there are professional terms such as plaintiff, defendant, third party in litigation, recusal, objection to jurisdiction, counterclaim, presentation and examination of evidence, burden of proof, exclusion of illegally obtained evidence, res judicata, appeal, enforcement, and litigation preservation.
[0042] It should be noted that the court hearing areas are determined according to the type of case, including criminal cases, civil cases, and administrative litigation cases.
[0043] Furthermore, the recognition deviation includes the number of times the court hearing speech recognition engine has misrecognized different professional terms in history.
[0044] Specifically, such as Figure 2 As shown, it was determined that the recognition accuracy of the court hearing speech recognition engine did not meet the requirements, specifically including:
[0045] Based on the recognition deviation under different professional terms, determine the number of recognition deviations of the court hearing speech recognition engine for different professional terms in history;
[0046] Specialized terms that exhibit a high number of identification biases are classified as defective terms.
[0047] Based on the number of defective words, determine whether the recognition accuracy of the court hearing speech recognition engine meets the requirements.
[0048] It is understood that the number of identification deviations refers to the number of times in history that the identification result does not belong to the professional terminology.
[0049] Specifically, when the number of defective words does not meet the requirements, in one possible embodiment, when the proportion of defective words in the professional vocabulary under the court hearing speech recognition engine is greater than 0.25, it is determined that the recognition accuracy of the court hearing speech recognition engine does not meet the requirements.
[0050] Specifically, when the recognition accuracy of the court hearing speech recognition engine meets the requirements, there is no need to determine the court hearing duration threshold, and all historical court hearing voices can be used as effective training data.
[0051] Optionally, it can be determined that the recognition accuracy of the court hearing speech recognition engine does not meet the requirements, specifically including:
[0052] Based on the recognition deviation under different professional terms, determine the number of recognition deviations of the court hearing speech recognition engine for different professional terms in history;
[0053] Based on the number of recognition deviations, identify the words with recognition deviations in the professional vocabulary;
[0054] Based on the number of words with recognition deviations, it is determined whether the recognition accuracy of the court hearing speech recognition engine meets the requirements.
[0055] It is understood that the identification deviation words are professional words whose identification deviation number does not meet the requirements. Specifically, professional words whose identification deviation number is greater than the preset deviation number threshold are identified as identification deviation words. In one possible embodiment, professional words whose identification deviation number in the court hearing speech recognition engine is more than 20 times in history are identified as identification deviation words.
[0056] Specifically, when the number of identified biased words does not meet the requirements, i.e., when it exceeds the preset threshold for the number of biased words, it is determined that the recognition accuracy of the court hearing speech recognition engine does not meet the requirements.
[0057] S2 determines the identification deviation words based on the identification deviation situation, and determines the court hearing duration threshold corresponding to the training data based on the composition data of the identification deviation words in the court hearing domain corresponding to the court hearing speech recognition engine, and combined with the similar speech words under different identification deviation words. The court hearing duration threshold is used to obtain effective training data.
[0058] Furthermore, the composition data of the corresponding identification deviation words in the court trial field is determined based on the proportion of the identification deviation words in the professional vocabulary corresponding to the court trial field.
[0059] Specifically, the similar speech words are words that are similar to the words spoken in the words with the recognition deviation, such as deposit and order, legal person and issuer, arbitration and final award, planting vegetables, litigation and lawsuit.
[0060] Specifically, such as Figure 3 As shown, the method for determining the trial duration threshold corresponding to the training data is as follows:
[0061] Based on the composition data of the recognition deviation words in the court hearing domain corresponding to the court hearing speech recognition engine, determine the composition ratio of the recognition deviation words.
[0062] Based on the similar speech words under different recognition deviation words, identify the recognition deviation words that have similar speech words;
[0063] The threshold for court hearing duration corresponding to the training data is determined based on the identification deviation words with similar pronunciations and the proportion of the number of identification deviation words.
[0064] It is understandable that the threshold for court hearing duration corresponding to the training data is determined based on the identification deviation words with similar pronunciations and the proportion of the number of identification deviation words, specifically including:
[0065] The identification risk value is determined based on the proportion of words with similar pronunciations that have identification deviations in the professional vocabulary corresponding to the court trial field, and the average proportion of words with identification deviations in the professional vocabulary corresponding to the court trial field.
[0066] The trial duration threshold is determined based on the preset trial duration threshold corresponding to the identified risk value.
[0067] Specifically, the higher the risk value, the more difficult the annotation process becomes because the court hearing speech recognition engine may not be able to effectively identify the time periods of court hearing speech containing professional vocabulary. Therefore, the lower the court hearing duration threshold is, the more likely it is to be determined. In one possible embodiment, the court hearing duration threshold is determined according to the one-to-one correspondence in the constructed preset table, or it can be determined according to the product of the average value and the difference of the existing court hearing speech duration that can be used as training data, wherein the difference is the difference between 1 and the risk value.
[0068] Optionally, the method for determining the trial duration threshold corresponding to the training data is as follows:
[0069] S21 uses the composition data of the identification deviation words in the court hearing domain corresponding to the court hearing speech recognition engine to determine the proportion of the number of identification deviation words and the number of identification deviation words.
[0070] Optionally, if the number of identified biased words is large or the proportion of the total number of words is large in the above steps, i.e., it exceeds the threshold, then a preset time threshold is used to determine the trial duration threshold. It can be understood that the preset time threshold is set to 1 hour or is determined from the training data of the top 10% of the shortest trial durations in the training data.
[0071] It should also be noted that if the number of words with recognition errors is small and their proportion is not large, then it is necessary to determine whether the sum of the number of recognition errors of different words with recognition errors in the court hearing speech recognition engine meets the requirements. It can be understood that when the sum of the number of recognition errors of different words with recognition errors in the court hearing speech recognition engine is greater than the preset number threshold, then the preset duration threshold is used to determine the court hearing duration threshold.
[0072] It should also be noted that if the sum of the number of recognition errors of different words in the court hearing speech recognition engine is not greater than the preset threshold, then proceed to the next step.
[0073] S22 determines the identification deviation words with similar speech words based on the similar speech words under different identification deviation words, and determines the number of similar speech words for different identification deviation words.
[0074] It should be noted that in the above steps, the number of words with similar pronunciations that have recognition deviations is obtained. When the number of words with similar pronunciations that have recognition deviations does not meet the requirements, that is, when it is greater than the preset threshold, the risk of recognition deviation is relatively high. Therefore, based on this, the preset duration threshold is used to determine the court hearing duration threshold.
[0075] In addition, it is understandable that when the number of similar pronunciation words with recognition deviation meets the requirements, it is still necessary to determine the number of similar pronunciation words of different recognition deviation words. When the total number of similar pronunciation words of different recognition deviation words does not meet the requirements, that is, when it is greater than the preset number threshold, the risk of recognition deviation is relatively large. Therefore, on this basis, the preset duration threshold is used to determine the court hearing duration threshold.
[0076] Furthermore, even if the total number of similar speech words of different identification deviation words meets the requirements, it is still necessary to determine identification deviation words with more than 2 similar speech words. When the proportion of identification deviation words with more than 2 similar speech words in the professional vocabulary corresponding to the court hearing field does not meet the requirements, in one possible embodiment, if it is greater than 0.05, then it is determined that the court hearing duration threshold is determined by using a preset duration threshold. In other cases, proceed to the next step.
[0077] S23 determines the court hearing duration threshold corresponding to the training data based on the number of similar speech words for different identification deviation words, combined with the proportion of the number of identification deviation words and the number of identification deviation words.
[0078] In one possible embodiment, the identification deviation weight value of different identification deviation words is determined by the number of similar speech words of different identification deviation words. The identification risk value is determined based on the sum of the identification deviation weight values of the number of identification deviations and the proportion of the number of identification deviation words. The trial duration threshold is determined based on the preset trial duration threshold corresponding to the identification risk value.
[0079] It is understood that, in one possible embodiment, the recognition deviation weight value is determined based on the number of similar speech words, specifically based on a preset ratio factor and the number of similar speech words, or it can be determined based on expert scoring methods or data fitting methods.
[0080] In another possible embodiment, the risk value is determined by constructing a mathematical function based on the sum of the identification deviation weights of the number of identification deviations and the proportion of the number of words constituting the identification deviations. In one possible embodiment, the mathematical function includes any one of the following: product, mean, or analytic hierarchy process (AHP) mathematical model.
[0081] Furthermore, the effective training data is training data whose trial duration is less than the trial duration threshold.
[0082] S3 uses the court hearing speech recognition engine to perform recognition processing on the effective training data, and uses the recognition data of professional words in different time periods in the recognition processing results, and combines the recognition deviation of different professional words to determine the training processing data in the effective training data.
[0083] Specifically, such as Figure 4 As shown, the method for determining the training processing data in the effective training data is as follows:
[0084] Using the identification data of professional terms in different time periods, the time periods containing professional terms are determined and used as the matching time periods;
[0085] Based on the recognition deviation of professional terms in different matching time periods, determine the matching time periods in which there are no recognition deviations;
[0086] Based on the number of matching time periods in the effective training data that do not contain words with identification bias, it is determined whether the effective training data is training processing data.
[0087] It is understood that when the number of matching time periods without identifying biased words does not meet the requirements, the effective training data is determined not to be training data. It is also understood that when the proportion of matching time periods without identifying biased words in the number of time periods in the effective training data is less than 0.1, the number of matching time periods without identifying biased words does not meet the requirements.
[0088] In another possible embodiment, when the number of matching time periods without identification bias words meets the requirement, the effective training data is determined as training processing data when the ratio of the number of matching time periods without identification bias words and the sum of the number of matching time periods in the unit time periods in the effective training data is greater than 0.7. The others are not considered training processing data.
[0089] Optionally, the method for determining the training processing data in the effective training data is as follows:
[0090] S31 uses the recognition data of professional terms in different time periods to determine the unit time period in which professional terms exist and uses it as the matching unit time period. Based on the recognition deviation of professional terms in different matching unit time periods, it determines the matching unit time period in which there are no recognition deviation words and uses it as the reliable recognition time period.
[0091] Optionally, if the number of matching time periods in the above steps does not meet the requirements, that is, if the number of matching time periods in the effective training data is small, then the effective data is determined not to belong to the training processing data.
[0092] Furthermore, when the number of matching unit time periods in the above steps meets the requirement, that is, when the number of matching unit time periods is greater than the preset threshold for the number of matching unit time periods, if the number of matching unit time periods without identification deviation words does not meet the requirement, then it is determined that the effective training data does not belong to the training processing data. It can be understood that when the proportion of the number of matching unit time periods without identification deviation words in the unit time periods of the effective training data is less than 0.1, it is determined that the number of matching unit time periods without identification deviation words does not meet the requirement.
[0093] In another possible embodiment, if the number of reliable time periods is sufficient, the effective training data is determined as training data based on the proportion of the number of matching time periods without identification bias words and the sum of the number of matching time periods in the effective training data. If the proportion is greater than 0.7, the effective training data is used for training processing. Otherwise, the process proceeds to the next step.
[0094] S32 determines the number of matching unit time periods corresponding to different identification deviation words based on the composition data of matching unit time periods of different identification deviation words, and determines the identification matching value of different identification deviation words by combining the composition data of professional words that do not belong to identification deviation words in different matching unit time periods.
[0095] Optionally, in the above steps, in one possible embodiment, the recognition reliability value of different matching unit time periods is determined based on the proportion of professional terms that do not belong to the recognition deviation words in different matching unit time periods, and the recognition matching value of different recognition deviation words is determined based on the sum of the recognition reliability values of the matching unit time periods corresponding to different recognition deviation words.
[0096] It is understandable that when there are words with identification deviations whose matching values are greater than the preset matching threshold, it indicates that the degree of need for annotation processing is also high in the matching unit time period in which the words with identification deviations exist. Therefore, based on this, the effective training data can be directly determined as training processing data.
[0097] Additionally, it should be noted that when there are no identification deviation words with matching values greater than the preset matching threshold, it is necessary to further determine whether there are identification deviation words with matching values that meet the requirements. Specifically, when there are no identification deviation words with matching values that meet the requirements, in one possible embodiment, if the identification matching values of different identification deviation words are all less than 0.15, it can be determined that the effective training data does not belong to the training processing data.
[0098] Furthermore, when there are identification deviation words whose identification matching values meet the requirements, it is necessary to further determine whether the sum of the identification matching values of different identification deviation words is greater than the matching threshold. If the sum of the identification matching values of different identification deviation words is greater than the matching threshold, then the effective training data can be determined to be training processing data. Otherwise, proceed to the next step.
[0099] S33 determines whether the effective training data is training data based on the number of reliable time periods and matching unit time periods identified in the effective training data, and in combination with the identification matching values of different identification deviation words.
[0100] It is understood that, in one possible embodiment, the proportion of the number of matching time periods without identification bias words and the sum of the number of matching time periods in the unit time period in the effective training data is used as the time period matching value. The data matching value of the effective training data is determined based on the average of the sum of the time period matching value and the identification matching values of different identification bias words. When the data matching value is greater than a preset matching threshold, the effective training data is determined to be training processing data.
[0101] Specifically, based on the recognition data of professional terms in the training data, the time periods containing professional terms are determined. After manually annotating the time periods containing professional terms, a training dataset is formed. The court hearing speech recognition engine is then incrementally trained based on the training dataset to obtain the trained court hearing speech recognition engine.
[0102] It is understood that the value range of the unit time period is 1 minute.
[0103] S4, based on the training data, performs training processing on the court hearing speech recognition engine after annotation, and determines the adjustment processing strategy for the court hearing duration threshold corresponding to the training data based on the usage data of the training data in the effective training data after different training processing times of the court hearing speech recognition engine.
[0104] Specifically, the method for determining the adjustment strategy for the trial duration threshold corresponding to the training data is as follows:
[0105] Based on the recognition results of the court hearing speech recognition engine in the effective training data after different training processing times, the changes in the training processing data in the effective training data are determined.
[0106] Based on the aforementioned changes, determine the number of training processes for which new valid training data exists, and use this number as the new processing count.
[0107] Based on the number of new processing steps and the new data from the effective training data in different number of new processing steps, determine the adjustment strategy for the court hearing duration threshold corresponding to the training data.
[0108] Optionally, if the number of new processing operations is large, that is, when the proportion of new processing operations in different training processing operations of the court hearing speech recognition engine is greater than the preset new processing proportion threshold and the number of new processing operations is greater than the preset new processing number threshold, in a possible embodiment, when it is greater than 0.7 and the number of new processing operations is more than 10, then it is not necessary to determine the court hearing duration threshold. Instead, all training data can be used as the basis to screen the training processing data and effective training data, and then the court hearing speech recognition engine can be updated and trained.
[0109] Additionally, it should be noted that if the proportion of new processing times in different training processing iterations of the court hearing speech recognition engine is not greater than a preset new processing proportion threshold or the number of new processing iterations is not greater than a preset new processing iteration threshold, then if the proportion of new processing times in different training processing iterations of the court hearing speech recognition engine is greater than a preset new processing proportion threshold or the number of new processing iterations is greater than a preset new processing iteration threshold, then as long as there are new processing iterations, the recognition results of the effective training data in the training data within different duration intervals will be used to determine whether an update of the court hearing duration threshold is needed. Specifically, when the proportion of effective training data meets the threshold... If the required duration interval includes more than one duration interval shorter than the trial duration threshold, then the maximum value of the endpoints in the duration interval where the proportion of effective training data satisfies the required duration interval is used as the new trial duration threshold. In a possible embodiment, if the duration interval includes duration intervals shorter than the preset duration threshold, duration intervals shorter than the second duration threshold, and duration intervals shorter than the trial duration threshold, then the largest preset duration threshold is used as the new trial duration threshold. If the proportion of effective training data in the training data within the duration interval is greater than 0.5, then the maximum value of the endpoints corresponding to the duration interval is determined as the new trial duration threshold.
[0110] If the proportion of new processing times in different training processing times of the court hearing speech recognition engine is not greater than the preset new proportion threshold and the number of new processing times is not greater than the preset new processing times threshold, and if it is less than the second proportion threshold and the number of new processing times is less than the preset processing times threshold, i.e. less than 0.2 and the number of new processing times is 2 times or less, then the court hearing duration threshold will not be updated for the time being.
[0111] Furthermore, if the number of new processing operations is not less than 0.2 or is not less than 2, then it is determined whether the court hearing duration threshold needs to be updated based on the proportion of new processing operations within the most recent preset number of operations. In addition, it can be understood that if the proportion of new processing operations within the most recent preset number of operations is above the preset new proportion threshold, then the maximum value of the endpoints in the duration interval where the proportion of effective training data meets the requirements is determined as the new court hearing duration threshold. In other cases, the court hearing duration threshold is not updated for the time being.
[0112] Example 2
[0113] In a second aspect, the present invention provides a computer device comprising: a memory and a processor connected in communication, and a computer program stored in the memory and capable of running on the processor, wherein the processor executes the above-described method for training a courtroom speech recognition engine that supports a legal language model when running the computer program.
[0114] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0115] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0116] The above description is merely one or more embodiments of this specification and is not intended to limit this specification. Various modifications and variations can be made to the one or more embodiments of this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of one or more embodiments of this specification should be included within the scope of the claims of this specification.
Claims
1. A method for training a courtroom speech recognition engine that supports a legal language model, characterized in that, Specifically, it includes: If the recognition accuracy of the court hearing speech recognition engine does not meet the requirements based on the recognition deviation of the engine under different professional terms, proceed to the next step. Based on the aforementioned recognition deviation, the recognition deviation words are determined. According to the composition data of the recognition deviation words in the court hearing domain corresponding to the court hearing speech recognition engine, and combined with the similar speech words under different recognition deviation words, the court hearing duration threshold corresponding to the training data is determined, and the court hearing duration threshold is used to obtain effective training data. The court hearing speech recognition engine is used to recognize and process the effective training data, and the recognition data of professional words in different time periods in the recognition and processing results are used to determine the training processing data in the effective training data, combined with the recognition deviation of different professional words. Based on the training data, the court hearing speech recognition engine is trained after annotation. Based on the usage data of the effective training data of the court hearing speech recognition engine after different training times, the adjustment strategy of the court hearing duration threshold corresponding to the training data is determined. The method for determining the adjustment strategy for the trial duration threshold corresponding to the training data is as follows: Based on the recognition results of the court hearing speech recognition engine in the effective training data after different training processing times, the changes in the training processing data in the effective training data are determined. Based on the changes, determine the number of training processes that have newly added valid training data, and use this as the number of new processing processes. Based on the number of new processing steps and the new data from the effective training data in different number of new processing steps, determine the adjustment strategy for the court hearing duration threshold corresponding to the training data.
2. The courtroom speech recognition engine training method supporting legal language models as described in claim 1, characterized in that, The terminology is determined based on the field of court proceedings and the analysis results of the terminology in legal documents within that field.
3. The courtroom speech recognition engine training method supporting legal language models as described in claim 2, characterized in that, The courtroom jurisdiction is determined based on the type of case, including criminal cases, civil cases, and administrative litigation cases.
4. The courtroom speech recognition engine training method supporting legal language models as described in claim 1, characterized in that, The recognition deviation includes the number of times the court hearing speech recognition engine has misrecognized different professional terms in history.
5. The courtroom speech recognition engine training method supporting legal language models as described in claim 1, characterized in that, The accuracy rate of the court hearing speech recognition engine was determined to be unsatisfactory, specifically including: Based on the recognition deviation under different professional terms, determine the number of recognition deviations of the court hearing speech recognition engine for different professional terms in history; Specialized terms that exhibit a high number of identification biases are classified as defective terms. Based on the number of defective words, determine whether the recognition accuracy of the court hearing speech recognition engine meets the requirements.
6. The courtroom speech recognition engine training method supporting legal language models as described in claim 5, characterized in that, If the number of defective words does not meet the requirements, then the recognition accuracy of the court hearing speech recognition engine is determined to be unsatisfactory.
7. The courtroom speech recognition engine training method supporting legal language models as described in claim 1, characterized in that, The composition data of the corresponding identification deviation words in the court trial field are determined based on the proportion of the number of identification deviation words in the professional vocabulary corresponding to the court trial field.
8. The courtroom speech recognition engine training method supporting legal language models as described in claim 1, characterized in that, The method for determining the trial duration threshold corresponding to the training data is as follows: Based on the composition data of the recognition deviation words in the court hearing domain corresponding to the court hearing speech recognition engine, determine the composition ratio of the recognition deviation words. Based on the similar speech words under different recognition deviation words, identify the recognition deviation words that have similar speech words; The threshold for court hearing duration corresponding to the training data is determined based on the identification deviation words with similar pronunciations and the proportion of the number of identification deviation words.
9. A computer device, comprising: A memory and processor connected by communication, and a computer program stored in the memory and capable of running on the processor, characterized in that, when the processor runs the computer program, it executes a courtroom speech recognition engine training method supporting a legal language model as described in any one of claims 1-8.
Citation Information
Patent Citations
Language domain model self-adaptive classification method oriented to court trial record automatic generation
CN116578739A
Automatic error correction method for real-time court hearing speech recognition, storage medium and computing device
CN108984529A
Court trial process real-time inspection method and system through voice recognition
CN112686782A