Multichannel court voice real-time recognition system and method based on anti-crosstalk algorithm
By monitoring the characteristics of the forensic voice lines, adjusting the voice recognition strategy and combining with the trial video verification, the accuracy and reliability problems caused by crosstalk and feature changes in court multi-channel voice recognition are solved, and more efficient real-time recognition is achieved.
Patent Information
- Application Number
- CN202510629784.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-05-16
AI Technical Summary
In a court environment, the multi-channel voice recognition system has reduced recognition accuracy and reliability due to changes in crosstalk and speech characteristics, making it difficult to achieve real-time accurate recognition.
By monitoring the voice characteristics changes of the voice lines, the voice lines are determined and updated, and the voice recognition strategy is adjusted based on the anti-crossing algorithm, and secondary verification is carried out in combination with the trial surveillance video to improve the recognition accuracy.
It effectively reduces the impact of speech feature changes on recognition results, improves the accuracy and reliability of court speech recognition, and ensures the real-time nature of multi-channel speech recognition.
Smart Images

Figure CN120148496B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of speech recognition, and particularly relates to a multi-channel court speech real-time recognition system and method based on an anti-crosstalk algorithm. Background Art
[0002] There are multiple parties in a court, and there are often certain interest conflicts between different parties. Therefore, crosstalk occurs from time to time. To reduce the impact of crosstalk on the accuracy of speech recognition, in the invention patent application CN202111654359.7 "Noise Processing Method for Multiple Sound Sources", feature analysis and mining are performed on the channel speech, the microphone channels collecting environmental noise are excluded, the crosstalk sound sources in the crosstalk channels are eliminated, and the normal sound sources are sent to the speech recognition system for recognition, reducing the influence of crosstalk noise.
[0003] As the trial time and speaking time increase, the speech features of different channel speeches may change accordingly. Therefore, if the changes in the above speech features are ignored, it may not be possible to accurately exclude crosstalk, and it is difficult to ensure the reliability of real-time recognition processing of multi-channel speeches.
[0004] To solve the above technical problems, the present application provides a multi-channel court speech real-time recognition system and method based on an anti-crosstalk algorithm. Summary of the Invention
[0005] To achieve the object of the present invention, the present invention adopts the following technical solutions:
[0006] Specifically, in the first aspect, the present application provides a multi-channel court speech real-time recognition method based on an anti-crosstalk algorithm, which specifically includes:
[0007] S1 When the similarity of the speech features of different speech lines in the court determines that the accuracy of speech recognition processing in the court meets the requirements, proceed to the next step;
[0008] S2 Use the speech distribution data of different speech lines at different time periods to determine the monitored speech lines in the speech lines;
[0009] S3 According to the distribution data of the crosstalk time periods between the monitored speech lines and different speech lines, determine the updated speech lines in the monitored speech lines;
[0010] S4 performs an update process on the voice features of the updated voice line, determines the change situation of the voice features of the updated voice line and the similarity situation with the voice features of other voice lines, and combines the change situations of the voice features of other updated voice lines to determine the update processing strategy for the voice features of different updated voice lines, and determines the voice recognition processing strategy for different voice lines in the courtroom based on the voice features after the update processing.
[0011] The beneficial effects of the present invention are as follows:
[0012] By using the voice distribution data in different time periods in different voice lines, the monitored voice line in the voice lines is determined, realizing the recognition of the monitored voice line with possible changes in voice features from multiple angles such as speech duration and the aggregation situation of speech time periods, avoiding the influence of the change in voice features caused by a long speech duration on the accuracy rate of the voice recognition result in the original method, and improving the accuracy of the voice recognition processing.
[0013] Based on the voice features after the update processing, the voice recognition processing strategy for different voice lines in the courtroom is determined, realizing the influence on the voice recognition accuracy rate in the courtroom from the change situations of the voice features of different updated voice lines, fully considering the influence of the relatively large change in voice features on the similarity degree of the voice features of different voice channels and the recognition processing accuracy rate, thereby ensuring the reliability and accuracy of the voice recognition processing in the courtroom.
[0014] A further technical solution is that the similarity situation of the voice features is determined according to the Euclidean distance function between the voice features of different voice lines.
[0015] A further technical solution is that determining that the voice recognition processing accuracy rate in the courtroom meets the requirements specifically includes:
[0016] Based on the similarity situation of the voice features of different voice lines in the courtroom, determine the similarity coefficient of the voice features of different voice lines;
[0017] According to the similarity coefficient of the voice features of different voice lines, determine whether the voice recognition processing accuracy rate in the courtroom meets the requirements.
[0018] A further technical solution is that when there is a voice line with a similarity coefficient of voice features greater than the preset feature similarity coefficient threshold, it is determined that the voice recognition processing accuracy rate in the courtroom does not meet the requirements.
[0019] A further technical solution is that when the voice recognition processing accuracy rate in the courtroom does not meet the requirements, the voice recognition results of all voice lines are subjected to a secondary verification process in combination with the trial monitoring video.
[0020] A further technical solution lies in that the method for determining the speech recognition processing strategy of the speech line is as follows:
[0021] To update the processed speech features, determine the average value of the variation coefficients of the speech features of different updated speech lines during each update process, and use it as the feature variation value;
[0022] According to the feature variation values of different updated speech lines, determine the speech recognition processing strategy of the speech line.
[0023] A further technical solution lies in that, according to the feature variation values of different updated speech lines, determining the speech recognition processing strategy of the speech line specifically includes:
[0024] When there is no updated speech line with a feature variation value greater than the preset feature variation threshold, perform real-time speech recognition processing on different speech lines using the updated processed speech features;
[0025] When there is an updated speech line with a feature variation value greater than the preset feature variation threshold, regard the updated speech line with a feature variation value greater than the preset feature variation threshold as a feature variation line. When the number of the feature variation lines is greater than the preset variation line number threshold, perform secondary verification processing on the speech recognition results of all speech lines in the remaining time period in combination with the court trial monitoring video;
[0026] When the number of the feature variation lines is not greater than the preset variation line number threshold, only perform secondary verification processing on the speech recognition results of the feature variation lines in the remaining time period in combination with the court trial monitoring video.
[0027] A further technical solution lies in that the secondary verification processing in combination with the court trial monitoring video specifically includes:
[0028] Based on the lip recognition result of the speaker in the speech line in the court trial monitoring video, determine the recognized speech text result of the speech line, and use the recognized speech text result to verify the speech recognition result of the speech line.
[0029] In the second aspect, the present application provides a multi-channel court speech real-time recognition system based on an anti-crosstalk algorithm, adopting the above-mentioned multi-channel court speech real-time recognition method based on an anti-crosstalk algorithm, specifically including:
[0030] A speech line positioning module, an update processing module, and an identification processing module;
[0031] Among them, the speech line positioning module is responsible for determining the updated speech lines in the monitored speech lines;
[0032] The update processing module is responsible for determining the update processing strategies for the voice features of different update voice lines;
[0033] The recognition processing module is responsible for determining the voice recognition processing strategies for different voice lines in the courtroom based on the updated voice features.
[0034] Other features and advantages will be described in the subsequent specification. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification and the drawings.
[0035] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following provides preferred embodiments in conjunction with the accompanying drawings and describes them in detail as follows. Brief Description of the Drawings
[0036] By referring to the accompanying drawings and describing its exemplary embodiments in detail, the above and other features and advantages of the present invention will become more obvious;
[0037] Figure 1 is a flowchart of a multi-channel courtroom voice real-time recognition method based on an anti-crosstalk algorithm;
[0038] Figure 2 is a flowchart for determining that the voice recognition processing accuracy rate in the courtroom meets the requirements;
[0039] Figure 3 is a flowchart of a method for determining a monitored voice line in a voice line;
[0040] Figure 4 is a framework diagram of a multi-channel courtroom voice real-time recognition system based on an anti-crosstalk algorithm. Detailed Embodiments
[0041] To enable those skilled in the art of the present technology to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.
[0042] In this application, by analyzing the update situation of the voice features in the voice line, the change situation of the voice features of different voice lines in the courtroom is determined, and a differentiated voice recognition processing strategy is generated according to the change situation, thereby reducing the impact of the change in voice features on the voice recognition accuracy rate.
[0043] Embodiment 1
[0044] Such as Figure 1As shown in the figure, the present application provides a real-time multi-channel court voice recognition method based on a crosstalk prevention algorithm, which specifically includes:
[0045] S1 When the similarity of the voice features of different voice lines in the court is used to determine that the accuracy rate of voice recognition processing in the court meets the requirements, proceed to the next step;
[0046] A further technical solution is that the similarity of the voice features is determined according to the Euclidean distance function between the voice features of different voice lines.
[0047] Further, as Figure 2 shown, determining that the accuracy rate of voice recognition processing in the court meets the requirements specifically includes:
[0048] Using the similarity of the voice features of different voice lines in the court to determine the similarity coefficient of the voice features of different voice lines;
[0049] According to the similarity coefficient of the voice features of different voice lines, determine whether the accuracy rate of voice recognition processing in the court meets the requirements.
[0050] Specifically, when there is a voice line with a similarity coefficient of voice features greater than the preset feature similarity coefficient threshold, it is determined that the accuracy rate of voice recognition processing in the court does not meet the requirements.
[0051] It should be noted that when the accuracy rate of voice recognition processing in the court does not meet the requirements, the voice recognition results of all voice lines are subjected to secondary verification processing in combination with the court trial monitoring video.
[0052] In another possible embodiment, determining that the accuracy rate of voice recognition processing in the court meets the requirements specifically includes:
[0053] Using the similarity of the voice features of different voice lines in the court to determine the similarity coefficient of the voice features of different voice lines;
[0054] According to the similarity coefficient of the voice features of different voice lines, determine the number of voice lines with the similarity coefficient of voice features within the preset similarity coefficient range, and use it as the number of similar lines;
[0055] Based on the number of similar lines, determine whether the accuracy rate of voice recognition processing in the court meets the requirements.
[0056] Further, when the number of similar lines is greater than the preset number of similar line threshold, it is determined that the accuracy rate of voice recognition processing in the court does not meet the requirements.
[0057] Optionally, it is determined that the accuracy rate of speech recognition processing in the court meets the requirements, specifically including:
[0058] Based on the similarity of the speech features of different speech lines in the court, the similarity coefficient of the speech features of different speech lines is determined. When there is a speech line with a similarity coefficient of speech features greater than the preset feature similarity coefficient threshold, it is determined that the accuracy rate of speech recognition processing in the court does not meet the requirements;
[0059] When there is no speech line with a similarity coefficient of speech features greater than the preset feature similarity coefficient threshold:
[0060] Based on the similarity coefficient of the speech features of different speech lines, the speech lines with the similarity coefficient of speech features within the preset similarity coefficient interval are determined. When there is no speech line with the similarity coefficient of speech features within the preset similarity coefficient interval, it is determined that the accuracy rate of speech recognition processing in the court meets the requirements;
[0061] When there is a speech line with a similarity coefficient of speech features within the preset similarity coefficient interval:
[0062] The number of speech lines with the similarity coefficient of speech features within the preset similarity coefficient interval is used as the number of similar lines. When either the number of similar lines or the proportion of the number of similar lines in the number of speech lines does not meet the requirements, it is determined that the accuracy rate of speech recognition processing in the court does not meet the requirements;
[0063] When both the number of similar lines and the proportion of the number of similar lines in the number of speech lines meet the requirements:
[0064] The speech lines with the similarity coefficient of speech features within the preset similarity coefficient interval are used as similar lines, and based on the line types of different similar lines, when the number of target line types belonging to similar lines does not meet the requirements, it is determined that the accuracy rate of speech recognition processing in the court does not meet the requirements;
[0065] When the number of target line types belonging to similar lines meets the requirements:
[0066] Based on the line types of different similar lines, the preset crosstalk probability between different similar lines is determined, and in combination with the similarity coefficient of the speech features between different similar lines, the speech recognition deviation amount of the court is determined. Based on the speech recognition deviation amount, it is determined whether the accuracy rate of speech recognition processing in the court meets the requirements.
[0067] Furthermore, the target line types include the parties and their litigation representatives on both sides.
[0068] Specifically, when the number of similar lines of the target line type is more than 2, it is determined that the number of similar lines of the target line type does not meet the requirements.
[0069] S2 uses the voice distribution data at different time periods in different voice lines to determine the monitored voice line in the voice lines;
[0070] Specifically, as Figure 3 shown, the method for determining the monitored voice line in the voice lines is:
[0071] Based on the voice distribution data at different time periods in the voice line, determine the cumulative speech duration of the voice line from the current moment;
[0072] According to the cumulative speech duration, determine whether the voice line is a monitored voice line.
[0073] Further, when the cumulative speech duration of the voice line is greater than the preset speech duration threshold, it is determined that the voice line is a monitored voice line.
[0074] Optionally, the method for determining the monitored voice line in the voice lines is:
[0075] Based on the voice distribution data at different time periods in the voice line, determine the speech time period between the voice line and the current moment;
[0076] According to the interval duration between different speech time periods, regard the speech time period with an interval duration less than the preset interval duration threshold as an aggregated speech time period;
[0077] Based on the average value of the proportion of the number of the speech time period in the time period between the voice line and the current moment and the proportion of the number of the aggregated speech time period in the time period between the voice line and the current moment, determine the speech anomaly coefficient of the voice line, and according to the speech anomaly coefficient, determine whether the voice line is a monitored voice line.
[0078] Further, when the speech anomaly coefficient of the voice line is greater than the preset speech anomaly coefficient threshold, it is determined that the voice line is a monitored voice line.
[0079] In another possible embodiment, the method for determining the monitored voice line in the voice lines is:
[0080] Based on the voice distribution data at different time periods in the voice line, determine the cumulative speech duration of the voice line from the current moment. When the cumulative speech duration is within the preset speech duration interval, determine whether the voice line is a monitored voice line based on the cumulative speech duration;
[0081] When the cumulative speech duration is not within the preset speech duration range:
[0082] Determine the speech period between the voice line and the current moment. When the number of such speech periods is less than the preset speech period quantity threshold, then determine the voice line as a monitored voice line;
[0083] When the number of the speech periods is not less than the preset speech period quantity threshold:
[0084] According to the interval duration between different speech periods, regard the speech periods with an interval duration less than the preset interval duration threshold as aggregated speech periods. When the number of the aggregated speech periods is greater than the preset aggregated period quantity ratio or the sum of the speech durations of the aggregated speech periods does not meet the requirements, then determine the voice line as a monitored voice line;
[0085] When the number of the aggregated speech periods is not greater than the preset aggregated period quantity ratio and the sum of the speech durations of the aggregated speech periods meets the requirements:
[0086] Determine the voice influence factors of different speech periods based on the duration of different speech periods, the interval duration from the current moment, and the quantity and duration of the speech periods within the interval duration. When there is a speech period with a voice influence factor greater than the preset influence factor threshold, then determine the voice line as a monitored voice line;
[0087] When there is no speech period with a voice influence factor greater than the preset influence factor threshold:
[0088] Based on the voice influence factors of different speech periods, determine the speech anomaly coefficient of the voice line. According to the speech anomaly coefficient, determine whether the voice line is a monitored voice line.
[0089] Furthermore, determining whether the voice line is a monitored voice line based on the cumulative speech duration specifically includes:
[0090] When the cumulative speech duration of the voice line is greater than the preset speech duration threshold, then determine the voice line as a monitored voice line;
[0091] When the cumulative speech duration of the voice line is not greater than the preset speech duration threshold, then determine that the voice line does not belong to the monitored voice lines.
[0092] S3 Determine the updated voice line in the monitored voice line according to the distribution data of the crosstalk periods in the monitored voice line and different voice lines;
[0093] Specifically, the crosstalk period is the period when the monitored voice line and the voice line speak simultaneously.
[0094] Specifically, the method for determining the updated voice line in the monitored voice line is as follows:
[0095] Based on the distribution data of the crosstalk periods between the monitored voice line and different voice lines, determine the number of crosstalk periods in different voice lines;
[0096] Based on the number of crosstalk periods, regard the voice lines with crosstalk periods as crosstalk voice lines;
[0097] Determine whether the monitored voice line is an updated voice line according to the number of crosstalk voice lines.
[0098] Further, when the number of crosstalk voice lines of the monitored voice line is greater than the preset crosstalk voice line number threshold, it is determined that the monitored voice line is an updated voice line.
[0099] In another possible embodiment, the method for determining the updated voice line in the monitored voice line is as follows:
[0100] S31 Based on the distribution data of the crosstalk periods between the monitored voice line and different voice lines, determine the number of crosstalk periods in different voice lines;
[0101] S32 Regard the voice lines with crosstalk periods as crosstalk voice lines, and determine the crosstalk interference coefficients for different crosstalk voice lines according to the number of crosstalk periods and the duration of different crosstalk periods;
[0102] S33 Determine the line interference coefficient of the monitored voice line according to the crosstalk interference coefficients for different crosstalk voice lines, and determine whether the monitored voice line is an updated voice line based on the line interference coefficient.
[0103] Further, when the line interference coefficient is greater than the preset line interference coefficient threshold, it is determined that the monitored voice line is an updated voice line.
[0104] Optionally, the following content is included in step S31 above:
[0105] S311 Based on the distribution data of the crosstalk periods between the monitored voice line and different voice lines, determine the number of crosstalk periods in different voice lines. When the total number of crosstalk periods does not meet the requirements, it is determined that the monitored voice line is an updated voice line. When the total number of crosstalk periods meets the requirements, proceed to step S312;
[0106] S312 determines the total duration of the crosstalk period based on the durations of different crosstalk periods. When the total duration of the crosstalk period does not meet the requirements, it determines that the monitored voice line is an updated voice line. When the total duration of the crosstalk period meets the requirements, it proceeds to step S312;
[0107] S313 When there is no crosstalk period with a duration greater than the crosstalk duration preset value, it proceeds to step S32. When there is a crosstalk period with a duration greater than the crosstalk duration preset value, it proceeds to step S314;
[0108] S314 When the number of crosstalk periods with a duration greater than the crosstalk duration preset value does not meet the requirements, it determines that the monitored voice line is an updated voice line. When the number of crosstalk periods with a duration greater than the crosstalk duration preset value meets the requirements, it proceeds to step S32.
[0109] Optionally, the above step S32 includes the following:
[0110] S321 regards the voice line with a crosstalk period as a crosstalk voice line. When the number of crosstalk voice lines does not meet the requirements, it determines that the monitored voice line is an updated voice line. When the number of crosstalk voice lines meets the requirements, it proceeds to step S322;
[0111] S322 When the total duration of the crosstalk periods with different voice lines all meet the requirements, it determines that the monitored voice line does not belong to the updated voice line. When there is a voice line with a total duration of the crosstalk period that does not meet the requirements, it proceeds to step S323;
[0112] S323 When the number of voice lines with a total duration of the crosstalk period that does not meet the requirements is greater than the line number preset value, it determines that the monitored voice line does not belong to the updated voice line. When the number of voice lines with a total duration of the crosstalk period that does not meet the requirements is not greater than the line number preset value, it proceeds to step S324;
[0113] S324 determines the crosstalk interference coefficients with different crosstalk voice lines according to the number of crosstalk periods with different crosstalk voice lines and the durations of different crosstalk periods. When the sum of the crosstalk interference coefficients with different crosstalk voice lines does not meet the requirements, it determines that the monitored voice line is an updated voice line. When the sum of the crosstalk interference coefficients with different crosstalk voice lines meets the requirements, it proceeds to step S33.
[0114] S4 performs an update process on the voice features of the updated voice line, determines the changes in the voice features of the updated voice line and the similarity with the voice features of other voice lines, and combines the changes in the voice features of other updated voice lines to determine the update processing strategy for the voice features of different updated voice lines. Based on the voice features after the update processing, the voice recognition processing strategy for different voice lines in the courtroom is determined.
[0115] It should be noted that the method for determining the update processing strategy for the voice features of the updated voice line is as follows:
[0116] Based on the changes in the voice features of the updated voice line, determine the updated voice features and the change coefficient of the voice features of the updated voice line;
[0117] Based on the similarity between the updated voice features of the updated voice line and the voice features of other voice lines, determine the similarity coefficient with the voice features of other voice lines, and use the other voice lines within the preset similarity coefficient range as the variable similarity voice lines;
[0118] According to the changes in the voice features of other updated voice lines, determine the change coefficient of the voice features of other updated voice lines, and use the other updated voice lines with a change coefficient greater than the preset change coefficient threshold as the variable voice lines;
[0119] Based on the average value of the change coefficient of the voice features of the updated voice line, the proportion of the number of variable similarity voice lines, and the proportion of the number of variable voice lines, determine the update necessity coefficient of the voice features of the updated voice line, and based on the update necessity coefficient, determine the update processing strategy for the voice features of the updated voice line.
[0120] Furthermore, the change coefficient of the voice features of the updated voice line is determined according to the deviation amount between the updated voice features of the updated voice line and the voice features.
[0121] Optionally, determining the update processing strategy for the voice features of the updated voice line based on the update necessity coefficient specifically includes:
[0122] When the update necessity coefficient is greater than the preset update necessity coefficient threshold, perform an update process on the voice features of each segment of the speech voice of the updated voice line;
[0123] When the update necessity coefficient is not greater than the preset update necessity coefficient threshold, judge whether the update necessity coefficient is within the preset necessity coefficient range. If so, there is no need to perform an update process on the voice features of the updated voice line. If not, perform an update process on the voice features of the updated voice line based on the preset pronunciation duration threshold.
[0124] It should be noted that, based on the preset pronunciation duration threshold, the update process of the voice features of the updated voice line is specifically as follows:
[0125] When the speaking duration of the updated voice line since the last update time of the voice features is greater than the preset speaking duration threshold, the update process of the voice features of the updated voice line is performed.
[0126] Furthermore, the method for determining the speech recognition processing strategy of the voice line is as follows:
[0127] Based on the updated voice features, determine the average value of the variation coefficients of the voice features of different updated voice lines at each update process, and use it as the feature variation value;
[0128] According to the feature variation values of different updated voice lines, determine the speech recognition processing strategy of the voice line.
[0129] Specifically, according to the feature variation values of different updated voice lines, determining the speech recognition processing strategy of the voice line specifically includes:
[0130] When there is no updated voice line with a feature variation value greater than the preset feature variation threshold, perform real-time speech recognition processing on different voice lines using the updated voice features;
[0131] When there is an updated voice line with a feature variation value greater than the preset feature variation threshold, regard the updated voice line with a feature variation value greater than the preset feature variation threshold as a feature variation line. When the number of the feature variation lines is greater than the preset variation line number threshold, perform secondary verification processing on the speech recognition results of all voice lines in the remaining time period in combination with the court trial monitoring video;
[0132] When the number of the feature variation lines is not greater than the preset variation line number threshold, only perform secondary verification processing on the speech recognition results of the feature variation lines in the remaining time period in combination with the court trial monitoring video.
[0133] Furthermore, the secondary verification processing in combination with the court trial monitoring video specifically includes:
[0134] Based on the lip movement recognition result of the speaker in the voice line in the court trial monitoring video, determine the recognized speech text result of the voice line, and use the recognized speech text result to verify the speech recognition result of the voice line.
[0135] Embodiment 2
[0136] In the second aspect, as Figure 4As shown, the present application provides a multi-channel court voice real-time recognition system based on a crosstalk prevention algorithm, which adopts the above-mentioned multi-channel court voice real-time recognition method based on a crosstalk prevention algorithm, and specifically includes:
[0137] A voice line positioning module, an update processing module, and an identification processing module;
[0138] Among them, the voice line positioning module is responsible for determining the updated voice line in the monitored voice line;
[0139] The update processing module is responsible for determining the update processing strategy of the voice characteristics of different updated voice lines;
[0140] The identification processing module is responsible for determining the voice recognition processing strategy of different voice lines in the court based on the updated voice characteristics.
[0141] Each embodiment in this specification is described in a progressive manner. The same or similar parts between each embodiment can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0142] The above specifically describes certain embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0143] The above is only one or more embodiments of this specification and is not used to limit this specification. For those skilled in the art, one or more embodiments of this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of the claims of this specification.
Claims
1. A multi-channel courtroom speech real-time recognition method based on an anti-crosstalk algorithm, characterized in that: Specifically include: When it is determined that the accuracy of the speech recognition processing in the courtroom meets the requirements based on the similarity of the speech features of different voice lines in the courtroom, proceed to the next step; Determining a monitoring voice line among the voice lines by using voice distribution data in different time periods in different voice lines; determining an updated voice line in the monitored voice line according to distribution data of crosstalk periods in the monitored voice line and different voice lines; performing an update process on the voice features of the updated voice line, determining changes in the voice features of the updated voice line and similarities with the voice features of other voice lines, and determining different voice feature update processing strategies for the updated voice lines based on changes in the voice features of other updated voice lines, and determining voice recognition processing strategies for different voice lines in the courtroom based on the updated voice features; Determining that the accuracy of speech recognition processing in the courtroom meets the requirements specifically includes: Determining similarity coefficients of the voice features of different voice lines in the courtroom based on similarities in voice features. If there is a voice line with a voice feature similarity coefficient greater than a preset feature similarity coefficient threshold, determining that the voice recognition processing accuracy in the courtroom does not meet the requirement. The method for determining the monitoring voice line in the voice line is: Determine the cumulative speech duration of the voice line from the current moment based on the voice distribution data of the voice line at different time periods, and determine the voice line as a monitored voice line when the cumulative speech duration of the voice line is greater than a preset speech duration threshold; The method for determining the updated voice line in the monitored voice line is: The number of crosstalk periods in the monitored voice line and different voice lines is determined using the distribution data of the crosstalk periods. Based on the number of crosstalk periods, the voice lines with crosstalk periods are regarded as crosstalk voice lines. When the number of crosstalk voice lines of the monitored voice line is greater than a preset crosstalk voice line number threshold, the monitored voice line is determined to be an updated voice line.
2. The multi-channel courtroom speech real-time recognition method based on the anti-crosstalk algorithm as claimed in claim 1, characterized in that: The similarity of the speech features is determined based on a Euclidean distance function between the speech features of different speech lines.
3. The multi-channel courtroom speech real-time recognition method based on the anti-crosstalk algorithm as claimed in claim 1, characterized in that: When the accuracy of the voice recognition processing in the court does not meet the requirements, the voice recognition results of all voice lines will be subject to secondary verification in combination with the court trial monitoring video.
4. The multi-channel courtroom speech real-time recognition method based on the anti-crosstalk algorithm as claimed in claim 1, characterized in that: The method for determining the voice recognition processing strategy of the voice line is: Determine an average value of the variation coefficients of the voice features of different updated voice lines during each update process using the updated voice features, and use the average value as the feature variation value; According to different updated characteristic change values of the voice line, a speech recognition processing strategy of the voice line is determined.
5. The multi-channel courtroom speech real-time recognition method based on the anti-crosstalk algorithm as claimed in claim 4 is characterized in that: Determining a speech recognition processing strategy for the voice line based on different updated feature change values of the voice line specifically includes: When there is no updated voice line with a feature change value greater than a preset feature change threshold, the updated voice features are used to perform real-time voice recognition processing on different voice lines; When there is an updated voice line with a feature change value greater than the preset feature change threshold, the updated voice line with a feature change value greater than the preset feature change threshold is regarded as a feature change line. When the number of such feature change lines exceeds the preset change line number threshold, the speech recognition results of all voice lines in the remaining time period are combined with the court trial surveillance video for secondary verification processing; When the number of the characteristic change lines is not greater than the preset change line number threshold, only the voice recognition results of the characteristic change lines in the remaining time period are subjected to secondary verification in combination with the trial monitoring video.
6. The multi-channel courtroom speech real-time recognition method based on the anti-crosstalk algorithm as claimed in claim 5, characterized in that: A secondary verification process is conducted in conjunction with the court trial surveillance video, specifically including: The speech text recognition result of the voice line is determined based on the lip shape recognition result of the speaker in the trial monitoring video, and the speech text recognition result is used to verify the speech recognition result of the voice line.
7. A multi-channel courtroom voice real-time recognition system based on an anti-crosstalk algorithm, using a multi-channel courtroom voice real-time recognition method based on an anti-crosstalk algorithm according to any one of claims 1 to 6, characterized in that: Specifically include: Voice line positioning module, update processing module, identification processing module; The voice line location module is responsible for determining the updated voice line in the monitored voice lines; The update processing module is responsible for determining different update processing strategies for updating the voice features of the voice lines; The recognition processing module is responsible for determining the speech recognition processing strategies for different voice lines in the court based on the updated speech features.
Citation Information
Patent Citations
Anti-crosstalk method based on conference live recording system, electronic device and storage medium
CN111429919A
Noise processing method for multiple sound sources
CN114613377A