Content recognition method based on multi-language continuous voice stream
By processing and predicting the coincident speech segments in multilingual continuous speech streams, combining coincidence degree and statement correlation verification, the problem of large recognition errors in traditional technology is solved, and more efficient and accurate content recognition is achieved.
Patent Information
- Application Number
- CN202510749882.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-08-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In traditional multilingual continuous speech stream recognition, the recognition error of overlapping speech segments is large and the verification processing is not effective, resulting in low content recognition accuracy.
By extracting the sonic data of the coincident voice segment from the multilingual continuous speech stream, the diffused sonic data prediction and adjustment of the main sound wave data, combined with verification of the overlap degree and the degree of statement correlation, the final content recognition results are obtained.
It reduces the chance of misjudgment of content recognition results and improves the efficiency and accuracy of identification work.
Smart Images

Figure CN120564718A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech recognition technology, and in particular to a content recognition method based on multilingual continuous speech streams. Background Art
[0002] With the acceleration of globalization, the demand for multilingual speech recognition is becoming increasingly urgent in scenarios such as cross-border communication, multinational business, and international travel. For example, at international conferences, where participants come from many different countries, real-time recognition of speech content in multiple languages can facilitate tasks such as meeting notes and translation.
[0003] In traditional technologies, the simultaneous recognition of continuous speech streams in multiple languages often results in overlaps of different speech segments, leading to errors in the recognition of the speech content of the overlapping segments. In addition, the differentiated recognition of the speech content of the overlapping segments simply involves content recognition after noise filtering of different dialects, without further verification, resulting in low accuracy of the content recognition results. Summary of the Invention
[0004] In order to overcome the above-mentioned deficiencies in the prior art, the present application provides a content recognition method based on multilingual continuous speech streams.
[0005] The present application provides a method for content recognition based on a multilingual continuous speech stream, the method comprising:
[0006] Step S1, extracting a first voice stream and a second voice stream having overlapping voice segments from a pre-recognized voice stream after noise removal processing, extracting concentrated main sound wave data from the main sound wave data after sound wave data conversion of the overlapping voice segments, and statistically analyzing the overall sound wave data after the first voice stream and the second voice stream are converted to obtain first whole wave data and second whole wave data, and extracting first auxiliary sound wave data 1 and first auxiliary sound wave data 2 belonging to the two voice segments from the first whole wave data according to the main sound wave data;
[0007] Step S2, based on the first auxiliary sound wave data 1 and the first auxiliary sound wave data 2, the concentrated main sound wave data is predicted and adjusted for diffuse sound wave data to obtain first overlapping segment sound wave data, and the second overlapping segment sound wave data is statistically calculated based on the first overlapping segment sound wave data, the concentrated main sound wave data and the second whole wave data;
[0008] Step S3: If the degree of overlap between the pre-processed sound wave data and the main sound wave data after the integration of the first overlapping sound wave data and the second overlapping sound wave data is greater than or equal to the preset overlap condition threshold, output verification condition one; if it is determined that the correction result one and the correction result two after the corresponding content recognition correction of the first overlapping sound wave data and the second overlapping sound wave data respectively meet the correlation degree relationship between the upper and lower sentences, output verification condition two; based on verification condition one and verification condition two, the voice stream content recognition result is comprehensively obtained.
[0009] Preferably, a multilingual continuous speech stream is received, and noise removal processing is performed on the multilingual continuous speech stream to obtain a pre-recognized speech stream. If overlapping speech segments exist in the pre-recognized speech stream, a first speech stream and a second speech stream containing the overlapping speech segments are extracted, where the second speech stream is other speech streams containing the remaining overlapping speech segments in addition to the first speech stream containing the overlapping speech segments in the multilingual continuous speech stream.
[0010] Preliminary content recognition is performed on the first voice stream and the second voice stream respectively to obtain a first preliminary recognition result and a second preliminary recognition result.
[0011] Preferably, the overlapping speech segments are converted into sound wave data to obtain main sound wave data, and concentrated sound wave data are extracted from the main sound wave data to obtain concentrated main sound wave data;
[0012] Counting the overall sound wave data of the first voice stream and the second voice stream after conversion to obtain first whole wave data and second whole wave data;
[0013] Continuous equidistant segments of sound wave data adjacent to the main sound wave data are extracted from the first whole wave data to obtain the first sound wave data to be measured and the second sound wave data to be measured.
[0014] Preferably, the fluctuation difference between the main sound wave data and the sound wave data to be measured, as well as between the main sound wave data and the sound wave data to be measured, is comprehensively counted to obtain the main measured fluctuation value;
[0015] Extract the first auxiliary sound wave data 1 and the first auxiliary sound wave data 2 belonging to the two speech segments which are closest in time sequence to the main sound wave data and whose fluctuation difference is the same as the main measured fluctuation value from the first whole wave data.
[0016] Preferably, concentrated sound wave data are extracted from the first auxiliary sound wave data 1 and the first auxiliary sound wave data 2 to obtain concentrated consonant sound wave data 1 and concentrated consonant sound wave data 2 respectively;
[0017] The first sound wave pre-adjustment value is obtained by averaging the difference between the first auxiliary sound wave data 1 and the concentrated consonant sound wave data 1 and the difference between the first auxiliary sound wave data 2 and the concentrated consonant sound wave data 2;
[0018] According to the first sound wave preset value, the concentrated main sound wave data is predicted and adjusted for the diffuse sound wave data to obtain the first overlapping segment sound wave data;
[0019] According to the first overlapping segment to be identified, the concentrated main sound wave data and the second whole wave data, a second sound wave pre-adjustment value is calculated, and according to the second sound wave pre-adjustment value, the concentrated main sound wave data is predicted and adjusted for diffuse sound wave data to obtain the second overlapping segment sound wave data.
[0020] Preferably, the first overlapping sound wave data and the second overlapping sound wave data are acoustically integrated to obtain pre-processed sound wave data, and an overlap condition threshold is preset. If the overlap between the pre-processed sound wave data and the main sound wave data is greater than or equal to the overlap condition threshold, the verification condition one is output.
[0021] Preferably, content recognition is performed on the first overlapping sound wave data segment and the second overlapping sound wave data segment to obtain a first overlapping recognition result and a second overlapping recognition result;
[0022] According to the first overlap recognition result, the corresponding overlap segment content recognition result in the first recognition initial result is corrected to obtain correction result one; according to the second overlap recognition result, the corresponding overlap segment content recognition result in the second recognition initial result is corrected to obtain correction result two.
[0023] Preferably, a first relevance determination threshold is preset based on the comprehensive relevance of the context sentences in the first recognition preliminary result, and a second relevance determination threshold is preset based on the comprehensive relevance of the context sentences in the second recognition preliminary result;
[0024] If the correlation degree between the correction result 1 and the context of the first overlapping recognition result is greater than or equal to the correlation determination threshold 1, and the correlation degree between the correction result 2 and the context of the second overlapping recognition result is greater than or equal to the correlation determination threshold 2, then output verification compliance condition 2;
[0025] According to the verification condition 1 and the verification condition 2, the correction result 1 and the first initial recognition result are content integrated, and the correction result 2 and the second initial recognition result are content integrated to comprehensively obtain the voice stream content recognition result.
[0026] Compared with the prior art, the present invention has the following characteristics and beneficial effects:
[0027] By mainly intercepting and analyzing the overlapping speech segments existing in the process of multi-language continuous speech stream recognition, the overlapping speech segments are converted into sound wave data, so as to extract the concentrated sound wave data in the main sound wave data. The concentrated sound wave data in the main sound wave data is the same pronunciation part of multiple languages. Subsequently, the diffuse sound wave data are predicted and adjusted on the concentrated sound wave data respectively. Because a speech stream has a stability law in the pronunciation diffuse sound wave of continuous time sequence, the sound wave data to be tested and measured are obtained by extracting continuous equidistant segments of sound wave data adjacent to the main sound wave data from the first whole wave data belonging to the first speech stream according to the main sound wave data, rather than extracting and comparing and analyzing several segments of sound wave data, so as to reduce the influence of large errors between the sound wave data with low interference and time sequence correlation. According to the fluctuation difference characteristics between the sound wave data to be tested and the sound wave data to be tested and the main sound wave data, two auxiliary sound wave data are refined and extracted from the first whole wave data. According to these two auxiliary sound wave data, the method is realized. Now, the diffuse sound wave data is predicted for the concentrated main sound wave data in the first whole wave data. Similarly, the second whole wave data is analyzed accordingly to predict the first overlapping segment sound wave data and the second overlapping segment sound wave data. In order to improve the reliability of the entire analysis process and the accuracy of the prediction results, it is necessary to verify the predicted first overlapping segment sound wave data and the second overlapping segment sound wave data, that is, one is to judge the degree of overlap between the pre-processed sound wave data and the main sound wave data after the two overlapping segment sound wave data are integrated, and the other is to judge the degree of correlation between the two results of the content recognition of the two overlapping segment sound wave data and the overall content recognition result initially recognized. If both verification conditions are met, the two verification results identified after the prediction adjustment are determined to be the final content recognition results of the overlapping voice segment, and the voice stream content recognition result is obtained comprehensively. Through the above-mentioned two-condition verification processing method, the probability of misjudgment of the content recognition result is reduced and the efficiency of the entire recognition work is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 This is a flowchart of a method for content recognition based on multilingual continuous speech streams, which is mainly embodied in this embodiment. DETAILED DESCRIPTION
[0029] The present invention is further described in detail below with reference to the following examples.
[0030] Reference Figure 1 A content recognition method based on a multilingual continuous speech stream comprises the following steps:
[0031] Step S1: extract the first voice stream and the second voice stream with overlapping voice segments from the pre-recognized voice stream after noise removal processing, extract concentrated main sound wave data from the main sound wave data after sound wave data conversion of the overlapping voice segments, and count the overall sound wave data after the conversion of the first voice stream and the second voice stream to obtain the first whole wave data and the second whole wave data, and according to the main sound wave data, extract the first auxiliary sound wave data 1 and the first auxiliary sound wave data 2 belonging to the two voice segments from the first whole wave data.
[0032] Step S2: Based on the first auxiliary sound wave data 1 and the first auxiliary sound wave data 2, the concentrated main sound wave data is predicted and adjusted for diffuse sound wave data to obtain the first overlapping segment sound wave data; based on the first overlapping segment sound wave data, the concentrated main sound wave data and the second whole wave data, the second overlapping segment sound wave data is statistically calculated.
[0033] Step S3: If the degree of overlap between the pre-processed sound wave data and the main sound wave data after the integration of the first overlapping sound wave data and the second overlapping sound wave data is greater than or equal to the preset overlap condition threshold, output verification condition one; if it is determined that the correction result one and the correction result two after the corresponding content recognition correction of the first overlapping sound wave data and the second overlapping sound wave data respectively meet the correlation degree relationship between the upper and lower sentences, output verification condition two; based on verification condition one and verification condition two, the voice stream content recognition result is comprehensively obtained.
[0034] Specifically, by mainly intercepting and analyzing the overlapping speech segments existing in the multi-language continuous speech stream recognition process, the overlapping speech segments are converted into sound wave data to facilitate the extraction of concentrated sound wave data in the main sound wave data. The concentrated sound wave data in the main sound wave data is the same pronunciation part of multiple languages. Subsequently, the diffuse sound wave data are predicted and adjusted on the concentrated sound wave data respectively. Since a speech stream has a stability law in the pronunciation diffuse sound wave of continuous time sequence, the sound wave data to be tested and measured are obtained by extracting continuous equidistant segments of sound wave data adjacent to the main sound wave data from the first whole wave data belonging to the first speech stream according to the main sound wave data, rather than extracting and comparing and analyzing several segments of sound wave data, so as to reduce the influence of large errors between the sound wave data with low interference and time sequence correlation, according to the fluctuation difference characteristics between the sound wave data to be tested and the sound wave data to be tested and the main sound wave data, two auxiliary sound wave data are refined and extracted from the first whole wave data, and according to these two auxiliary sound wave data, In order to realize the prediction of diffuse sound wave data for the concentrated main sound wave data in the first whole wave data, similarly, the second whole wave data is analyzed accordingly to predict the first overlapping segment sound wave data and the second overlapping segment sound wave data. In order to improve the reliability of the entire analysis process and the accuracy of the prediction results, it is necessary to verify the predicted first overlapping segment sound wave data and the second overlapping segment sound wave data, that is, one is to judge the degree of overlap between the pre-processed sound wave data and the main sound wave data after the two overlapping segment sound wave data are integrated, and the other is to judge the degree of correlation between the two results of the content recognition of the two overlapping segment sound wave data and the overall content recognition result initially recognized. If both verification conditions are met, the two verification results identified after the prediction adjustment are determined to be the final content recognition results of the overlapping voice segment, and the voice stream content recognition results are obtained comprehensively. Through the verification processing method of the two conditions mentioned above, the probability of misjudgment of the content recognition results is reduced and the efficiency of the entire recognition work is improved.
[0035] The specific step S1 includes the following sub-steps:
[0036] A multi-language continuous speech stream is received, and noise removal is performed on the multi-language continuous speech stream to obtain a pre-recognition speech stream. If there are overlapping speech segments in the pre-recognition speech stream, a first speech stream and a second speech stream with the overlapping speech segments are extracted, and the second speech stream is the other speech streams with the remaining overlapping speech segments other than the first speech stream with the overlapping speech segments in the multi-language continuous speech stream.
[0037] Preliminary content recognition is performed on the first voice stream and the second voice stream respectively to obtain a first preliminary recognition result and a second preliminary recognition result.
[0038] The overlapping speech segments are converted into sound wave data to obtain main sound wave data, and concentrated sound wave data are extracted from the main sound wave data to obtain concentrated main sound wave data.
[0039] The first whole wave data and the second whole wave data are obtained by counting the whole sound wave data of the first voice stream and the second voice stream respectively after conversion.
[0040] Continuous equidistant segments of sound wave data adjacent to the main sound wave data are extracted from the first whole wave data to obtain the first sound wave data to be measured and the second sound wave data to be measured.
[0041] The fluctuation differences between the main sound wave data and the sound wave data to be measured one, as well as between the main sound wave data and the sound wave data to be measured two, are comprehensively statistically analyzed to obtain the main measured fluctuation value.
[0042] Extract the first auxiliary sound wave data 1 and the first auxiliary sound wave data 2 belonging to the two speech segments which are closest in time sequence to the main sound wave data and whose fluctuation difference is the same as the main measurement fluctuation value from the first whole wave data.
[0043] Specifically, such as pre-recognition of voice stream (the noise removal processing here refers to the filtering of voice streams to remove other noisy sounds except the target voice stream, such as environmental noise: background noise, echo, etc., which will seriously affect the accuracy of voice recognition. In a noisy environment, the signal-to-noise ratio of the voice signal is reduced, resulting in an increase in recognition error. And when performing content recognition on a voice stream, other voice streams with dialects are filtered to reduce the sound interference between different voice streams. If the pre-recognition voice streams are A voice stream, B voice stream, C voice stream, and D voice stream respectively), the first voice stream and the second voice stream (if there is an overlapping voice segment Y between voice stream A and voice stream B (here the overlapping voice segment Y can be located at the initial end of the voice stream, the middle end of the voice stream, or the end of the voice stream), then the first voice stream is voice stream A, and the second voice stream is voice stream B), the first recognition initial The results and the second initial recognition results (that is, recognizing the voice stream into text, using existing technology - end-to-end recognition technology: using an end-to-end training method, directly mapping from the voice signal to the text sequence, avoiding the complex module division and error transmission problems in the traditional method, and being able to more efficiently handle the recognition task of multi-language continuous voice streams, improving the accuracy and efficiency of recognition. If the content result of voice stream A is recognized as N1, and the content result of voice stream B is recognized as N2), the main sound wave data (for example, normalizing and framing the pre-recognized voice stream, normalization: normalizing the amplitude of the voice stream so that its amplitude is within a certain range, usually it can be normalized to the interval [-1,1]. This ensures that voice signals of different intensities have the same weight in subsequent processing, which is convenient for analysis and processing. Framing: Since the voice signal has short-term stability, that is, in a short time (usually 10ms), - 30ms) can be considered stationary. Therefore, the speech stream is divided into several frames so that each frame can be processed independently. A common frame length is 256 samples, and a frame shift is 128 samples. Converting sound wave data involves extracting features and generating sound wave representations. Feature extraction involves extracting parameters that represent speech characteristics from preprocessed speech frames. Common features include Mel-Frequency Cepstral Coefficients (MFCCs) and Linear Prediction Cepstral Coefficients (LPCCs). These features reflect information such as the speech spectrum and vocal tract characteristics, and are used in subsequent tasks such as speech recognition and speaker identification. Taking MFCCs as an example, the extraction process involves performing a fast Fourier transform (FFT) on the speech frame to obtain a spectrum. The spectrum is then filtered through a triangular filter bank. The filtered result is then subjected to logarithmic operations and discrete cosine transforms (DCTs) to obtain MFCC feature vectors. Generating sound wave representations: Based on the extracted features, the speech stream can be converted into sound wave data.A common way is to plot the amplitude of the speech signal over time as a waveform graph, with the horizontal axis representing time and the vertical axis representing amplitude), concentrate the main sound wave data (concentrate the sound waves, such as the acoustic feature level: certain phonemes or pronunciation methods in different languages are similar, and they will show a certain concentration trend in acoustic features such as spectrum. For example, the pronunciation of some vowels in different languages may have overlapping parts in the frequency band. These similar acoustic features will cause the sound waves to be relatively concentrated in certain frequency ranges. Language feature level: Some languages may have commonalities in pronunciation characteristics, which makes the sound waves show concentration. For example, some languages emphasize specific syllables or intonation patterns when pronouncing, and these commonalities will be reflected in the sound wave data. Appear. If the main sound wave data is d, the concentrated sound wave data is extracted, which is the concentrated and identical sound wave data between voice stream A and voice stream B, if it is d1. Diffuse sound wave (that is, d-d1 is the diffuse sound wave of d2. This diffuse sound wave is most likely caused by the language pronunciation difference between voice stream A and voice stream B, and is caused by the comprehensive pronunciation diffusion of voice stream A and voice stream B. Subsequently, it is necessary to predict the diffuse sound wave data of voice stream A and the diffuse sound wave data of voice stream B), such as the language difference level: There are many differences between different languages in pronunciation, grammar, vocabulary, etc. These differences will lead to the diffusion of sound wave data. For example, some languages have unique consonants or The pronunciation of vowels, and its corresponding sound wave frequency and waveform are significantly different from those of other languages, making the sound waves diffuse as a whole. At the level of speech stream characteristics: In a continuous speech stream, sound waves will continue to change and diffuse due to factors such as the speaker's speaking speed, intonation, voice, and switching between languages. For example, in a sentence, the pronunciation of different words may cause rapid changes in the frequency and amplitude of sound waves, and when words of different languages are combined together, the transition and connection of sound waves will also be more complicated, causing the sound waves to exhibit diffusion characteristics in both time and frequency dimensions), the first whole wave data and the second whole wave data (if Z1 and Z2), the sound wave data to be tested one and the sound wave data to be tested two (such as Z1 and If the speech segments with the closest adjacent time sequences to Y are R1 and R2), the main measured fluctuation value (if the fluctuation difference between Y and R1 is c1, and the fluctuation difference between R2 and Y is c2, then the main measured fluctuation value is c1+c2 if it is C), the first auxiliary measured sound wave data one and the first auxiliary measured sound wave data two (if Z1 contains sound wave data segments with different fluctuation amplitudes: s1, s2, s3, s4, s5, if the comprehensive fluctuation value between s2 and s1 and s3 respectively is C, and the comprehensive fluctuation value between s4 and s3 and s5 respectively is also C, then s2 and s4 are extracted as the first auxiliary measured sound wave data one and the first auxiliary measured sound wave data two, because they have the same stable fluctuation law as Y, they can be used as auxiliary analysis data).
[0044] The specific step S2 includes the following sub-steps:
[0045] The concentrated sound wave data are extracted from the first auxiliary sound wave data 1 and the first auxiliary sound wave data 2 to obtain the concentrated consonant sound wave data 1 and the concentrated consonant sound wave data 2 respectively.
[0046] The first sound wave preset value is obtained by averaging the difference between the first auxiliary sound wave data 1 and the concentrated consonant sound wave data 1 and the difference between the first auxiliary sound wave data 2 and the concentrated consonant sound wave data 2.
[0047] According to the first sound wave preset value, the concentrated main sound wave data is predicted and adjusted to obtain the first overlapping segment sound wave data.
[0048] According to the first overlapping segment to be identified, the concentrated main sound wave data and the second whole wave data, a second sound wave pre-adjustment value is calculated, and according to the second sound wave pre-adjustment value, the concentrated main sound wave data is predicted and adjusted for diffuse sound wave data to obtain the second overlapping segment sound wave data.
[0049] Specifically, such as the concentrated consonant wave data one and the concentrated consonant wave data two (if they are F1 and F2 respectively), the first sound wave pre-adjustment value (such as the difference between the first auxiliary sound wave data one and the concentrated consonant wave data one: s2-F1, the difference between the first auxiliary sound wave data two and the concentrated consonant wave data two: S4-F2, the first sound wave pre-adjustment value: [(s2-F1) + (S4-F2)] / 2 if it is T1), the first overlapping segment sound wave data (that is, d1+T1 if it is L1), the second sound wave pre-adjustment value (if it is T2), and the second overlapping segment sound wave data (that is, d1+T2 if it is L2).
[0050] The specific step S3 includes the following sub-steps:
[0051] The first overlapping sound wave data and the second overlapping sound wave data are acoustically integrated to obtain pre-processed sound wave data, and a coincidence condition threshold is preset. If the degree of coincidence between the pre-processed sound wave data and the main sound wave data is greater than or equal to the coincidence condition threshold, the verification condition one is output.
[0052] Content recognition is performed on the first overlapping sound wave data segment and the second overlapping sound wave data segment to obtain a first overlapping recognition result and a second overlapping recognition result.
[0053] According to the first overlap recognition result, the corresponding overlap segment content recognition result in the first recognition initial result is corrected to obtain correction result one; according to the second overlap recognition result, the corresponding overlap segment content recognition result in the second recognition initial result is corrected to obtain correction result two.
[0054] According to the comprehensive correlation degree of the context sentences in the first recognition preliminary result, the correlation determination threshold 1 is preset, and according to the comprehensive correlation degree of the context sentences in the second recognition preliminary result, the correlation determination threshold 2 is preset.
[0055] If the degree of correlation between the correction result 1 and the context of the first overlapping recognition result is greater than or equal to the correlation determination threshold 1, and the degree of correlation between the correction result 2 and the context of the second overlapping recognition result is greater than or equal to the correlation determination threshold 2, then the verification condition 2 is output.
[0056] According to verification condition one and verification condition two, the correction result one and the first preliminary recognition result are content integrated, and the correction result two and the second preliminary recognition result are content integrated to comprehensively obtain the voice stream content recognition result.
[0057] Specifically, such as pre-processing sound wave data (i.e., if the sound waves of L1 and L2 are overlapped and the new sound wave data is b), presetting the overlap condition threshold (if it is W), verifying the compliance condition one (if the sound wave data overlap between b and d is w, if w is greater than or equal to W, it can be judged that the one-time verification condition of the prediction adjustment result of b is passed), the first overlap recognition result and the second overlap recognition result (if the content recognition results are g1 and g2 respectively), correction result one, correction result two (if the corresponding overlap content recognition results of the first recognition result are g3 and g4, then the part of the text content that differs between g1 and g3 is semantically comprehensively adjusted, and the same part of the text content is retained to obtain a correction result one G1, and the part of the text content that differs between g2 and g4 is semantically comprehensively adjusted, and the same part of the text content is retained to obtain a correction result one G2), comprehensive correlation degree (the comprehensive correlation degree here is to remove the overlapping speech in the first recognition result). The results of the recognition of the remaining content of the segment are counted, and can be counted based on semantic similarity: for example, using word vectors (such as Word2Vec, GloVe, etc.) to represent the words in the sentence as vectors. For the two sentences above and below, the sentence vector can be obtained by calculating the average of their word vectors. Then calculate the cosine similarity between the two sentence vectors to measure the degree of semantic similarity. For example, for the sentences "Apple is a kind of fruit" and "Banana is also a kind of fruit", after averaging the word vectors, their cosine similarity will be relatively high, because "apple" and "banana" both belong to the fruit category semantically, and the sentence structure is similar. Or dialogue scenes and topic tracking: If it is in a dialogue scene, it is important to understand the topic of the dialogue. For example, in a dialogue about travel, if the previous sentence mentions a tourist attraction, and the next sentence mentions the ticket price or surrounding facilities of this attraction, then the degree of correlation is very high. This can be achieved by establishing a topic model (such as LDA - Latent Dirichlet Allocation) to identify the topic of the conversation, and then determine the degree of correlation between the previous and next sentences based on the consistency of the topic), preset correlation determination threshold one (if it is M1), preset correlation determination threshold (if it is M2), verification condition two (if the degree of correlation between the previous and next sentences of the correction result one and the first overlapping recognition result is m1, if m1 is greater than or equal to M1, and at the same time, the degree of correlation between the previous and next sentences of the correction result two and the second overlapping recognition result is m2, if m2 is greater than or equal to M2, then it can be determined that the secondary verification condition of the prediction adjustment result b has passed), voice stream content recognition result (if both verification conditions are met, the correction result one is determined as the final content recognition result of the overlapping voice segment of voice stream A, and the correction result two is determined as the final content recognition result of the overlapping voice segment of voice stream B, and the correction result one and the first initial recognition result are content-integrated, and the correction result two and the second initial recognition result are content-integrated to obtain the content recognition result of the multilingual continuous voice stream).
[0058] The above are all preferred embodiments of the present application, and are not intended to limit the scope of protection of the present application. Therefore, any equivalent changes made based on the structure, shape, and principle of the present application should be included in the scope of protection of the present application.
Claims
1. A content recognition method based on multilingual continuous speech stream, characterized in that: The following steps are involved: Step S1, extracting a first voice stream and a second voice stream having overlapping voice segments from a pre-recognized voice stream after noise removal processing, extracting concentrated main sound wave data from the main sound wave data after sound wave data conversion of the overlapping voice segments, and statistically analyzing the overall sound wave data after the first voice stream and the second voice stream are converted to obtain first whole wave data and second whole wave data, and extracting first auxiliary sound wave data 1 and first auxiliary sound wave data 2 belonging to the two voice segments from the first whole wave data according to the main sound wave data; Step S2, based on the first auxiliary sound wave data 1 and the first auxiliary sound wave data 2, the concentrated main sound wave data is predicted and adjusted for diffuse sound wave data to obtain first overlapping segment sound wave data, and the second overlapping segment sound wave data is statistically calculated based on the first overlapping segment sound wave data, the concentrated main sound wave data and the second whole wave data; Step S3: If the degree of overlap between the pre-processed sound wave data and the main sound wave data after the integration of the first overlapping sound wave data and the second overlapping sound wave data is greater than or equal to the preset overlap condition threshold, output verification condition one; if it is determined that the correction result one and the correction result two after the corresponding content recognition correction of the first overlapping sound wave data and the second overlapping sound wave data respectively meet the correlation degree relationship between the upper and lower sentences, output verification condition two; based on verification condition one and verification condition two, the voice stream content recognition result is comprehensively obtained.
2. The method for content recognition based on multilingual continuous speech stream according to claim 1, characterized in that: Step S1 includes: receiving a multilingual continuous speech stream, performing noise removal on the multilingual continuous speech stream to obtain a pre-recognized speech stream, and if overlapping speech segments exist in the pre-recognized speech stream, extracting a first speech stream and a second speech stream containing the overlapping speech segments, wherein the second speech stream is other speech streams containing the remaining overlapping speech segments in addition to the first speech stream containing the overlapping speech segments in the multilingual continuous speech stream; Preliminary content recognition is performed on the first voice stream and the second voice stream respectively to obtain a first preliminary recognition result and a second preliminary recognition result.
3. The method for content recognition based on multilingual continuous speech stream according to claim 2, characterized in that: Step S1 further includes: Converting the overlapping speech segments into sound wave data to obtain main sound wave data, and extracting concentrated sound wave data from the main sound wave data to obtain concentrated main sound wave data; Counting the overall sound wave data of the first voice stream and the second voice stream after conversion to obtain first whole wave data and second whole wave data; Continuous equidistant segments of sound wave data adjacent to the main sound wave data are extracted from the first whole wave data to obtain the first sound wave data to be measured and the second sound wave data to be measured.
4. The method for content recognition based on multilingual continuous speech stream according to claim 3, characterized in that: Step S2 includes: The fluctuation differences between the main sound wave data and the sound wave data to be measured, as well as between the main sound wave data and the sound wave data to be measured, are comprehensively counted to obtain the main measured fluctuation value; Extract the first auxiliary sound wave data 1 and the first auxiliary sound wave data 2 belonging to the two speech segments which are closest in time sequence to the main sound wave data and whose fluctuation difference is the same as the main measured fluctuation value from the first whole wave data.
5. The method for content recognition based on multilingual continuous speech stream according to claim 4, characterized in that: Step S2 includes: Extracting concentrated sound wave data from the first auxiliary sound wave data 1 and the first auxiliary sound wave data 2 to obtain concentrated consonant sound wave data 1 and concentrated consonant sound wave data 2 respectively; The first sound wave pre-adjustment value is obtained by averaging the difference between the first auxiliary sound wave data 1 and the concentrated consonant sound wave data 1 and the difference between the first auxiliary sound wave data 2 and the concentrated consonant sound wave data 2; According to the first sound wave preset value, the concentrated main sound wave data is predicted and adjusted for the diffuse sound wave data to obtain the first overlapping segment sound wave data; According to the first overlapping segment to be identified, the concentrated main sound wave data and the second whole wave data, a second sound wave pre-adjustment value is calculated, and according to the second sound wave pre-adjustment value, the concentrated main sound wave data is predicted and adjusted for diffuse sound wave data to obtain the second overlapping segment sound wave data.
6. The method for content recognition based on multilingual continuous speech stream according to claim 5, characterized in that: Step S3 includes: The first overlapping sound wave data and the second overlapping sound wave data are acoustically integrated to obtain pre-processed sound wave data, and a coincidence condition threshold is preset. If the degree of coincidence between the pre-processed sound wave data and the main sound wave data is greater than or equal to the coincidence condition threshold, the verification condition one is output.
7. The method for content recognition based on multilingual continuous speech stream according to claim 6, characterized in that: Step S3 further includes: Performing content recognition on the first overlapping sound wave data segment and the second overlapping sound wave data segment to obtain a first overlapping recognition result and a second overlapping recognition result; According to the first overlap recognition result, the corresponding overlap segment content recognition result in the first recognition initial result is corrected to obtain correction result one; according to the second overlap recognition result, the corresponding overlap segment content recognition result in the second recognition initial result is corrected to obtain correction result two.
8. The method for content recognition based on multilingual continuous speech stream according to claim 7, characterized in that: Step S3 further includes: Presetting a first relevance determination threshold value based on the comprehensive relevance between the context sentences in the first recognition preliminary result, and a second relevance determination threshold value based on the comprehensive relevance between the context sentences in the second recognition preliminary result; If the correlation degree between the correction result 1 and the context of the first overlapping recognition result is greater than or equal to the correlation determination threshold 1, and the correlation degree between the correction result 2 and the context of the second overlapping recognition result is greater than or equal to the correlation determination threshold 2, then output verification compliance condition 2; According to the verification condition 1 and the verification condition 2, the correction result 1 and the first initial recognition result are content integrated, and the correction result 2 and the second initial recognition result are content integrated to comprehensively obtain the voice stream content recognition result.