A semantic analysis method for speech in a crowded environment
By combining lip movement video streams and voice signal data, and using deep learning and voiceprint separation technology, the noise interference problem of voice analysis in crowded environments is solved, and high-precision voice separation and semantic analysis are achieved.
Patent Information
- Application Number
- CN202510935465.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-07-08
AI Technical Summary
Existing speech analysis methods are easily affected by background noise in crowded environments, making it difficult to accurately distinguish the target user's voice from noise. They also fail to effectively utilize visual information for feature extraction, resulting in decreased speech recognition accuracy and semantic ambiguity.
Combining lip movement video streams and voice signal data, through deep learning and voiceprint separation technology, the target user's voice features are extracted, the initial and final consonant structure of Chinese pronunciation is used for time domain segmentation and pinyin tone analysis, and the contextual semantics are combined to optimize the understanding of speech segments.
It significantly improves the accuracy of speech recognition and semantic analysis in noisy environments, effectively separates the target user's voice, and improves the accuracy of speech separation and semantic understanding.
Smart Images

Figure CN120431913B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to speech analysis, in particular to a semantic analysis method of speech in a crowded environment. BACKGROUND
[0002] Semantic analysis of speech is widely used in intelligent assistants, which can help users control devices, query information, and manage daily tasks through voice. At the same time, this technology has been widely used in automatic customer service, voice translation, smart home, medical diagnosis, etc. With the continuous development of artificial intelligence technology, the accuracy and real-time performance of speech semantic analysis have gradually improved, making voice interaction a more natural and convenient way of human-computer interaction, greatly promoting the popularization and development of intelligent life.
[0003] Current speech analysis methods on the market often face significant limitations in crowded environments, mainly relying on the processing of pure speech signals, which are easily disturbed by background noise, leading to a decrease in speech recognition accuracy. In noisy environments, sound source separation and noise suppression are not effective, often making it difficult to accurately distinguish the target user's speech from other noise. In addition, existing speech analysis techniques focus on the processing of audio data, ignoring the auxiliary role of visual information, and cannot effectively extract features using additional signals such as user lip movement video streams. Therefore, in crowded or multi-speaker scenarios, existing methods often cannot provide efficient and accurate speech separation and semantic analysis. More importantly, many methods fail to combine context to correct semantics when segmenting and recognizing pinyin speech signals, leading to misidentification or semantic ambiguity. SUMMARY
[0004] To improve existing semantic analysis methods of speech, a semantic analysis method of speech in a crowded environment is provided, which combines lip movement video streams and speech signal data, uses deep learning and voiceprint separation technology to effectively extract the target user's speech, and optimizes speech segment understanding based on pinyin tone analysis and context semantics, significantly improving speech recognition accuracy in noisy environments.
[0005] To achieve the above purposes, the technical solution adopted by the present application is as follows:
[0006] A semantic analysis method of speech in a crowded environment, characterized in that it comprises:
[0007] Collecting speech signal data and lip movement video stream data of a target user;
[0008] Extracting lip movement feature sequences based on lip movement video stream data, and mapping the lip movement feature sequences to predicted speech feature vectors through a pre-trained convolutional neural network model;
[0009] Input the predicted speech feature vector into the voiceprint separation model to separate the speech segment of the target user from the mixed speech signal and generate pure speech data features;
[0010] Segment the pure speech data features in the time domain based on the initial and final consonant-vowel structure of Chinese pronunciation, and obtain the time-domain waveform of the speech signal of a single word;
[0011] Perform dynamic time warping matching on the time-domain waveform of the single-word speech signal and the initial and final consonant-vowel time-domain waveform library to obtain the corresponding pinyin expression of each word;
[0012] Perform tone combination correlation analysis based on the pinyin expression of consecutive single words, and obtain the meaning of the speech segment of the target user in combination with the context semantics.
[0013] Preferably, the lip movement feature sequence is extracted based on the lip movement video stream data, and the lip movement feature sequence is mapped into a predicted speech feature vector through a pre-trained convolutional neural network model, which specifically comprises:
[0014] Based on the obtained lip movement video stream data of the target user, the face region of each frame is located, the lip is accurately located through a key point regression algorithm, and pre-processing is performed;
[0015] The lip movement features are extracted through the primary, intermediate and high-level feature layers in the convolutional neural network model to obtain the lip movement feature sequence;
[0016] Based on the obtained lip movement feature sequence, it is aligned with the speech frame rate, and the lip movement feature sequence is mapped into a predicted speech feature vector through a pre-trained convolutional neural network encoder-decoder;
[0017] The predicted speech feature vector includes a joint representation of mel-frequency cepstral coefficients and fundamental frequency.
[0018] Preferably, the predicted speech feature vector is input into the voiceprint separation model to separate the speech segment of the target user from the mixed speech signal and generate pure speech data features, which specifically comprises:
[0019] Frame the mixed speech signal and convert it into a time-frequency spectrum, and enhance the acoustic feature dimension of lip movement prediction;
[0020] Generate a compact voiceprint vector based on the mel-frequency cepstral coefficients and fundamental frequency of the predicted speech feature vector, construct a voiceprint template matrix through time series expansion, and calibrate the speaker characteristics;
[0021] Based on the obtained voiceprint template matrix, input into the convolutional neural network encoder, and obtain the time-frequency mask of the target speech through similarity weighting;
[0022] The target complex spectrum is extracted based on a time-frequency mask, and a phase iteration is constrained according to a fundamental frequency predicted by lip movement, and a time-domain waveform is reconstructed to generate pure speech data features.
[0023] Preferably, the initial pure speech data features are segmented in the time domain based on the initial Chinese pronunciation initial and final structure, and the time-domain waveform of the single-character speech signal specifically comprises:
[0024] By monitoring the signal data with high zero-crossing rate and low energy in the pure speech data, a sharp pulse with a duration less than a threshold on the energy curve is obtained, and the initial is determined;
[0025] By monitoring the data with high energy and stable formant in the pure speech data, the final is determined through energy envelope analysis and formant tracking;
[0026] The silent section is determined based on the energy duration of the speech signal, and the inter-word interval is obtained;
[0027] For each candidate syllable, it is determined whether it contains an initial and a final, and the final section is directly extracted for the syllable without an initial;
[0028] The time-domain waveform of the single-character speech signal is obtained by segmenting the single-character syllable.
[0029] Preferably, the final is determined by monitoring the data with high energy and stable formant in the pure speech data, through energy envelope analysis and formant tracking, specifically comprising:
[0030] The logarithmic energy of each frame of speech is calculated to generate an energy envelope curve, and the starting boundary of the final is marked as the starting point of the rising edge of the envelope;
[0031] The first three formants F1-F3 are calculated by linear predictive coding, and when the fluctuation of F1 is less than a threshold for 10 consecutive frames, the stable final section is determined.
[0032] Preferably, the time-domain waveform of the single-character speech signal is matched with the initial and final time-domain waveform library by dynamic time warping to obtain the corresponding pinyin expression of each character, specifically comprising:
[0033] Based on 21 types of pure initial and 35 types of final pronunciations of a speaker, the time-domain waveform templates of each type of pronunciation are extracted to construct the initial and final time-domain waveform library;
[0034] Based on the priority of clear consonants, the average zero-crossing rate of the first 50ms section of the waveform of the single character to be matched is determined to determine the clear consonant initial;
[0035] Based on the non-clear consonant initial, the initial time-domain waveform library is matched to obtain the first three initial candidates with the minimum cost by DTW calculation;
[0036] The first three formant frequencies are obtained by performing linear predictive coding analysis on the to-be-matched coda segment, and the matched coda is obtained by performing nonlinear time scaling on the DTW path.
[0037] Based on the obtained initial consonants and coda, the splicing verification is performed, the combination legality is checked, and the corresponding pinyin expression of each character is generated.
[0038] Preferably, the pinyin expression based on the continuous single character is subjected to tone combination relevance analysis, and the meaning of the voice segment of the target user is obtained in combination with the context semantics, which specifically comprises:
[0039] The pinyin tones of the continuous single character are arranged and combined to form all possible tone sequences.
[0040] Based on the statistical tone combination frequency in the corpus, a relevance weight is assigned to each tone combination in combination with the word semantics, the higher the frequency, the greater the weight, the higher the semantic relevance, and the greater the weight.
[0041] The tone sequences are sorted from high to low based on the weight, and the tone combination with the highest weight is selected.
[0042] Based on the pinyin expression of the single character obtained by combining the obtained tone, the specific content of the entire voice segment is obtained, and the voice segment content is corrected according to the grammar structure of the words and sentences before and after the pinyin segment, to obtain the meaning of the voice segment of the target user.
[0043] Compared with the prior art, the advantages of the present application are:
[0044] By fusing the lip movement video stream data and the voice signal data, the method can effectively extract the voice of the target user from complex background noise, improving the accuracy of voice recognition. The use of lip movement video stream, combined with the deep learning processing of the convolutional neural network model, not only improves the extraction accuracy of voice features, but also better deals with the problem of voice separation in noisy environments under the support of visual and auditory information. Secondly, the method separates the voice of the target user from the mixed voice signal accurately through the voiceprint separation model, improving the clarity of the voice data. Through dynamic time warping technology and matching of initial consonants and coda, the accuracy of pinyin expression of Chinese single character voice is effectively improved. In addition, based on the analysis of pinyin tone combination, combined with the context semantics for intelligent correction, the accuracy of semantic analysis can be further improved to ensure more accurate understanding of the voice segment. BRIEF DESCRIPTION OF DRAWINGS
[0045] Figure 1 The method proposed in the present application is shown in the schematic diagram;
[0046] Figure 2 The predicted voice feature vector proposed in the present application is shown in the schematic diagram;
[0047] Figure 3 Pure speech data feature schematic diagram proposed for the present application;
[0048] Figure 4 Single word speech signal time domain waveform schematic diagram proposed for the present application;
[0049] Figure 5 Judgment of the present application for the final schematic diagram;
[0050] Figure 6 Pinyin table expression schematic diagram proposed for the present application;
[0051] Figure 7 Speech segment meaning schematic diagram proposed for the present application;
[0052] Figure 8 The architecture of the electronic device in the present application;
[0053] Figure 9 The structure of the computer readable storage medium in the present application. DETAILED DESCRIPTION
[0054] The following description is used to disclose the present application so that those skilled in the art can implement the present application. The preferred embodiments in the following description are only as examples, and other obvious modifications can be made by those skilled in the art.
[0055] Referring to Figure 1 The semantic analysis method of speech in a crowded environment, comprising:
[0056] Step one: collect and obtain the speech signal data and lip movement video stream data of the target user;
[0057] Step two: extract the lip movement feature sequence based on the lip movement video stream data, and map the lip movement feature sequence to the predicted speech feature vector through the pre-trained convolutional neural network model;
[0058] Step three: input the predicted speech feature vector into the voiceprint separation model, separate the speech segment of the target user from the mixed speech signal, and generate pure speech data features;
[0059] Step four: based on the initial and final structure of Chinese pronunciation, the pure speech data features are segmented in time domain to obtain the time domain waveform of single word speech signal;
[0060] Step five: dynamic time warping matching is performed between the single word speech signal time domain waveform and the initial and final time domain waveform library to obtain the corresponding pinyin expression of each word;
[0061] Step six: based on the continuous single word pinyin expression, the tone combination correlation analysis is carried out, and the speech segment meaning of the target user is obtained combined with the context semantics.
[0062] Referring to Figure 2 As shown, the lip movement feature sequence is extracted based on the lip movement video stream data, and the lip movement feature sequence is mapped to the predicted speech feature vector through the pre-trained convolutional neural network model, specifically comprising:
[0063] Based on the obtained lip movement video stream data of the target user, the face region of each frame is located, the lip is accurately located through the key point regression algorithm, and pre-processing is performed;
[0064] The lip movement feature is extracted through the primary, intermediate and high-level feature layers in the convolutional neural network model, and the lip movement feature sequence is obtained;
[0065] Based on the obtained lip movement feature sequence, it is aligned with the speech frame rate, and the lip movement feature sequence is mapped to the predicted speech feature vector through the pre-trained convolutional neural network encoder-decoder;
[0066] The predicted speech feature vector includes the joint representation of mel frequency cepstral coefficient and fundamental frequency.
[0067] Specifically, within the face boundary box, the key point regression algorithm is applied to accurately locate the lip contour, and usually the lip is represented as a plurality of key points;
[0068] According to the accurately located lip region, the image of the lip region is cropped to ensure that the region input to the network only contains the lip region, and a deep convolutional neural network is used to extract the features of the lip region. The network is composed of multiple convolutional layers, pooling layers and fully connected layers to extract features of different scales. The primary feature layer extracts basic information such as edges and textures from local details, the intermediate feature layer captures structural features of lip movement such as lip shape and opening and closing, and the high-level feature layer extracts high-level features of lip movement such as patterns and dynamic changes of lip movement from a high-level semantic level.
[0069] Due to the different frame rates of video and speech, time alignment is needed to ensure that each video frame corresponds to a speech frame. The aligned lip movement feature sequence is mapped to a latent space representation through the encoder, and the decoder generates the predicted speech feature according to the latent space representation.
[0070] Mel frequency cepstral coefficient is used to represent the spectral features of speech, which is usually calculated by short-time Fourier transform, Mel frequency filtering, logarithmic compression and discrete cosine transform on the speech signal. The formula is:
[0071] ;
[0072] Where, is the mel frequency cepstral coefficient, is the discrete cosine transform, is the mel filter bank, is a short-time Fourier transform;
[0073] The fundamental frequency represents the pitch of the speech, which can be extracted from the speech signal by a fundamental frequency estimation algorithm;
[0074] The Mel-frequency cepstral coefficients and the fundamental frequency features are spliced to form the final speech feature vector.
[0075] Referring to Figure 3 The predicted speech feature vector is input into the voiceprint separation model to separate the speech segment of the target user from the mixed speech signal, and generate pure speech data features, which specifically include:
[0076] The mixed speech signal is framed and converted into a time-frequency spectrum, and the acoustic feature dimension of the lip movement prediction is enhanced;
[0077] Based on the Mel-frequency cepstral coefficients and the fundamental frequency of the predicted speech feature vector, a compact voiceprint vector is generated, a voiceprint template matrix is constructed through time series expansion, and the speaker characteristics are calibrated;
[0078] Based on the obtained voiceprint template matrix, it is input into the convolutional neural network encoder, and the time-frequency mask of the target speech is obtained through similarity weighting;
[0079] Based on the time-frequency mask, the target complex spectrum is extracted, the phase iteration is constrained according to the fundamental frequency of the lip movement prediction, and the time-domain waveform is reconstructed to generate the pure speech data features.
[0080] Specifically, first, the speech signal is divided into short-time frames, and the frame length is usually 20-40 milliseconds. A short-time Fourier transform is applied to each frame to obtain a time-frequency spectrum, which represents the energy distribution of the signal at different times and frequencies, and the formula is:
[0081] ;
[0082] wherein, is a signal sequence, is a window function, is a frequency term in the Fourier transform, t is a time index, and f is a frequency index;
[0083] The voiceprint template matrix is input into the convolutional neural network encoder to learn the time-frequency mask, and the time-frequency mask is used to weight the complex spectrum. For each frequency point and time point, the mask is applied to remove the interfering speech signal, and based on the fundamental frequency of the lip movement prediction, the phase part of the complex spectrum is constrained. The weighted complex spectrum is converted back to the time-domain waveform using the inverse STFT, and finally the pure speech signal is reconstructed.
[0084] Referring to Figure 4As shown, the initial phonetic data features are segmented in time domain based on the initial phonetic data features, and the phonetic data features are segmented in time domain based on the initial phonetic data features. The time-domain waveform of the single-character phonetic signal specifically includes:
[0085] By monitoring the signal data with high zero-crossing rate and low energy in the pure phonetic data, the sharp pulse with energy curve less than the threshold is obtained, and the initial phonetic data features are segmented in time domain based on the initial phonetic data features. The initial phonetic data features are segmented in time domain based on the initial phonetic data features.
[0086] By monitoring the signal data with high energy and stable formant in the pure phonetic data, the energy envelope analysis and formant tracking are used to judge the final phonetic data features.
[0087] The energy of the signal is used to judge the silence segment, and the interval between the characters is obtained.
[0088] Whether the initial phonetic data features contain the initial phonetic data features is judged, and the final phonetic data features are directly extracted from the initial phonetic data features.
[0089] The time-domain waveform of the single-character phonetic signal is obtained by segmenting the single-character phonetic signal.
[0090] Specifically, the zero-crossing rate is the number of times the signal crosses zero in a unit of time, which is used to reflect the pitch and noise of the signal. High zero-crossing rate usually means that the signal is a part without initial phonetic data features or contains higher frequency components.
[0091] The energy of the signal is calculated, and the part with low energy is usually silent or weak noise. In the judgment of the initial phonetic data features, low energy can be used as the basis for ignoring a part of the signal.
[0092] According to the low value of the zero-crossing rate and the energy, the sharp pulse part with low energy and high zero-crossing rate is identified as the initial phonetic data features. For example, if the energy of the signal is less than a certain threshold and the zero-crossing rate is high, it is determined that the segment is the initial phonetic data features.
[0093] The energy envelope is the smooth change of the energy of the signal in time. By analyzing the energy envelope of the audio signal, different phonetic structures can be distinguished. The final phonetic data features usually have higher energy and more stable energy curve. The formula is:
[0094] ;
[0095] Wherein, is the energy envelope, is the window function or filter, is the value of the signal at time n;
[0096] The formant frequency of the signal is calculated using the formant extraction algorithm. By tracking the formant frequency, the existence of the final phonetic data features is judged. If there is a stable formant in the signal and the energy is high, it may be the final phonetic data features.
[0097] A silent segment is an area where the energy of the speech signal is continuously below a certain threshold. Energy detection is used to identify silent segments. If the energy is below a certain threshold, it is considered a silent segment.
[0098] Each candidate syllable can be checked to see if it contains an initial and a final. Energy, zero-crossing rate, and formant are used to determine whether the syllable is an initial or a final. If it contains both an initial and a final, the syllable is marked as a complete syllable.
[0099] Syllables are segmented using features such as silence segments, zero-crossing rate, and energy envelope. The start and end times of each candidate syllable can be calculated from the silence segments and the interval between words to obtain the single-word syllable signal after segmentation.
[0100] See Figure 5 As shown, by monitoring the data with high energy and stable formant in the clean speech data, and through energy envelope analysis and formant tracking, the finals are judged to include:
[0101] Calculate the logarithmic energy of each frame of speech, generate an energy envelope curve, and mark the starting point of the envelope rising edge as the starting boundary of the final;
[0102] The first three formants F1-F3 are calculated by linear predictive coding. When the fluctuation of F1 is less than the threshold for 10 consecutive frames, it is determined to be a stable final segment.
[0103] See Figure 6 As shown, the time domain waveform of the single-word speech signal is dynamically time-warped and matched with the initial and final time domain waveform libraries to obtain the pinyin expression corresponding to each word. Specifically, the following steps are performed:
[0104] Based on the speaker's 21 types of pure initials and 35 types of finals, the time domain waveform template of each type of pronunciation is extracted to build the initial and final time domain waveform library;
[0105] Based on the priority of voiceless consonants, the voiceless consonant initials are determined according to the average zero-crossing rate of the first 50ms segment of the waveform of the word to be matched;
[0106] Based on the non-voiceless consonant initials, match them from the initial time domain waveform library, calculate the minimum path cost through DTW, and obtain the top three initial candidates with the minimum cost;
[0107] By performing linear predictive coding analysis on the matching final segment, the first three formant frequencies are obtained, and the matching final is obtained by performing nonlinear time scaling on the DTW path;
[0108] Based on the obtained initials and finals, splicing verification is performed to check the legality of the combination and generate the corresponding pinyin expression for each character.
[0109] Specifically, for each type of initial and final, through recording or voice data collection, the time domain waveform of each type of pronunciation is extracted, and the time domain template signal with clear features can be obtained through frame division, denoising and other preprocessing steps;
[0110] Based on the average zero-crossing rate of the first 50 ms segment of the to-be-matched single character waveform, it is judged whether it is a clear consonant initial. The zero-crossing rate is the number of times the signal crosses zero in the speech signal, reflecting the frequency characteristics of the signal. Based on the size of the average zero-crossing rate, it is judged whether it is a clear consonant initial.
[0111] For non-clear consonant initials, match from the initial time domain waveform library. The DTW algorithm finds the most matched initial template by calculating the minimum matching path cost between two time series signals. By comparing the cost function value, the top three candidate initials with the smallest cost are selected.
[0112] The DTW path of the final signal is nonlinearly time-scaled to realize time alignment. By adjusting the time scale, the signal best matched with the final template in the library is obtained. According to the above steps, the most matched final is selected from the final library.
[0113] The matched initials and finals are spliced, and the legality of the splicing is checked. Usually, the legality of the splicing is verified by the smoothness of the speech splicing. According to the splicing result, the corresponding pinyin expression is generated.
[0114] Referring to Figure 7 The pinyin expression of the continuous single character is used to analyze the relevance of the tone combination. The meaning of the target user's speech segment is obtained by combining the context semantics, which includes:
[0115] The pinyin tones of the continuous single character are arranged and combined to form all possible tone sequences.
[0116] Based on the frequency of tone combination statistics in the corpus, a relevance weight is assigned to each tone combination based on the word semantics. The higher the frequency, the higher the weight, the higher the semantic relevance, and the higher the weight.
[0117] Based on the weight, the tone sequences are sorted from high to low, and the tone combination with the highest weight is selected.
[0118] Based on the pinyin expression of the obtained tone combination single character, the specific content of the entire speech segment is obtained. The content of the speech segment is corrected according to the grammar structure of the words and sentences before and after the pinyin segment, and the meaning of the target user's speech segment is obtained.
[0119] Specifically, for each pinyin, a tone is assigned according to its syllable composition, and all possible tone sequences are generated by arranging and combining the tones of each single character.
[0120] In the corpus, the frequency of different tone combinations is counted, the higher the frequency, the greater the possibility of the combination appearing in the actual context; each tone combination may represent different words or semantics, based on the context and grammar rules of different words in the corpus, the semantic relevance of each tone combination is calculated, and the frequency and semantic relevance are considered comprehensively to assign a weight to each tone combination;
[0121] According to the weight of each tone combination, all tone combinations are sorted in descending order. The sorted combinations are represented in order from high to low, and the tone combination with the highest weight is selected as the final tone combination of the target voice segment;
[0122] According to the grammar structure of the words and sentences before and after the pinyin segment, it is checked whether the pinyin segment conforms to the grammar and context, if the pinyin segment does not conform to the sentence structure or grammar, appropriate correction is performed, for example, some tone combinations may not be common or conform to the grammar rules in a specific context, and the pinyin may need to be adjusted through grammar rules, the correction method can include:
[0123] Vocabulary level correction: correcting the infrequent combination through a dictionary or grammar rules;
[0124] Sentence level correction: adjusting the tone combination of the pinyin according to the context before and after the sentence to ensure the correctness of the grammar.
[0125] Further, the method according to the embodiments of the present application can also be implemented by means of Figure 8 The architecture of the electronic device is shown. As Figure 8 shown, the electronic device 500 can include a bus 501, one or more CPUs 502, a read-only memory (ROM) 503, a random access memory (RAM) 504, a communication port connected to a network 505, an input / output component 506, a hard disk 507, etc. The storage device in the electronic device 500, such as the ROM 503 or the hard disk 507, can store a semantic analysis method for voice in a crowded environment provided by the present application. The electronic device 500 can also include a terminal interface 508. Of course, Figure 8 The architecture shown is only exemplary, and when implementing different devices, one or more components of the electronic device shown can be omitted according to actual needs. Figure 8
[0126] Figure 9 is a structural diagram of a computer readable storage medium provided by an embodiment of the present application. As Figure 9 Fig. 6 shows a computer readable storage medium 600 according to an embodiment of the present application. The computer readable storage medium 600 stores computer readable instructions. When the computer readable instructions are run by a processor, a method for semantic analysis of speech in a people-dense environment according to an embodiment of the present application described above with reference to the accompanying drawings can be implemented. The storage medium 600 includes, but is not limited to, volatile memory and / or non-volatile memory. The volatile memory can include, for example, random access memory (RAM), cache memory and the like. The non-volatile memory can include, for example, read only memory (ROM), hard disk, flash memory and the like.
[0127] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description, not representing the advantages and disadvantages of the embodiments. The above-mentioned specific embodiments of the present application are described. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are possible or can be advantageous.
[0128] Each of the embodiments in the present specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments.
[0129] The above-mentioned is only the preferred embodiment of the present application, and does not limit the present application. Any modification, equivalent replacement, improvement and the like within the principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for semantic parsing of speech in a crowded environment, characterized by: include: Collect and obtain the target user's voice signal data and lip movement video stream data; Extract lip movement feature sequences based on lip movement video stream data, and map them into predicted speech feature vectors using a pre-trained convolutional neural network model. The predicted speech feature vector is input into the voiceprint separation model to separate the target user's speech segment from the mixed speech signal and generate pure speech data features; Based on the initial and final structure of Chinese pronunciation, the clean speech data features are segmented in the time domain to obtain the time domain waveform of the speech signal of a single word; Based on the speaker's 21 types of pure initials and 35 types of finals, the time domain waveform template of each type of pronunciation is extracted to build the initial and final time domain waveform library; Based on the priority of voiceless consonants, the voiceless consonant initials are determined according to the average zero-crossing rate of the first 50ms segment of the waveform of the word to be matched; Based on the non-voiceless consonant initials, match them from the initial time domain waveform library, calculate the minimum path cost through DTW, and obtain the top three initial candidates with the minimum cost; By performing linear predictive coding analysis on the matching final segment, the first three formant frequencies are obtained, and the matching final is obtained by performing nonlinear time scaling on the DTW path; Based on the acquired initials and finals, perform splicing verification, check the legality of the combination, and generate the corresponding pinyin expression of each character; Based on the pinyin expression of consecutive single words, the tone combination correlation analysis is performed, and the meaning of the target user's speech segment is obtained by combining the contextual semantics.
2. The method for semantic analysis of speech in a crowded environment according to claim 1, characterized in that: The method of extracting a lip movement feature sequence based on the lip movement video stream data and mapping the lip movement feature sequence into a predicted speech feature vector through a pre-trained convolutional neural network model specifically includes: Based on the acquired lip movement video stream data of the target user, the facial region is located in each frame, the lips are accurately located using a key point regression algorithm, and preprocessing is performed; The lip motion features are extracted through the primary, intermediate and advanced feature layers in the convolutional neural network model to obtain the lip motion feature sequence; Based on the acquired lip movement feature sequence, it is aligned with the speech frame rate and mapped into a predicted speech feature vector through a pre-trained convolutional neural network encoder-decoder. The predicted speech feature vector includes a joint representation of Mel-frequency cepstral coefficients and fundamental frequency.
3. The method for semantic analysis of speech in a crowded environment according to claim 1, characterized in that: Inputting the predicted speech feature vector into the voiceprint separation model, separating the target user's speech segment from the mixed speech signal, and generating pure speech data features specifically includes: The mixed speech signal is divided into frames and converted into a time-frequency spectrum, and the acoustic feature dimension of lip movement prediction is enhanced; Generate a compact voiceprint vector based on the Mel-frequency cepstral coefficients and fundamental frequency of the predicted speech feature vector, construct a voiceprint template matrix through time series expansion, and calibrate the speaker characteristics; Based on the obtained voiceprint template matrix, it is input into the convolutional neural network encoder, and the time-frequency mask of the target speech is obtained by similarity weighting; The target complex spectrum is extracted based on the time-frequency mask, and the time domain waveform is reconstructed according to the fundamental frequency constraint phase of the lip movement prediction to generate the pure speech data features.
4. The method for semantic analysis of speech in a crowded environment according to claim 1, characterized in that: The time domain segmentation of the clean speech data features based on the initial and final structure of Chinese pronunciation to obtain the time domain waveform of the speech signal of a single word specifically includes: By monitoring the signal data with high zero-crossing rate and low energy in the clean speech data, the spike pulses on the energy curve that are continuously less than the threshold are obtained and determined to be initial consonants; By monitoring the data with high energy and stable formant in the clean speech data, the final vowel is determined through energy envelope analysis and formant tracking; Based on the energy of the speech signal, the silence segment is continuously determined to obtain the inter-word interval; For each candidate syllable, determine whether it contains an initial consonant and a final vowel structure, and directly extract the final vowel segment for syllables without initial consonants; By segmenting the single word syllables, the time domain waveform of the speech signal of the single word is obtained.
5. The method for semantic analysis of speech in a crowded environment according to claim 4, characterized in that: The method of monitoring data with high energy and stable formants in the clean speech data and determining the finals by energy envelope analysis and formant tracking specifically includes: Calculate the logarithmic energy of each frame of speech, generate an energy envelope curve, and mark the starting point of the envelope rising edge as the starting boundary of the final; The first three formants F1-F3 are calculated by linear predictive coding. When the fluctuation of F1 is less than the threshold for 10 consecutive frames, it is determined to be a stable final segment.
6. The method for semantic analysis of speech in a crowded environment according to claim 1, characterized in that: The tone combination correlation analysis based on the pinyin expression of consecutive words and the acquisition of the target user's speech segment meaning in combination with contextual semantics specifically include: Arrange and combine the pinyin tones of consecutive words to form all possible tone sequences; Based on the frequency of tone combinations in the corpus and the semantics of the words, a relevance weight is assigned to each tone combination. The higher the frequency, the greater the weight; the higher the semantic relevance, the greater the weight. Sort the tone sequences from high to low based on their weights, and select the tone combination with the highest weight; The specific content of the entire speech segment is obtained based on the pinyin expression of the acquired tone combination and single word, and the speech segment content is modified according to the grammatical structure of the words and sentences before and after the pinyin segment to obtain the meaning of the speech segment of the target user.
7. An electronic device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the semantic parsing method of speech in a crowded environment as described in any one of claims 1-6.
8. A computer-readable storage medium storing computer-readable instructions, characterized in that: When the computer-readable instructions are executed by a processor, a method for semantic parsing of speech in a crowded environment according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Chinese phonetic recognition method
CN102208186A
Multi-person voice separation method and device based on voiceprint features and medium
CN113990344A