A method and system for extracting voice data using artificial intelligence

Through artificial intelligence technology, combined with the multi-form fusion processing of voice and video data and dynamic configuration of weights, the accuracy and stability of voice data extraction in video communication is solved, and efficient speech recognition in complex environments is achieved.

CN120319224BActive Publication Date: 2025-09-02HANGZHOU ZHILIAO INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510820326.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-02
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

The existing voice data extraction technology in video communication scenarios has poor accuracy and completeness of voice data extraction, which is difficult to meet the needs of high-quality information extraction in complex application scenarios.

Method used

By monitoring the audio and video data in video communication, the speech speed and volume are recognized, combined with the proportion of lip-shaped image and user historical voice data, the weight is dynamically configured to realize the fusion processing of multi-form data, and improve the accuracy and stability of voice data extraction.

Benefits of technology

Under low-quality audio and complex speech speed conditions, the accuracy and stability of speech data extraction are improved, the recognition error rate is reduced, and environmental adaptability and application generalization are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120319224B_ABST
    Figure CN120319224B_ABST
Patent Text Reader

Abstract

This application relates to a method and system for extracting voice data using artificial intelligence, and relates to the field of intelligent voice extraction technology, including: monitoring and collecting video and audio data generated by video communications to obtain corresponding data sequences; performing speech rate and volume recognition on the audio sequence, obtaining feature parameters and respectively configuring lip shape and audio weights to obtain first and second sets of weights; performing lip shape image ratio analysis on the video sequence, combining it with vocabulary analysis of the user's historical voice data to obtain third and fourth sets of weights; fusing the four sets of lip shape and audio weights, performing voice content recognition and extraction on the audio and video sequence to obtain the final voice data. This invention solves the problems of existing technologies such as reliance on a single voice data extraction method, poor voice data extraction quality, and lack of personalized adaptation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent speech extraction, and in particular to a speech data extraction method and system using artificial intelligence. Background Art

[0002] With the widespread application of video communication technology, voice data extraction has become a core link in information exchange and processing in areas such as intelligent customer service and online conferencing. This data plays a key role in achieving efficient human-computer communication, content analysis and storage, and understanding user needs.

[0003] Most existing voice data extraction technologies rely on a single natural language recognition method. When faced with video communication scenarios, the accuracy and completeness of voice data extraction are often poor due to poor audio quality (such as background noise, fast speaking speed, low volume, etc.), making it difficult to meet the high-quality information extraction needs in complex application scenarios. Summary of the Invention

[0004] In order to solve the above technical problems, the present application provides a voice data extraction method and system using artificial intelligence, which overcomes the defects of the existing technology of single voice audio recognition method and poor extraction quality, improves the accuracy and stability of voice data extraction in complex scenarios, and realizes reliable extraction under poor conditions such as noisy environments and fast speaking speeds.

[0005] The embodiments of this application disclose the following technical solutions:

[0006] In a first aspect, an embodiment of the present application provides a method for extracting speech data using artificial intelligence, the method comprising:

[0007] Monitor and collect video data and audio data generated by video communication to obtain a video data sequence and an audio data sequence;

[0008] Performing speech rate recognition and volume recognition according to the audio data sequence to obtain speech rate feature parameters and volume feature parameters, and respectively configuring lip shape weights and audio weights to obtain a first lip shape weight, a first audio weight, a second lip shape weight, and a second audio weight;

[0009] Performing a lip shape image ratio analysis on the video data sequence to obtain a third lip shape weight and a third audio weight, and performing a vocabulary analysis based on historical speech data of the target user to obtain a fourth lip shape weight and a fourth audio weight;

[0010] According to the first lip shape weight, the first audio weight, the second lip shape weight, the second audio weight, the third lip shape weight, the third audio weight, the fourth lip shape weight, and the fourth audio weight, the fused lip shape weight and the fused audio weight are obtained by processing, and voice content recognition and extraction are performed on the video data sequence and the audio data sequence to obtain voice data.

[0011] In a second aspect, an embodiment of the present application provides a voice data extraction system using artificial intelligence, the system comprising:

[0012] The data acquisition module is used to monitor and collect video data and audio data generated by video communication, and obtain video data sequences and audio data sequences;

[0013] An audio feature analysis and weight configuration module is used to perform speech rate recognition and volume recognition based on the audio data sequence to obtain speech rate feature parameters and volume feature parameters, and to configure lip shape weights and audio weights respectively to obtain a first lip shape weight, a first audio weight, a second lip shape weight, and a second audio weight;

[0014] a video and vocabulary feature analysis and weight configuration module, configured to perform a lip shape image ratio analysis on the video data sequence, process the analysis to obtain a third lip shape weight and a third audio weight, and extract and record the target user's historical voice data, perform a vocabulary analysis, and process the analysis to obtain a fourth lip shape weight and a fourth audio weight;

[0015] The weight fusion and speech recognition module is used to process the first lip shape weight, the first audio weight, the second lip shape weight, the second audio weight, the third lip shape weight, the third audio weight, the fourth lip shape weight, and the fourth audio weight to obtain a fused lip shape weight and a fused audio weight, and perform speech content recognition and extraction on the video data sequence and the audio data sequence to obtain speech data.

[0016] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0017] This application proposes a method and system for extracting voice data using artificial intelligence. By collecting audio data from video communications, comprehensively analyzing audio features, video lip shape features, and the user's historical language habits, the audio and video recognition results are combined with dynamic fusion weights. It simultaneously processes multiple forms of data, including video and audio, and dynamically adjusts weights based on speech speed, volume, the proportion of lip shape images, and the user's historical vocabulary usage. It also adapts to the user's language habits, effectively improving the accuracy and stability of voice data extraction, achieving efficient voice recognition under poor conditions such as low-quality audio and complex speech speeds, and reducing the recognition error rate. Through the steps of audio feature analysis process, lip shape weight calculation, and user habit adaptation, it combines multiple data and accurately evaluates each influencing factor, effectively avoiding the problem of extraction failure caused by reliance on a single factor and environmental changes. At the same time, dynamic weight control effectively improves the environmental adaptability and application generalization of voice data extraction.

[0018] The technical solution provided by this application solves the problems existing in existing voice data extraction technologies, such as single reliance on voice data extraction, poor voice data extraction quality, and lack of personalized adaptation. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0020] Figure 1 A flowchart of a method for extracting voice data using artificial intelligence provided in an embodiment of the present application;

[0021] Figure 2 A schematic diagram of the structure of a voice data extraction system using artificial intelligence provided in an embodiment of the present application;

[0022] In the accompanying drawings, the components represented by the reference numerals are described as follows:

[0023] Data collection module 01, audio feature analysis and weight configuration module 02, video and vocabulary feature analysis and weight configuration module 03, weight fusion and speech recognition module 04. DETAILED DESCRIPTION

[0024] The present application provides a method and system for extracting voice data using artificial intelligence, which is used to solve technical problems existing in the prior art, such as single reliance on voice data extraction, poor voice data extraction quality, and lack of personalized adaptation.

[0025] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0026] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the described features. In the description of this application, "plurality" means two or more, unless otherwise specifically specified.

[0027] In the description of this application, the term "for example" is used to mean "used as an example, illustration or explanation". Any embodiment described in this application as "for example" is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in this application.

[0028] Example 1, as shown in the attached Figure 1 As shown, the present application provides a method for extracting speech data using artificial intelligence, the method comprising the following steps:

[0029] S100: monitoring and collecting video data and audio data generated by video communication to obtain a video data sequence and an audio data sequence;

[0030] In the embodiments of this application, during the voice data extraction process, in order to achieve synchronous processing and reliable analysis of video and audio information, it is necessary to obtain complete and valid raw video and audio data. The collected raw video and audio data are then converted into a data sequence form that can be used for subsequent analysis and processing. This lays the data foundation for subsequent audio feature-based weight calculation, voice content recognition, and fusion processing, and avoids problems such as low processing efficiency and information omissions caused by inconsistent data sequences and inconsistent order.

[0031] Specifically, video data collection is achieved by monitoring the video stream during video communication. Video images are captured using a video acquisition device (such as a mobile phone camera). Data conversion technology is then used to convert the captured video images into a standardized digital format video data sequence, providing the basic data support for subsequent voice data extraction and analysis.

[0032] Among them, data conversion technology is to capture the analog video signal in the communication through video acquisition equipment, and convert it into a digital signal through analog-to-digital conversion technology, that is, use standard compression such as H.264 encoding, and then organize it into a video data sequence according to storage formats such as MP4. After unifying parameters such as format and frame rate, the conversion of the original picture to standardized digital video is completed.

[0033] Furthermore, the audio data generated by video communication is collected by deploying audio collection equipment (such as microphones) on the communication equipment, monitoring the audio signals in the communication in real time, converting the analog audio into digital audio signals, and forming an audio data sequence according to a certain sampling frequency and quantization accuracy, so as to facilitate the subsequent unified processing and analysis of the video and audio data.

[0034] Video and audio data collection supports both real-time monitoring and post-recording extraction. For example, a 2-second audio and video clip is captured. In real-time mode, audio and video data is synchronously captured by the audio and video acquisition device deployed on the communication device. Recording mode allows for the extraction of relevant clips from stored videos or call logs. The collected audio and video data is converted into a standardized digital format sequence to ensure consistent processing in the subsequent voice extraction and analysis processes.

[0035] S200: Performing speech rate recognition and volume recognition based on the audio data sequence to obtain speech rate feature parameters and volume feature parameters, and configuring lip shape weights and audio weights respectively to obtain a first lip shape weight, a first audio weight, a second lip shape weight, and a second audio weight;

[0036] In the embodiment of the present application, during the voice data extraction process, in order to ensure the accuracy and reliability of the voice data extraction, the video and audio data sequences acquired in real time need to be processed for speech rate recognition and volume recognition respectively.

[0037] Specifically, the audio data sequence is identified and analyzed using preset recognition rules and algorithms to obtain speech rate characteristic parameters and volume characteristic parameters, which are then input into a weight calculation module.

[0038] Furthermore, the module analyzes and processes the input feature parameters and configures the lip shape weight and audio weight according to the preset weight configuration rules and algorithms, and outputs the first lip shape weight, the first audio weight, the second lip shape weight and the second audio weight.

[0039] The weight parameters obtained according to these steps can effectively optimize the speech data extraction process, ensure the accuracy of subsequent speech data extraction, and avoid problems such as speech data extraction errors or losses caused by factors such as fast speaking speed and low volume.

[0040] Step S200 of the method provided in the embodiment of the present application includes:

[0041] Inputting the audio data sequence into a speech rate recognizer to identify and obtain speech rate feature parameters;

[0042] Extracting the average volume in the audio data sequence as a volume feature parameter;

[0043] According to the speech rate characteristic parameters, the lip shape weight and the audio weight are configured to obtain a first lip shape weight and a first audio weight. According to the volume characteristic parameters, the lip shape weight and the audio weight are configured to obtain a second lip shape weight and a second audio weight.

[0044] In an embodiment of the present application, the acquired audio data sequence is transmitted to a pre-built speech rate recognizer, which parses the speech rate in the audio data sequence based on preset algorithms and rules, and calculates and outputs characteristic parameters that can characterize the speed of the speech by analyzing information such as the time interval and syllable distribution of the speech signal.

[0045] The step of “building a speech rate recognizer” in the method provided in the embodiment of the present application includes:

[0046] Using machine learning, a network structure of a speech rate identifier is constructed, wherein the input data of the speech rate identifier is an audio data sequence and the output data is a speech rate feature parameter;

[0047] According to the speech data recognition records in the historical time, a set of sample audio data sequences is collected, and the speech speed of different sample audio data sequences is marked to obtain a set of sample speech speed feature parameters;

[0048] The speech rate identifier is supervised trained and tested using the sample audio data sequence set and the sample speech rate feature parameter set until convergence.

[0049] In the embodiment of the present application, a speech rate recognizer network architecture is first constructed based on a machine learning method, and its input is set as an audio data sequence, and its output is a characteristic parameter representing the speech rate.

[0050] The speech rate identifier network extracts and analyzes hierarchical features from audio data through a multi-layer neural network structure. This network design includes a feature preprocessing module, a time-frequency conversion module, a feature extraction module, and a parameter output module. These modules work together to analyze the speech rate characteristics of audio data sequences.

[0051] Specifically, an audio data sequence is received through an input interface. First, the audio data is subjected to noise reduction, format unification and other processing through a feature preprocessing module; then, the audio signal is converted from the time domain to the frequency domain through a time-frequency conversion module; subsequently, according to the preset rules of the feature extraction module, detailed information such as the time interval and frequency distribution of the speech signal is analyzed, and its speech rate characteristics are further calculated; finally, the quantized speech rate feature parameters are output through a parameter output module.

[0052] After repeated debugging and optimization of the above steps and determining the optimal parameters and operating logic of each module, the construction of the speech rate recognizer was completed, thus forming a complete audio data processing flow, ensuring the stable extraction of speech rate features in the audio and the output of reliable speech rate feature parameters.

[0053] Furthermore, by collecting historical speech data recognition records, a sample audio data sequence set is screened and constructed, and the speaking speed of each sample audio data sequence is manually labeled to form a corresponding sample speaking speed feature parameter set.

[0054] The sample audio data sequence collection includes speech data from various scenarios, such as formal presentations and daily conversations, as well as data from men, women, young and old, with different accents and speaking styles. After audio detection and elimination of invalid audio segments, the valid audio is unified into a standard encoding format using format conversion tools and parameter adjustments. This ultimately creates a sample audio data sequence collection containing a variety of speech features, providing reliable data support for speech rate recognition.

[0055] Furthermore, according to a unified standard, the speaking speed of the sample audio data sequence is measured and labeled one by one, and a corresponding set of sample speaking speed feature parameters is generated, providing a reliable basis for subsequent training optimization.

[0056] The speech rate recognizer was trained and tested using the two sample sets as training data. Network parameters were adjusted and the model structure optimized through multiple iterations, while the training results were continuously verified until the recognition results stabilized.

[0057] Specifically, during the training process, the sample audio data sequence is input into the recognizer, the output speech rate feature parameters are compared with the labeled values, and after calculating the error, the processing parameters such as the feature preprocessing module and the time-frequency conversion module are adjusted in a targeted manner to optimize the feature extraction processes such as the time interval analysis of the speech signal and the frequency distribution calculation.

[0058] Furthermore, by repeatedly testing different module combinations and data processing sequences, the speech rate identifier architecture was effectively adjusted. After multiple rounds of iterative training, when the error fluctuations over multiple training sessions were kept to a very small range, such as an error rate below 0.5%, the model was considered to have reached a state of convergence, ensuring that it could accurately extract speech rate characteristic parameters from various types of audio data.

[0059] In the embodiment of the present application, the input audio data sequence is scanned frame by frame through an audio processing algorithm, the instantaneous volume value of each frame of the audio signal is calculated, and then the volume values ​​of all frames are arithmetic averaged. The resulting value is the average volume of the audio data sequence, and it is output as a volume feature parameter representing the audio loudness.

[0060] The audio processing method in the above steps may use a short-time energy algorithm to process the input audio data sequence.

[0061] Specifically, the collected audio data is first divided into multiple short frames. The data and energy value of each frame are calculated. This energy value is generally related to the volume level. The energy values ​​of each frame are then summed and divided by the total number of frames to calculate the average energy of the audio data sequence. This is then converted into a corresponding average volume value, which is then output as a volume feature parameter representing the audio loudness.

[0062] In the method provided in the embodiment of the present application, the steps of “configuring the lip shape weight and the audio weight according to the speech rate characteristic parameter to obtain the first lip shape weight and the first audio weight, and configuring the lip shape weight and the audio weight according to the volume characteristic parameter to obtain the second lip shape weight and the second audio weight” include:

[0063] According to the speech data recognition record in the historical time, the average speech speed characteristic parameter and the average volume characteristic parameter are obtained;

[0064] Calculating a ratio of the average speaking rate characteristic parameter to the speaking rate characteristic parameter, performing correction calculation on a preset audio weight to obtain a first audio weight, and calculating a first lip shape weight based on the first audio weight;

[0065] The ratio of the volume characteristic parameter to the average volume characteristic parameter is calculated, and the preset audio weight is corrected and calculated to obtain a second audio weight. The second lip shape weight is calculated based on the second audio weight.

[0066] In the embodiment of the present application, all voice data within a specified historical time period is filtered out based on the historically stored voice data recognition records.

[0067] Furthermore, the speech rate characteristic parameter and volume characteristic parameter in each data record are summed up respectively, and the summed result is divided by the total number of data records in the historical time period, thereby calculating the average speech rate characteristic parameter and the average volume characteristic parameter.

[0068] Furthermore, the ratio of the average speaking rate characteristic parameter to the current speaking rate characteristic parameter is calculated. Specifically, the average speaking rate characteristic parameter is divided by the current speaking rate characteristic parameter to obtain the ratio between the two, that is, ratio = average speaking rate characteristic parameter / current speaking rate characteristic parameter.

[0069] Furthermore, the preset audio weight is corrected and calculated using the ratio, specifically by multiplying the preset audio weight by the ratio to obtain the first audio weight, that is, the first audio weight=the preset audio weight×the ratio.

[0070] The default audio weight is a parameter pre-set in the speech data extraction system to measure the importance of audio information in speech content recognition. The value is usually between 0 and 1, representing the initial importance of audio information in the speech content recognition process. For example, when the default audio weight is 0.58, it means that before adjusting other factors such as speech speed and volume, the system defaults to audio information accounting for 58% of the proportion of speech content recognition, while other information such as lip shape accounts for 42%.

[0071] Furthermore, based on the constraint that the sum of the lip shape weight and the audio weight is 1, the first lip shape weight is calculated, specifically 1 minus the first audio weight, to obtain the first lip shape weight, that is: the first lip shape weight = 1 - the first audio weight.

[0072] For example, the preset audio weight is 0.58. If the average speaking speed feature parameter is 120 words / min and the current speaking speed feature parameter is 100 words / min, the ratio of the average speaking speed feature parameter to the current speaking speed feature parameter is 120 / 100=1.2. At this time, the first audio weight is 1.2×0.58=0.696 (approximate value 0.7), and the first lip shape weight is 1-0.696=0.304 (approximate value 0.3).

[0073] Similarly, the volume characteristic parameters in all data records within the historical time period are summarized, the sum is divided by the total number of data records, and the average volume characteristic parameter is obtained, which reflects the average level of historical volume.

[0074] Furthermore, the volume characteristic parameter of the current voice data is obtained and divided by the average volume characteristic parameter to obtain the ratio of the two, that is: ratio = average volume characteristic parameter / volume characteristic parameter of the current voice data. The ratio reflects the degree of change of the current volume compared with the historical average volume.

[0075] Furthermore, this ratio is used to correct the preset audio weight. The preset audio weight is multiplied by the ratio to obtain the second audio weight. If the ratio is greater than 1, it means that the current volume is higher than the historical average level, and the second audio weight will increase accordingly, otherwise it will decrease.

[0076] Furthermore, based on the constraint that the sum of the lip shape weight and the audio weight is always 1, the second lip shape weight can be calculated by subtracting the second audio weight from 1, so that the audio and lip shape weights can be adjusted to each other as the volume changes, ensuring the reasonable distribution of the weights of the two during speech recognition and improving recognition accuracy.

[0077] For example, if the preset audio weight is 0.5, and the average volume characteristic parameter is 60 decibels and the current volume characteristic parameter is 60 decibels, then the ratio of the current volume characteristic parameter to the average volume characteristic parameter is 60 / 60 = 1. In this case, the second audio weight is 1×0.5=0.5, and the second lip shape weight is 1-0.5=0.5.

[0078] The embodiment of the present application constructs a speech rate identifier and realizes dynamic adjustment between the lip shape of the speech rate feature and the audio weight through the weight calculation method of the above steps, effectively improving the accuracy and stability of voice data extraction in complex scenarios.

[0079] For example, when speech is spoken too quickly, the voice signal tends to blur or overlap, making audio recognition more difficult and quality deteriorating. The system automatically reduces the audio weight according to pre-set weight adjustment rules. When speech volume is low, the voice signal is more susceptible to background noise interference, causing poor audio quality. Similarly, the audio weight is reduced according to pre-set rules. These steps increase the weight of video lip-sync information in the voice data extraction process, ensuring accuracy.

[0080] S300: Analyzing the proportion of lip shape images in the video data sequence to obtain a third lip shape weight and a third audio weight, and extracting and recording historical speech data of the target user to perform vocabulary analysis to obtain a fourth lip shape weight and a fourth audio weight.

[0081] In an embodiment of the present application, when configuring lip shape and audio weights, in order to coordinately process video image features and voice data features and achieve accurate configuration, it is necessary to perform lip shape image ratio analysis on the video data sequence and perform vocabulary analysis on the historical voice data of the target user.

[0082] Furthermore, the above analysis results are converted into corresponding weight parameters through an established algorithm. Subsequently, the lip shape and audio weights are dynamically adjusted according to the actual picture conditions and user usage habits to achieve personalized voice interaction.

[0083] During the lip shape and audio weighting process, a lip shape ratio recognizer constructed using a convolutional neural network processes the video data sequence. Through the recognizer's multi-layer convolution, pooling, and fully connected operations, it outputs the lip shape image ratio coefficient for each video frame. These coefficients are averaged and compared with preset baseline values. The weight parameters are then adjusted according to the established proportional relationship to derive the third lip shape weight and the third audio weight.

[0084] Furthermore, a lexical analysis is performed on the historical speech data of the target user to calculate the fourth lip shape weight and the fourth audio weight, which are then combined with the third set of weights to form a final configuration solution.

[0085] Step S300 in the method provided in the embodiment of the present application includes:

[0086] Inputting each video data in the video data sequence into a lip shape ratio recognizer to identify and obtain a plurality of lip shape image ratio coefficients, wherein the lip shape ratio recognizer is constructed using a convolutional neural network and is trained using a sample video data set and a set of labeled sample lip shape image ratio coefficients;

[0087] Get the average lip shape image ratio;

[0088] Calculating a ratio of a mean of the plurality of lip shape image proportion coefficients to the average lip shape image proportion coefficient, and performing a correction calculation on the preset lip shape weight to obtain a third lip shape weight;

[0089] A third audio weight is calculated based on the third lip shape weight.

[0090] In an embodiment of the present application, when configuring lip shape and audio weights, in order to achieve coordinated and precise configuration of video image features and voice data features, a lip shape ratio recognizer constructed based on a convolutional neural network is used to process the video data sequence frame by frame.

[0091] Among them, the recognizer performs computational analysis on each frame of video image through multi-layer convolution, pooling and fully connected structure, and outputs multiple lip image proportion coefficients.

[0092] Specifically, when constructing a lip shape ratio identifier, a convolutional neural network architecture is utilized to perform sliding convolutions on video frames using convolution kernels, extracting detailed features from bottom-level to high-level layers, including the edge contours, geometric shapes, and dynamic changes of the lip shape area in the video. A pooling layer further reduces the dimensionality of the feature map, preserving key information while reducing computational complexity. This identifier can identify key information about lip shape and size changes in a video frame by frame, as well as the dynamic trends of lip shape across consecutive frames, according to established program rules, clearly analyzing the lip shape information in the video.

[0093] Furthermore, to ensure the accurate operation of the recognizer, it is necessary to construct a sample video dataset that includes different scenes, characters, and lip shape changes (including different scenes such as formal reports and daily conversations, as well as video data of men, women, young and old, with different accents and speaking styles). At the same time, manual labeling is used to calculate the proportion of the lip shape image in each video in the picture, and collect and form a standard sample dataset.

[0094] During the recognizer training phase, sample video data is fed into the recognizer network segment by segment, and the backpropagation algorithm gradually adjusts the recognizer's internal calculation parameters and processing logic. Using the manually labeled percentage coefficients as a reference, the recognizer's calculation process and judgment criteria are continuously optimized to minimize the error between the predicted results and the actual percentage.

[0095] Furthermore, after multiple iterations of training and parameter optimization, the lip shape ratio recognizer can stably and efficiently recognize and process the input video data sequence, accurately calculate the lip shape image ratio coefficient of each video frame, and provide key data support for subsequent lip shape weight configuration and lip shape and audio collaborative processing.

[0096] After completing the construction and annotation of the sample video dataset, the lip image ratio coefficients of all sample videos are summarized, and the average lip image ratio coefficient is obtained by summing up and dividing by the total number of samples.

[0097] Among them, this coefficient serves as a benchmark value for measuring the proportion of lip-sync images in the video screen, and can be used to compare the subsequent analysis results of the lip-sync image proportion with the actual video data sequence, providing an important reference basis for the correction and configuration of weight parameters.

[0098] Furthermore, a plurality of lip image proportion coefficients outputted from the video data sequence after being processed by the lip proportion recognizer are summarized, and the mean of the group of coefficients is obtained by arithmetic averaging.

[0099] Furthermore, the mean obtained in the above steps is ratioed with the pre-calculated average lip image ratio coefficient to obtain a correction coefficient, that is: correction coefficient = average lip image ratio coefficient after recognition processing / pre-calculated average lip image ratio coefficient.

[0100] Furthermore, based on the correction coefficient, a multiplication operation is performed on the preset lip shape weight, thereby completing the adjustment and optimization of the preset lip shape weight, and finally obtaining a third lip shape weight suitable for the current video data feature.

[0101] The preset lip shape weight is a pre-set parameter used to measure the importance of lip shape information in speech content recognition. It represents the initial importance of lip shape information in the speech content recognition process, and its value is usually between 0 and 1.

[0102] Furthermore, after obtaining the third lip shape weight, based on the constraint that the sum of the lip shape weight and the audio weight is one, the third audio weight can be obtained by subtraction, that is, the third audio weight=1-third audio weight.

[0103] For example, if the preset lip shape weight is 0.25, and the lip shape image ratio coefficients output after the lip shape ratio recognizer processes the video data sequence are 0.1, 0.2, 0.3, and 0.2, respectively, their average is (0.1 + 0.2 + 0.3 + 0.2) / 4 = 0.2. The pre-calculated average lip shape image ratio coefficient is 0.25, so the correction factor is 0.2 / 0.25 = 0.8. In this case, the third lip shape weight is 0.25 × 0.8 = 0.2, and the third audio weight is 1 - 0.2 = 0.8.

[0104] The calculation method in the above steps ensures that the weight distribution of lip sync and audio playback is always balanced during video playback, thereby achieving coordinated optimization of the two and providing users with a more natural and smooth audio-visual experience.

[0105] For example, when a video focuses on a close-up of a person, with their mouth occupying a large portion of the frame, if the calculated lip shape image area accounts for 70%, this indicates that the lip shape information is clear and complete, and a higher lip shape weight will be assigned. If, in a video shot from a distance, the person's mouth area only occupies 5% or even less of the frame, or if the mouth is completely out of frame due to a camera switch (a percentage of 0), the corresponding lip shape weight will also be significantly reduced. By comparing and associating the lip shape image area percentage value with the weight, differentiated processing of lip shape information in different video scenes can be achieved.

[0106] The method provided in the embodiment of the present application includes the steps of “extracting records based on the historical speech data of the target user, performing vocabulary analysis, and processing to obtain the fourth lip shape weight and the fourth audio weight”.

[0107] Collect historical extracted voice data sets from users who have extracted voice data in historical time;

[0108] Counting the number of occurrences of uncommon words on the historically extracted speech data set to obtain the number of occurrences of uncommon words;

[0109] Collect the number of occurrences of uncommon words from multiple sample users and calculate the mean to obtain the average number of occurrences of uncommon words;

[0110] According to the ratio of the average number of occurrences of uncommon words to the number of occurrences of uncommon words, the preset lip shape weight is modified and calculated to obtain a fourth lip shape weight;

[0111] A fourth audio weight is calculated based on the fourth lip shape weight.

[0112] In an embodiment of the present application, multiple users of different ages, occupations and language habits are selected as sample objects, and the voice data generated by them in scenarios such as voice chatting, voice command input, and voice course learning in historical time are captured through a built-in voice data acquisition device (such as a microphone, etc.).

[0113] During the collection process, the acquired voice data is processed according to established data collection standards and operating procedures. This involves converting data storage formats uniformly, such as converting voice data from different sources to common formats like MP3 and WAV, and adjusting parameters like audio sampling rate and bit depth to ensure consistent formatting across all voice data. Furthermore, duplicate voice segments are identified and removed to ensure the accuracy and validity of the collected data.

[0114] Furthermore, valid voice data that has been formatted and deduplicated is categorized based on data generation scenarios, usage time, user type, and other dimensions. For example, voice data can be divided into categories such as voice commands, voice chat, and voice learning based on voice interaction scenarios; or categorized chronologically, using monthly or quarterly cycles.

[0115] After classification is complete, the data is aggregated and integrated, archived and stored according to standardized naming rules and storage paths, ultimately forming a complete historical voice data set. This set provides reliable data support for further analysis of user voice usage preferences, calculation of lip shape and audio weight parameters, and more.

[0116] Furthermore, all speech content in the historically extracted speech data set is converted into text format and compared against a pre-organized dictionary of rare words. During the verification process, each word in the text is carefully identified. If a word matches an uncommon word in the dictionary, it is marked and counted in a dedicated counting table or system, and the corresponding count value is increased, thus accurately counting the number of occurrences of the uncommon word.

[0117] Furthermore, after completing the verification of all text content, the final cumulative value in the counting table is the number of occurrences of uncommon words in the historically extracted speech data set. This value reflects the frequency of uncommon words used by users during speech expression and provides important data for the subsequent accurate calculation of weight parameters.

[0118] We then aggregate the number of rare word occurrences across all sample users, summing the total. This total is then divided by the total number of sample users to arrive at the average rare word occurrence. The specific calculation formula is: Average rare word occurrence = Sum of rare word occurrences / Total number of sample users. This calculation provides a direct reflection of the overall frequency of rare word use across the sample user population.

[0119] Furthermore, by calculating the ratio of the number of occurrences of uncommon words for a single sample user to the average number of occurrences of uncommon words, the relative level of the user's frequency of using uncommon words in the group is evaluated, which is specifically expressed as: ratio = number of occurrences of uncommon words for a single sample user / average number of occurrences of uncommon words.

[0120] Specifically, if the ratio is greater than 1, it indicates that the user's frequency of using uncommon words is higher than the group average. Since there is less lip shape training data corresponding to uncommon words, in order to reduce the lip shape recognition bias, the preset lip shape weight is lowered according to the preset rules; if the ratio is less than 1, it means that the user's frequency of using uncommon words is lower than the average, and the preset lip shape weight can be appropriately adjusted.

[0121] Through the above steps, the preset lip shape weights are corrected and calculated to obtain a fourth lip shape weight that fits the user's personalized language usage habits. This weight parameter fully considers the impact of the user's use of uncommon words on lip shape recognition, closely fits the user's personalized language usage habits, and can more accurately adapt to the user's voice characteristics.

[0122] Furthermore, based on the constraint that the sum of the lip shape weight and the audio weight is 1, after determining the fourth lip shape weight, a subtraction operation, that is, 1 minus the fourth lip shape weight, can be used to obtain the corresponding fourth audio weight.

[0123] For example, the preset lip shape weight is 0.5. If there are 5 sample users, and their uncommon words appear 3 times, 4 times, 2 times, 5 times, and 6 times respectively, the total number of uncommon word appearances is 20, and the average number of uncommon word appearances is 20 / 5 = 4. A target user's uncommon word appears 2 times, and the ratio is 2 / 4 = 0.5, which is less than 1. According to the rule of "upward adjustment coefficient = 0.5 × (1-ratio)", the upward adjustment coefficient is 0.5 × (1-0.5) = 0.25. Therefore, the fourth lip shape weight is 0.5 × (1+0.25) = 0.625 (approximately 0.6), and the fourth audio weight is 1-0.6 = 0.4.

[0124] Through the calculation method of the above steps, the fourth audio weight and the fourth lip shape weight obtained are adapted to each other. When the voice and lip shape data are subsequently processed, resources can be reasonably allocated, the audio data processing efficiency can be improved, and better coordination between voice and lip shape can be achieved.

[0125] For example, in the statistical analysis of user voice data, if the proportion of uncommon characters in the user's historical voice data reaches 30%, the system will assign a lower lip shape weight because there are few standard lip shape samples corresponding to uncommon characters, and recognition errors are prone to occur when matching lip shape with voice; on the contrary, if the proportion of uncommon characters used by the user is only 5%, it indicates that the user mostly uses common words, has rich corresponding lip shape training samples, and has high lip shape recognition accuracy, and the system will assign a higher lip shape weight.

[0126] S400: Based on the first lip shape weight, the first audio weight, the second lip shape weight, the second audio weight, the third lip shape weight, the third audio weight, the fourth lip shape weight, and the fourth audio weight, processing is performed to obtain a fused lip shape weight and a fused audio weight, and voice content recognition and extraction are performed on the video data sequence and the audio data sequence to obtain voice data.

[0127] In the embodiments of this application, during the collaborative processing of voice and video, to achieve precise adaptation of the video image and audio content, it is necessary to integrate and optimize multiple sets of lip-sync and audio weight parameters, and further analyze the video data sequence and audio data sequence. By combining the integrated weight parameters with the extracted voice data, a basis is provided for the subsequent intelligent matching of voice and video, providing a personalized audio-visual experience.

[0128] Among them, in the process of weight parameter fusion, the first to fourth lip shape weights and audio weights are summarized, and weighted calculation and optimization processing are performed through a preset algorithm to obtain the fused lip shape weight and fused audio weight respectively.

[0129] Furthermore, based on the obtained fusion lip-sync and audio weights, speech content recognition and extraction are performed on the video and audio data sequences. By analyzing the lip-sync changes in the video frame by frame and parsing the audio content sentence by sentence, the speech information is identified and extracted. The extracted speech data is combined with the fusion weights to develop a final voice and video collaborative processing solution, achieving coordinated matching of voice and video.

[0130] Step S400 in the method provided in the embodiment of the present application includes:

[0131] According to the first lip shape weight, the first audio weight, the second lip shape weight, the second audio weight, the third lip shape weight, the third audio weight, the fourth lip shape weight, and the fourth audio weight, a fusion calculation is performed to obtain a fusion lip shape weight and a fusion audio weight;

[0132] Performing lip-sync speech content recognition and extraction and audio speech content recognition and extraction according to the video data sequence and the audio data sequence respectively to obtain a lip-sync speech content set and an audio speech content set;

[0133] According to the fused lip shape weight and the fused audio weight, weighted fusion processing is performed on the lip shape speech content set and the audio speech content set to obtain speech data, wherein the weighted fusion processing includes weighting each speech content and screening the speech content with the largest total weight sum as the speech data.

[0134] In an embodiment of the present application, the first to fourth lip shape weights and audio weights obtained are summarized, and according to a preset weight ratio scheme, through weighted calculation and numerical optimization, the fused lip shape weight and fused audio weight are respectively obtained, providing core parameter basis for subsequent voice and video collaborative processing.

[0135] Among them, when voice and video are processed collaboratively, the calculation of the fusion lip shape weight and the fusion audio weight needs to be completed in multiple steps.

[0136] The first step is the data preparation stage, which collects four sets of lip shape and audio weights calculated based on different algorithms such as video frame rate, speech clarity, lip shape ratio recognition, and rare word frequency analysis. At the same time, based on system requirements and application scenarios, a priority coefficient is assigned to each set of weights to ensure that the sum of all coefficients is 1.

[0137] Next, the weighted calculation phase begins. The weights for the fused lip sync and audio are calculated by multiplying the corresponding group weights by the coefficients and then adding them up. After the weighted calculations are complete, the numerical optimization phase begins. If the sum of the fused lip sync and audio weights is not 1, normalization adjustments are performed to ensure that the final weights are within the valid range of 0 to 1.

[0138] Specifically, the calculated fusion weights are applied to voice and video processing. The fusion lip-sync weights effectively adjust the detail and accuracy of the lip-sync image; the fusion audio weights optimize speech recognition accuracy and enhance background noise filtering. By coordinating the two weights in real time, we further ensure that speech playback and lip-sync changes remain consistent.

[0139] For example, during the data preparation phase, four sets of lip shape and audio weights were obtained: (0.3, 0.7), (0.5, 0.5), (0.2, 0.8), and (0.6, 0.4), corresponding to priority coefficients of 0.2, 0.3, 0.2, and 0.3. During the weighted calculation phase, each set of lip shape and audio weights was multiplied by the coefficients and then added together, resulting in an initial fused lip shape weight of 0.43 and a fused audio weight of 0.57. Since the sum of the two is 1, no normalization adjustment is required. The final fused lip shape weight is 0.43 (approximately 0.4) and the fused audio weight is 0.57 (approximately 0.6). These can be used to optimize lip shape details, speech recognition accuracy, and video and audio synchronization in speech and video processing.

[0140] In the method provided in the embodiment of the present application, the step of “performing lip-sync speech content recognition and extraction and audio speech content recognition and extraction based on the video data sequence and the audio data sequence to obtain a lip-sync speech content set and an audio speech content set” includes:

[0141] Inputting the video data sequence into a plurality of lip-synced speech content recognizers respectively, and recognizing and outputting a lip-synced speech content set, wherein the input data of each lip-synced speech content recognizer is the video data sequence, the output data is the lip-synced speech content, and the training data of each lip-synced speech content recognizer is different;

[0142] The audio data sequence is input into a plurality of audio and speech content recognizers, and the recognition output is used to obtain an audio and speech content set. The input data of each audio and speech content recognizer is the audio data sequence, and the output data is the audio and speech content. The training data for each audio and speech content recognizer is different. Each audio and speech content recognizer is constructed based on a natural language recognition model and is trained by collecting different sample audio data sequences and recognizing and annotating the sample audio and speech content.

[0143] In an embodiment of the present application, in order to more accurately identify the voice content, a multi-angle, multi-scene analysis method is adopted, that is, the same video data sequence is input into multiple lip-sync voice content recognizers in sequence, and each recognizer uses the complete video data sequence as the analysis object, analyzes the lip-sync changes in the video data, and outputs the corresponding voice content.

[0144] Among them, since each recognizer used video samples from different scenes during the training stage, including diverse scenes such as daily conversations, news broadcasts, film and television dialogues, and contained video data with different speaking styles, accent characteristics, and speaking speed rhythm, the output results of each recognizer are different.

[0145] For example, when a recognizer trained with daily conversation video samples processes video data sequences of life scenes, it tends to recognize colloquial expressions and common vocabulary; when a recognizer trained with news broadcast video samples, it tends to recognize standardized terms and current political terms; and when a recognizer trained with film and television dialogue video samples, it tends to capture emotional expressions and the content of lines in specific contexts.

[0146] Furthermore, these different recognition results are summarized and integrated to construct a collection of lip-synced speech content containing multiple possibilities, which further provides rich data reference for subsequent speech content screening and analysis based on weight analysis.

[0147] Similarly, to achieve accurate recognition of audio data, the same audio data sequence is simultaneously input into multiple audio speech content recognizers. Each audio speech content recognizer takes the complete audio data sequence as input and outputs the corresponding speech content recognition result by analyzing the audio waveform features, voiceprint features, and semantic features.

[0148] Among them, since each audio speech content recognizer used audio samples from different sources during the training stage, including news broadcasts, conference recordings, film and television soundtracks, telephone calls and other diverse scenarios, and contained audio data with different accents, speaking speeds, emotional expressions and background noise environments, each recognizer showed significant differences in recognition methods and accuracy distribution.

[0149] For example, when a recognizer is trained using news broadcast audio samples, it tends to recognize clear, standard Mandarin, current affairs vocabulary, and professional terms; when a recognizer is trained using conference recording samples, it tends to recognize scenes where multiple people speak alternately and there is slight noise, and can accurately distinguish the speech content of different speakers; when a recognizer is trained using film and television soundtrack samples, it is more sensitive to speech recognition of lines and dialect expressions.

[0150] Furthermore, these different recognition results are integrated and summarized to form an audio voice content collection containing multiple possibilities, providing a reliable data basis for subsequent voice content verification, error correction and optimization.

[0151] Furthermore, based on the obtained fused lip shape weights and fused audio weights, weighted fusion processing is performed on the lip shape speech content set and the audio speech content set to obtain speech data.

[0152] Specifically, the corresponding fusion weight is first marked for each content in the lip-shaped speech content set and the audio speech content set; then, the lip-shaped weight and audio weight of the candidate speech content with similar or identical pronunciation in the two sets are added together; finally, the content with the largest total weight is selected from all the candidate speech content and determined as the final speech data, so as to achieve accurate matching of speech and lip-shaped information.

[0153] For example, suppose the lip-synced speech content set identifies two "doctors" and the audio speech content set identifies two "clothes." Since "doctor" and "clothes" have similar pronunciations, there is ambiguity in recognition. At this time, the determined fused lip-synced weight of 0.4 is assigned to "doctor," and the fused audio weight of 0.6 is assigned to "clothes." Through calculation, the total weight of "doctor" is its lip-synced weight of 0.4, and the total weight of "clothes" is its audio weight of 0.6 plus the lip-synced weight of 0.6 (because both sets identify "clothes"), which is 1.2. By comparison, it can be seen that the total weight of "clothes" of 1.2 is greater than the total weight of "doctor" of 0.4. Therefore, "clothes" with the largest total weight is selected as the final speech data, thereby effectively solving the recognition problem caused by similar pronunciations and achieving accurate matching of speech and lip-synced information.

[0154] The embodiments of the present application achieve the following technical effects through the above specific implementation methods:

[0155] The embodiments of the present application provide a method and system for extracting voice data using artificial intelligence. By synchronously collecting video and audio data sequences and analyzing them using multiple recognizers with different training data, a diverse voice content collection is constructed, which effectively solves the problem that traditional voice extraction relies on a single data source and obtains information in a one-sided manner. Secondly, by collecting various information such as speech speed, volume, proportion of lip images, and user's historical vocabulary usage, a dynamic weight configuration mechanism is constructed. By calculating and fusing multiple sets of weight parameters, the accuracy of voice content recognition is effectively improved, and misjudgment caused by similar pronunciation or environmental interference is avoided. Finally, by analyzing the frequency of use of uncommon words in the user's historical voice data, personalized weight adaptation is achieved, and training optimization is performed on a large number of samples to ensure that the matching of voice and lip information can be completed stably and accurately in complex application scenarios such as different speaking styles, different accents, and video picture features.

[0156] The method and system provided in the embodiments of the present application solve the problems of existing voice data extraction technology, such as single reliance, poor voice data extraction quality, and lack of personalized adaptation. They effectively avoid recognition errors caused by environmental factors or individual differences, improve the accuracy and reliability of voice data extraction, provide support for the application of voice recognition technology in multiple scenarios, and promote the development of voice processing technology towards intelligence and precision.

[0157] Example 2, as shown in the attached Figure 2 As shown, based on the inventive concept of a voice data extraction method using artificial intelligence provided in Example 1, the present application also provides a voice data extraction system using artificial intelligence, specifically comprising:

[0158] The data acquisition module 01 is used to monitor and collect video data and audio data generated by video communication, and obtain video data sequences and audio data sequences;

[0159] The audio feature analysis and weight configuration module 02 is used to perform speech rate recognition and volume recognition based on the audio data sequence to obtain speech rate feature parameters and volume feature parameters, and respectively configure lip shape weights and audio weights to obtain a first lip shape weight, a first audio weight, a second lip shape weight, and a second audio weight;

[0160] The video and vocabulary feature analysis and weight configuration module 03 is used to analyze the lip shape image ratio of the video data sequence to obtain a third lip shape weight and a third audio weight, and to extract and record the historical voice data of the target user, perform vocabulary analysis, and obtain a fourth lip shape weight and a fourth audio weight.

[0161] The weight fusion and speech recognition module 04 is used to process the first lip shape weight, the first audio weight, the second lip shape weight, the second audio weight, the third lip shape weight, the third audio weight, the fourth lip shape weight, and the fourth audio weight to obtain a fused lip shape weight and a fused audio weight, and perform speech content recognition and extraction on the video data sequence and the audio data sequence to obtain speech data.

[0162] In one embodiment, the audio feature analysis and weight configuration module 02 is further configured to:

[0163] Inputting the audio data sequence into a speech rate recognizer to identify and obtain speech rate feature parameters;

[0164] Extracting the average volume in the audio data sequence as a volume feature parameter;

[0165] According to the speech rate characteristic parameters, the lip shape weight and the audio weight are configured to obtain a first lip shape weight and a first audio weight. According to the volume characteristic parameters, the lip shape weight and the audio weight are configured to obtain a second lip shape weight and a second audio weight.

[0166] In one embodiment, the video and vocabulary feature analysis and weight configuration module 03 is further configured to:

[0167] Inputting each video data in the video data sequence into a lip shape ratio recognizer to identify and obtain a plurality of lip shape image ratio coefficients, wherein the lip shape ratio recognizer is constructed using a convolutional neural network and is trained using a sample video data set and a set of labeled sample lip shape image ratio coefficients;

[0168] Get the average lip shape image ratio;

[0169] Calculating a ratio of a mean of the plurality of lip shape image proportion coefficients to the average lip shape image proportion coefficient, and performing a correction calculation on the preset lip shape weight to obtain a third lip shape weight;

[0170] A third audio weight is calculated based on the third lip shape weight.

[0171] Collect historical extracted voice data sets from users who have extracted voice data in historical time;

[0172] Counting the number of occurrences of uncommon words on the historically extracted speech data set to obtain the number of occurrences of uncommon words;

[0173] Collect the number of occurrences of uncommon words from multiple sample users and calculate the mean to obtain the average number of occurrences of uncommon words;

[0174] According to the ratio of the average number of occurrences of uncommon words to the number of occurrences of uncommon words, the preset lip shape weight is modified and calculated to obtain a fourth lip shape weight;

[0175] A fourth audio weight is calculated based on the fourth lip shape weight.

[0176] In one embodiment, the weight fusion and speech recognition module 04 is further configured to:

[0177] According to the first lip shape weight, the first audio weight, the second lip shape weight, the second audio weight, the third lip shape weight, the third audio weight, the fourth lip shape weight, and the fourth audio weight, a fusion calculation is performed to obtain a fusion lip shape weight and a fusion audio weight;

[0178] Performing lip-sync speech content recognition and extraction and audio speech content recognition and extraction according to the video data sequence and the audio data sequence respectively to obtain a lip-sync speech content set and an audio speech content set;

[0179] According to the fused lip shape weight and the fused audio weight, weighted fusion processing is performed on the lip shape speech content set and the audio speech content set to obtain speech data, wherein the weighted fusion processing includes weighting each speech content and screening the speech content with the largest total weight sum as the speech data.

[0180] It should be noted that the order in which the embodiments of the present application are presented is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. Furthermore, the foregoing descriptions of specific embodiments of this specification are provided. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential sequence shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0181] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.

[0182] This specification and drawings are merely illustrative of the present application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Obviously, those skilled in the art may make various modifications and variations to this application without departing from the scope of this application. Thus, this application is intended to include such modifications and variations as fall within the scope of this application and its equivalents.

Claims

1. A method for extracting speech data using artificial intelligence, characterized in that: The method comprises: Monitor and collect video data and audio data generated by video communication to obtain a video data sequence and an audio data sequence; Performing speech rate recognition and volume recognition according to the audio data sequence to obtain speech rate feature parameters and volume feature parameters, and respectively configuring lip shape weights and audio weights to obtain a first lip shape weight, a first audio weight, a second lip shape weight, and a second audio weight; Performing a lip shape image ratio analysis on the video data sequence to obtain a third lip shape weight and a third audio weight, and performing a vocabulary analysis based on historical speech data of the target user to obtain a fourth lip shape weight and a fourth audio weight; The method includes: processing the first lip shape weight, the first audio weight, the second lip shape weight, the second audio weight, the third lip shape weight, the third audio weight, the fourth lip shape weight, and the fourth audio weight to obtain a fused lip shape weight and a fused audio weight; and performing voice content recognition and extraction on the video data sequence and the audio data sequence to obtain voice data, including: According to the first lip shape weight, the first audio weight, the second lip shape weight, the second audio weight, the third lip shape weight, the third audio weight, the fourth lip shape weight, and the fourth audio weight, a fusion calculation is performed to obtain a fusion lip shape weight and a fusion audio weight; Performing lip-sync speech content recognition and extraction and audio speech content recognition and extraction according to the video data sequence and the audio data sequence respectively to obtain a lip-sync speech content set and an audio speech content set; According to the fused lip shape weight and the fused audio weight, weighted fusion processing is performed on the lip shape speech content set and the audio speech content set to obtain speech data, wherein the weighted fusion processing includes weighting each speech content and screening the speech content with the largest total weight sum as the speech data.

2. The method for extracting speech data using artificial intelligence according to claim 1, wherein: According to the audio data sequence, speech rate recognition and volume recognition are performed to obtain speech rate feature parameters and volume feature parameters, and lip shape weights and audio weights are configured respectively to obtain a first lip shape weight, a first audio weight, a second lip shape weight, and a second audio weight, including: Inputting the audio data sequence into a speech rate recognizer to identify and obtain speech rate feature parameters; Extracting the average volume in the audio data sequence as a volume feature parameter; According to the speech rate characteristic parameters, the lip shape weight and the audio weight are configured to obtain a first lip shape weight and a first audio weight. According to the volume characteristic parameters, the lip shape weight and the audio weight are configured to obtain a second lip shape weight and a second audio weight.

3. The method for extracting speech data using artificial intelligence according to claim 2, wherein: The steps of constructing the speech rate identifier include: Using machine learning, a network structure of a speech rate identifier is constructed, wherein the input data of the speech rate identifier is an audio data sequence and the output data is a speech rate feature parameter; According to the speech data recognition records in the historical time, a set of sample audio data sequences is collected, and the speech speed of different sample audio data sequences is marked to obtain a set of sample speech speed feature parameters; The speech rate identifier is supervised trained and tested using the sample audio data sequence set and the sample speech rate feature parameter set until convergence.

4. The method for extracting speech data using artificial intelligence according to claim 2, wherein: According to the speech rate characteristic parameter, the lip shape weight and the audio weight are configured to obtain a first lip shape weight and a first audio weight; according to the volume characteristic parameter, the lip shape weight and the audio weight are configured to obtain a second lip shape weight and a second audio weight, including: According to the speech data recognition record in the historical time, the average speech speed characteristic parameter and the average volume characteristic parameter are obtained; Calculating a ratio of the average speaking rate characteristic parameter to the speaking rate characteristic parameter, performing correction calculation on a preset audio weight to obtain a first audio weight, and calculating a first lip shape weight based on the first audio weight; The ratio of the volume characteristic parameter to the average volume characteristic parameter is calculated, and the preset audio weight is corrected and calculated to obtain a second audio weight. The second lip shape weight is calculated based on the second audio weight.

5. The method for extracting speech data using artificial intelligence according to claim 1, wherein: Performing a lip shape image ratio analysis on the video data sequence to obtain a third lip shape weight and a third audio weight includes: Inputting each video data in the video data sequence into a lip shape ratio recognizer to identify and obtain a plurality of lip shape image ratio coefficients, wherein the lip shape ratio recognizer is constructed using a convolutional neural network and is trained using a sample video data set and a set of labeled sample lip shape image ratio coefficients; Get the average lip shape image ratio; Calculating a ratio of a mean of the plurality of lip shape image proportion coefficients to the average lip shape image proportion coefficient, and performing a correction calculation on the preset lip shape weight to obtain a third lip shape weight; A third audio weight is calculated based on the third lip shape weight.

6. The method for extracting speech data using artificial intelligence according to claim 1, wherein: Extract records based on the target user's historical voice data, perform vocabulary analysis, and process to obtain the fourth lip shape weight and the fourth audio weight, including: Collect historical extracted voice data sets from users who have extracted voice data in historical time; Counting the number of occurrences of uncommon words on the historically extracted speech data set to obtain the number of occurrences of uncommon words; Collect the number of occurrences of uncommon words from multiple sample users and calculate the mean to obtain the average number of occurrences of uncommon words; According to the ratio of the average number of occurrences of uncommon words to the number of occurrences of uncommon words, the preset lip shape weight is modified and calculated to obtain a fourth lip shape weight; A fourth audio weight is calculated based on the fourth lip shape weight.

7. The method for extracting speech data using artificial intelligence according to claim 1, wherein: Performing lip-sync speech content recognition and extraction and audio speech content recognition and extraction according to the video data sequence and the audio data sequence respectively to obtain a lip-sync speech content set and an audio speech content set, including: Inputting the video data sequence into a plurality of lip-synced speech content recognizers respectively, and recognizing and outputting a lip-synced speech content set, wherein the input data of each lip-synced speech content recognizer is the video data sequence, the output data is the lip-synced speech content, and the training data of each lip-synced speech content recognizer is different; The audio data sequence is input into a plurality of audio speech content recognizers respectively, and the recognition output is used to obtain an audio speech content set, wherein the input data of each audio speech content recognizer is the audio data sequence, the output data is the audio speech content, and the training data of each audio speech content recognizer is different.

8. A voice data extraction system using artificial intelligence, characterized in that: The system is used to perform the method according to any one of claims 1 to 7, and the system includes: The data acquisition module is used to monitor and collect video data and audio data generated by video communication, and obtain video data sequences and audio data sequences; An audio feature analysis and weight configuration module is used to perform speech rate recognition and volume recognition based on the audio data sequence to obtain speech rate feature parameters and volume feature parameters, and to configure lip shape weights and audio weights respectively to obtain a first lip shape weight, a first audio weight, a second lip shape weight, and a second audio weight; a video and vocabulary feature analysis and weight configuration module, configured to perform a lip shape image ratio analysis on the video data sequence, process the analysis to obtain a third lip shape weight and a third audio weight, and extract and record the target user's historical voice data, perform a vocabulary analysis, and process the analysis to obtain a fourth lip shape weight and a fourth audio weight; The weight fusion and speech recognition module is used to process the first lip shape weight, the first audio weight, the second lip shape weight, the second audio weight, the third lip shape weight, the third audio weight, the fourth lip shape weight, and the fourth audio weight to obtain a fused lip shape weight and a fused audio weight, and perform speech content recognition and extraction on the video data sequence and the audio data sequence to obtain speech data, including: According to the first lip shape weight, the first audio weight, the second lip shape weight, the second audio weight, the third lip shape weight, the third audio weight, the fourth lip shape weight, and the fourth audio weight, a fusion calculation is performed to obtain a fusion lip shape weight and a fusion audio weight; Performing lip-sync speech content recognition and extraction and audio speech content recognition and extraction according to the video data sequence and the audio data sequence respectively to obtain a lip-sync speech content set and an audio speech content set; According to the fused lip shape weight and the fused audio weight, weighted fusion processing is performed on the lip shape speech content set and the audio speech content set to obtain speech data, wherein the weighted fusion processing includes weighting each speech content and screening the speech content with the largest total weight sum as the speech data.

Citation Information

Patent Citations

  • Vehicle-mounted terminal equipment, vehicle-mounted interaction system and interaction method

    CN109941231A

  • Cross-modal speech recognition method

    CN115938367A