Model training method and device, lip language recognition method and device, electronic equipment and medium
By constructing a multimodal corpus and training a lip recognition model, the communication difficulties of patients with vocal loss after surgery is solved, the efficiency and accuracy of lip recognition are improved, and the quality of postoperative life of patients is promoted.
Patent Information
- Application Number
- CN202510113675.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-24
AI Technical Summary
Patients with middle and advanced laryngeal cancer and hypopharyngeal cancer lose voice after surgery, resulting in difficulty in communication and affecting postoperative recovery and quality of life.
By building a multimodal corpus, obtaining and labeling the audio and video data of patients before and after the operation, training the lip recognition model, and achieving automatic lip recognition of patients' audio and video.
It improves the efficiency and accuracy of lip recognition, solves the problem of difficulty in communication in a state of loss of voice, and improves the convenience of life after surgery.
Smart Images

Figure CN119993156A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of visual speech recognition, and in particular to a model training method, a lip reading recognition method, a device, an electronic device and a medium. Background Art
[0002] Every year, a large number of patients with advanced laryngeal cancer and hypopharyngeal cancer undergo partial laryngectomy, complete laryngectomy or tracheotomy. These patients are often able to speak and communicate normally before surgery, but may suddenly lose their voice after surgery, which may bring the following negative effects:
[0003] 1) Failure to communicate normally with the accompanying personnel: This causes the patient to be unable to express his or her wishes in a timely manner. Whether it is physiological needs or physical discomfort, failure to deal with them in a timely manner may affect the patient's postoperative recovery;
[0004] 2) Unable to communicate normally with medical staff: The decline in medical work efficiency is one problem; in addition, if the patient's wishes are not conveyed in a timely and accurate manner, it may cause misunderstandings, affect the doctor-patient relationship and even medical decision-making;
[0005] 3) Unable to communicate normally with family members: Patients will feel the inconvenience of postoperative life to the greatest extent, which is difficult to accept psychologically, and is not conducive to mental health and postoperative recovery.
[0006] Therefore, the above problems need to be solved urgently. Summary of the invention
[0007] The main purpose of the embodiments of the present application is to propose a model training method, a lip reading recognition method, a device, an electronic device and a medium, which can automatically recognize the lip reading of the patient's audio and video, improve the efficiency and accuracy of lip reading recognition, and solve the problem of communication difficulties of patients in a voiceless state.
[0008] On the one hand, an embodiment of the present application proposes a model training method, the method comprising the following steps:
[0009] Acquire a sample data set from a multimodal corpus; the sample data set includes a plurality of patient audio and video data with text labels;
[0010] Dividing the sample data set to obtain a training set and a test set;
[0011] Constructing a lip reading recognition model, and training the lip reading recognition model using the training set;
[0012] The test set is used to perform model evaluation on the trained lip reading recognition model to determine whether to continue training the lip reading recognition model.
[0013] In some embodiments, obtaining a sample data set from a multimodal corpus specifically includes:
[0014] constructing the multimodal corpus;
[0015] Collecting pre- and post-operative audio and video data corresponding to a plurality of medical surgery patients, and performing data pre-processing on the pre- and post-operative audio and video data; the pre- and post-operative audio and video data include pre-operative audio and video data and post-operative audio and video data; the medical surgery patients include patients undergoing partial laryngectomy, complete laryngectomy or tracheotomy;
[0016] Perform text recognition and annotation on the pre-processed audio and video data before and after the operation to obtain text labels corresponding to the audio and video data before and after the operation, and store the audio and video data before and after the operation including the text labels in the multimodal corpus.
[0017] In some embodiments, the performing of text recognition and labeling on the pre-processed audio and video data before and after the surgery to obtain text labels corresponding to the audio and video data before and after the surgery specifically includes:
[0018] Using a speech recognition model to perform text recognition on the audio and video data before and after the operation, and determining a plurality of text annotation contents and a video timestamp and a dialect type corresponding to each of the text annotation contents;
[0019] Using an emotional computing model to perform emotional recognition on each of the text annotation contents, and determining the emotional feature labels corresponding to each of the text annotation contents;
[0020] Performing medical entity annotation on each of the text annotation contents using a medical entity annotation model to determine the medical entity annotation corresponding to each of the text annotation contents;
[0021] Dividing the audio and video data before and after the operation into video frames to obtain a plurality of video frames and a time stamp corresponding to each of the video frames;
[0022] Determine a text description label corresponding to each of the text annotation contents according to the dialect type, the emotion feature label, and the medical entity label corresponding to each of the text annotation contents;
[0023] According to the video timestamp corresponding to each of the text annotation contents and the timestamp corresponding to each of the video frames, pairing the text annotation contents with each of the video frames to generate a video frame text pairing label corresponding to each of the text annotation contents;
[0024] The text label is determined according to the video frame text pairing label corresponding to each of the text annotation contents.
[0025] In some embodiments, pairing the text annotation content with each video frame according to the video timestamp corresponding to each text annotation content and the timestamp corresponding to each video frame to generate a video frame text pairing label corresponding to each text annotation content specifically includes:
[0026] Determining a plurality of target video frames corresponding to each of the text annotation contents from the plurality of video frames according to the video timestamp corresponding to each of the text annotation contents and the timestamp corresponding to each of the video frames;
[0027] The text description tags corresponding to each of the text annotation contents are associated and bound to each of the target video frames to generate the video frame text pairing tags corresponding to each of the text annotation contents.
[0028] In some embodiments, the using the test set to perform model evaluation on the trained lip reading recognition model to determine whether to continue training the lip reading recognition model specifically includes:
[0029] Obtaining multiple model evaluation indicators and model evaluation indicator thresholds corresponding to each of the model evaluation indicators;
[0030] Inputting the test set into the lip reading recognition model, and using the lip reading recognition model to output corresponding lip reading recognition results;
[0031] According to each of the model evaluation indicators and the lip reading recognition result, the lip reading recognition model is evaluated to determine the current evaluation value of the model corresponding to each of the model evaluation indicators;
[0032] Determine a model evaluation result according to the model evaluation indicator threshold corresponding to each model evaluation indicator and the current evaluation value of the model, and determine whether to continue training the lip reading recognition model according to the model evaluation result;
[0033] When the model evaluation result is that the current evaluation values of the model corresponding to each of the model evaluation indicators exceed the model evaluation indicator threshold, the training of the lip reading recognition model is stopped; otherwise, the training of the lip reading recognition model continues.
[0034] On the other hand, an embodiment of the present application provides a lip reading recognition method, the method comprising the following steps:
[0035] Acquire the audio and video data to be identified corresponding to the medical operation patient, and perform data preprocessing on the audio and video data to be identified;
[0036] The lip reading recognition model is used to perform lip reading recognition on the pre-processed audio and video data to be recognized, and the lip reading recognition result is output; the lip reading recognition model is trained by the model training method described above.
[0037] In some embodiments, the lip reading recognition model is used to perform lip reading recognition on the pre-processed audio and video data to be recognized, and the lip reading recognition result is output, specifically including:
[0038] Using the lip reading recognition model, the preprocessed audio and video data to be recognized are divided into video frames and lip reading recognition is performed, and a plurality of text recognition information corresponding to the video data to be recognized and a plurality of video frames to be paired corresponding to each of the text recognition information are output;
[0039] The lip reading recognition model is used to pair each of the text recognition information with the corresponding multiple video frames to be paired, and the video frame text pairing results corresponding to each of the text recognition information are determined. According to each of the video frame text pairing results, each of the text recognition information is inserted into the audio and video data to be recognized, and the corresponding lip reading recognition audio and video data is output.
[0040] On the other hand, an embodiment of the present application provides a lip reading recognition device, the device comprising:
[0041] The first module is used to obtain the audio and video data to be identified corresponding to the medical operation patient, and perform data preprocessing on the audio and video data to be identified;
[0042] The second module is used to perform lip reading recognition on the pre-processed audio and video data to be recognized by using a lip reading recognition model, and output the lip reading recognition result; the lip reading recognition model is trained by the model training method described above.
[0043] On the other hand, an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the model training method or the lip reading recognition method described above when executing the computer program.
[0044] On the other hand, an embodiment of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the model training method or the lip reading recognition method described above.
[0045] The embodiments of the present application include at least the following beneficial effects: the present application provides a model training method, lip reading recognition method, device, electronic device and medium, which obtains the audio and video data to be recognized corresponding to the medical surgery patient, performs data preprocessing on the audio and video data to be recognized, uses the trained lip reading recognition model to perform lip reading recognition on the preprocessed audio and video data to be recognized, and outputs the lip reading recognition result. The present application can automatically recognize the lip reading of the patient's audio and video, improve the efficiency and accuracy of lip reading recognition, solve the problem of communication difficulties of patients in the state of aphonia, and improve the convenience of patients' life after surgery. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a flow chart of a model training method provided in an embodiment of the present application;
[0047] Figure 2 is a flow chart of a lip reading recognition model provided in an embodiment of the present application;
[0048] Figure 3 is a structural schematic diagram of a lip reading recognition device provided in an embodiment of the present application;
[0049] Figure 4 It is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the attached claims.
[0051] It is understood that the terms "first", "second", etc. used in this application can be used to describe various concepts in this article, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another concept. For example, without departing from the scope of the embodiment of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein can be interpreted as "at the time of" or "when" or "in response to determination".
[0052] The terms "at least one", "multiple", "each", "any", etc. used in this application, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.
[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0054] It should be noted that in each specific implementation of the present application, when it comes to the need to perform relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0055] At present, a large number of patients with advanced laryngeal cancer and hypopharyngeal cancer undergo partial laryngectomy, complete laryngectomy or tracheotomy every year. These patients can often speak and communicate normally before the operation, but may suddenly lose their voice after the operation, which brings the following negative effects:
[0056] 1) Failure to communicate normally with the accompanying personnel: This causes the patient to be unable to express his or her wishes in a timely manner. Whether it is physiological needs or physical discomfort, failure to deal with them in a timely manner may affect the patient's postoperative recovery;
[0057] 2) Unable to communicate normally with medical staff: The decline in medical work efficiency is one problem; in addition, if the patient's wishes are not conveyed in a timely and accurate manner, it may cause misunderstandings, affect the doctor-patient relationship and even medical decision-making;
[0058] 3) Unable to communicate normally with family members: Patients will feel the inconvenience of postoperative life to the greatest extent, which is difficult to accept psychologically, and is not conducive to mental health and postoperative recovery.
[0059] "Lip recognition technology" is also known as "lip reading" (Lip Reading, LR), which belongs to the category of visual speech recognition technology (Visual Speech Recognition, VSR). It is a technology that uses visual information to infer and understand the speaker's language or pronunciation. As an emerging topic in the intersection of computer vision and natural language processing, with the development of deep learning in recent years, lip reading has broad development prospects in the fields of speech recognition, human-computer interaction and public safety; in the field of healthcare, lip reading systems have been used for auxiliary communication of hearing-impaired people.
[0060] At present, the application of lip reading technology in medical scenarios still faces many challenges, which are specifically reflected in the following aspects:
[0061] 1) Lack of special corpora adapted to dialect environments. Most existing lip-reading technologies rely on standardized open-source corpora, which are mostly concentrated in languages such as standard Mandarin or English. As a result, existing models lack adaptability to dialect features and have poor recognition effects.
[0062] 2) There is a lack of corpora designed for special scenarios. For example, hospital scenarios often involve a large number of medical terms, and existing models lack such personalized training data, so the recognition effect is limited;
[0063] 3) Multimodal fusion problem. Lip reading recognition not only relies on visual information (lip shape, facial expressions, etc.), but also requires the support of audio information such as voice and intonation. Most existing lip reading models rely only on a single modality (such as video or voice), and fail to effectively combine the multimodal features of video and audio data, affecting the accuracy and robustness of the model.
[0064] The above problems seriously restrict the application of lip reading technology in medical scenarios.
[0065] Based on this, the embodiments of the present application propose a model training method, a lip reading recognition method, a device, an electronic device and a medium, which combine the dialect characteristics of different regions (such as Cantonese, Hakka, Minnan, etc.) with otolaryngology medical scenarios to construct a multimodal corpus, aiming to provide high-quality and diversified data support for lip reading recognition technology used in postoperative aphonia patients. The corpus is used to train the lip reading recognition model to complete the lip reading work, improve the efficiency and accuracy of lip reading recognition, and solve the problem of patients' communication difficulties in the aphonia state.
[0066] Reference Figure 1 , Figure 1 This is an optional flowchart of a model training method provided in an embodiment of the present application. The method may include but is not limited to steps S101 to S104:
[0067] Step S101, obtaining a sample data set from a multimodal corpus; the sample data set includes patient audio and video data with multiple text labels;
[0068] Step S102, dividing the sample data set to obtain a training set and a test set;
[0069] Step S103, constructing a lip reading recognition model, and training the lip reading recognition model using the training set;
[0070] Step S104, using the test set to perform model evaluation on the trained lip reading recognition model to determine whether to continue training the lip reading recognition model.
[0071] In some embodiments, step S101 may include but is not limited to steps S201 to S203:
[0072] Step S201, constructing a multimodal corpus;
[0073] Step S202, collecting pre- and post-operative audio and video data corresponding to a plurality of medical surgery patients, and performing data pre-processing on the pre- and post-operative audio and video data; the pre- and post-operative audio and video data include pre-operative audio and video data and post-operative audio and video data; the medical surgery patients include patients undergoing partial laryngectomy, complete laryngectomy or tracheotomy;
[0074] Step S203 , performing text recognition and annotation on the pre-processed audio and video data before and after the operation, obtaining text labels corresponding to the audio and video data before and after the operation, and storing the audio and video data before and after the operation including the text labels in the multimodal corpus.
[0075] In some embodiments, the above-mentioned pre- and post-operative audio and video data include audio and video data recorded by the patient before surgery (such as the above-mentioned pre-operative audio and video data), as well as audio and video data or video (lip shape) data of each recovery stage after surgery (such as the above-mentioned post-operative audio and video data). For example, by using the pre-operative audio and video data to perform personalized small sample training on the lip reading recognition model, the lip reading recognition model can be quickly adjusted in the subsequent data training of the input post-operative audio and video data, thereby greatly improving the accuracy and robustness of lip reading recognition.
[0076] In some embodiments, a multimodal corpus is constructed, which combines the dialect characteristics of different regions (such as Cantonese, Hakka, Minnan, etc.) and otolaryngology medical scenarios. After collecting pre- and post-operative audio and video data, it is also stored in the multimodal corpus. Then, the pre- and post-operative audio and video data that have not been text-annotated are obtained from the multimodal corpus and text-annotated.
[0077] Specifically, the audio and video data before and after the operation may include but are not limited to:
[0078] 1) Preoperative audio and video data and postoperative audio and video data: During the patient's hospitalization, audio and video data are collected before and after surgery, and the conversations between the patient and medical staff, family members and caregivers are collected; the recording must be agreed to by both parties in advance, and the recording process must comply with ethical approval requirements. Note that audio and video data with unclear lip shape or pronunciation cannot be excluded, because postoperative aphonia patients may have certain deformation and weakening of facial movements due to changes in laryngeal anatomy and wound swelling. In order to learn and adapt to lip reading recognition in such situations, this part of the data also needs to be collected for training;
[0079] 2) Postoperative outpatient follow-up data: After the patient is discharged from the hospital, outpatient follow-up data can be collected to obtain feedback on the patient's long-term recovery after surgery;
[0080] 3) Online health consultation data: For Internet medical care, patients can record videos and upload them;
[0081] 4) Postoperative simulated dialogue data: Design common problems that patients may encounter after surgery, and organize medical experts, doctors and volunteers to simulate different postoperative dialogue scenarios, such as postoperative consultation, health guidance, and discussion of follow-up treatment plans.
[0082] At the same time, the multimodal corpus also stores public resources and open data sets, including some public data sets related to medical care and Cantonese, which can serve as a supplement to the preliminary corpus and guide subsequent speech recognition models to identify dialect types in audio and video data before and after surgery.
[0083] In some embodiments, optionally, since whether a medical surgery patient can make a sound for a period of time after surgery varies from person to person, which is related to the recovery of the medical surgery patient, if the tracheotomy tube of the medical surgery patient is blocked smoothly, the patient can start to try to make a sound. Based on this, the collected postoperative audio-visual data includes:
[0084] 1) Video data before postoperative tube blocking: Postoperative patients cannot speak before tube blocking, so video data of patients in this recovery period is collected;
[0085] 2) Audio and video data of successful postoperative tube occlusion: After the patient's tube is successfully blocked after surgery, he or she can try to speak. Therefore, the patient's audio and video data during the recovery period is collected. The patient's audio and video data includes video data and audio data.
[0086] The above-mentioned lip reading recognition model can skillfully handle the problem of "modality missing" in the data set. Therefore, the missing audio data in the above-mentioned video data before postoperative tube occlusion will not interfere with the construction of the multimodal corpus and model training.
[0087] In some embodiments, step S202 may include, but is not limited to, steps S301 to S307:
[0088] Step S301, using a speech recognition model to perform text recognition on the audio and video data before and after the operation, and determining a plurality of text annotation contents and the video timestamp and dialect type corresponding to each text annotation content;
[0089] Step S302, using the sentiment computing model to perform sentiment recognition on each text annotation content, and determine the sentiment feature label corresponding to each text annotation content;
[0090] Step S303, using the medical entity annotation model to perform medical entity annotation on each text annotation content, and determine the medical entity annotation corresponding to each text annotation content;
[0091] Step S304, dividing the video data before and after the operation into video frames to obtain multiple video frames and timestamps corresponding to each video frame;
[0092] Step S305, determining a text description label corresponding to each text annotation content according to the dialect type, emotional feature label and medical entity label corresponding to each text annotation content;
[0093] Step S306, pairing the text annotation content with each video frame according to the video timestamp corresponding to each text annotation content and the timestamp corresponding to each video frame, and generating a video frame text pairing label corresponding to each text annotation content;
[0094] Step S307: determining text labels according to the video frame text pairing labels corresponding to each text annotation content.
[0095] In some embodiments, data preprocessing is performed on the audio and video data before and after the operation, which may include but is not limited to preliminary preprocessing and data enhancement. The preliminary preprocessing may include but is not limited to video editing, audio noise reduction, video frame extraction, and resampling to improve data quality and consistency. A pre-trained model based on deep learning is introduced to enhance and repair the image quality of the audio and video data before and after the operation, and to optimize the clips with blurred image quality or insufficient light.
[0096] Data enhancement is performed on the pre- and post-operative audio and video data after preliminary preprocessing, including rotation, zooming in, and zooming out, to improve the comprehensiveness of the data and enhance the robustness of the subsequent model.
[0097] In some embodiments, a speech recognition model is used to transcribe the audio data of the data-enhanced pre- and post-operative audio and video data (such as the above-mentioned speech recognition). Specifically, the speech recognition model (such as Whisper or Wav2Vec2.0) is pre-trained through a multimodal corpus, and the audio signal is transcribed into text data through the pre-trained speech recognition model, and multiple text annotation contents and the video timestamps corresponding to each text annotation content are determined, and each text annotation content is annotated with the dialect type and accent characteristics, and the dialect type corresponding to each text annotation content is determined, and manual proofreading is performed to improve accuracy. The unclear speech that may exist in the audio of postoperative patients is processed by a specific audio enhancement system, such as a speech enhancement system based on a generative adversarial network.
[0098] Based on multiple preset dialect types (such as Cantonese, Hakka and Minnan, etc.), dialect recognition is performed on the audio and video data before and after the operation, and the dialect type corresponding to the audio and video data before and after the operation is determined and labeled. This can provide a new language environment data for lip reading technology and solve the problem of insufficient adaptability of traditional lip reading technology to dialects.
[0099] In some embodiments, a medical entity annotation model (such as MedGPT or BioBERT) is used to perform medical entity annotation on the text annotation contents obtained after transcribing the audio data, including disease names (such as laryngeal cancer, hypopharyngeal cancer, etc.), drug names (such as anti-inflammatory drugs, analgesics, chemotherapy drugs, targeted drugs, etc.), medical operations (such as surgery, laryngoscopy, pathology, follow-up visits, etc.), symptom descriptions (such as pain, bleeding, nausea, etc.) and other information, to determine the medical entity information contained in each text annotation content and annotate it, to generate corresponding medical entity annotations, and at the same time combine manual proofreading to improve the effectiveness of medical entity annotations, so that the relevant medical terms can be accurately captured during the subsequent lip reading recognition model training.
[0100] In some embodiments, a sentiment computing model (such as a Transformer-based sentiment analysis framework) adapted to medical scenarios is trained and fine-tuned in a targeted manner. Based on the sentiment computing model, sentiment recognition is performed on each text annotation content, including sentiment (anxiety, concern, etc.) and tone (affirmation, question, request, etc.) annotation, so as to capture the emotional characteristics of patients in different contexts and generate sentiment feature labels corresponding to each text annotation content.
[0101] In some embodiments, step S306 may include but is not limited to steps S401 to S402:
[0102] Step S401, determining a number of target video frames corresponding to each text annotation content from a plurality of video frames according to the video timestamp corresponding to each text annotation content and the timestamp corresponding to each video frame;
[0103] Step S402 : Associating and binding the text description tags corresponding to each text annotation content with each target video frame to generate video frame text pairing tags corresponding to each text annotation content.
[0104] In some embodiments, the video frames are aligned with the corresponding text annotation content according to the timestamp to obtain the corresponding video-audio-text label pairs, that is, the above-mentioned video frame text pairing labels.
[0105] Exemplarily, assuming that the video timestamps of the text annotation content are T1 to T2, according to the timestamps of each video frame, it is determined that video frame X to video frame Y are multiple target video frames corresponding to the text annotation content, the text description tags of the text annotation content are associated and bound to video frames X to video frame Y, and a video frame text pairing tag is generated. According to the video frame text pairing tag, the text annotation content containing the text description tag is inserted into the video segment corresponding to video frame X to video frame Y, wherein the timestamp of video frame X is T1, the timestamp of video frame Y is T2, and multiple video frames are included between video frame X and video frame Y.
[0106] In some embodiments, the video frame text pairing tags corresponding to each text annotation content constitute the text tags of the pre- and post-operative audio and video data.
[0107] In some embodiments, the sample data set is divided into a training set, a validation set and a test set. The training set is used to train the lip reading recognition model, the validation set is used to verify the lip reading recognition model, and the test set is used to evaluate the lip reading recognition model. The evaluated lip reading recognition model is used for subsequent lip reading recognition of the audio and video data to be recognized.
[0108] The lip reading recognition model may include, but is not limited to, a deep neural network (DNN) model, an End2End model, an Encoder-Decoder model, and a 3DCNN audio-visual synchronization auxiliary model.
[0109] In some embodiments, optionally, the pre- and post-operative audio and video data containing text labels are converted into a corresponding preset format and stored in the multimodal corpus, and a sample data set is obtained from the multimodal corpus, the sample data set including the patient audio and video data with multiple text labels.
[0110] In some embodiments, step S104 may include but is not limited to steps S501 to S505:
[0111] Step S501, obtaining multiple model evaluation indicators and model evaluation indicator thresholds corresponding to each model evaluation indicator;
[0112] Step S502, input the test set into the lip reading recognition model, and use the lip reading recognition model to output the corresponding lip reading recognition result;
[0113] Step S503, performing model evaluation on the lip reading recognition model according to each model evaluation index and the lip reading recognition result, and determining the current model evaluation value corresponding to each model evaluation index;
[0114] Step S504, determining the model evaluation result according to the model evaluation indicator threshold corresponding to each model evaluation indicator and the current evaluation value of the model, and determining whether to continue training the lip reading recognition model according to the model evaluation result.
[0115] Step S505, when the model evaluation result is that the current model evaluation value corresponding to each model evaluation index exceeds the model evaluation index threshold, stop training the lip reading recognition model, otherwise, continue training the lip reading recognition model.
[0116] In some embodiments, the model evaluation indicator may be accuracy, recall rate, F1 value, AUC value, etc., and each model evaluation indicator is provided with a corresponding model evaluation indicator threshold.
[0117] Reference Figure 2 , Figure 2 This is an optional flow chart of a lip reading recognition method provided in an embodiment of the present application. The method may include but is not limited to steps S1 to S2:
[0118] Step S1, obtaining the audio and video data to be identified corresponding to the medical operation patient, and performing data preprocessing on the audio and video data to be identified;
[0119] Step S2, using the lip reading recognition model to perform lip reading recognition on the pre-processed audio and video data to be recognized, and outputting the lip reading recognition result; the lip reading recognition model is trained by the above-mentioned model training method.
[0120] In some embodiments, data preprocessing is performed on the audio and video data to be recognized, including video editing, audio noise reduction, video frame extraction, resampling and data enhancement, so as to improve the quality of the audio and video data to be recognized and improve the accuracy and efficiency of subsequent lip reading recognition of the audio and video data to be recognized.
[0121] In some embodiments, step S2 may include but is not limited to steps S21 to S22:
[0122] Step S21, using the lip reading recognition model to perform video frame division and lip reading recognition on the pre-processed video data to be recognized, and output multiple text recognition information corresponding to the video data to be recognized and multiple video frames to be paired corresponding to each text recognition information;
[0123] In step S22, the lip reading recognition model is used to pair each text recognition information with a corresponding plurality of video frames to be paired, and the video frame text pairing result corresponding to each text recognition information is determined. According to each video frame text pairing result, each text recognition information is inserted into the audio and video data to be recognized, and the corresponding lip reading recognition audio and video data is output.
[0124] In some embodiments, a lip reading recognition model is used to pair each text recognition information with a corresponding plurality of video frames to be paired, a plurality of target paired video frames are determined, the text recognition information is inserted into the video segments corresponding to the plurality of target paired video frames, and the corresponding lip reading recognition audio and video data is output to achieve lip reading recognition.
[0125] Reference Figure 3 , Figure 3 : is an optional structural diagram of a lip reading recognition device provided in an embodiment of the present application. The device is used to implement the above-mentioned lip reading recognition method. The device may include:
[0126] The first module is used to obtain the audio and video data to be identified corresponding to the medical operation patient, and perform data preprocessing on the audio and video data to be identified;
[0127] The second module is used to use the lip reading recognition model to perform lip reading recognition on the pre-processed audio and video data to be recognized, and output the lip reading recognition results; the lip reading recognition model is trained by the above-mentioned model training method.
[0128] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0129] The embodiment of the present application also provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the above-mentioned model training method or lip reading recognition method when executing the computer program. The electronic device can be any smart terminal including a tablet computer.
[0130] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0131] See also Figure 4 , Figure 4 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:
[0132] The processor 901 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0133] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solution provided in the embodiments of this specification is implemented by software or firmware, the relevant program code is stored in the memory 902, and the processor 901 calls and executes the model training method or lip reading recognition method of the embodiments of this application;
[0134] Input / output interface 903, used to implement information input and output;
[0135] Communication interface 904, used to realize communication interaction between the device and other devices, which can be realized by wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.);
[0136] A bus 905 that transmits information between various components of the device (e.g., the processor 901, the memory 902, the input / output interface 903, and the communication interface 904);
[0137] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .
[0138] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned model training method or lip reading recognition method is implemented.
[0139] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0140] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0141] The embodiments of the present application provide a model training method, a lip reading recognition method, a device, an electronic device and a medium, which can automatically recognize the lip reading of the patient's audio and video, improve the efficiency and accuracy of lip reading recognition, solve the problem of communication difficulties of patients in the aphonic state, and improve the convenience of patients' life after surgery.
[0142] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0143] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0144] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0145] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.
[0146] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0147] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0148] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0149] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0150] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0151] It should be appreciated that embodiments of the present invention may be implemented or enforced by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable memory. The method may be implemented in a computer program using standard programming techniques including a non-transitory computer-readable storage medium configured with a computer program, wherein the storage medium so configured causes the computer to operate in a specific and predefined manner - according to the methods and drawings described in the specific embodiments. Each program may be implemented in a high-level procedural or object-oriented programming language to communicate with a computer system. However, if desired, the program may be implemented in an assembly or machine language. In any case, the language may be a compiled or interpreted language. In addition, the program may be run on a programmed ASIC for this purpose.
[0152] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.
[0153] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.
Claims
1. A model training method, characterized in that: The method comprises the following steps: Acquire a sample data set from a multimodal corpus; the sample data set includes a plurality of patient audio and video data with text labels; Dividing the sample data set to obtain a training set and a test set; Constructing a lip reading recognition model, and training the lip reading recognition model using the training set; The test set is used to perform model evaluation on the trained lip reading recognition model to determine whether to continue training the lip reading recognition model.
2. The model training method according to claim 1, characterized in that: The step of obtaining a sample data set from a multimodal corpus specifically includes: constructing the multimodal corpus; Collecting pre- and post-operative audio and video data corresponding to a plurality of medical surgery patients, and performing data pre-processing on the pre- and post-operative audio and video data; the pre- and post-operative audio and video data include pre-operative audio and video data and post-operative audio and video data; the medical surgery patients include patients undergoing partial laryngectomy, complete laryngectomy or tracheotomy; Perform text recognition and annotation on the pre-processed audio and video data before and after the operation to obtain text labels corresponding to the audio and video data before and after the operation, and store the audio and video data before and after the operation including the text labels in the multimodal corpus.
3. The model training method according to claim 2, characterized in that: The performing of text recognition and labeling on the pre-processed audio and video data before and after the surgery to obtain text labels corresponding to the audio and video data before and after the surgery specifically includes: Using a speech recognition model to perform text recognition on the audio and video data before and after the operation, and determining a plurality of text annotation contents and a video timestamp and a dialect type corresponding to each of the text annotation contents; Using an emotional computing model to perform emotional recognition on each of the text annotation contents, and determining the emotional feature labels corresponding to each of the text annotation contents; Performing medical entity annotation on each of the text annotation contents using a medical entity annotation model to determine the medical entity annotation corresponding to each of the text annotation contents; Dividing the audio and video data before and after the operation into video frames to obtain a plurality of video frames and a time stamp corresponding to each of the video frames; Determine a text description label corresponding to each of the text annotation contents according to the dialect type, the emotion feature label, and the medical entity label corresponding to each of the text annotation contents; According to the video timestamp corresponding to each of the text annotation contents and the timestamp corresponding to each of the video frames, pairing the text annotation contents with each of the video frames to generate a video frame text pairing label corresponding to each of the text annotation contents; The text label is determined according to the video frame text pairing label corresponding to each of the text annotation contents.
4. The model training method according to claim 3, characterized in that: The pairing of the text annotation content and each video frame according to the video timestamp corresponding to each text annotation content and the timestamp corresponding to each video frame to generate a video frame text pairing label corresponding to each text annotation content specifically includes: Determining a plurality of target video frames corresponding to each of the text annotation contents from the plurality of video frames according to the video timestamps corresponding to each of the text annotation contents and the timestamps corresponding to each of the video frames; The text description tags corresponding to each of the text annotation contents are associated and bound to each of the target video frames to generate the video frame text pairing tags corresponding to each of the text annotation contents.
5. The model training method according to claim 1, characterized in that: The step of using the test set to perform model evaluation on the trained lip reading recognition model to determine whether to continue training the lip reading recognition model specifically includes: Obtaining multiple model evaluation indicators and model evaluation indicator thresholds corresponding to each of the model evaluation indicators; Inputting the test set into the lip reading recognition model, and using the lip reading recognition model to output corresponding lip reading recognition results; According to each of the model evaluation indicators and the lip reading recognition result, the lip reading recognition model is evaluated to determine the current evaluation value of the model corresponding to each of the model evaluation indicators; Determine a model evaluation result according to the model evaluation indicator threshold corresponding to each model evaluation indicator and the current evaluation value of the model, and determine whether to continue training the lip reading recognition model according to the model evaluation result; When the model evaluation result is that the current evaluation values of the model corresponding to each of the model evaluation indicators exceed the model evaluation indicator threshold, the training of the lip reading recognition model is stopped; otherwise, the training of the lip reading recognition model continues.
6. A lip reading recognition method, characterized in that: The method comprises the following steps: Acquire the audio and video data to be identified corresponding to the medical operation patient, and perform data preprocessing on the audio and video data to be identified; The lip reading recognition model is used to perform lip reading recognition on the preprocessed audio and video data to be recognized, and the lip reading recognition result is output; the lip reading recognition model is trained by the model training method described in any one of claims 1 to 5.
7. The lip reading recognition method according to claim 6, characterized in that: The lip reading recognition model is used to perform lip reading recognition on the pre-processed audio and video data to be recognized, and the lip reading recognition result is output, specifically including: Using the lip reading recognition model, the preprocessed audio and video data to be recognized are divided into video frames and lip reading recognition is performed, and a plurality of text recognition information corresponding to the video data to be recognized and a plurality of video frames to be paired corresponding to each of the text recognition information are output; The lip reading recognition model is used to pair each of the text recognition information with the corresponding multiple video frames to be paired, and the video frame text pairing results corresponding to each of the text recognition information are determined. According to each of the video frame text pairing results, each of the text recognition information is inserted into the audio and video data to be recognized, and the corresponding lip reading recognition audio and video data is output.
8. A lip reading recognition device, characterized in that: The device comprises: The first module is used to obtain the audio and video data to be identified corresponding to the medical operation patient, and perform data preprocessing on the audio and video data to be identified; The second module is used to use a lip reading recognition model to perform lip reading recognition on the preprocessed audio and video data to be recognized, and output the lip reading recognition result; the lip reading recognition model is trained by the model training method described in any one of claims 1 to 5.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the model training method described in any one of claims 1 to 5 or the lip reading recognition method described in any one of claims 6 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the model training method described in any one of claims 1 to 5 or the lip reading recognition method described in any one of claims 6 to 7 is implemented.
Citation Information
Patent Citations
Lip language recognition-based artificial intelligence sound production system and method
CN108831472A
Cantonese lip reading recognition method and device and storage medium
CN114299418A
Lip language recognition method for patient with language disorder in hospital environment
CN116959060A
Mouth shape recognition method based on lip language recognition model and application system
CN117133291A
Model training method, antibacterial peptide prediction method and system
CN117875444A