Speech Recognition Method, Device, Medium and Equipment Based on User Facial Expressions

By combining facial expression recognition and voice data processing, the speech recognition model is trained to identify user intentions, and the problem of inaccurate speech recognition in the prior art is solved, and the accuracy and user experience of the interaction between smart devices and users is improved.

CN115440196BActive Publication Date: 2025-06-24SHENZHEN YIJIAN INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211163199.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-23
Publication Date
2025-06-24
Estimated Expiration
2042-09-23

AI Technical Summary

Technical Problem

In the prior art, speech recognition is inaccurate, resulting in inaccurate intention recognition when the smart device interacts with the user, affecting the user experience.

Method used

The infrared acquisition device collects facial thermal images, recognizes facial feature points, and generates dynamic feature images based on facial expression changes, combines speech data for semantic recognition, and trains the speech recognition model through emotional tags.

Benefits of technology

Improve the accuracy of speech recognition, enable smart devices to more accurately identify user intentions, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115440196B_ABST
    Figure CN115440196B_ABST
Patent Text Reader

Abstract

The present disclosure provides a voice recognition method, device, medium and equipment based on the facial expressions of users. The method includes: determining the change situation of the facial feature points of a target user in a monitoring environment within a preset time period according to an identification model to generate a facial dynamic feature image; matching multiple dynamic sub-images of feature regions with preset dynamic sub-images of corresponding feature regions to determine the emotion label corresponding to the target user; collecting audio data of the target user in the monitoring environment within the preset time period to generate user voice corresponding to the target user; training a voice recognition model according to the emotion label; and performing semantic recognition on the user voice through the trained voice recognition model to generate semantic information corresponding to the target user. Thereby, the intelligent device can more accurately recognize the user intention corresponding to the user voice, improving the accuracy of voice recognition and bringing a better product experience to users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition, and in particular to a speech recognition method, device, medium and equipment based on a user's facial expression. Background Art

[0002] In the prior art, only the text recognition of the user's speech is performed, the text of the user's language is extracted, and the user's intention is recognized through text semantics. However, in the voice dialogue scenario of human-computer interaction, the user's intention obtained only by analyzing the semantics is inaccurate, which seriously affects the subsequent interaction process between the intelligent device and the user, and brings a bad experience to the user. Summary of the Invention

[0003] In view of this, the purpose of the present disclosure is to provide a speech recognition method, device, medium and equipment based on a user's facial expression, so as to solve the technical problem of inaccurate speech recognition in the related art.

[0004] Based on the above invention purpose, the first aspect of the present disclosure provides a speech recognition method based on a user's facial expression, and the method includes:

[0005] Collect a thermal image in a monitoring environment through an infrared acquisition device, and when it is confirmed that there is a human face in the monitoring environment based on an image recognition model, determine the facial feature points of the target user corresponding to the human face according to a feature recognition algorithm, and based on the preset distribution rule of the facial feature points, repeatedly execute the following steps until it is determined that the facial feature points of the target user in the monitoring environment change: select a frame image with a corresponding duration from the initial dynamic image according to a preset target duration to generate a facial dynamic feature image corresponding to the target user, and match the facial dynamic feature image with a preset standard dynamic image to generate a comparison result, and determine whether the matching is successful according to the comparison result. If the matching is successful, it is determined that the target user does not have an emotional fluctuation in the frame image with the corresponding duration, then extend the used target duration, and re-obtain the frame image with the corresponding duration. If the matching is not successful, extract the frame image with the duration, generate a facial dynamic feature image corresponding to the target user, and based on the preset distribution rule and multiple feature regions corresponding to the human face, segment the facial dynamic feature image to generate multiple feature region dynamic sub-images corresponding to the multiple feature regions, where the multiple feature regions at least include an eye feature region, a nose feature region, and a mouth feature region;

[0006] Match the multiple feature region dynamic sub-images with the multiple preset dynamic sub-images corresponding to the multiple feature regions, determine the multiple expression recognition results corresponding to the multiple feature region dynamic sub-images, and fuse the multiple expression recognition results according to preset weights to determine the emotion label corresponding to the target user. Among them, the expression recognition result is used to represent the emotion label corresponding to the target user, and the preset weights are set according to the strength relationship of each feature region representing the emotion label;

[0007] Collect the audio data of the target user in the monitoring environment within the preset time period to generate target audio data, identify the user voice frequency band corresponding to the target user in the target audio data, perform noise reduction processing on the target audio data according to the user voice frequency band, and perform voice extraction on the noise-reduced target audio data according to the set voice features to generate the user voice corresponding to the target user. Among them, control instructions are issued to the intelligent terminal through the user voice collected by the microphone;

[0008] Screen out the initial sample voice data corresponding to the emotion label from the initial database, add the initial sample voice data to the sample training set of the voice recognition model, perform recognition training on the voice recognition model based on the sample training set, and perform semantic recognition on the user voice through the trained voice recognition model to generate the semantic information corresponding to the target user. Among them, the initial database includes the mapping relationship between multiple initial sample voice data and multiple emotion labels.

[0009] Further, the performing noise reduction processing on the target audio data according to the user voice frequency band, performing voice extraction on the noise-reduced target audio data according to the set voice features to generate the user voice corresponding to the target user includes:

[0010] Analyze the user voice in the target audio data according to the historical user voice corresponding to the target user, so as to generate the user voice frequency band and environmental audio according to the target audio data;

[0011] Perform noise reduction processing on the target audio data based on the user voice frequency band to remove the environmental audio in the target audio data, and perform topology restoration on the processed target audio data to generate the user voice corresponding to the target user.

[0012] Further, screening out the initial sample voice data corresponding to the emotion label from the initial database, adding the initial sample voice data to the sample training set of the speech recognition model, and performing recognition training on the speech recognition model based on the sample training set, and performing semantic recognition on the user voice through the trained speech recognition model to generate the semantic information corresponding to the target user, including:

[0013] Screening the initial sample voice data in the initial database based on the emotion label to obtain a preset number of first sample voice data and corresponding first emotion semantics, where the first emotion semantics is the semantic information of the first sample voice data under the emotion label;

[0014] Extracting features from the first sample voice data through the feature extraction network of the speech recognition model to generate a feature vector corresponding to the first sample voice data, and performing semantic recognition on the feature vector through the fully connected neural network of the speech recognition model to generate target semantic information. In the case where the target semantic information is determined to be inconsistent with the first emotion semantics, updating the speech recognition model according to the first emotion semantics;

[0015] Performing semantic recognition on the user voice based on the updated speech recognition model to generate the semantic information corresponding to the target user.

[0016] Further, matching the multiple feature region dynamic sub-images with the multiple preset dynamic sub-images corresponding to the multiple feature regions, determining the multiple expression recognition results corresponding to the multiple feature region dynamic sub-images, and fusing the multiple expression recognition results according to preset weights to determine the emotion label corresponding to the target user, including:

[0017] Performing normalization processing on any one of the feature region dynamic sub-images to generate a dynamic grayscale sub-image of the same size;

[0018] Identifying the grayscale sub-image to determine the feature region corresponding to the grayscale sub-image;

[0019] Obtaining the multiple preset dynamic sub-images corresponding to the feature region, and matching the multiple preset dynamic sub-images with the grayscale sub-image to determine the similarity between the multiple preset dynamic sub-images and the grayscale sub-image, where each preset dynamic sub-image corresponds to a preset expression recognition result;

[0020] Determining the target expression recognition result corresponding to the target preset dynamic sub-image with the maximum similarity as the expression recognition result.

[0021] The second aspect of the present disclosure provides a voice recognition device based on a user's facial expression. The device includes:

[0022] A first generation module, configured to collect a thermal image in a monitoring environment through an infrared acquisition device, and when it is confirmed that there is a human face in the monitoring environment based on an image recognition model, determine the facial feature points of the target user corresponding to the human face according to a feature recognition algorithm, and based on a preset distribution rule of the facial feature points, repeatedly execute the following steps until it is determined that the facial feature points of the target user in the monitoring environment change: Select frame images with a corresponding duration from the initial dynamic image according to a preset target duration to generate a facial dynamic feature image corresponding to the target user, and match the facial dynamic feature image with a preset standard dynamic image to generate a comparison result. Determine whether the match is successful according to the comparison result. If the match is successful and it is determined that the target user does not have an emotional fluctuation in the frame images with the corresponding duration, extend the used target duration and obtain frame images with the corresponding duration again. If the match is not successful, extract the frame images with the duration, generate a facial dynamic feature image corresponding to the target user, and segment the facial dynamic feature image based on the preset distribution rule and multiple feature regions corresponding to the target user to generate multiple feature region dynamic sub-images corresponding to the multiple feature regions; where the multiple feature regions at least include an eye feature region, a nose feature region, and a mouth feature region;

[0023] A determination module, configured to match the multiple feature region dynamic sub-images with multiple preset dynamic sub-images corresponding to the multiple feature regions, determine multiple expression recognition results corresponding to the multiple feature region dynamic sub-images, and fuse the multiple expression recognition results according to a preset weight to determine an emotion label corresponding to the target user. Wherein, the expression recognition result is used to represent the emotion label corresponding to the target user, and the preset weight is set according to the strength relationship of each feature region representing the emotion label;

[0024] A second generation module, configured to collect audio data of the target user in the monitoring environment within the preset time period, generate target audio data, identify the user voice frequency band corresponding to the target user in the target audio data, perform noise reduction processing on the target audio data according to the user voice frequency band, and perform voice extraction on the noise-reduced target audio data according to a set voice feature to generate the user voice corresponding to the target user. Wherein, a control instruction is issued to the intelligent terminal through the user voice collected by the microphone;

[0025] A third generation module, configured to screen out initial sample voice data corresponding to the emotion label from an initial database, add the initial sample voice data to a sample training set of a speech recognition model, and perform recognition training on the speech recognition model based on the sample training set, and perform semantic recognition on the user voice through the trained speech recognition model to generate semantic information corresponding to the target user, wherein the initial database includes a mapping relationship between multiple groups of initial sample voice data and multiple emotion labels.

[0026] Further, the second generation module may further be configured to:

[0027] Analyze the user voice in the target audio data according to the historical user voice corresponding to the target user, so as to generate the user voice frequency band and the environmental audio according to the target audio data;

[0028] Perform noise reduction processing on the target audio data based on the user voice frequency band, remove the environmental audio in the target audio data, and perform topological restoration on the processed target audio data to generate the user voice corresponding to the target user.

[0029] A third aspect of the present disclosure provides a computer storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the steps of the speech recognition method based on the user's facial expression as described in any one of the first aspects.

[0030] A fourth aspect of the present disclosure provides an electronic device, including a computer program, and when the computer program is executed by a processor, it implements the steps of the speech recognition method based on the user's facial expression as described in any one of the first aspects.

[0031] The present disclosure can at least achieve the following beneficial effects:

[0032] By collecting the changes of the facial feature points of the target user in the monitoring environment within a preset time period, a facial dynamic feature image corresponding to the target user is generated. The facial dynamic feature image is segmented to generate multiple feature region dynamic sub-images. The multiple feature region dynamic sub-images are matched with the preset dynamic sub-images of the corresponding feature regions to determine the emotion label corresponding to the target user. The audio data of the target user in the monitoring environment within the preset time period is collected. According to the voice features, the target audio data after noise reduction is subjected to voice extraction to generate the user voice corresponding to the target user. The initial sample voice data corresponding to the emotion label is screened out from the initial database, and the initial sample voice data is added to the sample training set of the voice recognition model. The voice recognition model is trained based on the sample training set. The user voice is semantically recognized by the trained voice recognition model to generate the semantic information corresponding to the target user. Thus, by judging the facial emotion of the user, an emotion label is generated and the voice recognition model is trained according to the emotion label. The user voice is semantically recognized by the trained voice recognition model, enabling the intelligent device to more accurately recognize the user intention corresponding to the user voice, improving the accuracy of voice recognition and bringing a better product experience to the user. Description of the Drawings

[0033] Figure 1 It is a flowchart of a voice recognition method based on the facial expression of a user in an embodiment of the present disclosure.

[0034] Figure 2 It is a structural diagram of a voice recognition device based on the facial expression of a user in an embodiment of the present disclosure. Detailed Embodiments

[0035] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention is provided in conjunction with the accompanying drawings. Many specific details are set forth in the following description to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein. Those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0036] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation on the present invention.

[0037] In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined.

[0038] In the present invention, unless otherwise clearly specified and limited, the terms "mounted", "connected", "coupled", "fixed", etc. shall be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two elements or the interaction relationship between two elements, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0039] In the present invention, unless otherwise clearly specified and limited, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature may mean that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "beneath" and "underneath" the second feature may be that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.

[0040] It should be noted that when an element is referred to as "fixed to" or "disposed on" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "upper", "lower", "left", "right" and similar expressions used herein are only for illustrative purposes and do not represent the only implementation.

[0041] Figure 1 The following is a flowchart of a voice recognition method based on a user's facial expression in an embodiment of the present disclosure. As Figure 1 shown, the method includes the following steps:

[0042] In step S11, according to the image recognition model, determine the change of the facial feature points of the target user in the monitoring environment within a preset time period to generate a facial dynamic feature image;

[0043] Among them, the thermal image in the monitoring environment is collected by an infrared acquisition device, and when it is confirmed that there is a human face facial feature in the monitoring environment based on the image recognition model, the facial feature points of the target user corresponding to the human face are determined according to the feature recognition algorithm, and based on the preset distribution rule of the facial feature points, the following steps are cyclically executed until it is determined that the facial feature points of the target user in the monitoring environment have changed: select frame images with a corresponding duration from the initial dynamic image according to a preset target duration to generate a facial dynamic feature image corresponding to the target user, and match the facial dynamic feature image with a preset standard dynamic image to generate a comparison result, and judge whether the matching is successful according to the comparison result. If the matching is successful, judge that the target user has not had an emotional fluctuation in the frame images with the corresponding duration, then extend the used target duration and obtain frame images with the corresponding duration again. If the matching is not successful, extract the frame images with the duration, generate a facial dynamic feature image corresponding to the target user, and based on the preset distribution rule and the multiple feature regions corresponding to the human face, segment the facial dynamic feature image to generate multiple feature region dynamic sub-images corresponding to the multiple feature regions, where the multiple feature regions at least include an eye feature region, a nose feature region, and a mouth feature region.

[0044] In step S12, match the multiple feature region dynamic sub-images with multiple preset dynamic sub-images corresponding to the multiple feature regions to determine the emotion label corresponding to the target user;

[0045] Among them, the facial expression recognition result is used to represent the emotion label corresponding to the target user. The preset weight is set according to the strength relationship of the emotion labels represented by each feature region. The emotion label is determined by fusing multiple facial expression recognition results according to the preset weight based on the facial expression recognition results corresponding to the dynamic sub-images of each feature region.

[0046] In step S13, the audio data of the target user in the monitoring environment within the preset time period is collected to generate the user voice corresponding to the target user.

[0047] Among them, by collecting the audio data of the target user in the monitoring environment within the preset time period, target audio data is generated, the user voice frequency band in the target audio data is identified, the target audio data is denoised according to the user voice frequency band, and voice extraction is performed on the denoised target audio data according to voice features to generate the user voice corresponding to the target user. A control instruction is issued to the intelligent terminal through the user voice collected by the microphone.

[0048] In step S14, the speech recognition model is trained according to the emotion label, and the semantic recognition of the user voice is performed through the trained speech recognition model to generate the semantic information corresponding to the target user.

[0049] Among them, by screening out the initial sample voice data corresponding to the emotion label from the initial database, the initial sample voice data is added to the sample training set of the speech recognition model, and the speech recognition model is trained for recognition based on the sample training set. The semantic recognition of the user voice is performed through the trained speech recognition model to generate the semantic information corresponding to the target user. Among them, the initial database includes the mapping relationship between multiple groups of initial sample voice data and multiple emotion labels.

[0050] Adopting the above technical solution, by collecting the change situation of the facial feature points of the target user in the monitoring environment within a preset time period, a facial dynamic feature image corresponding to the target user is generated, and the facial dynamic feature image is segmented to generate multiple feature region dynamic sub-images; the multiple feature region dynamic sub-images are matched with the preset dynamic sub-images of the corresponding feature regions to determine the emotion label corresponding to the target user, the audio data of the target user in the monitoring environment within the preset time period is collected, voice extraction is performed on the denoised target audio data according to voice features to generate the user voice corresponding to the target user, initial sample voice data corresponding to the emotion label is screened out from the initial database, the initial sample voice data is added to the sample training set of the voice recognition model, and the voice recognition model is recognized and trained based on the sample training set, and the semantic recognition of the user voice is performed through the trained voice recognition model to generate the semantic information corresponding to the target user. Thus, by judging the facial emotion of the user, an emotion label is generated and the voice recognition model is trained according to the emotion label, and the semantic recognition of the user voice is performed through the trained voice recognition model, so that the intelligent device can more accurately recognize the user intention corresponding to the user voice, improving the accuracy of voice recognition and bringing a better product experience to the user.

[0051] Further, step S14 above includes:

[0052] Analyze the user voice in the target audio data according to the historical user voice corresponding to the target user, so as to generate the user voice frequency band and environmental audio according to the target audio data;

[0053] Perform noise reduction processing on the target audio data based on the user voice frequency band, remove the environmental audio in the target audio data, and perform topology restoration on the processed target audio data to generate the user voice corresponding to the target user.

[0054] Further, step S14 above includes:

[0055] Screen the initial sample voice data in the initial database based on the emotion label to obtain a preset number of first sample voice data and corresponding first emotion semantics, and the first emotion semantics is the semantic information of the first sample voice data under the emotion label;

[0056] Feature extraction is performed on the first sample speech data through the feature extraction network of the speech recognition model to generate a feature vector corresponding to the first sample speech data. Semantic recognition is performed on the feature vector through the fully connected neural network of the speech recognition model to generate target semantic information. In the case where it is determined that the target semantic information is inconsistent with the first emotional semantics, the speech recognition model is updated according to the first emotional semantics;

[0057] Semantic recognition is performed on the user speech based on the updated speech recognition model to generate the semantic information corresponding to the target user.

[0058] Further, step S12 above includes:

[0059] Normalization processing is performed on any of the feature region dynamic sub-images to generate dynamic grayscale sub-images of the same size;

[0060] The grayscale sub-image is recognized to determine the feature region corresponding to the grayscale sub-image;

[0061] Multiple preset dynamic sub-images corresponding to the feature region are obtained, and the multiple preset dynamic sub-images are matched with the grayscale sub-image to determine the similarity between the multiple preset dynamic sub-images and the grayscale sub-image, where each frame of the preset dynamic sub-image corresponds to a preset expression recognition result;

[0062] The target expression recognition result corresponding to the target preset dynamic sub-image with the maximum similarity is determined as the expression recognition result.

[0063] Figure 2 It is a structural diagram of a speech recognition device based on a user's facial expression in an embodiment of the present disclosure. The recognition device 100 includes: a first generation module 110, a determination module 120, a second generation module 130, and a third generation module 140.

[0064] The first generation module 110 is configured to collect thermal images in the monitoring environment through an infrared acquisition device. When it is confirmed that there is a human face in the monitoring environment based on an image recognition model, it determines the facial feature points of the target user corresponding to the human face according to a feature recognition algorithm, and based on the preset distribution rule of the facial feature points, repeatedly executes the following steps until it is determined that the facial feature points of the target user in the monitoring environment have changed: Select frame images of a corresponding duration from the initial dynamic image according to a preset target duration to generate a facial dynamic feature image corresponding to the target user, and match the facial dynamic feature image with a preset standard dynamic image to generate a comparison result. Determine whether the match is successful according to the comparison result. If the match is successful and it is determined that the target user has not had an emotional fluctuation in the frame images of the corresponding duration, extend the used target duration and obtain frame images of the corresponding duration again. If the match is not successful, extract the frame images of the duration, generate a facial dynamic feature image corresponding to the target user, and segment the facial dynamic feature image based on the preset distribution rule and multiple feature regions corresponding to the target user to generate multiple feature region dynamic sub-images corresponding to the multiple feature regions; where the feature regions at least include an eye feature region, a nose feature region, and a mouth feature region.

[0065] The determination module 120 is configured to match the multiple feature region dynamic sub-images with the preset dynamic sub-images of the multiple feature regions, determine multiple expression recognition results corresponding to the multiple feature region dynamic sub-images, and fuse the multiple expression recognition results according to preset weights to determine the emotion label corresponding to the target user. Among them, the expression recognition result is used to represent the emotion label corresponding to the target user, and the preset weights are set according to the strength relationship of the emotion labels represented by each feature region.

[0066] The second generation module 130 is configured to collect audio data of the target user in the monitoring environment within the preset time period, generate target audio data, identify the user voice frequency band corresponding to the target user in the target audio data, perform noise reduction processing on the target audio data according to the user voice frequency band, and perform voice extraction on the noise-reduced target audio data according to set voice features to generate the user voice corresponding to the target user. Among them, the user voice collected by the microphone issues a control instruction to the intelligent terminal.

[0067] A third generation module 140, configured to screen out initial sample voice data corresponding to the emotion label from an initial database, add the initial sample voice data to a sample training set of a speech recognition model, and perform recognition training on the speech recognition model based on the sample training set, and perform semantic recognition on the user voice through the trained speech recognition model to generate semantic information corresponding to the target user, where the initial database includes a mapping relationship between a plurality of initial sample voice data and a plurality of emotion labels.

[0068] The above device collects the change situation of the facial feature points of the target user in the monitoring environment within a preset time period to generate a facial dynamic feature image corresponding to the target user, segments the facial dynamic feature image to generate a plurality of feature region dynamic sub-images; matches the plurality of feature region dynamic sub-images with preset dynamic sub-images of corresponding feature regions to determine an emotion label corresponding to the target user, collects audio data of the target user in the monitoring environment within the preset time period, performs speech extraction on the denoised target audio data according to speech features to generate the user voice corresponding to the target user, screens out initial sample voice data corresponding to the emotion label from an initial database, adds the initial sample voice data to a sample training set of a speech recognition model, and performs recognition training on the speech recognition model based on the sample training set, and performs semantic recognition on the user voice through the trained speech recognition model to generate semantic information corresponding to the target user. Thus, by judging the facial emotion of the user, an emotion label is generated and the speech recognition model is trained according to the emotion label, and the semantic recognition of the user voice is performed through the trained speech recognition model, so that the intelligent device can more accurately recognize the user intention corresponding to the user voice, improves the accuracy of speech recognition, and brings a better product experience to the user.

[0069] Further, the first generation module 110 may further be configured to:

[0070] Based on a preset distribution rule of facial feature points, the following steps are cyclically executed until it is determined that the facial feature points of the target user in the monitoring environment change;

[0071] Select frame images with a corresponding duration from the initial dynamic image according to the preset target duration to generate a facial dynamic feature image corresponding to the target user; and match the facial dynamic feature image with a preset standard dynamic image to generate a comparison result, and judge whether the matching is successful according to the comparison result;

[0072] If the matching is successful and it is determined that the target user does not have an emotional fluctuation in the frame images with the corresponding duration, extend the used target duration and re-obtain the frame images with the corresponding duration;

[0073] If the matching is unsuccessful, extract the frame images of the duration to generate a facial dynamic feature image corresponding to the target user.

[0074] Furthermore, the second generation module 130 can also be used for:

[0075] Analyze the user voice in the target audio data according to the historical user voice corresponding to the target user, so as to generate the user voice frequency band and ambient audio according to the target audio data;

[0076] Perform noise reduction processing on the target audio data based on the user voice frequency band, remove the ambient audio in the target audio data, and perform topological restoration on the processed target audio data to generate the user voice corresponding to the target user.

[0077] Furthermore, the third generation module 140 can also be used for:

[0078] Based on the emotion label, screen the initial sample voice data in the initial database to obtain a preset number of first sample voice data and corresponding first emotion semantics, and the first emotion semantics is the semantic information of the first sample voice data under the emotion label;

[0079] Extract features from the first sample voice data through the feature extraction network of the speech recognition model to generate a feature vector corresponding to the first sample voice data, perform semantic recognition on the feature vector through the fully connected neural network of the speech recognition model to generate target semantic information, and in the case where it is determined that the target semantic information is inconsistent with the first emotion semantics, update the speech recognition model according to the first emotion semantics;

[0080] Perform semantic recognition on the user voice based on the updated speech recognition model to generate the semantic information corresponding to the target user.

[0081] Furthermore, the third generation module 140 can also be used for:

[0082] Perform normalization processing on any feature region dynamic sub-image to generate dynamic grayscale sub-images of the same size;

[0083] Identify the grayscale sub-image to determine the feature region corresponding to the grayscale sub-image;

[0084] Obtain the multiple preset dynamic sub-images corresponding to the feature region, and match the multiple preset dynamic sub-images with the grayscale sub-image to determine the similarity between the multiple preset dynamic sub-images and the grayscale sub-image, wherein each preset dynamic sub-image corresponds to a preset facial expression recognition result;

[0085] Determine the target facial expression recognition result corresponding to the target preset dynamic sub-image with the maximum similarity as the facial expression recognition result.

[0086] The present disclosure also provides a computer storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the voice recognition method based on the user's facial expression as described in any one of the foregoing.

[0087] The present disclosure also provides an electronic device, including a computer program, which implements the steps of the voice recognition method based on the user's facial expression as described in any one of the foregoing when executed by a processor.

[0088] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0089] The above embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent should be subject to the appended claims.

Claims

1. A speech recognition method based on the user's facial expression, characterized in that, The method includes: Collecting a thermal image in a monitoring environment through an infrared acquisition device, and when it is confirmed that there is a human face in the monitoring environment based on an image recognition model, determining facial feature points of a target user corresponding to the human face according to a feature recognition algorithm, and based on a preset distribution rule of the facial feature points, cyclically executing the following steps until it is determined that the facial feature points of the target user in the monitoring environment change: Selecting frame images of a corresponding duration from an initial dynamic image according to a preset target duration to generate a facial dynamic feature image corresponding to the target user, and matching the facial dynamic feature image with a preset standard dynamic image to generate a comparison result, and judging whether the matching is successful according to the comparison result. If the matching is successful and it is judged that there is no emotional fluctuation of the target user in the frame images of the corresponding duration, then extending the used target duration and obtaining frame images of the corresponding duration again. If the matching is not successful, extracting the frame images of the duration, generating a facial dynamic feature image corresponding to the target user, and segmenting the facial dynamic feature image based on the preset distribution rule and multiple feature regions corresponding to the human face to generate multiple feature region dynamic sub-images corresponding to the multiple feature regions, where the multiple feature regions at least include an eye feature region, a nose feature region, and a mouth feature region; Matching the multiple feature region dynamic sub-images with multiple preset dynamic sub-images corresponding to the multiple feature regions, determining multiple expression recognition results corresponding to the multiple feature region dynamic sub-images, and fusing the multiple expression recognition results according to preset weights to determine an emotional label corresponding to the target user, where the expression recognition result is used to represent the emotional label corresponding to the target user, and the preset weights are set according to the strength relationship of each feature region representing the emotional label; Collecting audio data of the target user in the monitoring environment within a preset time period to generate target audio data, identifying a user voice frequency band corresponding to the target user in the target audio data, performing noise reduction processing on the target audio data according to the user voice frequency band, and performing voice extraction on the noise-reduced target audio data according to set voice features to generate a user voice corresponding to the target user, where a control instruction is issued to an intelligent terminal through the user voice collected by a microphone; Screening out initial sample voice data corresponding to the emotional label from an initial database, adding the initial sample voice data to a sample training set of a voice recognition model, performing recognition training on the voice recognition model based on the sample training set, and performing semantic recognition on the user voice through the trained voice recognition model to generate semantic information corresponding to the target user, where the initial database includes a mapping relationship between multiple initial sample voice data and multiple emotional labels; Among them, the noise reduction process for the target audio data according to the user voice frequency band, and the voice extraction for the denoised target audio data according to the set voice features to generate the user voice corresponding to the target user, includes: Analyze the user voice in the target audio data according to the historical user voice corresponding to the target user, so as to generate the user voice frequency band and the environmental audio according to the target audio data; Perform noise reduction processing on the target audio data based on the user voice frequency band to remove the environmental audio in the target audio data, and perform topology restoration on the processed target audio data to generate the user voice corresponding to the target user.

2. The recognition method according to claim 1, wherein The screening of the initial sample voice data corresponding to the emotion label from the initial database, adding the initial sample voice data to the sample training set of the voice recognition model, performing recognition training on the voice recognition model based on the sample training set, and performing semantic recognition on the user voice through the trained voice recognition model to generate the semantic information corresponding to the target user, includes: Screen the initial sample voice data in the initial database based on the emotion label to obtain a preset number of first sample voice data and corresponding first emotion semantics, where the first emotion semantics is the semantic information of the first sample voice data under the emotion label; Extract features from the first sample voice data through the feature extraction network of the voice recognition model to generate a feature vector corresponding to the first sample voice data, perform semantic recognition on the feature vector through the fully connected neural network of the voice recognition model to generate target semantic information, and in the case where it is determined that the target semantic information is inconsistent with the first emotion semantics, update the voice recognition model according to the first emotion semantics; Perform semantic recognition on the user voice based on the updated voice recognition model to generate the semantic information corresponding to the target user.

3. The recognition method according to claim 1, wherein The matching of the multiple feature region dynamic sub-images with the multiple preset dynamic sub-images corresponding to the multiple feature regions, and determining the multiple expression recognition results corresponding to the multiple feature region dynamic sub-images, includes: Perform normalization processing on any one of the feature region dynamic sub-images to generate a dynamic grayscale sub-image of the same size; Identify the grayscale sub-image to determine the feature region corresponding to the grayscale sub-image; Obtain the multiple preset dynamic sub-images corresponding to the feature region, and match the multiple preset dynamic sub-images with the grayscale sub-image to determine the similarity between the multiple preset dynamic sub-images and the grayscale sub-image, where each preset dynamic sub-image corresponds to a preset expression recognition result; Determine that the target expression recognition result corresponding to the target preset dynamic sub-image with the maximum similarity is the expression recognition result.

4. A speech recognition device based on the user's facial expression, characterized in that, Includes: The first generation module is used to collect thermal images in the monitoring environment through an infrared acquisition device. When it is confirmed that there is a human face in the monitoring environment based on an image recognition model, it determines the facial feature points of the target user corresponding to the human face according to a feature recognition algorithm, and based on the preset distribution rules of the facial feature points, repeatedly executes the following steps until it is determined that the facial feature points of the target user in the monitoring environment have changed: Select frame images with a corresponding duration from the initial dynamic image according to a preset target duration to generate a facial dynamic feature image corresponding to the target user, and match the facial dynamic feature image with a preset standard dynamic image to generate a comparison result. Determine whether the match is successful according to the comparison result. If the match is successful and it is determined that the target user has not had an emotional fluctuation in the frame images with the corresponding duration, extend the used target duration and obtain frame images with the corresponding duration again. If the match is unsuccessful, extract the frame images with the duration, generate a facial dynamic feature image corresponding to the target user, and segment the facial dynamic feature image based on the preset distribution rules and multiple feature regions corresponding to the target user to generate multiple feature region dynamic sub-images corresponding to the multiple feature regions; Wherein the multiple feature regions at least include an eye feature region, a nose feature region, and a mouth feature region; The determination module is used to match the multiple feature region dynamic sub-images with multiple preset dynamic sub-images corresponding to the multiple feature regions, determine multiple expression recognition results corresponding to the multiple feature region dynamic sub-images, and fuse the multiple expression recognition results according to preset weights to determine an emotion label corresponding to the target user. Among them, the expression recognition result is used to represent the emotion label corresponding to the target user, and the preset weights are set according to the strength relationship of each feature region representing the emotion label; The second generation module is used to collect audio data of the target user in the monitoring environment within a preset time period to generate target audio data, identify the user voice frequency band corresponding to the target user in the target audio data, perform noise reduction processing on the target audio data according to the user voice frequency band, and perform voice extraction on the noise-reduced target audio data according to set voice features to generate user voice corresponding to the target user. Among them, a control instruction is issued to the intelligent terminal through the user voice collected by the acquisition microphone; The third generation module is used to screen out initial sample voice data corresponding to the emotion label from an initial database, add the initial sample voice data to the sample training set of a voice recognition model, perform recognition training on the voice recognition model based on the sample training set, and perform semantic recognition on the user voice through the trained voice recognition model to generate semantic information corresponding to the target user. Among them, the initial database includes a mapping relationship between multiple initial sample voice data and multiple emotion labels; Among them, the second generation module is further used for: Analyze the user speech in the target audio data according to the historical user speech corresponding to the target user, so as to generate the user speech frequency band and environmental audio according to the target audio data; Perform noise reduction processing on the target audio data based on the user speech frequency band, remove the environmental audio in the target audio data, and perform topology restoration on the processed target audio data to generate the user speech corresponding to the target user.

5. A computer storage medium, characterized in that, A computer program is stored on the computer storage medium, and when the computer program is run by a processor, it executes the steps of the speech recognition method based on user facial expressions according to any one of claims 1-3.

6. An electronic device, including a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the speech recognition method based on user facial expressions according to any one of claims 1-3.

Citation Information

Patent Citations

  • Emotion recognition method, system and equipment and medium

    CN113380271A

  • Emotion recognition method and device, electronic equipment and storage medium

    CN114120425A