A method, device and equipment for analyzing facial expressions and target object states
By generating a facial expression recognition model and using image data to analyze the facial expression changes of the target object, the problem of the inability to continuously collect and accurately analyze facial expressions in the prior art is solved, real-time and accurate analysis of the emotional changes of the target object is achieved, and the accuracy and objectivity of the lie detection are improved.
Patent Information
- Application Number
- CN202111561312.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-15
AI Technical Summary
The existing technology cannot continuously collect facial expression information and cannot accurately and objectively analyze the real emotions of the target object, resulting in unsatisfactory lie detection results.
By obtaining the image data of the target object, facial expression analysis is performed based on the preset facial expression recognition model. The process of generating facial expression recognition model includes obtaining initial state images, adjusting color space attributes, obtaining emotional change images, and amplitude and frequency of associated image information, and generating facial expression recognition model through training of neural network models.
Real-time and accurate analysis of the facial expression changes of the target object is achieved, and the emotional changes of the target object can be objectively judged, thereby improving the accuracy and objectivity of the lie detection.
Smart Images

Figure CN114241565B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human face recognition, and particularly to a method, device and equipment for analyzing facial expressions and target object states. Background Art
[0002] Currently, in the process of solving criminal cases, the most widely used and mature lie detection technology is the polygraph technique. The tester asks a series of standardized questions to the subject, and records the physiological reactions of the subject to each question through a wearable lie detector, and judges whether the subject lies based on this. However, the existing technology generally uses a variety of devices to collect the bioelectrical signals of the subject, such as using a sphygmomanometer to detect blood pressure, etc. In this way, using various devices to detect the bioelectrical signals of the subject separately requires frequent wearing of relevant instruments, and thus it is impossible to continuously and long-term detect the subject, and only the relevant detection data at each stage of the entire interrogation process can be obtained. The lie detection effect is not ideal, and the interrogation results are mostly based on subjective judgment to analyze the behavioral reactions and self-rationalizing expressions of the subject during the interview / statement, and it is impossible to accurately and objectively analyze the true emotions of the subject. Summary of the Invention
[0003] Therefore, in order to solve the defects in the prior art that it is impossible to continuously collect the required information and impossible to accurately and objectively analyze the true emotions, the present invention provides a method, device and equipment for analyzing facial expressions and target object states.
[0004] According to a first aspect, an embodiment of the present invention provides a method for analyzing facial expressions, including: obtaining image data of a target object; performing facial expression analysis based on a preset facial expression recognition model and the image data to obtain an analysis result, where the preset facial expression recognition model is generated by training based on an initial state image of the target object and a sample image including state changes.
[0005] Optionally, the process of training and generating the preset facial expression recognition model includes: obtaining the initial state image of the target object as a reference sample image; adjusting the color space attribute of the reference sample image to generate a reference sample image; obtaining a change image of the target object in the emotional change stage; associating the state change of the target object based on the amplitude and frequency of the image information in the change image to generate a comparison sample image; training a neural network model based on the reference sample image and the comparison sample image to generate the preset facial expression recognition model.
[0006] Optionally, adjusting the color space attributes of the reference sample image to generate a reference sample image includes: adjusting the brightness and / or contrast of the region of interest of the reference sample image; replacing the region of interest in the original reference sample image with the adjusted region of interest to generate the reference sample image.
[0007] Optionally, adjusting the color space attributes of the reference sample image to generate a reference sample image further includes: inputting the adjusted reference sample image into a deep generation model, and outputting a reference sample image whose similarity to the adjusted reference sample image is less than a preset threshold;
[0008] Optionally, the training process of the deep generation model includes: generating a fake sample image according to random noise and the adjusted reference sample image through the generator of the deep generation model; judging the similarity between the fake sample image and the original reference sample image through the discriminator of the deep generation model; adjusting the parameters of the deep generation model based on the judgment result until the similarity between the generated fake sample image and the original reference sample image is greater than a preset threshold.
[0009] According to a second aspect, an embodiment of the present invention provides a target state analysis method, including: acquiring image information of a target object, collecting audio information of the target object, and determining the correspondence between the image information and the audio information based on time information; performing facial expression analysis based on a preset facial expression recognition model and the image information to obtain a facial expression analysis result; performing state analysis according to the correspondence between the image information and the audio information and the facial expression analysis result to obtain a state analysis result.
[0010] Optionally, the preset facial expression recognition model is trained and generated using the facial expression analysis method described in the first aspect or any one of the embodiments.
[0011] Optionally, performing state analysis according to the correspondence between the image information and the audio information and the facial expression analysis result to obtain a state analysis result includes: determining the correspondence between the audio information and the facial expression analysis result based on the correspondence between the image information and the audio information; using three-dimensional control to analyze the facial expression analysis result based on the correspondence between the audio information and the facial expression analysis result, and calculating and analyzing to obtain the state analysis result.
[0012] According to a second aspect, a facial expression analysis device includes: an acquisition module configured to acquire image data of a target object; a training module configured to perform facial expression analysis based on a preset facial expression recognition model and the image data to obtain an analysis result, where the preset facial expression recognition model is generated by training based on an initial state image of the target object and sample images including state changes.
[0013] According to a third aspect, a target object state analysis device includes: a collection module configured to acquire image information of a target object, acquire audio information of the target object, and determine a correspondence between the image information and the audio information based on time information; an analysis module configured to perform facial expression analysis based on a preset facial expression recognition model and the image information to obtain a facial expression analysis result; and a communication module configured to perform state analysis based on the correspondence between the image information and the audio information and the facial expression analysis result to obtain a state analysis result.
[0014] According to a fourth aspect, a computer device includes: a communication unit, a memory, and a processor, where the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the steps of the method described in the first aspect, the second aspect, or any optional implementation manner.
[0015] According to a fifth aspect, a computer-readable storage medium is characterized in that the computer-readable storage medium stores computer instructions for causing the computer to perform the steps of the method described in the first aspect, the second aspect, or any optional implementation manner.
[0016] The technical solution of the present invention has the following advantages:
[0017] A facial expression analysis method, device, and equipment provided by an embodiment of the present invention, the method includes the following steps: acquiring facial expression image data of a target object in real time, obtaining a facial expression recognition model by training the image data of the initial state of the target object and sample data during state changes, and performing facial expression analysis based on the facial expression recognition model and the acquired facial expression image data. By comprehensively analyzing the real-time acquired facial expression image data and the preset facial expression recognition model, the embodiment of the present invention can objectively and accurately analyze the true emotions of the target object from the changes in the facial expressions of the target object.
[0018] A method, device and equipment for analyzing the state of a target object provided by an embodiment of the present invention. The method includes the following steps: obtaining image data of the target object in each time period in real time, and collecting audio information of the target object in each time period at the same time. Determine the correspondence between the image information and the audio information based on the time information, and perform facial analysis based on the image information and the facial expression recognition model generated by the above facial expression analysis method to obtain a facial analysis result. According to the correspondence between the image information and the audio information and the facial analysis result, obtain the state analysis result of the target object. Through the correspondence between the image data and the audio information of the target object in each time period in the embodiment of the present invention, combined with the facial expression recognition model for comprehensive facial analysis, the state analysis result of the target object is obtained, which can more scientifically obtain the analysis result of the correspondence between the facial expression and the audio information of the target object, and can more objectively obtain the state changes of the target object in each stage. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0020] Figure 1 It is a specific example flowchart of a facial expression analysis method according to an embodiment of the present invention;
[0021] Figure 2 It is another specific example flowchart of a facial expression analysis method according to an embodiment of the present invention;
[0022] Figure 3 It is a schematic diagram of the calm state of the target object in a facial expression analysis method according to an embodiment of the present invention;
[0023] Figure 4 It is a schematic diagram of the tense state of the target object in a facial expression analysis method according to an embodiment of the present invention;
[0024] Figure 5 It is a specific example flowchart of a method for analyzing the state of a target object according to an embodiment of the present invention;
[0025] Figure 6 It is another specific example flowchart of a method for analyzing the state of a target object according to an embodiment of the present invention;
[0026] Figure 7 It is a schematic diagram showing the analysis result of a method for analyzing the state of a target object according to an embodiment of the present invention;
[0027] Figure 8 Schematic structural diagram of an apparatus for a facial expression analysis method according to an embodiment of the present invention;
[0028] Figure 9 Schematic structural diagram of an apparatus for a target object state analysis method according to an embodiment of the present invention;
[0029] Figure 10 Schematic structural diagram of a computer according to an embodiment of the present invention. Detailed implementation manners
[0030] The technical solutions of the present invention will be described clearly and completely below with reference to the accompanying drawings. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0031] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0032] The embodiments of the present invention take the criminal case detection technical field where the human face recognition technology is most widely applied as an example. Specifically, by obtaining the facial expression image information of a target object, facial expression analysis is performed according to a preset facial expression recognition model and image data. The following embodiments are only the preferred embodiments of the present invention, and the present invention can also be applied to other technical fields and is not limited thereto.
[0033] Figure 1 The flowchart of a facial expression analysis method according to an embodiment of the present invention is shown. The facial expression analysis method specifically includes the following steps:
[0034] S100: Obtain the image data of the target object;
[0035] S200: Perform facial expression analysis based on the preset facial expression recognition model and the image data to obtain an analysis result. The preset facial expression recognition model is generated by training based on the initial state image of the target object and sample images including state changes.
[0036] In an embodiment of the present invention, video recording devices such as law enforcement cameras are used to collect non-contact facial expression image information in real time. The image data of the target object in the initial state collected is trained with the sample images of the state changes at each stage during the entire collection process to obtain a facial expression recognition model. By analyzing the facial expression recognition model and the image data of the target object, the facial expression changes of the target object can be accurately obtained, so as to accurately and objectively judge the emotional changes caused by the facial expression changes of the target object.
[0037] Figure 2 An alternative embodiment of the present invention is shown, and the process of training and generating the preset facial expression recognition model specifically includes the following steps:
[0038] S201: Obtain the initial state image of the target object as the reference sample image.
[0039] Specifically, to obtain the initial state image of the target object, that is, to collect the image data of the target object in a calm state and use this as the reference sample image. In the actual application process, when obtaining the initial state image of the target object, the target object can be in states such as sadness (Sad), happiness (Happy), fear (Fear), disgust (Disgust), surprise (Surprise), and anger (Angry), etc. The present invention is not limited thereto.
[0040] S202: Adjust the color space attributes of the reference sample image to generate a reference sample image.
[0041] Specifically, adjust the color space attributes such as brightness and contrast of the reference sample image, adjust the color space attributes such as brightness and contrast of the reference sample image to an appropriate degree, and generate a reference sample based on the adjusted image data.
[0042] S203: Obtain the changing image of the target object during the emotional change stage.
[0043] Specifically, collect the changing images of the target object during the emotional changes in the entire stage. The emotional changes are the emotional stages other than the initial state, and the changing images during the emotional changes are obtained in real time.
[0044] S204: Correlate the state changes of the target object based on the amplitude and frequency of the image information in the changing image to generate a comparison sample image.
[0045] Specifically, extract the frequency and amplitude of the movement around the head of the target object in the changed image, obtain the parameter values of each subtle vibration variable of the frequency and amplitude, and based on the parameter values, perform correlation discrimination on various mental states of the target object such as anger, tension, and aggressiveness, and correlate with the state of the target object, and generate a comparison sample image based on the corresponding relationship. In practical applications, the maximum frequency and average amplitude of the movement around the head are displayed in the form of a vibration halo. Among them, as Figure 3 , Figure 4 shown, the vibration frequency is displayed as the color of the vibration halo, red represents the color of activity and aggressiveness, yellow represents the color of annoyance and tension, green represents the color of the normal state, purple represents the color of rest and calm, the average amplitude is displayed as the size of the vibration halo, and through the color and size changes of the vibration halo, the emotional changes of the target object are visually displayed. The present invention is not limited thereto.
[0046] S205: Train the neural network model based on the reference sample image and the comparison sample image to generate the preset facial expression recognition model.
[0047] Specifically, perform three-dimensional image processing based on the image information of the reference sample image to obtain a dynamic image in three-dimensional space, and analyze the head and neck movement and facial muscle fluctuations caused by emotional changes of the target object through three-dimensional control. During the entire image information acquisition process, collect multiple video images, collect frame differences and accumulate frames in multiple video images, realize the digital visualization of each parameter of psychology and physiology and display it through the amplitude pixels and frequency pixels of the changed image, and generate the preset facial expression recognition model.
[0048] In the embodiment of the present invention, an initial state image of the target object in a calm state is obtained through a network or a camera, and this is used as a reference sample image. The color space attributes of the reference sample image are adjusted, and the adjusted reference sample image is used to generate a reference sample image. During the entire image acquisition process, the micro-motion data of the facial muscles of the target object is collected in real time through a network or a camera, that is, the change image data of the facial expressions when the target object shows emotional changes at each stage during the entire image acquisition process is collected. The vibration parameters of all pixel points in the change image data are extracted. The vibration parameters are the frequency change and amplitude change data of the change image data. Based on the frequency change and amplitude change data of each frame in the change image data, they are associated with the state change of the target object to correspondingly form a comparison sample image. Based on the reference sample image and the comparison sample image, a neural network model is trained to generate the preset facial expression recognition model. Generating a reference sample image from the image data of the initial state of the target object, generating a comparison sample image from the transformed image data, and training through the reference sample image and the comparison sample image to generate a facial expression recognition model can comprehensively analyze the state change of the target object at each stage during the entire interrogation process, so as to more objectively and accurately analyze the true change of the target object, further assist in solving cases, and improve work efficiency.
[0049] In an embodiment of the present invention, the entire image acquisition process is divided into five stages. The first stage acquires image data of the initial state of the target object; the second stage is used to acquire facial expression data of the target object when it is relatively calm; the third stage acquires the facial expression image of the target object in real time, and analyzes the psychological state and emotional changes of the target object in real time based on the vibration image generated by the facial expression image; the fourth stage arouses strong emotional changes in the target object, acquires the facial expression of the target object in real time, and extracts the frequency and amplitude in the image information of the facial image of the target object, associates the parameter values of the frequency and amplitude with the emotional changes of the target object, and judges the authenticity of the acquired audio information of the target object through the emotional changes; the fifth stage acquires the facial information of the target object and compares it with the facial expression image acquired in the first stage, and comprehensively draws the psychological changes of the target object in the entire image acquisition process based on the comparison results. In practical applications, the embodiment of the present invention combines a five-stage inquiry method to assist criminal case detection and interrogation work. Specifically, in the first stage, the state tracking stage (ST), the video recording device collects image data of the subject in a sitting state and analyzes the basic psychological state of the subject at this time; the second stage, the adaptive stage (AP), continuously collects image data of the subject when calm, and collects audio data of the subject speaking freely, and analyzes the psychological state of the subject at this time based on the image data and audio information; the third stage, the recall stage (IR), continuously collects audio information and facial expression images of the subject, establishes a correlation between the audio information and the facial expression images, analyzes the generated vibration image, and judges the emotional change state of the subject; the fourth stage, the deepening problem stage (Concrete Clarification), actively arouses the emotional change of the subject through inquiry, analyzes the facial expression image data of the subject at this stage in real time, and judges the psychological state of the subject through the change of the vibration image, and adjusts the inquiry method in a targeted manner through the change of the vibration image; the fifth stage, the adjustment stage (AR) The facial expression image data of the subjects in this stage are collected by video recording equipment, and the facial expression image data are compared with the facial expression image data of the subjects collected in the first stage to obtain a comparison result, based on which the psychological changes of the subjects in the whole process of questioning are analyzed, and the psychological changes of the subjects in the process of the five-stage questioning method are comprehensively obtained, which plays an auxiliary reference role in the interrogation of criminal cases.
[0050] In an alternative embodiment of the present invention, the method for adjusting the color space attributes of the reference sample image in step S202 to generate a reference sample image specifically includes the following steps:
[0051] (1) Adjust the brightness and / or contrast of the region of interest in the reference sample image.
[0052] Specifically, in this embodiment, the parameter values of the frequency and amplitude in the reference sample image are extracted. Based on these parameter values, three-dimensional space processing is performed on the reference sample image to obtain a dynamic image of the three-dimensional space of the reference sample image. After comprehensive processing of the parameter values, different color images are displayed in the dynamic image. The maximum frequency and evaluation amplitude obtained from the movement around the head of the target object are regarded as a halo. The vibration frequency is used to represent the color of the vibration halo, and the average amplitude is used to represent the size of the halo. Adjust the brightness and / or contrast of the halo to more clearly display the region of interest in the reference sample image. In practical applications, the halos in the reference sample image are uniform, and the region of interest is the region where the facial features of the target object are more prominent. The present invention is not limited thereto.
[0053] (2) Replace the region of interest in the original reference sample image with the adjusted region of interest to generate the reference sample image.
[0054] Specifically, in this embodiment, using the expression data enhancement method based on the region of interest, the facial features and chin of the target object, etc., which have more prominent amplitude and frequency characteristics in the study of expression recognition, are set as the region of interest. By replacing the image data of the initial state of the target object with a face image with a more prominent region of interest after adjusting the color space attributes, the reference sample image is generated.
[0055] In the embodiment of the present invention, the color space attributes of the reference sample image are adjusted. By adjusting the color space attributes such as brightness and contrast of the reference sample image, the facial image data of the target object is divided into regions of interest based on the expression data enhancement method in the region of interest. The adjusted region of interest is used to replace the region of interest in the original reference sample image to generate the reference sample image. By adjusting the color space attributes of the original reference sample image, the influence of light on expression recognition is reduced, and further, the accuracy and scientificity of expression recognition in the process of assisting interrogation in the present invention are improved.
[0056] In an alternative embodiment of the present invention, the method for adjusting the color space attributes of the reference sample image in step S202 to generate a reference sample image may further include the following steps:
[0057] (1) Input the adjusted reference sample image into the deep generation model, and output a reference sample image whose similarity to the original reference sample image is greater than a preset threshold.
[0058] Specifically, input the adjusted reference sample image into the deep generation model. The deep generation model includes a generator and a discriminator. The generator receives the random noise data of the reference sample image and generates a fake sample image according to the random noise data. The fake sample image refers to a noise image generated based on a random noise in the reference sample image data. The discriminator receives the fake sample image and the original reference sample image, compares the similarity between the fake sample image and the original reference sample image, adjusts the parameters of the fake sample image based on the similarity until the similarity is greater than the preset threshold, and outputs a reference sample image whose similarity to the original reference sample image is greater than the preset threshold.
[0059] In this embodiment, a fake sample image is generated based on the random noise data generated by the deep generation model receiving the adjusted reference sample image. The fake sample image is compared with the original reference sample image by the discriminator, and finally a fake sample image with a similarity greater than the preset threshold is output as the reference sample image. By continuously generating the fake sample image through the deep generation model and continuously discriminating it from the original reference sample image, an image with the same style as the initial state image is finally generated, thereby greatly improving the robustness of the facial expression analysis method of the present invention.
[0060] The above embodiments are only preferred embodiments of the present invention. The present invention can be applied to a variety of different application scenarios, such as in the technical fields of computer vision, social emotion analysis, medical diagnosis, etc., and can also be applied to analyzing the mental state of target objects in various fields.
[0061] Figure 5 A method for analyzing the state of a target object according to an embodiment of the present invention is shown. The method for analyzing the state of the target object specifically includes the following steps:
[0062] S300: Obtain the image information of the target object, collect the audio information of the target object, and determine the correspondence between the image information and the audio information based on the time information;
[0063] S400: Perform facial expression analysis based on a preset facial expression recognition model and the image information to obtain a facial expression analysis result;
[0064] S500: Perform state analysis according to the correspondence between the image information and the audio information and the facial expression analysis result to obtain a state analysis result.
[0065] Specifically, in this embodiment, video recording devices such as law enforcement cameras and cameras are used to obtain the facial expression image information of the target object, and audio information of the target object is collected through audio devices such as microphones. During the entire image acquisition process, different audio information collected at different time stages is corresponded to the image information of the target object at the corresponding stage to establish a corresponding relationship, and the facial expression of the image information is analyzed according to a preset facial expression recognition model. Relying on medical imaging and biometric recognition technologies, new visual image information is formed. According to different curves reflected by the new image information and the corresponding audio information, the state of the target object in different emotions such as aggression, stress, and false reaction is analyzed and recognized to obtain a state analysis result.
[0066] Specifically, the facial expression recognition model used in this embodiment is generated by training using the facial expression analysis method described in the above embodiment. For detailed content, reference can be made to the relevant descriptions of any method embodiment above.
[0067] In this embodiment, by combining the corresponding relationship between the image information and audio information of the target object, the change of the image information of the target object at each time stage during the entire image acquisition process is comprehensively analyzed and recognized, so as to more scientifically and objectively obtain the emotional changes and state changes of the target object at different stages during the entire image acquisition process.
[0068] Figure 6 An optional embodiment of the present invention is shown. The method for performing state analysis according to the corresponding relationship between the image information and audio information and the facial expression analysis result in step S500 above to obtain a state analysis result specifically includes the following steps:
[0069] S501: Based on the corresponding relationship between the image information and audio information, determine the corresponding relationship between the audio information and the facial expression analysis result;
[0070] S502: Based on the corresponding relationship between the audio information and the facial expression analysis result, use three-dimensional control to analyze the facial expression analysis result, calculate and analyze to obtain the state analysis result.
[0071] Specifically, in this embodiment, based on the correspondence between the image information and the audio information, the correspondence between the audio information and the facial expression analysis result is determined. The association between the facial micro-vibration of the target object and the vestibular organ (emotional reflex VER) of the target object is analyzed through three-dimensional control (3D), and the head and neck movements and fluctuations caused by emotional changes are analyzed through three-dimensional control. Frame differences are collected from multiple image data in the changed images, and frames are accumulated simultaneously. A dedicated mathematical formula is used for calculation and the corresponding results are analyzed to achieve digital visualization of each parameter of the mental and physiological states, and technical level classification and discrimination are performed. For example, Figure 7 as shown, this new type of image displays the unique information of the target object, so various mental states can be analyzed, and it is clearer to distinguish that the target object is in mental states such as anger, tension, aggression, etc. In practical applications, the new type of image is an emotion-energy change diagram, which can generate unique information like a color map, a heat map, or an X-ray Figure 1 and can perceive the emotional ups and downs of the target object throughout the image acquisition process. Combining the correspondence between the image information and the audio information, the credibility of the audio information of the target object during the image acquisition process can be inferred therefrom.
[0072] In this embodiment, through three-dimensional control (3D), the correspondence between the audio information and the facial expression analysis result is analyzed to obtain the state analysis result of the target object, so that in the auxiliary interrogation process, objective and accurate analysis conclusions can be generated, and the data generated by the emotion-energy change diagram, such as the excitement and sensitivity parameters, vibration amplitude parameters, comprehensive parameters of the emotion value, lying probability parameters, and emotion change parameters of the target object, can be used. By the changes of these parameters and the vibration image generated by analyzing the amplitude and frequency changes of the facial image of the target object, the emotional changes of the target object can be scientifically judged, and based on the emotional changes, the mental changes of the target object can be judged, which establishes a scientific basis for the association between the audio information and the image information. In practical applications, the mental and physiological changes of the target object are reflected in real time through the vibration graph, and the noise picture of the target object is obtained through a video recording device. The data generated by the emotion-energy change diagram can be one of the above parameter data, or any combination of the above parameter data, or other relevant parameter data, and the present invention is not limited thereto.
[0073] It should be understood that although Figure 1 、 Figure 2 、 Figure 5 、 Figure 6The steps in the flowchart are shown sequentially according to the arrows, but these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 and Figure 2 and Figure 5 and Figure 6 at least some of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least some of the steps or stages in other steps.
[0074] As Figure 8 shown, this embodiment provides a facial expression analysis device, including: an acquisition module 1 and a training module 2, where:
[0075] The acquisition module 1 is used to acquire image data of a target object. For detailed content, reference can be made to the relevant description of step S100 in any of the above method embodiments;
[0076] The training module 2 is used to perform facial expression analysis based on a preset facial expression recognition model and the image data to obtain an analysis result. The preset facial expression recognition model is generated by training based on the initial state image of the target object and sample images including state changes. For detailed content, reference can be made to the relevant description of step S200 in any of the above method embodiments.
[0077] In the embodiments of the present invention, video recording devices such as law enforcement cameras and ordinary cameras are used to collect facial expression image information in real time in an unconscious and non-contact manner. The image data of the target object in the initial state collected is trained with sample images of state changes in each stage during the entire collection process to obtain a facial expression recognition model. Through the analysis of the facial expression recognition model and the image data of the target object, the facial expression change situation of the target object can be accurately obtained, so as to accurately and objectively judge the emotional changes caused by the facial expression changes of the target object.
[0078] As Figure 9 shown, this embodiment of the present invention provides a target object state analysis device, including: a collection module 3, an analysis module 4, and a communication module 5, where:
[0079] The collection module 3 is used to acquire image information of a target object, acquire audio information of the target object, and determine the correspondence between the image information and the audio information based on time information. For detailed content, reference can be made to the relevant description of step S300 in any of the above method embodiments;
[0080] An analysis module 4, configured to perform facial expression analysis based on a preset facial expression recognition model and the image information to obtain a facial expression analysis result. For detailed content, reference can be made to the relevant description in step S400 of any of the above method embodiments;
[0081] A communication module 5, configured to perform state analysis according to the correspondence between the image information and the audio information and the facial expression analysis result to obtain a state analysis result. For detailed content, reference can be made to the relevant description in step S500 of any of the above method embodiments.
[0082] In this embodiment, by combining the correspondence between the image information and the audio information of the target object, the change situation of the image information of the target object at each time stage during the entire process of assisted interrogation is comprehensively analyzed and identified, so as to more scientifically and objectively obtain the emotional changes and state changes of the target object at different stages during the interrogation process, and at the same time improve the interrogation efficiency.
[0083] For the specific limitations and beneficial effects of the facial expression analysis device and the target object state analysis device, reference can be made to the limitations of the facial expression analysis method and the target object state analysis method in the above text, which will not be elaborated here. Each module of the above facial expression analysis device and the target object state analysis device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of the processor in the electronic device in the form of hardware, or stored in the memory in the electronic device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above modules.
[0084] An embodiment of the present invention further provides a computer device, such as Figure 10 shown, Figure 10 is a schematic structural diagram of a computer device provided by an optional embodiment of the present invention. The computer device may include at least one processor 41, at least one communication interface 42, at least one communication bus 43, and at least one memory 44. Among them, the communication interface 42 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the communication interface 42 may further include a standard wired interface and a wireless interface. The memory 44 may be a high-speed RAM memory (Random Access Memory, volatile random access memory), or a non-volatile memory, such as at least one disk memory. Optionally, the memory 44 may further be at least one storage device located far from the foregoing processor 41. Wherein the processor 41 may be combined with Figure 8 , Figure 9For the described device, an application program is stored in the memory 44, and the processor 41 calls the program code stored in the memory 44 to execute the steps of the facial expression analysis method and the target object state analysis method in any of the above method embodiments.
[0085] Among them, the communication bus 43 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus 43 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 10 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0086] Among them, the memory 44 can include volatile memory, such as random-access memory (RAM); the memory can also include non-volatile memory, such as flash memory, hard disk drive (HDD) or solid-state drive (SSD); the memory 44 can also include a combination of the above types of memories.
[0087] Among them, the processor 41 can be a central processing unit (CPU), a network processor (NP), or a combination of a CPU and an NP.
[0088] Among them, the processor 41 may further include a hardware chip. The above-mentioned hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above-mentioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0089] Optionally, the memory 44 is further configured to store program instructions. The processor 41 may call the program instructions to implement the methods described in any embodiment of the present invention.
[0090] An embodiment of the present invention further provides a non-transitory computer storage medium. The computer storage medium stores computer-executable instructions, and the computer-executable instructions can execute the methods described in any of the above method embodiments. Among them, the storage medium may be a magnetic disk, an optical disc, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium may also include a combination of the above types of memories.
[0091] Obviously, the above embodiments are merely examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or variations derived therefrom are still within the protection scope of the present invention.
Claims
1. A method for facial expression analysis, characterized in that, it includes: Obtain the image data of the target object; Perform facial expression analysis based on a preset facial expression recognition model and the image data to obtain an analysis result, and the preset facial expression recognition model is generated by training based on the initial state image of the target object and sample images including state changes; Among them, the process of training and generating the preset facial expression recognition model includes: Obtain the initial state image of the target object as a reference sample image; Adjust the color space attributes of the reference sample image to generate a reference sample image; Obtain the change image of the target object in the emotional change stage; Associate the state change of the target object based on the amplitude and frequency of the image information in the change image to generate a comparison sample image; Train a neural network model based on the reference sample image and the comparison sample image to generate the preset facial expression recognition model; The adjusting the color space attributes of the reference sample image to generate a reference sample image includes: Adjust the brightness and / or contrast of the region of interest part of the reference sample image; Replace the region of interest part in the original reference sample image with the adjusted region of interest part to generate the reference sample image.
2. The facial expression analysis method according to claim 1, characterized in that, the adjusting the color space attributes of the reference sample image to generate a reference sample image further includes: Input the adjusted reference sample image into a deep generation model, and output a reference sample image whose similarity to the adjusted reference sample image is less than a preset threshold; The training process of the deep generation model includes: Through the generator of the deep generation model, generate a fake sample image according to random noise and the adjusted reference sample image; Through the discriminator of the deep generation model, judge the similarity between the fake sample image and the original reference sample image; Adjust the parameters of the deep generation model based on the judgment result until the similarity between the generated fake sample image and the original reference sample image is greater than a preset threshold.
3. A method for analyzing the state of a target object, characterized in that, it includes: Obtain the image information of the target object, collect the audio information of the target object, and determine the correspondence between the image information and the audio information based on time information; Perform facial expression analysis based on a preset facial expression recognition model and the image information to obtain a facial expression analysis result, and the preset facial expression recognition model is generated by using the facial expression analysis method described in claim 1 or 2; Perform state analysis based on the correspondence between the image information and the audio information and the facial expression analysis result to obtain a state analysis result.
4. The method for analyzing the state of a target object according to claim 3, characterized in that, the performing state analysis based on the correspondence between the image information and the audio information and the facial expression analysis result to obtain a state analysis result includes: Determine the correspondence between the audio information and the facial expression analysis result based on the correspondence between the image information and the audio information; Based on the correspondence between the audio information and the facial expression analysis result, use three-dimensional control to analyze the facial expression analysis result, and calculate and analyze to obtain the state analysis result.
5. A facial expression analysis device, Characterized in that, Comprising: An acquisition module, configured to acquire image data of a target object; A training module, configured to perform facial expression analysis based on a preset facial expression recognition model and the image data to obtain an analysis result, where the preset facial expression recognition model is generated by training based on an initial state image of the target object and a sample image including state changes; Among them, the process of training and generating the preset facial expression recognition model includes: Acquire the initial state image of the target object as a reference sample image; Adjust the color space attributes of the reference sample image to generate a reference sample image; Acquire the change image of the target object in the emotional change stage; Based on the amplitude and frequency of the image information in the change image, associate the state change of the target object to generate a comparison sample image; Train a neural network model based on the reference sample image and the comparison sample image to generate the preset facial expression recognition model; The adjusting the color space attributes of the reference sample image to generate a reference sample image includes: Adjust the brightness and / or contrast of the region of interest part of the reference sample image; Use the adjusted region of interest part to replace the region of interest part in the original reference sample image to generate the reference sample image.
6. A target object state analysis device, Characterized in that, Comprising: An acquisition module, configured to acquire image information of a target object, acquire audio information of the target object, and determine the correspondence between the image information and the audio information based on time information; An analysis module, configured to perform facial expression analysis based on a preset facial expression recognition model and the image information to obtain a facial expression analysis result, where the preset facial expression recognition model is generated by training using the facial expression analysis method according to claim 1 or 2; A communication module, configured to perform state analysis according to the correspondence between the image information and the audio information and the facial expression analysis result to obtain a state analysis result.
7. A computer device, Characterized in that, Comprising: A communication unit, a memory, and a processor, where the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the steps of the facial expression analysis method according to claim 1 or 2, or execute the steps of the target object state analysis method according to claim 3 or 4.
8. A computer-readable storage medium, Characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the steps of the method according to claim 1 or 2, or for causing the computer to execute the steps of the target object state analysis method according to claim 3 or 4.
Citation Information
Patent Citations
Method and device for realizing face micro-expression change recognition based on AI technology
CN111178151A
Emotion early warning method based on camera
CN113143274A
Emotion detection method and apparatus, electronic device, and storage medium
WO2020248376A1