Emotion detection method, device, computer equipment and storage medium
Through multi-angle monitoring image fitting, the user's three-dimensional spatial position is determined, and the emotion score is combined with image and recording information, which solves the problem of high equipment costs and realizes low-cost emotion detection.
Patent Information
- Application Number
- CN202310101312.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-20
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2043-01-20
AI Technical Summary
Existing emotion detection methods require the arrangement of depth of field cameras, resulting in high equipment upgrade costs.
By acquiring multiple monitoring images for image fitting, determining the user's position information in three-dimensional space, combining behavioral actions and recording information, using image prediction models and speech prediction networks for emotion scores, reducing equipment costs.
It realizes the accurate evaluation of user emotions in three-dimensional space while reducing the cost of emotion detection.
Smart Images

Figure CN115990017B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to an emotion detection method, apparatus, computer equipment, storage medium, and computer program product. Background Art
[0002] When performing services for users, it's crucial to constantly monitor their emotions so that we can determine their satisfaction with the services. This allows us to adjust the service execution process and methods based on user satisfaction. Currently, emotion detection typically involves using depth-of-field cameras to collect user profile information at various locations and then analyzing this information to determine the user's emotions. However, current emotion detection requires deploying depth-of-field cameras, which incurs high equipment upgrade costs.
[0003] Therefore, current emotion detection methods have the defect of high detection cost. Summary of the Invention
[0004] Based on this, it is necessary to provide an emotion detection method, apparatus, computer equipment, computer-readable storage medium and computer program product that can reduce detection costs in response to the above technical problems.
[0005] In a first aspect, the present application provides an emotion detection method, the method comprising:
[0006] Obtaining at least two to-be-detected disparity images for an area to be detected, performing image fitting on the at least two to-be-detected disparity images to obtain a target disparity image for the area, and determining position information of a user to be detected in the area in three-dimensional space based on the target disparity image; the to-be-detected disparity image is determined based on multiple surveillance images of the area;
[0007] Acquiring behavioral action features of the user to be detected based on the location information and the multiple surveillance images;
[0008] Inputting the behavior action features into an image prediction model, and having the image prediction model output a behavior action emotion score of the user to be detected based on the behavior action features;
[0009] Acquire recording information in the area, and determine target recording information corresponding to the user to be detected from the recording information based on the location information; input the target recording information into a speech prediction network, and obtain a speech emotion score of the user to be detected output by the speech prediction network;
[0010] The emotion score of the user to be detected is determined according to the behavior emotion score and the voice emotion score.
[0011] In one embodiment, the step of acquiring at least two parallax images to be detected for the area to be detected includes:
[0012] Acquiring at least three surveillance images captured by at least three image acquisition devices;
[0013] Determining baseline distance information between a plurality of image pairs in the at least three surveillance images, wherein each image pair includes two surveillance images acquired by different image acquisition devices;
[0014] At least two to-be-detected parallax images are determined according to at least two groups of image pairs whose baseline distance information satisfies a preset distance condition.
[0015] In one embodiment, determining baseline distance information between multiple image pairs in the at least three surveillance images includes:
[0016] Obtaining coordinate values of the at least three image acquisition devices in the area, and an angle between directions of every two image acquisition devices;
[0017] For every two image acquisition devices, the corresponding Euclidean distance is determined based on the coordinate values of the two image acquisition devices; based on the preset optimal baseline distance, the Euclidean distance, the angle between the orientations of the two image acquisition devices and the preset optimal angle, the baseline distance information corresponding to the two image acquisition devices is determined, and the baseline distance information corresponding to the two image acquisition devices is used as the baseline distance information between the image pairs corresponding to the two image acquisition devices.
[0018] In one embodiment, determining at least two to-be-detected disparity images based on at least two image pairs whose baseline distance information satisfies a preset distance condition includes:
[0019] Acquire two first target image acquisition device sets corresponding to the first minimum baseline distance information, and acquire a second target image acquisition device set including any one target image acquisition device in the first target image acquisition device set and corresponding to the second minimum baseline distance information;
[0020] Determine a first disparity image based on a plurality of monitoring images corresponding to the first target image acquisition device set, and determine a second disparity image based on a plurality of monitoring images corresponding to the second target image acquisition device set;
[0021] At least two to-be-detected disparity images are obtained according to the first disparity image and the second disparity image.
[0022] In one embodiment, obtaining the behavioral characteristics of the user to be detected based on the location information and the multiple surveillance images includes:
[0023] Acquire, from the plurality of surveillance images, a user image corresponding to the user to be detected according to the location information;
[0024] Acquire posture features corresponding to the user to be detected in at least two frames of the user image, and posture optical flow features corresponding to the posture features;
[0025] Obtaining gait features corresponding to the user image; obtaining action features based on the posture features, the posture optical flow features, and the gait features;
[0026] Obtaining facial information corresponding to at least two frames of the user image, and extracting expression features corresponding to the facial information and expression optical flow features corresponding to the expression features;
[0027] The behavioral action features of the user to be detected are obtained according to the action features, the expression features and the expression optical flow features.
[0028] In one embodiment, the image prediction model includes a posture prediction network, a gait prediction network, and an expression prediction network;
[0029] Inputting the behavior action features into an image prediction model, and having the image prediction model output a behavior action emotion score of the user to be detected based on the behavior action features, includes:
[0030] Inputting the posture feature and the posture optical flow feature into the posture prediction network, and outputting a first posture emotion score based on the posture feature of each frame and a second posture emotion score based on the posture optical flow feature through the posture prediction network;
[0031] Inputting the gait features into a gait prediction network, and outputting a gait emotion score through the gait prediction network;
[0032] Obtaining an action emotion score according to the first posture emotion score, the second posture emotion score, and the gait emotion score;
[0033] Inputting the facial expression features and the facial expression optical flow features into an facial expression prediction network, and outputting a first facial expression emotion score based on the facial expression features and a second facial expression emotion score based on the facial expression optical flow features through the facial expression prediction network;
[0034] The behavioral action emotion score is obtained according to the action emotion score, the first expression emotion score and the second expression emotion score.
[0035] In one embodiment, determining target recording information corresponding to the user to be detected from the recording information according to the location information includes:
[0036] Obtaining mouth shape features and gender features corresponding to the user to be detected based on the user image and the facial information corresponding to the user image;
[0037] Determine the user corresponding to each frame of recording information based on the Euclidean distance between the voice position of each frame of recording information and the position of the mouth shape feature of the corresponding frame, and the Euclidean distance between the voice position of each frame of recording information and the position of the gender feature of the corresponding frame, and obtain the recording information of each user in the area to be detected;
[0038] According to the location information, target recording information corresponding to the user to be detected is determined from the recording information of each user.
[0039] In one embodiment, determining the emotion score of the user to be detected based on the behavior emotion score and the voice emotion score includes:
[0040] Sorting multiple behavioral emotion scores and multiple voice emotion scores corresponding to the user to be detected in chronological order to obtain an emotion score sequence corresponding to the user to be detected; each behavioral emotion score and each voice emotion score is determined based on at least one frame of the user image to be detected;
[0041] The emotion score sequence is input into an emotion score network based on long short-term memory, and an emotion score of the user to be detected is obtained after the emotion score network performs weighted superposition on multiple action emotion scores and multiple voice emotion scores.
[0042] In one embodiment, the method further comprises:
[0043] Acquire multiple frames of continuous disparity image samples and multiple frames of continuous audio recording samples within a preset time period, and input the multiple frames of continuous disparity image samples into the image prediction model to be trained;
[0044] Training a tracking and re-identification network in the image prediction model to be trained based on the multiple frames of continuous disparity image samples to obtain first model parameters corresponding to the tracking and re-identification network when training is completed; the tracking and re-identification network is used to determine position information of the user to be detected in three-dimensional space;
[0045] Training a behavior action prediction network in the image prediction model to be trained based on the multiple frames of continuous disparity image samples, and obtaining second model parameters corresponding to the behavior action prediction network when training is completed;
[0046] Obtaining the image prediction model according to the first model parameters and the second model parameters;
[0047] The speech prediction network to be trained is trained according to the multiple frames of continuous recording samples until the speech prediction network is obtained when the training conditions are met.
[0048] In a second aspect, the present application provides an emotion detection device, comprising:
[0049] a fitting module, configured to obtain at least two to-be-detected disparity images for an area to be detected, perform image fitting on the at least two to-be-detected disparity images to obtain a target disparity image for the area, and determine, based on the target disparity image, position information of a user to be detected in the area in three-dimensional space; the to-be-detected disparity image is determined based on multiple surveillance images of the area;
[0050] an acquisition module, configured to acquire the behavioral action features of the user to be detected based on the location information and the plurality of surveillance images;
[0051] A first detection module is configured to input the behavior action features into an image prediction model, and the image prediction model outputs a behavior action emotion score of the user to be detected based on the behavior action features;
[0052] a second detection module configured to obtain recording information in the area and determine target recording information corresponding to the user to be detected from the recording information based on the location information; input the target recording information into a speech prediction network, and obtain a speech emotion score of the user to be detected output by the speech prediction network;
[0053] The output module is used to determine the emotion score of the user to be detected based on the behavior emotion score and the voice emotion score.
[0054] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0055] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when executed by a processor.
[0056] In a fifth aspect, the present application provides a computer program product, comprising a computer program, which implements the steps of the above method when executed by a processor.
[0057] The above-mentioned emotion detection method, apparatus, computer equipment, storage medium and computer program product obtain a target disparity image of the area to be detected by fitting at least two disparity maps of the area to be detected, thereby determining the user's position information in three-dimensional space, determining the user's behavior and action characteristics based on the position information and multiple monitoring images, performing a behavior and action emotion score on the behavior and action characteristics based on an image prediction model, and performing a voice emotion score on the user's recording information based on a voice prediction network to obtain the user's emotion score. Compared with the traditional method of collecting and analyzing the characteristic information of users at various positions through a depth-of-field camera, this solution fits disparity maps through monitoring images from multiple angles to determine the user's three-dimensional spatial position, thereby extracting the user's image features from the image based on the position and determining the recording of the user at the position, and then performing an emotion score on the user's behavior and action and voice, thereby reducing the cost of emotion detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 A diagram showing an application environment of an emotion detection method according to an embodiment;
[0059] Figure 2 1 is a flow chart of an emotion detection method according to an embodiment;
[0060] Figure 3 Schematic diagram of the fitting step in one embodiment;
[0061] Figure 4 is a flowchart of an emotion detection method according to another embodiment;
[0062] Figure 5 is a structural block diagram of an emotion detection device in one embodiment;
[0063] Figure 6 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0064] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0065] The emotion detection method provided in the embodiment of the present application can be applied to Figure 1In the application environment shown, terminal 102 communicates with image acquisition device 104 located in the area to be detected via a network. Terminal 102 can obtain multiple surveillance images captured by the image acquisition device in the area to be detected, and generate at least two parallax images to be detected based on the multiple surveillance images. Based on the at least two parallax images, the user's position in three-dimensional space is determined, and by detecting the user's behavioral characteristics and recording information, corresponding emotion scores are output, thereby obtaining the user's emotion score. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, and tablet computers.
[0066] In one embodiment, Figure 2 As shown, a method for emotion detection is provided, which is applied to Figure 1 The following steps are used as an example to illustrate the terminal in the figure:
[0067] Step S202: Obtain at least two disparity images to be detected for the area to be detected, perform image fitting on the at least two disparity images to be detected to obtain a target disparity image of the area, and determine the position information of the user to be detected in the area in three-dimensional space based on the target disparity image; the disparity image to be detected is determined based on multiple monitoring images of the area.
[0068] The area to be detected can be an area where a user is located, such as a business location. The area can contain multiple users, and the terminal can identify the user to be detected from among them. The user to be detected can also be a user for whom emotion detection is required. The terminal can determine the user's satisfaction with the business process by detecting the user's emotions. To more accurately construct feature information of the user to be detected in the three-dimensional world, the terminal can first determine the user's position in three-dimensional space. The terminal can obtain at least two parallax images to be detected for the area to be detected. The parallax images can be obtained based on images of the same target from different perspectives. For example, the area can be provided with multiple image acquisition devices, which can be installed at different locations within the area and facing different directions. Each image acquisition device can capture a corresponding surveillance image, allowing the terminal to obtain multiple surveillance images of the area. These surveillance images can be images of the area taken from different angles. Since the area contains multiple users, the terminal can obtain images of each user from different perspectives, thereby determining at least two parallax images to be detected based on the multiple surveillance images. Each parallax image to be detected can be obtained by fitting multiple surveillance images from different angles.
[0069] After obtaining the at least two disparity images to be detected, the terminal can perform image fitting on the at least two disparity images to be detected, thereby obtaining the target disparity image of the above-mentioned area. The terminal can input the at least two disparity images to be detected into the target fitting model and obtain the target disparity image output by the target fitting model. Specifically, Figure 3 As shown, Figure 3 The figure is a flow chart of the fitting step in one embodiment. Taking the case that there are two disparity images to be detected, the terminal can obtain the two disparity images to be detected by a triplet disparity calculation method. The triplet disparity calculation method can be a method for obtaining the above-mentioned disparity images to be detected after inverse mapping based on three surveillance images from different perspectives. After the terminal obtains the above-mentioned two disparity images to be detected, both of the two disparity images to be detected are relatively rough disparity estimates of the same scene. Therefore, the terminal can input the above-mentioned two disparity images to be detected into a pre-trained target fitting model to fit the two disparity images to be detected, thereby obtaining a more accurate target disparity image. The terminal can use the real disparity map as a training label to supervise the training of the fitting model to obtain the above-mentioned target fitting model. The real disparity map can be a disparity map of the three-dimensional spatial position of each user in the known map. The terminal can compare the predicted position information of each user in the real disparity map output by the fitting model to be trained in each training with the corresponding known three-dimensional spatial position, and adjust the model parameters of the fitting model to be trained based on the similarity until the similarity exceeds a preset threshold or the number of training times reaches a preset number, thereby obtaining a target fitting model that has completed training.
[0070] Step S204: obtaining the behavior and action features of the user to be detected based on the location information and the multiple surveillance images.
[0071] The location information may be the location information of the user to be detected in three-dimensional space. After determining the location information of the user to be detected, the terminal may combine the multiple surveillance images to obtain the behavioral characteristics of the user to be detected. The behavioral characteristics include the user's motion characteristics, expression characteristics, and expression optical flow characteristics. Motion characteristics represent characteristics generated by the user's body movements. Expression characteristics represent characteristics extracted from the user's expression in each surveillance image frame. Expression optical flow characteristics may represent characteristics generated by the continuous changes in the user's expression. The terminal may determine the location of the user to be detected in each surveillance image using the location information and extract the user image at that location. The terminal may then extract the behavioral characteristics from the user image. The multiple surveillance images may be images captured by multiple image acquisition devices with different perspectives. The images from each image acquisition device may be multiple frames of images over a period of time. For each of the multiple frames of images from each image acquisition device, the terminal may select the user to be detected using a rectangular frame to obtain multiple frames of user images. The behavioral characteristics may then be extracted based on the multiple frames of user images.
[0072] Step S206 , inputting the behavior action features into the image prediction model, and the image prediction model outputs the behavior action emotion score of the user to be detected based on the behavior action features.
[0073] Among them, the image prediction model can be a model for predicting the emotional score of the user to be detected in the image, and the image prediction model can be obtained by training the image prediction model to be trained based on multiple continuous disparity image samples. The terminal can input the above-mentioned behavioral action features into the trained image prediction model, and the image prediction model recognizes the behavioral action features, and after determining the behavioral action emotional score corresponding to the user to be detected based on the above-mentioned each behavioral action feature, it outputs the above-mentioned behavioral action emotional score. Among them, the behavioral action emotional score represents the score of the emotion represented by the behavioral action of the user to be detected. The higher the behavioral action score, the more positive the emotion of the user to be detected, and thus it can be determined that the user to be detected has a higher satisfaction with the service; conversely, the lower the behavioral action score, the more negative the emotion of the user to be detected, and thus it can be determined that the user to be detected has a lower satisfaction with the service.
[0074] Step S208: Acquire the recording information in the area, and determine the target recording information corresponding to the user to be detected from the recording information based on the location information; input the target recording information into the speech prediction network, and obtain the speech emotion score of the user to be detected output by the speech prediction network.
[0075] The terminal can also obtain the recording information in the above-mentioned area, and the recording information may include the human voice information emitted by each user in the above-mentioned area. The terminal can collect the sound information in the area by deploying a multi-microphone array in the above-mentioned area, and obtain the spatial coordinates of the sound source in the real world through the sound localization algorithm. The terminal can determine the target recording information of the user to be detected from the recording information based on the above-mentioned position information. For example, the terminal can determine the three-dimensional spatial position of the user to be detected in the image based on the above-mentioned position information, and extract the mouth shape features and gender features of the user to be detected from the monitoring image based on the position, and determine the target recording information corresponding to the user to be detected by matching the mouth shape features and gender features of multiple frames of user images with multiple frames of recording information. For example, the terminal can match the above-mentioned position information with the spatial coordinates determined based on the sound localization algorithm, and determine the speaker corresponding to the above-mentioned position in combination with the mouth shape features and gender features, thereby determining the target recording information of the user to be detected.
[0076] After the terminal obtains the above-mentioned target recording information, it can input the above-mentioned target recording information into the speech prediction network. The speech prediction network can be obtained by training the speech prediction network to be trained based on multiple frames of continuous recording samples and corresponding position information samples. For example, the above-mentioned recording samples can be samples of the speech emotion scores of known users. The terminal can input the above-mentioned recording samples into the speech prediction network to be trained, and the speech prediction network can identify the predicted speech emotion score corresponding to the user's recording sample based on the above-mentioned recording samples, and output the predicted speech emotion score, so that the terminal can compare the predicted speech emotion score output by the speech prediction network to be trained with the speech emotion score sample corresponding to the above-mentioned recording sample, adjust the model parameters of the speech prediction network to be trained according to the similarity obtained by the comparison, and conduct the next training until the similarity of the above-mentioned comparison is greater than the preset similarity threshold, or the number of training times reaches the preset number, and obtain a trained speech prediction network.
[0077] Step S210 , determining the emotion score of the user to be detected based on the action emotion score and the voice emotion score.
[0078] After obtaining the behavioral emotion score based on the image prediction model and the voice emotion score based on the voice prediction network, the terminal can determine the emotion score of the user to be detected based on the behavioral emotion score and the voice emotion score. Since the user image may include multiple frames of user images, the behavioral emotion score and the voice emotion score will also include scores corresponding to the multiple frames of user images. Therefore, the terminal can determine the final emotion score of the user to be detected by weighted superposition.
[0079] For example, in one embodiment, the above-mentioned multiple frames of user images may be user images in consecutive frames in a surveillance video, and the terminal may sort the multiple behavioral action emotion scores and multiple voice emotion scores corresponding to the user to be detected in chronological order to obtain an emotion score sequence corresponding to the user to be detected. Wherein, each of the above-mentioned behavioral action emotion scores and each voice emotion score is determined based on at least one frame of the user image to be detected, and the multiple behavioral action scores and multiple voice emotion scores in the above-mentioned emotion score sequence may be scores corresponding to consecutive frame user images. The terminal may input the above-mentioned emotion score sequence into an emotion score network based on LSTM (Long short-term memory), and obtain the emotion score of the user to be detected outputted by the emotion score network after weighted superposition of the multiple behavioral action emotion scores and the multiple voice emotion scores.
[0080] Specifically, the aforementioned emotion scores for each action and each voice can be a satisfaction prediction indicator. The terminal sorts the emotion scores for each action and voice at a single moment in chronological order, thereby obtaining the aforementioned emotion score sequence and inputting it into the emotion scoring network, which is also referred to as a satisfaction detection module. The emotion scoring network can be composed of a stack of LSTM units. The emotion scoring network predicts the emotion scores for each action and voice, and the prediction results are passed through a softmax layer to obtain the final predicted satisfaction result, namely, the emotion score of the user to be tested. In other words, the emotion score can be used as a user satisfaction evaluation of the service. A higher emotion score indicates a higher user satisfaction with the service, while a lower emotion score indicates a lower user satisfaction with the service.
[0081] In the above-mentioned emotion detection method, the target disparity image of the area to be detected is obtained by fitting at least two disparity maps of the area to be detected, thereby determining the user's position information in three-dimensional space, determining the user's behavioral action characteristics based on the position information and multiple monitoring images, performing a behavioral action emotion score on the behavioral action characteristics based on the image prediction model, and performing a voice emotion score on the user's recording information based on the voice prediction network to obtain the user's emotion score. Compared with the traditional method of collecting and analyzing the characteristic information of users at various positions through a depth-of-field camera, this solution fits disparity maps through monitoring images from multiple angles to determine the user's three-dimensional spatial position, thereby extracting the user's image features from the image based on the position and determining the recording of the user at the position, and then performing an emotion score on the user's behavior, action, and voice, thereby reducing the cost of emotion detection.
[0082] In one embodiment, obtaining at least two parallax images to be detected for the area to be detected includes: obtaining at least three surveillance images captured by at least three image acquisition devices; determining baseline distance information between multiple groups of image pairs in the at least three surveillance images; each group of image pairs includes two surveillance images captured by different image acquisition devices; and determining at least two parallax images to be detected based on at least two groups of image pairs whose baseline distance information satisfies a preset distance condition.
[0083] In this embodiment, at least three image acquisition devices may be installed in the area, each capable of capturing corresponding surveillance images. The terminal can then obtain at least three surveillance images. Furthermore, the image acquisition devices may be installed in different positions and orientations, so the at least three surveillance images obtained by the terminal may be from different perspectives. The terminal may assemble a set of image pairs based on two surveillance images captured by different image acquisition devices. The terminal may then assemble multiple sets of image pairs based on the at least three image acquisition devices. The terminal may determine baseline distance information between multiple sets of image pairs within the at least three surveillance images. This baseline distance information may be baseline distance information between two different image acquisition devices within each set of image pairs. Each image acquisition device may have corresponding coordinate values in three-dimensional space and may have different orientations. The terminal may determine the baseline distance information based on these coordinate values and orientations. For example, in one embodiment, the terminal may obtain the coordinate values of the at least three image acquisition devices within the area. These coordinate values may be the coordinate values of each image acquisition device in three-dimensional space. Furthermore, the terminal may obtain the angle between the orientations of each pair of image acquisition devices. There may be multiple image acquisition devices. For each pair of image acquisition devices, the terminal can determine the corresponding Euclidean distance based on the coordinate values of the two image acquisition devices. The terminal can determine a preset optimal baseline distance and a preset optimal angle by conducting field measurements in the aforementioned area. Thus, the terminal can determine the baseline distance information corresponding to the two image acquisition devices based on the preset optimal baseline distance, the Euclidean distance, the angle between the orientations of the two image acquisition devices, and the preset optimal angle. The terminal can use the baseline distance information corresponding to the two image acquisition devices as the baseline distance information between the image pairs corresponding to the two image acquisition devices.
[0084] Thus, the terminal can filter out at least two groups of image pairs whose baseline distance information meets the preset distance condition, determine a parallax image to be detected based on each group of image pairs, and then the terminal can determine at least two parallax images to be detected. Among them, the at least two groups of image pairs can include the same surveillance image, that is, the at least two groups of image acquisition devices corresponding to the at least two groups of image pairs have the same image acquisition device. Taking two groups of image pairs as an example, the two groups of image pairs can include surveillance images of three different perspectives acquired by three different image acquisition devices, and the terminal can determine at least two parallax images to be detected based on the three different surveillance images. For example, taking two parallax images as an example, in one embodiment, the terminal can filter out the minimum baseline distance information and obtain two first target image acquisition device sets corresponding to the first minimum baseline distance information. The terminal can also obtain other image acquisition device sets including any one target image acquisition device in the first target image acquisition device set from the remaining image acquisition device sets, and filter out the second target image acquisition device set whose baseline distance information is the second minimum baseline distance information from the other image acquisition device sets. Thus, the terminal can determine the first disparity image based on the multiple surveillance images corresponding to the first target image acquisition device set, and determine the second disparity image based on the multiple surveillance images corresponding to the second target image acquisition device set. After obtaining the first and second disparity images, the terminal can obtain at least two to-be-detected disparity images based on the first and second disparity images. For example, the first disparity image can be used as one of the to-be-detected disparity images, and the second disparity image can be used as the other to-be-detected disparity image.
[0085] Specifically, disparity calculation is crucial for accurately determining the user's depth information within an image. The terminal can determine the user's depth information, i.e., the user's position in three-dimensional space, using a triplet disparity calculation method. This process requires first selecting three images with the closest baseline distances under three different perspectives in the same scene. This involves selecting the first target image acquisition device set and the second target image acquisition device set. The terminal can then select the two closest images from the three surveillance images and calculate a disparity map using an inverse mapping method. Using the same method, the terminal can then select the two closest images from the remaining surveillance images, formed by combining any one of these two closest images with the other surveillance images, and generate a disparity map based on these two images. These two closest images represent the two closest images excluding the two images corresponding to the first disparity map generated. The terminal can then input these two different disparity maps into a trained target fitting model to fit them into a final target disparity image, thereby obtaining the user's depth information.
[0086] The calculation formula for the terminal to determine the two images with the closest baseline distance is as follows: Among them, z is the index value, B is the optimal baseline distance, D is the Euclidean distance, α represents the influence weight coefficient, θ represents the camera's orientation angle, x and y represent the plane coordinates of the camera in the venue, β represents the angle between the optimal viewing angles formed by the orientations of the two cameras, and p represents the temperature coefficient used to control the degree of value change. The above-mentioned image acquisition device can be a camera. The terminal can select the monitoring images taken by the two groups of cameras with the largest z values as the left and right images with the closest baseline distances, and solve the corresponding disparity map according to the inverse mapping method. For example, each group has two cameras, and one of the two cameras in each group is the same, that is, the terminal can select the nearest three cameras A, B, and C. Assuming that camera B is the central camera, the terminal can calculate the disparity map determined by the three groups of cameras A, B, and C, and then calculate the disparity of A, B and B, C respectively. Figure 1 、 2 , and then fit the final disparity map. Different z values exist in different combinations. Assuming all cameras are a set, after finding A and B with B as the center, A will be excluded from the current set. Then, the z values of B and C are found, thus obtaining two disparity maps to be detected.
[0087] Through the above embodiment, the terminal can determine the two groups of image pairs with the closest baseline distances based on the baseline distance information between multiple groups of image pairs, and obtain two disparity images to be detected based on the two groups of image pairs and the inverse mapping method, so that the terminal can determine the user's position information in three-dimensional space based on the two disparity images to be detected, thereby reducing the cost of emotion detection on the user.
[0088] In one embodiment, behavioral action features of a user to be detected are obtained based on location information and multiple surveillance images, including: obtaining a user image corresponding to the user to be detected from multiple surveillance images based on the location information; obtaining posture features corresponding to the user to be detected in at least two frames of user images, and posture optical flow features corresponding to the posture features; obtaining gait features corresponding to the user image; obtaining action features based on the posture features, the posture optical flow features, and the gait features; obtaining facial information corresponding to at least two frames of user images, and extracting expression features corresponding to the facial information and expression optical flow features corresponding to the expression features; obtaining behavioral action features of the user to be detected based on the action features, the expression features, and the expression optical flow features.
[0089] In this embodiment, the terminal can obtain the behavioral action features of the user to be detected from the above-mentioned surveillance images based on the location information. For example, the terminal can first obtain the user image corresponding to the user to be detected from multiple surveillance images based on the location information. The above-mentioned surveillance images may include surveillance images of multiple time points, and the user image may include multiple frames. The terminal can obtain the posture features corresponding to the user to be detected in at least two frames of user images, and obtain the posture optical flow features corresponding to at least two posture features. Among them, the posture feature represents the feature corresponding to the user's posture, and the posture optical flow feature represents the feature generated by the change of the user's posture in at least two consecutive frames of user images. The terminal can also obtain the gait feature corresponding to the user image. Among them, the gait feature represents the feature generated by the user's gait. The terminal can obtain the action feature based on the posture feature, the posture optical flow feature and the gait feature.
[0090] The terminal can also obtain facial information corresponding to at least two frames of user images, and extract expression features corresponding to the facial information, as well as expression optical flow features corresponding to the expression features. Expression features represent features corresponding to the expression in the facial information of each frame of the user image, and expression optical flow features represent features resulting from changes in expression across at least two consecutive frames of user image information. Thus, the terminal can derive behavioral action features of the user to be detected based on these action features, expression features, and expression optical flow features.
[0091] Specifically, the terminal can use the location information to generate a positioning prediction frame corresponding to the user's position in the image. The terminal then extracts features using the positioning prediction frame. The positioning prediction frame can capture the user's posture information in each frame of the user image, and the user's posture in the current frame is compared with the posture in the previous frame to calculate optical flow information to obtain posture optical flow features. The terminal can also use the positioning detection frame to capture the user's ROI (Region of Interest) and perform face recognition on the ROI region using a face recognition module to obtain the face area. However, since the actual captured face is small and difficult to extract features, the terminal can reconstruct the face area using a super-resolution network to improve the resolution of the face area and thus obtain facial information. The terminal can then perform expression recognition on the reconstructed face area to extract expression features. It can also calculate expression optical flow information based on the expression features of the current frame and the expression features of the previous frame. In addition, since real scenes may have blind spots where faces cannot be captured, to improve the accuracy of satisfaction detection in scenes lacking facial features, the terminal can also determine the user's emotional score by detecting gait features.
[0092] Through this embodiment, the terminal can determine the user's various behavioral action features based on location information and multiple frames of user images, and then the terminal can determine the user's emotion score based on these behavioral action features, reducing the cost of emotion scoring.
[0093] In one embodiment, behavioral action features are input into an image prediction model, and the image prediction model outputs a behavioral action emotion score of the user to be detected based on the behavioral action features, including: inputting posture features and posture optical flow features into a posture prediction network, outputting a first posture emotion score based on the posture features of each frame through the posture prediction network, and outputting a second posture emotion score based on the posture optical flow features; inputting gait features into a gait prediction network, and outputting a gait emotion score through the gait prediction network; obtaining an action emotion score based on the first posture emotion score, the second posture emotion score, and the gait emotion score; inputting expression features and expression optical flow features into an expression prediction network, outputting a first expression emotion score based on the expression features through the expression prediction network, and outputting a second expression emotion score based on the expression optical flow features; and obtaining a behavioral action emotion score based on the action emotion score, the first expression emotion score, and the second expression emotion score.
[0094] In this embodiment, the above-mentioned image prediction model includes a posture prediction network, a gait prediction network, and an expression prediction network. The terminal can input different features into different networks respectively, and then obtain the emotion score of the corresponding feature through specific network prediction. For example, the terminal can input the above-mentioned posture features and posture optical flow features into the posture prediction network, and the posture prediction network outputs a first posture emotion score based on the posture features of each frame, and the posture prediction network outputs a second posture emotion score based on the posture optical flow features. The terminal can input the gait features into the gait prediction network, and the above-mentioned gait prediction network outputs a gait emotion score. In this way, the terminal can obtain the action emotion score based on the above-mentioned first posture emotion score, the second posture emotion score, and the gait emotion score.
[0095] The terminal can also input the above-mentioned facial expression features and facial expression optical flow features into an facial expression prediction network, which then outputs a first facial expression emotion score based on the facial expression features and a second facial expression emotion score based on the facial expression optical flow features. Thus, the terminal can obtain a behavioral action score based on the above-mentioned action emotion score, the first facial expression emotion score, and the second facial expression emotion score.
[0096] Specifically, in the image prediction model, the position of the user to be detected in the next frame of the video sequence can be relocated based on the position information of the user image in each frame, based on the relocation module, and the user's posture features and posture optical flow features can be intercepted through the positioning prediction frame. The first posture emotion score, also known as the posture satisfaction prediction index, is obtained through posture prediction network detection. The terminal can also input the posture optical flow features into the potential energy recognition module in the posture prediction network to capture the emotional signals contained in the user's limb movement information and obtain a second posture emotion score, also known as the posture potential energy satisfaction prediction index. The terminal can also intercept the user's ROI through the positioning detection frame and extract the expression features in the facial information to obtain the first expression emotion score based on the expression prediction network. The terminal can also input the expression optical flow features into the expression potential energy recognition module in the expression prediction network to obtain the emotional features contained in the user's expression movement information, and then output the second expression emotion score, also known as the expression potential energy satisfaction prediction index. In addition, since a person's emotional information is partially included in their gait, the terminal can obtain a gait emotion dataset by collecting and labeling relevant data, and then train the corresponding network on the dataset to obtain a gait emotion classification network, namely the above-mentioned gait prediction network. The terminal can use the gait prediction network for gait emotion classification in actual scenarios to obtain the customer's corresponding emotional score under a certain gait, also called a gait emotion score.
[0097] Through this embodiment, the terminal can determine the user's behavioral action emotion score based on multiple behavioral action features such as the user's static features and optical flow features, so that the terminal can determine the user's satisfaction with the service based on the behavioral action emotion score, reducing the cost of user emotion scoring.
[0098] In one embodiment, target recording information corresponding to the user to be detected is determined from the recording information based on the location information, including: obtaining the mouth shape features and gender features corresponding to the user to be detected based on the user image and the facial information corresponding to the user image; determining the user corresponding to each frame of recording information based on the Euclidean distance between the voice position of each frame of recording information and the position of the mouth shape features of the corresponding frame, and the Euclidean distance between the voice position of each frame of recording information and the position of the gender features of the corresponding frame, and obtaining the recording information of each user in the area to be detected; and determining the target recording information corresponding to the user to be detected from the recording information of each user based on the location information.
[0099] In this embodiment, the terminal can match the user with the recording information based on the aforementioned location information. The recording information can include the location of each sound. The terminal can extract the user image based on the aforementioned location information and extract the corresponding facial information from the user image, thereby obtaining the mouth shape features and gender features corresponding to the user to be detected. The gender feature can be determined based on at least one of the following features in the user image: the body shape, facial shape, and attire of the user to be detected. The recording information can include multiple frames of recording, and since the user image has multiple frames, there can also be multiple mouth shape features and gender features. The terminal can obtain the Euclidean distance between the location of the sound in each frame of the recording information and the location of the mouth shape feature in the corresponding frame, as well as the Euclidean distance between the location of the sound in each frame of the recording information and the location of the gender feature in the corresponding frame. Based on these two Euclidean distances, the terminal can determine the user corresponding to each frame of the recording information, thereby obtaining the recording information of each user in the aforementioned area. Since the location information of the user to be detected is known, the terminal can determine the target recording information corresponding to the user to be detected from the recording information of each user based on the aforementioned location information. For example, the recording information corresponding to the aforementioned location information is used as the target recording information of the user to be detected.
[0100] Specifically, because sound and image processing are not uniform, the terminal needs to match the current positioning target with the sound source target, that is, match the user with the sound source. During the sound source acquisition phase, the terminal already obtains rough information about the sound source, including the sound source's location. The terminal can match the sound source and user's localization box using Euclidean distance in three-dimensional space, but this can introduce uncertainty in scenarios where the targets overlap. Taking into account the imaging quality and imaging distance of the image acquisition device, the terminal uses the mouth movement detection module and the gender detection module to match the sound source corresponding to the sound source location with the gender characteristics and mouth shape characteristics of each user at that location based on Euclidean distance. This allows the sound source to be matched with the target with the closest Euclidean distance, matching the predicted gender and predicted mouth shape with motion, with higher confidence, thereby improving the matching accuracy between the sound source and the user to be detected. The terminal can then input the target recording information into the speech prediction network, which identifies the emotional state embedding contained in the target recording. This embedding is then input into a classifier to obtain a speech emotion score, also known as a voice satisfaction prediction indicator. Among them, since the speech recognition network needs to complete semantic understanding and classification tasks, the terminal can use a combination of CTC (Connectionist temporal classification, temporal classification based on neural network) + classifier to build a speech prediction network.
[0101] Through this embodiment, the terminal can determine the user's mouth shape features and gender characteristics based on the user's position information in three-dimensional space, and determine the voice of the user to be detected based on the mouth shape features and gender characteristics, as well as the sound position of each voice, thereby improving the accuracy of determining the sound object of the voice.
[0102] In one embodiment, it also includes: obtaining multiple frames of continuous disparity image samples and multiple frames of continuous audio samples within a preset time period, and inputting the multiple frames of continuous disparity image samples into the image prediction model to be trained; training the tracking re-identification network in the image prediction model to be trained based on the multiple frames of continuous disparity image samples, and obtaining the first model parameters corresponding to the tracking re-identification network when the training is completed; the tracking re-identification network is used to determine the position information of the user to be detected in the three-dimensional space; training the behavior action prediction network in the image prediction model to be trained based on the multiple frames of continuous disparity image samples, and obtaining the second model parameters corresponding to the behavior action prediction network when the training is completed; obtaining the image prediction model based on the first model parameters and the second model parameters; training the speech prediction network to be trained based on the multiple frames of continuous audio samples until the speech prediction network is obtained when the training conditions are met.
[0103] In this embodiment, the terminal can pre-train the image prediction model and speech prediction network. The terminal can obtain multiple frames of continuous disparity image samples and multiple frames of continuous audio recording samples within a preset time period, and input the multiple frames of continuous disparity image samples into the image prediction model to be trained. The terminal trains the tracking and re-identification network in the image prediction model to be trained based on the multiple frames of continuous disparity image samples, obtaining first model parameters corresponding to the tracking and re-identification network upon completion of training. The tracking and re-identification network determines the position of the user to be detected in three-dimensional space. For example, if the surveillance image comprises multiple frames, the terminal can use the tracking and re-identification network to determine the position of the user image in each surveillance image based on the position information corresponding to each surveillance image frame. The terminal can also train the behavior and action prediction network in the image prediction network to be trained based on the multiple frames of continuous disparity image samples, obtaining second model parameters corresponding to the behavior and action prediction network upon completion of training. The terminal can then obtain the image prediction model based on the first and second model parameters. Since the image prediction model also includes a posture prediction network, a gait prediction network, and an expression prediction network, the terminal can also train each prediction network separately based on the disparity image samples, and obtain second model parameters based on the model parameters of each prediction network.
[0104] The terminal can also train the speech prediction network to be trained based on multiple frames of continuous recording samples until a training condition is met, thereby obtaining a trained speech prediction network. The training condition to be met can be that the similarity between the predicted speech emotion score output by the speech prediction network to be trained and the actual speech emotion score corresponding to the input recording sample is greater than or equal to a preset similarity threshold, or the number of training cycles reaches a preset number.
[0105] Specifically, the image prediction model can use a 3D Yolo network as the baseline network. Sub-detection networks, such as the posture prediction network, gait prediction network, and expression prediction network, can be built based on 3D convolutional neural networks. Sound processing components, such as the speech prediction network, can be composed of a CTC+GRU (gated recurrent neural network) classifier. Since the satisfaction detection modules in each prediction network are fed with the temporal characteristics of the satisfaction detection indicator, the terminal can use an LSTM to build them. During training, the terminal can adopt a segmented training approach, training the image processing and sound processing components independently. For the image processing network, the terminal can first train the tracking-re-identification module and, after model convergence, fix the model parameters. It then trains the posture satisfaction detection module (i.e., the posture prediction network), the gait satisfaction detection module (i.e., the gait prediction network), the super-resolution module, and the face satisfaction detection module (i.e., the expression prediction network). After model convergence, the parameters are also fixed. The terminal then trains the posture potential satisfaction detection module (i.e., the posture potential prediction module within the posture prediction network) and the expression potential satisfaction detection module (i.e., the expression potential prediction module within the expression prediction network). These two modules share the posture recognition and face recognition subnetworks with the posture satisfaction detection and face satisfaction detection modules, respectively. For the sound processing network, the terminal can jointly train the CTC network and the classification network. The CTC network will clip the trained CTC network parameters to accelerate model convergence. Finally, fix the network parameters and train the satisfaction detection modules within each network. For the classification supervision component, the terminal can use the cross-entropy loss function for training; for the regression supervision component, the terminal can use the IoU (Intersection over Union) loss function for training.
[0106] Through this embodiment, the terminal can train the image prediction model and the speech prediction network respectively in a variety of different ways, so that the terminal can perform emotion scoring based on the trained image prediction model and speech prediction network, reducing the cost of emotion scoring and improving the accuracy of emotion scoring.
[0107] In one embodiment, Figure 4 As shown, Figure 4The figure is a flow chart of an emotion detection method according to another embodiment. In this embodiment, the terminal can input a preprocessed image and video sequence into an image prediction model. The preprocessed image and video sequence includes target disparity images at multiple time points and surveillance images corresponding to each target disparity image. These images can be RGBD images. The terminal can track the detector in the re-identification module and determine the detection box corresponding to each user based on the aforementioned position information, thereby obtaining the user image. The terminal can remove noise from the user image using a Kalman filter. The terminal can use a cosine metric to calculate the cosine angle between the target features in the current frame and the features in the previous frame, match the targets in the current and previous frames, and thus determine the user image of the same user in each surveillance image frame, thereby improving the accuracy of re-identification. The terminal can intercept relevant features from the user image in each frame and, based on each feature, perform satisfaction tests on the behavior and action features of the user to be detected, including posture satisfaction prediction, gait satisfaction prediction, posture potential satisfaction prediction, expression satisfaction prediction, and expression potential satisfaction prediction. The output of each of these predictions can be the emotion score corresponding to the feature.
[0108] The terminal can also detect the mouth movement features of the user to be detected, match the target sound source, determine the target recording information of the user to be detected, and obtain a pre-processed sound sequence. The terminal can use the CTC model and classifier to predict the voice satisfaction of the user to be detected and obtain a voice emotion score. Among them, the above-mentioned expression features, expression optical flow features and mouth movement features can be extracted from the facial information with improved resolution after reconstructing the facial area of the user image using a super-resolution network. After the terminal obtains the above-mentioned emotional scores, it can input the satisfaction detection module and output the final satisfaction prediction result, i.e., the final emotional score, through the softmax layer. The terminal can then determine the user's satisfaction with the service based on the size of the emotional score. For example, the larger the value of the emotional score, the higher the user's satisfaction with the service.
[0109] Through the above-described embodiment, the terminal fits a disparity map to surveillance images from multiple angles, thereby determining the user's three-dimensional spatial position. Based on this position, the terminal can extract the user's image features from the image and determine the user's audio recording at that location. This allows for an emotional score based on the user's actions, actions, and voice, thus reducing the cost of emotion detection. Furthermore, the use of contactless detection technology can reduce customer disruption and enhance the customer experience. The comprehensive detection metrics based on deep networks can significantly improve the accuracy of existing customer satisfaction measurement systems.
[0110] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0111] Based on the same inventive concept, embodiments of the present application also provide an emotion detection device for implementing the aforementioned emotion detection method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more emotion detection device embodiments provided below can be found in the above-described limitations of the emotion detection method and will not be further elaborated here.
[0112] In one embodiment, Figure 5 As shown, an emotion detection device is provided, including: a fitting module 500, an acquisition module 502, a first detection module 504, a second detection module 506 and an output module 508, wherein:
[0113] The fitting module 500 is used to obtain at least two disparity images to be detected for the area to be detected, perform image fitting on the at least two disparity images to be detected to obtain a target disparity image of the area, and determine the position information of the user to be detected in the area in three-dimensional space based on the target disparity image; the disparity image to be detected is determined based on multiple monitoring images of the area.
[0114] The acquisition module 502 is used to acquire the behavior and action features of the user to be detected based on the location information and multiple monitoring images.
[0115] The first detection module 504 is configured to input the behavior action features into an image prediction model, and the image prediction model outputs a behavior action emotion score of the user to be detected based on the behavior action features.
[0116] The second detection module 506 is used to obtain the recording information in the area and determine the target recording information corresponding to the user to be detected from the recording information based on the location information; input the target recording information into the speech prediction network, and obtain the speech emotion score of the user to be detected output by the speech prediction network.
[0117] The output module 508 is used to determine the emotion score of the user to be detected based on the action emotion score and the voice emotion score.
[0118] In one embodiment, the fitting module 500 is specifically used to obtain at least three surveillance images captured by at least three image acquisition devices; determine baseline distance information between multiple groups of image pairs in the at least three surveillance images; each group of image pairs includes two surveillance images captured by different image acquisition devices; and determine at least two parallax images to be detected based on at least two groups of image pairs whose baseline distance information meets a preset distance condition.
[0119] In one embodiment, the fitting module 500 is specifically used to obtain the coordinate values of at least three image acquisition devices in the area, and the angle between the orientations of each two image acquisition devices; for each two image acquisition devices, determine the corresponding Euclidean distance based on the coordinate values of the two image acquisition devices; determine the baseline distance information corresponding to the two image acquisition devices based on the preset optimal baseline distance, the Euclidean distance, the angle between the orientations of the two image acquisition devices, and the preset optimal angle, and use the baseline distance information corresponding to the two image acquisition devices as the baseline distance information between the image pairs corresponding to the two image acquisition devices.
[0120] In one embodiment, the fitting module 500 is specifically used to obtain two first target image acquisition device sets corresponding to the first minimum baseline distance information, and obtain a second target image acquisition device set that includes any one target image acquisition device in the first target image acquisition device set and has the second minimum baseline distance information; determine a first disparity image based on a plurality of monitoring images corresponding to the first target image acquisition device set, and determine a second disparity image based on a plurality of monitoring images corresponding to the second target image acquisition device set; and obtain at least two disparity images to be detected based on the first disparity image and the second disparity image.
[0121] In one embodiment, the above-mentioned acquisition module 502 is specifically used to obtain a user image corresponding to the user to be detected from multiple surveillance images based on the position information; obtain the posture features corresponding to the user to be detected in at least two frames of user images, and the posture optical flow features corresponding to the posture features; obtain the gait features corresponding to the user image; obtain the action features based on the posture features, the posture optical flow features and the gait features; obtain the facial information corresponding to at least two frames of user images, and extract the expression features corresponding to the facial information and the expression optical flow features corresponding to the expression features; obtain the behavioral action features of the user to be detected based on the action features, the expression features and the expression optical flow features.
[0122] In one embodiment, the above-mentioned first detection module 504 is specifically used to input posture features and posture optical flow features into a posture prediction network, and output a first posture emotion score based on the posture features of each frame through the posture prediction network, and output a second posture emotion score based on the posture optical flow features; input gait features into a gait prediction network, and output a gait emotion score through the gait prediction network; obtain an action emotion score based on the first posture emotion score, the second posture emotion score and the gait emotion score; input expression features and expression optical flow features into an expression prediction network, and output a first expression emotion score based on the expression features through the expression prediction network, and output a second expression emotion score based on the expression optical flow features; obtain a behavior action emotion score based on the action emotion score, the first expression emotion score and the second expression emotion score.
[0123] In one embodiment, the above-mentioned second detection module 506 is specifically used to obtain the mouth shape features and gender features corresponding to the user to be detected based on the user image and the facial information corresponding to the user image; determine the user corresponding to each frame of recording information based on the Euclidean distance between the sound position of each frame of recording information and the position of the mouth shape features of the corresponding frame, and the Euclidean distance between the sound position of each frame of recording information and the position of the gender features of the corresponding frame, and obtain the recording information of each user in the area to be detected; and determine the target recording information corresponding to the user to be detected from the recording information of each user based on the position information.
[0124] In one embodiment, the above-mentioned output module 508 is specifically used to sort multiple behavioral action emotion scores and multiple voice emotion scores corresponding to the user to be detected in chronological order to obtain an emotion score sequence corresponding to the user to be detected; each behavioral action emotion score and each voice emotion score are determined based on at least one frame of user image of the user to be detected; the emotion score sequence is input into an emotion scoring network based on long short-term memory, and the emotion score of the user to be detected is obtained after the emotion scoring network performs weighted superposition on the multiple behavioral action emotion scores and the multiple voice emotion scores.
[0125] In one embodiment, the above-mentioned device also includes: a training module, which is used to obtain multiple frames of continuous disparity image samples and multiple frames of continuous audio samples within a preset time period, and input the multiple frames of continuous disparity image samples into the image prediction model to be trained; training the tracking re-identification network in the image prediction model to be trained based on the multiple frames of continuous disparity image samples, and obtaining the first model parameters corresponding to the tracking re-identification network when the training is completed; the tracking re-identification network is used to determine the position information of the user to be detected in the three-dimensional space; training the behavior action prediction network in the image prediction model to be trained based on the multiple frames of continuous disparity image samples, and obtaining the second model parameters corresponding to the behavior action prediction network when the training is completed; obtaining the image prediction model based on the first model parameters and the second model parameters; training the speech prediction network to be trained based on the multiple frames of continuous audio samples until the speech prediction network is obtained when the training conditions are met.
[0126] Each module in the above-mentioned emotion detection device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0127] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, an emotion detection method is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse.
[0128] Those skilled in the art will understand that Figure 6The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0129] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the above-mentioned emotion detection method when executing the computer program.
[0130] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned emotion detection method is implemented.
[0131] In one embodiment, a computer program product is provided, comprising a computer program, which implements the above-mentioned emotion detection method when executed by a processor.
[0132] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0133] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.
[0134] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0135] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A method for emotion detection, characterized in that: The method comprises: Acquiring at least three surveillance images captured by at least three image acquisition devices; For every two image acquisition devices, determining a corresponding Euclidean distance according to the coordinate values of the two image acquisition devices; Determining baseline distance information corresponding to the two image acquisition devices based on a preset optimal baseline distance, the Euclidean distance, the angle between the orientations of the two image acquisition devices, and the preset optimal angle, and using the baseline distance information corresponding to the two image acquisition devices as baseline distance information between the image pairs corresponding to the two image acquisition devices; Acquire two first target image acquisition device sets corresponding to the first minimum baseline distance information, and acquire a second target image acquisition device set including any one target image acquisition device in the first target image acquisition device set and corresponding to the second minimum baseline distance information; Determine a first disparity image based on a plurality of monitoring images corresponding to the first target image acquisition device set, and determine a second disparity image based on a plurality of monitoring images corresponding to the second target image acquisition device set; Obtaining at least two to-be-detected disparity images of an area to be detected based on the first disparity image and the second disparity image, performing image fitting on the at least two to-be-detected disparity images to obtain a target disparity image of the area, and determining position information of a user to be detected in the area in three-dimensional space based on the target disparity image; Obtaining behavioral action features of the user to be detected based on the location information and the multiple surveillance images; the behavioral action features include posture features, posture optical flow features, gait features, expression features, and expression optical flow features; Inputting the behavior action features into an image prediction model, and having the image prediction model output a behavior action emotion score of the user to be detected based on the behavior action features; Acquire recording information in the area, and determine target recording information corresponding to the user to be detected from the recording information based on the location information; input the target recording information into a speech prediction network, and obtain a speech emotion score of the user to be detected output by the speech prediction network; The emotion score of the user to be detected is determined according to the behavior emotion score and the voice emotion score.
2. The method according to claim 1, characterized in that Before determining the corresponding Euclidean distance based on the coordinate values of the two image acquisition devices; and determining the baseline distance information corresponding to the two image acquisition devices based on a preset optimal baseline distance, the Euclidean distance, the angle between the orientations of the two image acquisition devices, and the preset optimal angle, the method further includes: The coordinate values of the at least three image acquisition devices in the area and the angle between the directions of every two image acquisition devices are obtained.
3. The method according to claim 1, characterized in that The step of obtaining the behavioral characteristics of the user to be detected based on the location information and the plurality of surveillance images includes: Acquire, from the plurality of surveillance images, a user image corresponding to the user to be detected according to the location information; Acquire posture features corresponding to the user to be detected in at least two frames of the user image, and posture optical flow features corresponding to the posture features; Obtaining gait features corresponding to the user image; obtaining action features based on the posture features, the posture optical flow features, and the gait features; Acquire facial information corresponding to at least two frames of the user image, and extract expression features corresponding to the facial information and expression optical flow features corresponding to the expression features; The behavioral action features of the user to be detected are obtained according to the action features, the expression features and the expression optical flow features.
4. The method according to claim 3, characterized in that The image prediction model includes a posture prediction network, a gait prediction network and an expression prediction network; Inputting the behavior action features into an image prediction model, and having the image prediction model output a behavior action emotion score of the user to be detected based on the behavior action features, includes: Inputting the posture feature and the posture optical flow feature into the posture prediction network, and outputting a first posture emotion score based on the posture feature of each frame and a second posture emotion score based on the posture optical flow feature through the posture prediction network; Inputting the gait features into a gait prediction network, and outputting a gait emotion score through the gait prediction network; Obtaining an action emotion score according to the first posture emotion score, the second posture emotion score, and the gait emotion score; Inputting the facial expression features and the facial expression optical flow features into an facial expression prediction network, and outputting a first facial expression emotion score based on the facial expression features and a second facial expression emotion score based on the facial expression optical flow features through the facial expression prediction network; The behavioral action emotion score is obtained according to the action emotion score, the first expression emotion score and the second expression emotion score.
5. The method according to claim 3, characterized in that The determining target recording information corresponding to the user to be detected from the recording information according to the location information includes: Obtaining mouth shape features and gender features corresponding to the user to be detected based on the user image and the facial information corresponding to the user image; Determine the user corresponding to each frame of recording information based on the Euclidean distance between the voice position of each frame of recording information and the position of the mouth shape feature of the corresponding frame, and the Euclidean distance between the voice position of each frame of recording information and the position of the gender feature of the corresponding frame, and obtain the recording information of each user in the area to be detected; According to the location information, target recording information corresponding to the user to be detected is determined from the recording information of each user.
6. The method according to claim 1, characterized in that Determining the emotion score of the user to be detected based on the behavior emotion score and the voice emotion score includes: Sorting multiple behavioral emotion scores and multiple voice emotion scores corresponding to the user to be detected in chronological order to obtain an emotion score sequence corresponding to the user to be detected; each behavioral emotion score and each voice emotion score is determined based on at least one frame of the user image to be detected; The emotion score sequence is input into an emotion score network based on long short-term memory, and an emotion score of the user to be detected is obtained after the emotion score network performs weighted superposition on multiple action emotion scores and multiple voice emotion scores.
7. The method according to claim 1, characterized in that The method further comprises: Acquire multiple frames of continuous disparity image samples and multiple frames of continuous audio recording samples within a preset time period, and input the multiple frames of continuous disparity image samples into the image prediction model to be trained; Training a tracking and re-identification network in the image prediction model to be trained based on the multiple frames of continuous disparity image samples to obtain first model parameters corresponding to the tracking and re-identification network when training is completed; the tracking and re-identification network is used to determine position information of the user to be detected in three-dimensional space; Training a behavior action prediction network in the image prediction model to be trained based on the multiple frames of continuous disparity image samples, and obtaining second model parameters corresponding to the behavior action prediction network when training is completed; Obtaining the image prediction model according to the first model parameters and the second model parameters; The speech prediction network to be trained is trained according to the multiple frames of continuous recording samples until the speech prediction network is obtained when the training conditions are met.
8. An emotion detection device, characterized in that: The device comprises: a fitting module, configured to obtain at least two parallax images to be detected for an area to be detected, perform image fitting on the at least two parallax images to be detected to obtain a target parallax image for the area, and determine position information of a user to be detected in the area in three-dimensional space based on the target parallax image; the parallax image to be detected is determined based on multiple surveillance images of the area; an acquisition module, configured to acquire the behavior and action features of the user to be detected based on the location information and the plurality of surveillance images; the behavior and action features include posture features, posture optical flow features, gait features, expression features, and expression optical flow features; A first detection module is configured to input the behavior action features into an image prediction model, and the image prediction model outputs a behavior action emotion score of the user to be detected based on the behavior action features; a second detection module configured to obtain recording information in the area and determine target recording information corresponding to the user to be detected from the recording information based on the location information; input the target recording information into a speech prediction network, and obtain a speech emotion score of the user to be detected output by the speech prediction network; An output module, configured to determine an emotion score of the user to be detected based on the behavior emotion score and the voice emotion score; wherein at least three surveillance images captured by at least three image acquisition devices are obtained; for each pair of image acquisition devices, a corresponding Euclidean distance is determined based on the coordinate values of the two image acquisition devices; baseline distance information corresponding to the two image acquisition devices is determined based on a preset optimal baseline distance, the Euclidean distance, the angle between the orientations of the two image acquisition devices, and the preset optimal angle, and the baseline distance information corresponding to the two image acquisition devices is used as the baseline distance information between the image pairs corresponding to the two image acquisition devices; Acquire two first target image acquisition device sets corresponding to the first minimum baseline distance information, and acquire a second target image acquisition device set including any one target image acquisition device in the first target image acquisition device set and corresponding to the second minimum baseline distance information; Determine a first disparity image based on a plurality of monitoring images corresponding to the first target image acquisition device set, and determine a second disparity image based on a plurality of monitoring images corresponding to the second target image acquisition device set; At least two to-be-detected disparity images are obtained according to the first disparity image and the second disparity image.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Intelligent monitoring real-time processing method, device and apparatus and storage medium
CN109543513A
Three-dimensional target detection method and device
CN111462096A
Customer satisfaction identification method and device, equipment and medium
CN114973419A
Micro-expression detection method, electronic equipment and computer readable storage medium
CN115424315A