IP network visual intercom emergency emotion recognition method and system based on artificial intelligence
By integrating artificial intelligence technology in emergency equipment, collecting video data and building an emotional feature correlation matrix, the problem of difficulty in accurately identifying users' emergency emotions in the existing technology is solved, and more efficient emergency response is achieved.
Patent Information
- Application Number
- CN202510063104.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-15
AI Technical Summary
Existing emergency response technologies are difficult to accurately identify users' emotions in emergency situations, resulting in insufficient targeted and effective emergency responses.
Using an IP network visual intercom emergency emotion recognition method based on artificial intelligence, video data is collected through cameras, emotional feature correlation matrix of face images is constructed, and user emotions are identified.
It realizes more accurately identifying users' emotions in emergency situations, provides strong support for subsequent emergency measures, and improves the pertinence and effectiveness of emergency responses.
Smart Images

Figure CN120032409A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to an artificial intelligence-based IP network visual intercom emergency emotion recognition method and system. Background Art
[0002] In emergency scenarios, timely and accurate identification of user emotions is crucial for taking appropriate emergency measures. Traditional emergency response methods often lack accurate judgment of user emotions, resulting in insufficient pertinence and effectiveness of emergency response. Summary of the invention
[0003] The purpose of the present invention is to provide an artificial intelligence-based IP network visual intercom emergency emotion recognition method and system.
[0004] In a first aspect, an embodiment of the present invention provides an artificial intelligence-based IP network video intercom emergency emotion recognition method, comprising:
[0005] In response to an alarm instruction of the IP network video intercom emergency device triggered by the current user, starting a camera of the IP network video intercom emergency device;
[0006] Collecting current video data through the camera, and determining the current face image from the video data;
[0007] An emotion feature association matrix of the current face image is constructed, and the emotion type corresponding to the current face image is determined according to the emotion feature association matrix.
[0008] In a second aspect, an embodiment of the present invention provides a server system, including a server, wherein the server is used to execute the method described in the first aspect.
[0009] Compared with the prior art, the beneficial effects provided by the present invention include: using an artificial intelligence-based IP network video intercom emergency emotion recognition method and system disclosed by the present invention, after the user triggers the alarm command of the IP network video intercom emergency device, the camera is started to collect video data and determine the current face image, and then the emotional feature association matrix of the face image is constructed, and the corresponding emotion type is determined according to the matrix. By constructing the association matrix, this method can more accurately identify the user's emotions in emergency situations, and provide strong support for the subsequent targeted emergency measures. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative work.
[0011] Figure 1 A schematic diagram of the steps of an artificial intelligence-based IP network video intercom emergency emotion recognition method provided by an embodiment of the present invention;
[0012] Figure 2 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0013] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0014] The specific implementation modes of the present invention are described in detail below in conjunction with the accompanying drawings.
[0015] In order to solve the technical problems in the aforementioned background technology, Figure 1 A flow chart of an artificial intelligence-based IP network video intercom emergency emotion recognition method provided in an embodiment of the present disclosure is provided. The artificial intelligence-based IP network video intercom emergency emotion recognition method is introduced in detail below.
[0016] Step S201, in response to an alarm instruction of the IP network video intercom emergency device triggered by the current user, starting a camera of the IP network video intercom emergency device;
[0017] Step S202, collecting current video data through the camera, and determining the current face image from the video data;
[0018] Step S203, constructing an emotion feature association matrix of the current face image, and determining the emotion type corresponding to the current face image according to the emotion feature association matrix.
[0019] In an embodiment of the present invention, for example, in a community in a city, a gas leak suddenly occurred in a resident's home. After the resident (i.e., the current user) sensed the danger, he quickly ran to the IP network video intercom emergency device installed at the door, pressed the alarm button on the device, and triggered the alarm command. This alarm command is quickly transmitted to the server through the network inside the community. After receiving the alarm command, the server immediately sends a command to start the camera to the IP network video intercom emergency device. After receiving the command from the server, the IP network video intercom emergency device quickly starts its built-in camera. At this time, the camera begins to prepare to collect relevant video data to provide a basis for subsequent analysis and processing. After the camera is started, it starts to collect current video data. In the gas leak scene, the camera captures the image of the resident standing in front of the emergency device with an anxious and slightly frightened expression. These video data are transmitted to the server through the network in the form of digital signals. After receiving the video data, the server uses advanced image recognition algorithms to analyze and process the video data. The server first scans each frame of the video data and detects whether there is a face in the image through face recognition technology. In this process, the server determines whether it is a human face based on the features of the face, such as the position and shape of the eyes, nose, mouth, etc. When the server accurately identifies a face in a certain frame of image, it extracts the image segment containing the face and determines it as the current face image. For example, the server recognizes the face of the resident from the 10th to the 20th frame of the video data, so the image composed of these 10 frames is determined as the current face image. After the server obtains the current face image determined above, it extracts multiple images from the image. For example, the 12th and 18th frames are selected from different frames of the current face image. Then, the server divides the two images into two image sets based on the features of the image, such as subtle differences in facial expressions, lighting conditions, etc. Suppose the eyes of the resident in the 12th frame are slightly wide open, and it is divided into an image set representing the "slightly nervous" feature; the resident's brows are furrowed in the 18th frame, and it is divided into an image set representing the "relatively nervous" feature. The server selects a representative target image from the above two image sets. For example, the 18th frame image is selected as the target image from the image set representing the "relatively nervous" feature. Then, the server uses a deep learning algorithm to enhance the image feature extraction of the target image. The server performs a convolution operation on the target image through a convolutional neural network (CNN) to extract features such as edges and textures in the image to obtain the first image features. For example, in the 18th frame image, CNN detects the texture features of the resident's frown and the features of slightly bloodshot eyes, which are extracted as the first image features. The server assigns the first emotion identifier to the current face image based on the pre-set emotion classification criteria and the overall performance of the current face image.In the gas leak scenario, the server assigns the first emotion identifier of "tension and fear" to the current face image based on the characteristics of the resident's frowning brows and anxious expressions. Then, the server extracts semantic features from this first emotion identifier. The server decomposes the emotion identifier "tension and fear" into semantic units, such as "tension" and "fear", through natural language processing (NLP) technology, and further analyzes the relevant features of each semantic unit. For example, "tension" may be related to physiological reactions such as accelerated heartbeat and muscle tension, and "fear" may be related to psychological reactions such as perception of danger and escape tendency. These analyzed features constitute the first semantic feature. The server fuses the first image feature and the first semantic feature to obtain the first multidimensional feature. This first multidimensional feature contains feature information of multiple dimensions. For example, one dimension corresponds to the facial expression feature in the image, such as the degree of frowning; another dimension corresponds to the psychological reaction in the semantic feature, such as the degree of fear. Assume that in this example, the first multidimensional feature includes multiple dimensional features such as "the degree of frowning is 8 (full score 10)" in the facial expression dimension and "the degree of fear is 7 (full score 10)" in the psychological reaction dimension. The server obtains the preset enhanced image feature matching pool and semantic feature matching pool. The enhanced image feature matching pool stores the correspondence between a large number of pending face images and their corresponding enhanced image features, and the semantic feature matching pool stores the correspondence between pending face images and their corresponding semantic features. The server matches the two matching pools according to the target enhanced image feature and the first semantic feature in the first multidimensional feature. For example, the server compares the target enhanced image feature in the first multidimensional feature (such as the texture feature of frowning) with the enhanced image features of each pending face image in the enhanced image feature matching pool, and compares the first semantic feature (such as the semantic feature related to "tension and fear") with the semantic features of each pending face image in the semantic feature matching pool. Through the comparison, the server finds at least two target pending face images that are close to the current face image. Assume that the two target facial images found are from the images of the parties involved in other gas leakage incidents and fire incidents recorded previously. Then, the server determines the reference facial image from the target facial images based on the image matching coefficients between the at least two target facial images and the current facial image. Specifically, the server confirms the matching coefficient between the current facial image and each target facial image matched by each matching strategy. For example, for the matching strategy based on facial expression features, the server calculates that the matching coefficient between the current facial image and the first target facial image is 0.85, and the matching coefficient between the current facial image and the second target facial image is 0.82. Then, for each target facial image, the server confirms the average matching coefficient between the target facial image and the current facial image in each matching strategy.Assume that the average matching coefficient of the first target face image under multiple matching strategies is 0.84, and the average matching coefficient of the second target face image is 0.80. Finally, the server determines the reference face image from the target face images according to the matching coefficient average. In this example, the first target face image with a higher matching coefficient average is determined as the reference face image. Perform a feature integration operation on at least two of the first image features to obtain a target enhanced image feature: the server performs a feature integration operation on at least two first image features extracted from the current face image. For example, the server fuses the first image feature representing "frown" and the first image feature representing "slightly bloodshot eyes" to obtain a comprehensive target enhanced image feature, which more comprehensively reflects the features of the current face image. The target enhanced image features are synchronized to the preset semantic feature domain through the pre-trained feature synchronization network to obtain the first image semantic identification set, and the second image features corresponding to each of the reference face images are feature mapped to obtain the second image semantic identification set corresponding to each of the reference face images: the server uses the pre-trained feature synchronization network to synchronize the target enhanced image features to the preset semantic feature domain. For example, the "frown" feature in the target enhanced image features is mapped to the "nervous" semantic identification in the semantic feature domain, and the "slightly bloodshot eyes" feature is mapped to the "excited" semantic identification, thereby obtaining the first image semantic identification set. At the same time, the server performs a similar feature mapping operation on the second image features corresponding to the reference face image to obtain the second image semantic identification set corresponding to the reference face image. Obtain the collection location of the current face image and the voice content recognized by the voice data of the current face image: the server obtains the collection location information of the current face image from the IP network visual intercom emergency device. In this example, the collection location is a unit and a room in a building in the community. At the same time, the server recognizes the voice data attached to the current facial image, and identifies that the resident's voice content is "Come and help me, there is a gas leak at home." The emergency scene information of the current facial image is generated based on the background image of the current facial image: the server analyzes the background image of the current facial image, such as seeing smoke from a gas leak in the background, and combines relevant knowledge and algorithms to generate the emergency scene information of the current facial image, and determines that this is an emergency scene of a gas leak. The first emotion identifier is generated based on the collection location, the voice content, and the emergency scene information: the server comprehensively collects the location, voice content, and emergency scene information to further confirm and refine the first emotion identifier. In this example, based on the scene of a gas leak in the resident's home and the content of the voice requesting help, the first emotion identifier is confirmed again as "tension and fear."According to the first semantic identifier used to distinguish different information in the same facial image, and the second image semantic identifier set corresponding to the same reference facial image, the second emotion identifier and the emotion target value, an emotion feature association parameter is established to obtain at least two emotion feature association parameters corresponding to at least two reference facial images: the server establishes the emotion feature association parameter according to the first semantic identifier (such as an identifier for distinguishing different expression features in the current facial image), the second image semantic identifier set of the reference facial image, the second emotion identifier (the emotion identifier corresponding to the reference facial image) and the emotion target value (the emotion quantization value pre-set for the reference facial image). For example, for the first reference facial image, its second image semantic identifier set includes identifiers such as "tension" and "worry", the second emotion identifier is "highly nervous and fearful", and the emotion target value is 9 (out of 10). The server establishes an emotion feature association parameter according to this information and the first semantic identifier of the current facial image. The at least two emotion feature association parameters are merged to obtain an initial emotion association matrix, wherein at least two emotion feature association parameters in the initial emotion association matrix are separated by a second semantic identifier for distinguishing information of different face images: the server merges the emotion feature association parameters corresponding to multiple reference face images to form an initial emotion association matrix. During the merging process, the second semantic identifier is used to distinguish information of different face images to ensure that the data in the matrix is clear and orderly. An emotion classification identifier is established according to the first semantic identifier, the first image semantic identifier set and the first emotion identifier: the server establishes an emotion classification identifier according to the first semantic identifier, the first image semantic identifier set and the first emotion identifier. This emotion classification identifier is used to classify and identify the emotion of the current face image, for example, the established emotion classification identifier is "tension and fear in the gas leak scene". The emotion feature association matrix is generated according to the initial emotion association matrix and the emotion classification identifier: the server merges the initial emotion association matrix and the emotion classification identifier to generate a final emotion feature association matrix. This matrix comprehensively reflects the emotion feature association relationship between the current face image and the reference face image. The server performs emotion element feature conversion on the emotion feature association matrix. For example, the server standardizes the various emotion feature parameters in the matrix according to the requirements of the target value recognition model, converts the feature values of different dimensions into a unified format and range, and obtains an emotion feature set. The data in this emotion feature set meets the input requirements of the target value recognition model, so that accurate emotion type recognition can be performed later. The server inputs the emotion feature set into the target value recognition model. The target value recognition model is carefully trained. On the basis of keeping the basic model parameters of the preset multimodal large model unchanged, it optimizes the updated model parameters by training the emotion feature set, and adds an attention mechanism to enhance the focus on key features.For example, in this gas leak scenario, the attention mechanism will pay more attention to features related to the emotion of "tension and fear", such as the tension of facial expressions and the anxious tone of voice. After analyzing and processing the input emotional feature set, the target value recognition model generates the target emotional target value of the current face image. In this example, the server obtains the target emotional target value of the current face image from the target value recognition model as 8 (out of 10), indicating that the resident's emotions are at a high level of tension and fear. The server determines the emotional type corresponding to the current face image based on the size of the target emotional target value and the pre-set emotional type classification criteria. For example, when the target emotional target value is between 7-10, the corresponding emotional type is "high tension and fear". Therefore, the server determines that the current resident's emotional type is "high tension and fear". In this way, the server successfully identifies the emotional type of the current user in an emergency situation, providing an important basis for taking corresponding emergency measures in the future.
[0020] In the embodiments of the present invention, the following implementation modes are also provided.
[0021] When the emotion type belongs to the first-level emotion type, the intercom function of the IP network visual intercom emergency device is turned on to realize manual real-time reply interaction;
[0022] When the emotion type belongs to the secondary emotion type, current speech data from the current user is loaded into the first generative dialogue model of the integrated model to obtain the basic goal of the current user; the urgency of the secondary emotion type is greater than that of the primary emotion type in an emergency scenario;
[0023] Loading the past speech contents of all users within a preset period range and the basic goal into the second generative dialogue model of the integrated model to obtain the final goal of the current user;
[0024] The final goal is adapted to a plurality of preset emergency strategies in a preset emergency strategy database to obtain a target emergency strategy adapted to the current user, and the target emergency strategy is provided to the current user.
[0025] In an embodiment of the present invention, for example, suppose that in a large shopping mall, a customer accidentally has a slight collision with others during shopping, and the customer is slightly excited, but has not yet reached a very urgent level. At this time, the server determines that the customer's emotion type belongs to the first-level emotion type through the previous emotion recognition method. After identifying the emotion type, the server immediately sends an instruction to the IP network video intercom emergency device to turn on its intercom function. After receiving the instruction, the IP network video intercom emergency device establishes a communication connection with the customer service center of the mall. The staff of the customer service center can then interact with the customer in real time through the intercom device. For example, the customer said through the emergency device: "I was hit by someone just now, and I feel a little uncomfortable." The customer service staff responded through the intercom device: "I am very sorry to hear that you have encountered such a situation. Please calm down first. May I ask if you are injured now?" Through this real-time interactive method, the customer service staff can further understand the specific situation of the customer, appease the customer's emotions, and avoid further escalation of the situation. If the customer subsequently makes some reasonable demands, such as hoping to find the person who bumped into him to negotiate and resolve, the customer service staff can also assist in handling in a timely manner. In a factory workshop, a worker was operating a machine when he suddenly found that the machine had a serious fault and there was a great safety hazard. This made him extremely panic and anxious. The server determined that the worker's emotion type belonged to the secondary emotion type. At this time, the worker sent the current voice data to the server through the IP network video intercom emergency device. He said: "The machine suddenly lost control. What can I do?" After receiving this voice data, the server loaded it into the first generative dialogue model of the integrated model. The first generative dialogue model analyzes and understands the voice data. It first converts the voice data into text form, and then analyzes the semantic information in the text through deep learning algorithms and a large amount of training data. In this example, the model recognizes that the key information described by the worker is "machine failure" and "seeking solutions", so that the basic goal of the worker is to solve the current machine failure problem and ensure production safety. After obtaining the basic goal of the worker, the server will retrieve the past voice content of all users in the preset period range (for example, the past month) from the database. These past voice contents may contain information such as descriptions, solution processes, and expectations of other workers when they encountered similar machine failures. Assume that there are some cases of machine failure of the same model in the past voice content. Some workers mentioned that they hope to repair the machine and resume production as soon as possible, while others emphasized that they must ensure the safety of personnel first. The server loads these past voice contents and the basic goals of the current workers into the second generative dialogue model of the integrated model. The second generative dialogue model will comprehensively analyze this information and further explore the potential needs and ultimate expectations of workers in combination with the current actual situation.For example, after model analysis, it is found that the current production task of the worker is relatively urgent, and the faulty machine has a great impact on the entire production process, so the ultimate goal of the worker is not only to solve the machine failure, but also to resume production in the shortest possible time, while ensuring the safety of himself and other colleagues. After determining the worker's ultimate goal, the server will adapt it to multiple preset emergency strategies in the preset emergency strategy database. The preset emergency strategy database stores various response strategies for different types of emergencies, including machine failure, fire, personal injury, etc. For emergencies such as machine failure, there may be multiple emergency strategies in the database, such as immediately stopping the machine operation, notifying professional maintenance personnel, starting backup equipment, etc. The server uses a matching algorithm to compare the worker's ultimate goal with these preset emergency strategies one by one. In this example, the server adapts the following target emergency strategies from the preset emergency strategy database based on the workers' ultimate goal of resuming production as soon as possible and ensuring safety: First, stop the operation of the faulty machine immediately to avoid further expansion of the danger; at the same time, notify professional maintenance personnel to bring relevant tools and parts to the site for emergency repairs; during the repair process, arrange other workers to do a good job of on-site safety protection to prevent accidents; in addition, start the backup equipment to temporarily maintain normal production. After determining the target emergency strategy, the server will send it to the workers through the IP network video intercom emergency device. After receiving the target emergency strategy, the workers operate according to the guidance of the strategy, so as to effectively respond to the current emergency and minimize losses. Through the above steps, the server can provide targeted emergency strategies according to the user's emotional type and specific needs, and improve the efficiency and effectiveness of emergency handling.
[0026] In the embodiment of the present invention, the step of constructing an emotion feature association matrix of the current facial image and determining the emotion type corresponding to the current facial image according to the emotion feature association matrix may be implemented through the following examples.
[0027] Acquire the current face image, and perform multi-dimensional feature extraction on the current face image to obtain a first multi-dimensional feature, where the first multi-dimensional feature includes a plurality of dimensional features corresponding to a plurality of dimensions respectively;
[0028] Matching at least two reference facial images close to the current facial image according to the first multi-dimensional feature, each of the reference facial images being preset with an emotional target value;
[0029] Establishing the emotion feature association matrix according to the second multidimensional features corresponding to each of the reference facial images, the preset emotion target value of each of the reference facial images, and the first multidimensional features of the current facial image;
[0030] According to the emotion feature association matrix, identifying the emotion target value of the current face image;
[0031] The emotion type corresponding to the current facial image is determined according to the recognition result of the emotion target value.
[0032] In an embodiment of the present invention, illustratively, in a campus environment, an IP network video intercom emergency device is installed in the school's teaching building. One day, a student suddenly felt unwell in the classroom, so he went to the emergency device and triggered an alarm command. After receiving the alarm command, the server starts the camera of the emergency device and collects video data containing the student's face, thereby determining the current face image. The server begins to extract multi-dimensional features from the current face image. First, feature extraction is performed from the image dimension. The server uses advanced computer vision algorithms to analyze the features of various parts of the face in the current face image. For example, it is observed that the student's eyebrows are slightly wrinkled, the eyes are slightly dull, and the lips are a little pale. By analyzing these facial details, the server extracts dimensional features related to facial expressions, such as the degree of wrinkling of the eyebrows is quantified as 3 (out of 10), the degree of dullness of the eyes is quantified as 4 (out of 10), etc. At the same time, the server also extracts features from the voice dimension. When collecting video data, the emergency device may also simultaneously collect the student's weak voice, such as the student's hoarse voice and weak tone. The server converts the sound into text through speech recognition technology, and further analyzes the characteristics of the voice, such as intonation, speaking speed, and volume. For example, the student's speaking speed is slow, the volume is low, and the tone is relatively gentle. The server quantifies these voice features, such as the speaking speed is quantified to 2 (full score 10), the volume is quantified to 3 (full score 10), etc., as the characteristics of the voice dimension. In addition, the server also takes into account the characteristics of the behavior dimension. From the current face image, the student's body posture and movement are analyzed, and it is found that the student is standing unsteadily and the body is shaking slightly. The server quantifies these behavioral features, such as the degree of body shaking is quantified to 5 (full score 10), etc. Combining the above features of multiple dimensions such as image, voice, and behavior, the server obtains the first multidimensional feature. This first multidimensional feature contains multiple dimensional features corresponding to multiple dimensions, which comprehensively reflects the current state of the student in the emergency moment. The server obtains a pre-established benchmark face image library, which stores a large number of face images in different scenes and their corresponding multidimensional features and preset emotional target values. The server performs a matching search in the benchmark face image library based on the various dimensional features in the first multidimensional feature. For example, the server first searches for facial images with similar features such as the degree of frowning of eyebrows and dullness of eyes in the benchmark facial image library based on the features of the facial expression dimension. After finding some benchmark facial images that meet the conditions, further screening is performed based on the features of the voice dimension and the behavior dimension. Assume that in the benchmark facial image library, there are two benchmark facial images that are relatively close to the first multi-dimensional feature of the current student.One is an image of another student who was sick before. Its facial expression, voice and behavior characteristics have a high similarity with the current student, and the preset emotional target value is 7 (out of 10), indicating weakness and discomfort; the other is an image of a student after being slightly frightened. Some features match the current student, and the preset emotional target value is 6 (out of 10), indicating mild nervousness. For the two reference face images found, the server obtains their corresponding second multidimensional features respectively. Taking the reference face image of the sick student as an example, its second multidimensional features may include a more obvious pale face feature with a quantitative value of 8 (out of 10) in the facial expression dimension, a cough sound feature with a quantitative value of 6 (out of 10) in the voice dimension, and a slow walking feature with a quantitative value of 7 (out of 10) in the behavior dimension. The server associates the second multidimensional features of the reference face image, the preset emotional target value and the first multidimensional features of the current face image. First, the feature values of each dimension are arranged correspondingly to form a preliminary framework of a matrix. For example, the first row places the first multidimensional feature value of the current face image, the second row places the second multidimensional feature value of the sick student benchmark face image, and the third row places the second multidimensional feature value of the frightened student benchmark face image. Then, add a column to the matrix to record the emotional target value, and fill the preset emotional target values of the two benchmark face images into the corresponding rows. At the same time, in order to more clearly represent the relationship between each feature, the server may assign weights to the elements in the matrix according to the category and importance of the feature. For example, the feature weight of the facial expression dimension is set to 0.3, the feature weight of the voice dimension is set to 0.3, and the feature weight of the behavior dimension is set to 0.4. In this way, the server establishes an emotional feature association matrix. This matrix comprehensively shows the association relationship between the current face image and the benchmark face image in each dimensional feature and the corresponding emotional target value. The server uses complex data analysis algorithms and machine learning models to analyze and process the emotional feature association matrix to identify the emotional target value of the current face image. First, the server calculates the similarity between the first multidimensional feature of the current face image and the second multidimensional feature of the benchmark face image according to the weight of each dimensional feature in the matrix. For example, in the facial expression dimension, the similarity between the current student and the sick student benchmark face image is calculated to be 0.8, and the similarity with the frightened student benchmark face image is 0.6; in the voice dimension, the similarities are 0.7 and 0.5 respectively; in the behavior dimension, the similarities are 0.9 and 0.4 respectively. Then, the server multiplies the similarity of each dimension by the corresponding weight and adds them together to obtain the comprehensive similarity. For the sick student benchmark face image, the comprehensive similarity is: (0.8×0.3+0.7×0.3+0.9×0.4)=0.81; for the frightened student benchmark face image, the comprehensive similarity is: (0.6×0.3+0.5×0.3+0.4×0.4)=0.49.Next, the server calculates the emotional target value of the current face image by linear interpolation and other methods based on the comprehensive similarity and the preset emotional target value of the benchmark face image. Assume that the calculated emotional target value of the current student's face image is 6.8 (out of 10). The server determines the emotional type corresponding to the current face image according to the emotional target value obtained by recognition based on the pre-set emotional type classification rules. For example, the emotional type classification rules set in the server are: emotional target values between 0-3 are calm, between 3-6 are mildly uneasy, between 6-9 are moderately uncomfortable or nervous, and between 9-10 are highly urgent. Since the emotional target value of the current student's face image is 6.8, which falls within the range of 6-9, the server determines that the student's current emotional type is moderately uncomfortable or nervous. After determining the emotional type, the server can perform subsequent processing according to the corresponding strategy, such as turning on the intercom function for inquiry and comfort, or calling relevant emergency resources to provide help. Through the above detailed steps, the server can accurately construct the emotion feature association matrix and determine the emotion type corresponding to the current face image based on the matrix, providing an important basis for reasonable response in emergency situations.
[0033] In an embodiment of the present invention, the first multidimensional feature includes a first image feature, and the second multidimensional feature includes a second image feature; the emotional feature association matrix is established based on the second multidimensional feature corresponding to each of the benchmark facial images and the preset emotional target value of each of the benchmark facial images, and the first multidimensional feature of the current facial image, which can be implemented through the following example.
[0034] Performing feature mapping on the first image features to obtain a first image semantic identification set, and performing feature mapping on the second image features corresponding to each of the reference face images to obtain a second image semantic identification set corresponding to each of the reference face images;
[0035] Acquire a first emotion identifier corresponding to the current face image and a second emotion identifier corresponding to each of the reference face images;
[0036] The emotion feature association matrix is established based on the second image semantic identifier set, the second emotion identifier and the emotion target value corresponding to each of the reference facial images, as well as the first image semantic identifier set and the first emotion identifier.
[0037] In an embodiment of the present invention, for example, in a hospital emergency room scenario, a patient comes to the IP network video intercom emergency device due to sudden abdominal pain and triggers an alarm command. After the server starts the camera to collect video data and determines the current face image, it extracts multidimensional features to obtain the first multidimensional features, which include the first image features. The server uses a pre-trained feature mapping model to perform feature mapping on the first image features. For example, the first image feature extracted from the current face image shows that the patient's face is frowning and the forehead is slightly sweating. The server's feature mapping model maps the image feature of "frowning" to the semantic space to obtain the corresponding semantic identifier "pain"; and maps "slightly sweating on the forehead" to "discomfort". By mapping multiple first image features one by one, the server obtains a set of first image semantic identifiers, such as {"pain", "discomfort"}. At the same time, the server finds two reference face images close to the current face image in the reference face image library. One of the reference face images is an image of another patient with abdominal pain before, and its second image features include pale face and chapped lips. The server also performs feature mapping on these second image features, mapping "facial paleness" to "weakness" and "cracked lips" to "dehydration", thereby obtaining a second image semantic identifier set corresponding to the benchmark face image, such as {"weakness", "dehydration"}. Another benchmark face image is an image of a patient who is physically unwell due to anxiety. Its second image features include wandering eyes and frequent blinking. After feature mapping, the second image semantic identifier set obtained is {"anxiety", "uneasiness"}. The server determines the first emotion identifier corresponding to the current face image based on the overall performance of the current face image and the first image semantic identifier set, combined with the pre-set emotion classification rules. In this emergency room scene, the server determines the first emotion identifier as "pain and discomfort" based on the patient's frowning brows, sweating forehead, and the first image semantic identifier set {"pain", "discomfort"}. For the two benchmark face images found, the server also determines the second emotion identifier based on their respective second image semantic identifier sets and overall performance. For the baseline facial image of the patient with abdominal pain, combined with its second image semantic identifier set {"weak", "dehydration"} and the overall morbid facial expression, the server determines its second emotion identifier as "pain and suffering". For the baseline facial image of the patient with physical discomfort caused by anxiety, based on the second image semantic identifier set {"anxiety", "uneasiness"} and the overall tense expression, the server determines its second emotion identifier as "anxiety and uneasiness". The server starts to build an emotional feature association matrix. First, the first image semantic identifier set and the first emotion identifier of the current facial image are used as a row of the matrix. For example, the first row of data is {"pain", "discomfort", "pain and discomfort"}. Then, the second image semantic identifier set, the second emotion identifier and the emotional target value corresponding to the first baseline facial image (the image of the patient with abdominal pain) are used as another row of the matrix.Assuming that the preset emotional target value of the benchmark face image is 8 (full score 10), this row of data is {"weak", "dehydration", "sickness", 8}. Then take the relevant information corresponding to the second benchmark face image (anxious patient image) as another row of the matrix. Assuming that its emotional target value is 7 (full score 10), the data in this row is {"anxiety", "uneasiness", "anxious and uneasy", 7}. In the process of constructing the matrix, the server will also consider the importance and relevance of each feature and assign corresponding weights to different elements. For example, the weight of the image semantic identifier is 0.3, the weight of the emotional identifier is 0.4, and the weight of the emotional target value is 0.3. In this way, the server successfully establishes the emotional feature association matrix. This matrix comprehensively reflects the association between the current face image and the benchmark face image in terms of image semantics, emotional identifiers, and emotional target values. Through further analysis and processing of the matrix, the server can more accurately identify the emotional state of the current patient and provide more powerful support for subsequent emergency treatment. For example, based on the matrix information, the server can determine that the current patient's emotions are closer to the baseline image of patients with abdominal pain, and thus give priority to emergency treatment measures related to abdominal pain, such as arranging medical staff to conduct preliminary examinations and preparing corresponding treatment equipment.
[0038] In the embodiment of the present invention, the number of the first image features includes at least two; and feature mapping is performed on the first image features to obtain a first image semantic identification set, which can be implemented through the following example.
[0039] Performing a feature integration operation on at least two of the first image features to obtain a target enhanced image feature;
[0040] The target enhanced image features are synchronized to a preset semantic feature domain through a pre-trained feature synchronization network to obtain the first image semantic identification set.
[0041] In an embodiment of the present invention, for example, in a busy train station waiting hall, a passenger suddenly feels unwell, and he walks to the IP network video intercom emergency device to trigger an alarm command. After receiving the alarm command, the server starts the camera to collect video data, and determines the current face image from it, and then extracts multi-dimensional features of the image to obtain multiple first image features. Assume that two key first image features are extracted from the current face image: one is the feature of frowning in facial expression, and the other is the feature of slightly pale face. The server first performs a feature integration operation on these two first image features. When performing the feature integration operation, the server uses advanced image processing algorithms and deep learning technologies. For the feature of frowning, the server converts it into a set of quantitative data by analyzing the detailed information such as the degree of wrinkling, angle, and duration of wrinkling, such as the degree of wrinkling is 7 (full score 10), the angle is 30 degrees, etc. For the feature of slightly pale face, the server also analyzes the degree of pallor, distribution area, etc., and converts it into corresponding quantitative data, such as the degree of pallor is 6 (full score 10), mainly distributed in the cheek and forehead areas, etc. Then, the server fuses and comprehensively processes these quantified data. For example, different weights are assigned to different features according to their importance in reflecting physical discomfort by weighted averaging. Assume that the weight of the frown feature is 0.4 and the weight of the pale face feature is 0.6. Then, after calculation, a comprehensive target enhanced image feature value is obtained. The calculation process is as follows: target enhanced image feature value = frown feature quantization value × weight + pale face feature quantization value × weight = 7 × 0.4 + 6 × 0.6 = 2.8 + 3.6 = 6.4 (full score 10). This 6.4 target enhanced image feature value comprehensively reflects the degree of physical discomfort reflected by the two first image features of frown and pale face. It can more comprehensively and accurately describe the current state of the passenger than a single feature, and becomes the target enhanced image feature. After obtaining the target enhanced image feature, the server will synchronize it to the preset semantic feature domain through the pre-trained feature synchronization network. This feature synchronization network is trained with a large amount of data, and it can understand the correspondence between image features and semantic information. In the train station scenario, the server inputs the target enhanced image feature value of 6.4 into the feature synchronization network. The feature synchronization network first analyzes and interprets this value. It knows what semantic information similar values are usually associated with in previous training data. For example, when the target enhanced image feature value is between 6-8, the corresponding semantic information may include "physical discomfort" and "possible health problems". Then, based on these associations, the feature synchronization network maps the target enhanced image features to the preset semantic feature domain. In the preset semantic feature domain, there is a series of predefined semantic identifiers.For the current target enhanced image feature value of 6.4, the feature synchronization network maps it to the two semantic identifiers of "physical discomfort" and "possible health problems". Finally, the server obtains these mapped semantic identifiers from the feature synchronization network to form a first image semantic identifier set, namely {"physical discomfort", "possible health problems"}. This first image semantic identifier set can more intuitively express the semantic information embodied in the current face image, and provides an important basis for the subsequent establishment of the emotional feature association matrix. Through this process, the server successfully converts multiple first image features extracted from the current face image into a first image semantic identifier set with clear semantic meanings, making the analysis and judgment of the passenger's emotional state more accurate and in-depth, and providing strong support for the subsequent appropriate emergency measures. For example, based on this first image semantic identifier set, the server can preliminarily judge that the passenger may need medical assistance, so as to promptly contact the medical staff at the station to assist.
[0042] In an embodiment of the present invention, the emotion feature association matrix is established based on the second image semantic identifier set, the second emotion identifier and the emotion target value corresponding to each of the reference facial images, as well as the first image semantic identifier set and the first emotion identifier, which can be implemented through the following examples.
[0043] Establishing emotion feature association parameters according to a first semantic identifier for distinguishing different information in the same facial image, and the second image semantic identifier set corresponding to the same reference facial image, the second emotion identifier, and the emotion target value, so as to obtain at least two emotion feature association parameters corresponding to at least two reference facial images;
[0044] Merging the at least two emotion feature association parameters to obtain an initial emotion association matrix, wherein the at least two emotion feature association parameters in the initial emotion association matrix are separated by a second semantic identifier for distinguishing information of different facial images;
[0045] Establishing an emotion classification identifier according to the first semantic identifier, the first image semantic identifier set, and the first emotion identifier;
[0046] The emotion feature association matrix is generated according to the initial emotion association matrix and the emotion classification identifier.
[0047] In an embodiment of the present invention, exemplarily, in an office area of an office building, an IP network video intercom emergency device is triggered by an employee. After receiving the alarm command, the server obtains the first image semantic identifier set and the first emotion identifier of the current face image through a series of processing, and also finds two reference face images close to it, and determines their respective corresponding second image semantic identifier sets, second emotion identifiers and emotion target values. Assume that the first semantic identifier is used to distinguish different information such as facial expressions and body postures in the same face image. For the first reference face image, its second image semantic identifier set is {"tired", "dull eyes"}, the second emotion identifier is "work tired", and the emotion target value is 7 (full score 10). The server associates this information according to the first semantic identifier. For example, for information related to facial expressions, "tired" is associated with "facial state" in the first semantic identifier, and "dull eyes" is also associated with "facial state"; for the overall emotional state, "work tired" is associated with "emotional state" in the first semantic identifier; the emotional target value 7 is associated with "emotional quantification". In this way, an emotional feature association parameter is established, which can be expressed as {(facial state: {"fatigue", "dull eyes"}), (emotional state: "work tired"), (emotional quantification: 7)}. For the second benchmark face image, its second image semantic identifier set is {"frown", "slightly trembling"}, the second emotional identifier is "overstressed", and the emotional target value is 8 (out of 10). Similarly, according to the first semantic identifier, the established emotional feature association parameter is {(facial state: {"frown", "slightly trembling"}), (emotional state: "overstressed"), (emotional quantification: 8)}. In this way, the server obtains two emotional feature association parameters corresponding to the two benchmark face images. After obtaining the two emotional feature association parameters, the server merges them to form an initial emotional association matrix. Assuming that the second semantic identifier is used to distinguish different face images, "base image 1" and "base image 2" are used to represent two different benchmark face images respectively. The server arranges the two emotional feature association parameters in a certain order and separates them using the second semantic identifier. For example, the initial sentiment correlation matrix can be expressed as:
[0048]
[0049] In this way, the emotional feature association parameters of different reference face images are clearly separated by the second semantic identifier, forming an ordered initial emotional association matrix. After obtaining the first semantic identifier, the first image semantic identifier set and the first emotional identifier, the server starts to establish an emotional classification identifier. Assume that the first image semantic identifier set of the current face image is {"slightly tired", "slightly dull eyes"}, and the first emotional identifier is "mild fatigue". The server integrates and classifies this information according to the first semantic identifier. For example, according to the "facial state" in the first semantic identifier, "slightly tired" and "slightly dull eyes" in the first image semantic identifier set are classified into one category; according to the "emotional state", the first emotional identifier "mild fatigue" is classified into one category. Then, the server establishes an emotional classification identifier based on these classification information, for example, it can be expressed as {(facial state: {"slightly tired", "slightly dull eyes"}), (emotional state: "mild fatigue")}. This emotional classification identifier can clearly reflect the characteristics and emotional state of the current face image. After the server has the initial emotion association matrix and emotion classification identifier, it merges them to generate the final emotion feature association matrix. The server adds the emotion classification identifier as a new row to the initial emotion association matrix. At the same time, in order to distinguish, a specific identifier, such as "current image", is used to indicate that this row corresponds to the information of the current face image. The generated emotion feature association matrix is as follows:
[0050]
[0051] The "-" here means that the emotional target value of the current face image has not been determined, and further analysis and calculation are required based on this emotional feature association matrix. In this way, the server successfully established the emotional feature association matrix, which combines the relevant information of the baseline face image and the current face image, and provides an important data basis for the subsequent accurate identification of the emotional target value and emotional type of the current face image. For example, the server can use specific algorithms and models to calculate the emotional target value of the current face image based on the relationship and difference between the elements in the matrix, and then determine its emotional type so that appropriate emergency measures can be taken.
[0052] In the embodiment of the present invention, the step of obtaining the first emotion identifier corresponding to the current facial image and the second emotion identifier corresponding to each of the reference facial images may be implemented through the following example.
[0053] Obtaining the location where the current face image was collected and the voice content recognized by the voice data of the current face image;
[0054] Generate emergency scene information of the current face image according to the background image of the current face image;
[0055] The first emotion identifier is generated according to the collection location, the voice content and the emergency scenario information.
[0056] In an embodiment of the present invention, for example, in a large shopping mall, IP network video intercom emergency equipment is distributed in various areas. At this time, a customer suddenly triggered the alarm command of the emergency equipment in the clothing area of the mall. After receiving the alarm command, the server immediately starts to process the relevant information. First, the server obtains the collection location information of the current face image from the IP network video intercom emergency equipment. The emergency equipment has a built-in positioning system, which can accurately send its own location information to the server. In this example, the collection location obtained by the server is near a store in the clothing area on the third floor of the mall. At the same time, the server recognizes the voice data attached to the current face image. When the emergency equipment collects video data, it will also synchronously collect the surrounding sound information. When the customer triggers the alarm command, it may be accompanied by some voice expressions, such as he said: "I suddenly feel dizzy and feel very uncomfortable." After receiving these voice data, the server uses advanced voice recognition technology to convert them into text form, that is, "I suddenly feel dizzy and feel very uncomfortable", thereby obtaining the voice content. After obtaining the current face image, the server will analyze the background image to generate emergency scene information. In this mall scene, the background of the current face image shows that there are many hangers and clothing display racks around, the lighting in the store is normal, the customer flow is moderate, but there is no obvious sign of chaos or danger. The server interprets the background image through the image recognition algorithm and the pre-trained scene analysis model. It recognizes that this is a normal clothing sales area. Combined with the "dizziness and discomfort" mentioned in the customer's voice content previously obtained, the server determines that the current emergency scene is that the customer suddenly feels unwell while shopping in the mall. In order to generate emergency scene information in more detail, the server may also analyze other details in the background image. For example, it is observed that there are no obvious obstacles or dangerous items around the customer, and other customers around do not show abnormal behavior. Combining this information, the emergency scene information generated by the server is: "In a store in the clothing area on the third floor of the mall, a customer suddenly feels dizzy and unwell during normal shopping, and there are no obvious dangerous factors in the surrounding environment." After obtaining the collection location, voice content and emergency scene information, the server begins to combine this information to generate the first emotion identifier. The collection location is the clothing area of the mall, which is a relatively safe and normal shopping environment that generally does not cause people to have strong emotions such as fear and panic. In the voice content, the customer clearly stated that he was "dizzy and felt very uncomfortable", which reflects that the customer's current physical condition is not good, and may be accompanied by some anxiety and uneasiness. The emergency scenario information shows that there are no obvious risk factors in the surrounding environment, and the main problem is the customer's own physical discomfort. The server conducts a comprehensive analysis and judgment of this information based on the pre-set emotion classification rules and algorithms. In this example, based on the customer's voice content and emergency scenario information, the server determines that the customer's current main emotion is anxiety and uneasiness caused by physical discomfort.Therefore, the first emotion identifier generated by the server is "anxiety caused by physical discomfort". This first emotion identifier accurately summarizes the emotional state of the current customer at a specific collection location and a specific emergency scenario. After generating the first emotion identifier, the server can further take corresponding measures based on this identifier. For example, the server can turn on the intercom function of the IP network video intercom emergency equipment, contact the mall staff or medical personnel to go to the collection location, provide help and support to the customer, and at the same time soothe the customer's emotions through voice and inform the customer that relevant personnel are on the way. In this way, the server can respond to emergencies more effectively and ensure the safety and health of customers.
[0057] In an embodiment of the present invention, the first multidimensional feature includes a target enhanced image feature and a first semantic feature; matching at least two reference face images close to the current face image according to the first multidimensional feature can be implemented through the following example.
[0058] Obtaining a preset enhanced image feature matching pool and a semantic feature matching pool, wherein the enhanced image feature matching pool includes a correspondence between each pending face image and the enhanced image feature corresponding to the pending face image, and the semantic feature matching pool includes a correspondence between each pending face image and the semantic feature corresponding to the pending face image;
[0059] According to the target enhanced image feature, the enhanced image feature matching pool and the semantic feature matching pool are matched respectively, and according to the first semantic feature, the enhanced image feature matching pool and the semantic feature matching pool are matched respectively, to obtain at least two target pending facial images close to the current facial image;
[0060] The reference facial image is determined from the at least two target facial images to be determined according to the image matching coefficients between the at least two target facial images to be determined and the current facial image.
[0061] In an embodiment of the present invention, exemplarily, in a smart community, a server stores a large amount of facial image data for assisting emergency emotion recognition. Among them, it includes a preset enhanced image feature matching pool and a semantic feature matching pool. The enhanced image feature matching pool is a huge database that stores many pending facial images and their corresponding enhanced image features. For example, a pending facial image shows the state of a resident after participating in intense exercise in a community activity. The corresponding enhanced image features may include a high degree of facial redness (quantitative value of 8, full score 10), obvious sweat beads on the forehead (quantitative value of 7, full score 10), etc. These features are obtained through advanced image processing algorithms and feature extraction techniques, and can more accurately describe the detailed information of the facial image. The semantic feature matching pool is also a rich data set, which records the corresponding relationship between each pending facial image and the corresponding semantic feature. For example, for the facial image of the resident after participating in intense exercise, the corresponding semantic features may be "fatigue after exercise" and "certain consumption of the body". These semantic features are obtained by analyzing and understanding multiple factors such as the behavior, expression and surrounding environment of the person in the image. When the server needs to match the face image, it will first obtain the data of the two matching pools from the storage system to prepare for the subsequent matching work. One day, a resident in the community triggered the alarm command of the IP network video intercom emergency device at his door. After receiving the command, the server starts the device camera to collect video data, and determines the current face image, and then extracts the target enhanced image features and the first semantic features from it. Assume that the target enhanced image features in the current face image show that the resident's face is slightly red (quantitative value is 5, full score 10), and there are a few beads of sweat on the forehead (quantitative value is 4, full score 10). The server will match these target enhanced image features with the enhanced image features of each pending face image in the enhanced image feature matching pool one by one. During the matching process, the server will calculate the similarity between the two. For example, through a certain similarity calculation algorithm, it is found that the similarity with the enhanced image features of the face image of the resident who participated in intense exercise mentioned above is 0.6 (the similarity value range is 0-1, and the higher the value, the more similar it is). At the same time, the server will also match according to the first semantic feature of the current face image. Assuming that the first semantic feature is "mild discomfort", the server will match this semantic feature with the semantic features of each pending face image in the semantic feature matching pool. When the semantic feature "post-exercise fatigue" of the face image of a resident who has participated in intense exercise is matched, it is found that the two have a certain semantic correlation. After calculation by the semantic analysis algorithm, the semantic similarity is 0.5. Based on the matching results of the comprehensive target enhanced image feature and the first semantic feature, the server will screen out the target pending face image that is close to the current face image.In addition to the above-mentioned facial images of residents after intense exercise, another pending facial image may be found, such as a facial image of a resident after working outdoors for a period of time. The enhanced image features of this image have a similarity of 0.55 with the target enhanced image features of the current facial image, and the semantic similarity of the semantic feature "fatigue after work" with the first semantic feature "slight discomfort" of the current facial image is 0.45. In this way, the server obtains two target pending facial images that are close to the current facial image. After obtaining the two target pending facial images, the server needs to further determine the reference facial image. This requires calculating the image matching coefficient between the target pending facial image and the current facial image. The calculation of the image matching coefficient will comprehensively consider the similarity of the target enhanced image features and the similarity of the first semantic feature. The server will assign different weights to these two similarities, assuming that the weight of the target enhanced image feature similarity is 0.6 and the weight of the first semantic feature similarity is 0.4. For the facial image of residents after participating in intense exercise, the image matching coefficient is calculated as follows: Image matching coefficient = target enhanced image feature similarity × weight + first semantic feature similarity × weight = 0.6 × 0.6 + 0.5 × 0.4 = 0.36 + 0.2 = 0.56. For the facial image of residents after working outdoors for a period of time, the image matching coefficient is calculated as follows: Image matching coefficient = 0.55 × 0.6 + 0.45 × 0.4 = 0.33 + 0.18 = 0.51. By comparing the image matching coefficients of the two target undetermined facial images, the server finds that the image matching coefficient (0.56) of the facial image of residents after participating in intense exercise is higher. Therefore, the server will determine the facial image of residents after participating in intense exercise as the benchmark facial image. After determining the benchmark facial image, the server can use the relevant information of the benchmark facial image, such as the preset emotional target value, to further analyze the emotional state of the current facial image, and provide a more accurate basis for subsequent emergency treatment. For example, if the emotional target value of the baseline facial image is expressed as "moderate fatigue", the server can combine the specific situation of the current facial image to preliminarily judge that the current resident's emotional state also tends to be moderate fatigue, and thus take corresponding measures, such as contacting community medical personnel for a simple examination.
[0062] In an embodiment of the present invention, the at least two target facial images include at least two target facial images obtained by matching each matching strategy; the determination of the reference facial image from the at least two target facial images based on the image matching coefficients between the at least two target facial images and the current facial image can be implemented through the following example.
[0063] For each matching strategy, the at least two target facial images to be determined are matched, and a matching coefficient between the current facial image and each target facial image to be determined is determined;
[0064] For each target undetermined face image, determining the average value of the matching coefficient between the target undetermined face image and the current face image in each matching strategy;
[0065] The reference face image is determined from the at least two target face images according to the matching coefficient mean.
[0066] In an embodiment of the present invention, exemplarily, in a large shopping mall, the server is responsible for processing various information transmitted by the IP network video intercom emergency equipment. One day, a customer suddenly felt unwell in the mall and triggered the alarm instruction of the emergency equipment. After receiving the instruction, the server starts the relevant process, obtains the target enhanced image features and the first semantic features through the previous steps, and obtains multiple target pending face images by matching in the enhanced image feature matching pool and the semantic feature matching pool through different matching strategies. It is assumed here that there are two matching strategies: a matching strategy based on image feature similarity and a matching strategy based on semantic feature relevance. The matching strategy based on image feature similarity mainly focuses on the similarity between the target enhanced image features and the enhanced image features of the pending face image; the matching strategy based on semantic feature relevance focuses on the relevance between the first semantic feature and the semantic features of the pending face image. Through the matching strategy based on image feature similarity, the server found two target pending face images. The first target face image shows a customer who is slightly tired after walking for a long time in the mall. The calculated similarity between its enhanced image features and the target enhanced image features of the current face image is 0.7 (the similarity range is 0-1, and the higher the value, the more similar it is). The second target face image shows a customer who has suffered from heatstroke in the mall. The similarity between its target enhanced image features and the target enhanced image features of the current face image is 0.6. Using the matching strategy based on the correlation of semantic features, the server found two more target face images. One of them is an image of a customer who is anxious because he can't find his child in the mall. The correlation between its semantic features and the first semantic features of the current face image is determined to be 0.4 after analysis (the correlation range is 0-1, and the higher the value, the stronger the correlation). The other is an image of a customer who feels tired after shopping in the mall. The correlation between its semantic features and the first semantic features of the current face image is 0.5. In this way, the server confirms the matching coefficient between the current face image and each target face image for each matching strategy. After determining the matching coefficient between each target face image and the current face image under different matching strategies, the server needs to calculate the average matching coefficient of each target face image in all matching strategies. For the face image of a customer who looks a little tired after walking for a long time, its matching coefficient is 0.7 under the matching strategy based on image feature similarity. Since it is not matched in the matching strategy based on semantic feature relevance, the matching coefficient under this strategy defaults to 0. Then its matching coefficient mean is calculated as follows: matching coefficient mean = (0.7 + 0) / 2 = 0.35.For the facial image of the customer who has suffered from heatstroke, the matching coefficient is 0.6 under the matching strategy based on image feature similarity, and the matching coefficient defaults to 0 under the matching strategy based on semantic feature relevance, and the mean matching coefficient is: matching coefficient mean = (0.6 + 0) / 2 = 0.3. For the facial image of the customer who is anxious because he can't find his child, the matching coefficient is 0.4 under the matching strategy based on semantic feature relevance, and the matching coefficient defaults to 0 under the matching strategy based on image feature similarity, and the mean matching coefficient is: matching coefficient mean = (0 + 0.4) / 2 = 0.2. For the facial image of the customer who feels tired after shopping, the matching coefficient is 0.5 under the matching strategy based on semantic feature relevance, and the matching coefficient defaults to 0 under the matching strategy based on image feature similarity, and the mean matching coefficient is: matching coefficient mean = (0 + 0.5) / 2 = 0.25. Through such calculations, the server obtains the mean matching coefficient of each target pending facial image in all matching strategies. After calculating the mean values of the matching coefficients of each target face image to be determined, the server will determine the benchmark face image based on these mean values. The higher the mean value of the matching coefficient, the higher the overall matching degree between the target face image to be determined and the current face image. In the above example, the mean value of the matching coefficient of the face image of the customer who is slightly tired after walking for a long time is 0.35, which is the highest mean value of the matching coefficient among the four target face images to be determined. Therefore, the server will determine the face image of the customer who is slightly tired after walking for a long time as the benchmark face image. After determining the benchmark face image, the server can use the relevant information of the benchmark face image, such as the preset emotional target value, to further analyze the emotional state of the current customer. For example, if the preset emotional target value of the benchmark face image is expressed as "mild fatigue", the server can preliminarily judge that the emotional state of the customer who currently triggers the alarm instruction is also inclined to mild fatigue based on the specific situation of the current face image. Based on this judgment, the server can take corresponding emergency measures, such as sending a notification to the mall staff through the IP network video intercom emergency equipment, informing them that a customer may be unwell and needs help, such as guiding the customer to a rest area, providing drinking water, etc. At the same time, the server can also send some warm reminders and comforting messages to customers through emergency equipment, informing customers that relevant personnel are on the way and asking customers not to worry too much.
[0067] In this way, the server can more accurately determine the baseline facial image, thereby more effectively analyzing the emotional state of the current facial image and providing strong support for emergency processing.
[0068] In the embodiment of the present invention, the multi-dimensional feature extraction of the current face image to obtain the first multi-dimensional feature can be implemented through the following example.
[0069] Extracting at least two images from the current face image, and dividing the at least two images into at least two image sets;
[0070] Selecting a target image from the at least two image sets, and performing enhanced image feature extraction on the target image to obtain a first image feature;
[0071] Acquire a first emotion marker corresponding to the current facial image, and perform semantic feature extraction on the first emotion marker to obtain a first semantic feature;
[0072] The first multidimensional feature is obtained according to the first image feature and the first semantic feature.
[0073] In an embodiment of the present invention, exemplarily, in a terminal hall of an airport, a passenger triggers an alarm command of an IP network video intercom emergency device. After receiving the alarm command, the server starts the camera to collect video data and determines the current face image from it. The server starts to extract images from the current face image. Assuming that the duration of the face image is 10 seconds, the server extracts two images from the image at a certain time interval, which are the images at the 3rd second and the 7th second. The two images capture the state of the passenger at different times. Next, the server segments the two images to form at least two image sets. The segmentation is based on features such as the passenger's facial expression and body posture in the image. In the image at the 3rd second, the passenger's facial expression is slightly anxious, with slightly wrinkled brows; the body posture is relatively relaxed, and there is no obvious nervous movement. Based on these features, the server divides this part of the image into an image set representing "mild anxiety but relaxed body". In the image at the 7th second, the passenger's facial expression becomes more anxious, his brows are wrinkled more tightly, and his eyes reveal a hint of uneasiness; at the same time, his body posture becomes somewhat stiff, and his hands are unconsciously clenched. Based on these features, the server divides this image into another image set representing "increased anxiety and physical tension". Through such a segmentation operation, the server classifies the two extracted images into two different image sets, providing a more targeted data basis for subsequent feature extraction and analysis. After completing the segmentation of the image set, the server needs to select the target image from these image sets. In this example, since the passenger state reflected by the image set of "increased anxiety and physical tension" is more prominent and critical, the server selects the 7th second image in this set as the target image. After selecting the target image, the server begins to extract enhanced image features. The server uses advanced image processing algorithms and deep learning technology to perform multi-faceted feature analysis on the target image. First, the server focuses on the facial features of the passenger in the target image. Through detailed facial analysis, the server detects that the passenger's frown reaches 6 (out of 10), the pupil of the eye is slightly enlarged, the quantitative value is 5 (out of 10), and the lips are slightly trembling, the quantitative value is 4 (out of 10). The quantitative values of these facial features reflect the passenger's anxious emotional state. Then, the server analyzes the body posture features of the passenger in the target image. It is found that the passenger's body is slightly leaning forward, with a quantitative value of 3 (out of 10), and the degree of clenching his hands is relatively high, with a quantitative value of 7 (out of 10). These body posture features further indicate the passenger's nervousness. Based on the above analysis, the server extracts the first image features from the target image, which are presented in the form of quantitative data, such as {(degree of frowning: 6), (degree of pupil dilation: 5), (degree of lip trembling: 4), (degree of body leaning forward: 3), (degree of clenching his hands: 7)}. These first image features can more accurately and detailedly describe the passenger's current state.After obtaining the first image feature, the server also needs to obtain the first emotion identifier corresponding to the current face image. Based on the previous analysis of the image and the scene that the passenger may be in (waiting for the flight in the airport terminal), the server comprehensively determines that the first emotion identifier of the current passenger is "flight-related anxiety". After determining the first emotion identifier, the server extracts semantic features for it. The server uses natural language processing technology and semantic analysis algorithms to conduct an in-depth analysis of the identifier "flight-related anxiety". From a semantic level, the server parses that "flight-related" means that this anxiety is closely related to the situation of the flight, which may be a worry about flight delays, lost luggage and other problems; "anxiety" reflects that the passenger's current psychological state is uneasy and worried. The server further refines and quantifies this semantic information. For example, the semantic features of "flight-related" can be quantified as {(flight delay worry: 0.8), (luggage problem worry: 0.2)}, indicating that the passenger's main worry is flight delay; the semantic features of "anxiety" can be quantified as {(uneasiness level: 0.7), (worry level: 0.9)}. Through such semantic feature extraction, the server obtains the first semantic features, such as {(flight delay worry: 0.8), (luggage problem worry: 0.2), (uneasiness: 0.7), (worry level: 0.9)}. These first semantic features describe the passenger's emotional state from a semantic perspective. After obtaining the first image feature and the first semantic feature, the server merges them to obtain the first multidimensional feature. The first multidimensional feature is an information set that combines image features and semantic features, which can more comprehensively and deeply reflect the passenger state represented by the current face image. In this example, the server merges the first image feature and the first semantic feature to obtain the first multidimensional feature as follows: {(brow wrinkling degree: 6), (pupil dilation degree: 5), (lip trembling degree: 4), (body leaning forward degree: 3), (hands clenched degree: 7), (flight delay worry: 0.8), (luggage problem worry: 0.2), (uneasiness: 0.7), (worry level: 0.9)}. Through such first multi-dimensional features, the server can more accurately understand the current state of the passenger, providing rich and detailed data support for subsequent matching of benchmark facial images, determining emotion types, and taking corresponding emergency measures. For example, the server can further analyze the severity of the passenger's emotions based on these features. If the emotion is judged to be more serious, the airport staff can be arranged to assist the passenger in time, understand the specific situation and provide corresponding help and comfort.
[0074] In the embodiment of the present invention, the identification of the emotional target value of the current facial image according to the emotional feature association matrix can be implemented through the following examples.
[0075] Performing emotional element feature conversion on the emotional feature association matrix to obtain an emotional feature set, wherein the emotional element feature is standardized data corresponding to a target value recognition model;
[0076] The emotion feature set is loaded into the target value recognition model, wherein the target value recognition model is obtained by maintaining the basic model parameters of the preset multimodal large model unchanged and optimizing the updated model parameters of the multimodal large model according to the training emotion feature set, wherein the updated model parameters are related to the attention mechanism added to the multimodal large model;
[0077] Obtain the target emotion target value of the current facial image generated by the target value recognition model.
[0078] In an embodiment of the present invention, for example, in a busy commercial district, an IP network video intercom emergency device is triggered by a passerby. After a series of complex processing, the server has constructed an emotional feature association matrix. This matrix contains various feature information of the current face image and multiple benchmark face images, such as facial expressions, semantic identifiers, and emotional target values. Now, the server will perform emotional element feature conversion on this emotional feature association matrix. Assume that the elements in the emotional feature association matrix contain feature values of different dimensions, such as the quantitative value of facial expression (0-10), the correlation degree of semantic features (0-1), and the emotional target value (0-10). The value ranges and data types of these feature values are different, and they need to be standardized to meet the input requirements of the target value recognition model. The server first standardizes the facial expression quantization values in the matrix. For example, the facial expression quantization value of the current face image is 8, and the server converts it into standardized data that can be accepted by the target value recognition model through a standardized algorithm. Assume that the standardization algorithm is to subtract the mean from the original value and then divide it by the standard deviation. After calculation, the value of 8 is converted to 0.6 (this is just an example, and the actual calculation will be performed according to the specific mean and standard deviation). For the correlation of semantic features, the server also performs similar processing. For example, the correlation between the current face image and a semantic feature is 0.7, which becomes 0.5 after standardization. The emotional target value is also standardized accordingly. For example, the emotional target value of a benchmark face image is 6, which becomes 0.4 after standardization. After the server performs such emotional element feature conversion on each element in the emotional feature correlation matrix, the converted data is organized into an emotional feature set. All data in this emotional feature set are standardized data corresponding to the target value recognition model, such as {(facial expression standardized value: 0.6), (semantic feature correlation standardized value: 0.5), (emotion target value standardized value: 0.4)}. This data form is more conducive to the processing and analysis of the subsequent target value recognition model. After obtaining the emotional feature set, the server loads it into the target value recognition model. This target value recognition model is obtained through careful training and optimization. The target value recognition model is built based on the preset multimodal large model. The multimodal large model itself has powerful processing capabilities and can process various types of data, such as images, text, etc. When training the target value recognition model, the server maintains the basic model parameters of the multimodal large model unchanged, in order to retain the general knowledge and feature representation capabilities learned by the multimodal large model on large-scale data. At the same time, the server optimizes the updated model parameters of the multimodal large model based on the training emotion feature set. The training emotion feature set is obtained from a large amount of facial image data with known emotion target values through a feature extraction and conversion process similar to the previous one.By letting the multimodal large model learn the data patterns and rules in these training emotion feature sets, the model parameters are adjusted and updated so that it can more accurately identify the emotion target value. In this process, the attention mechanism added to the multimodal large model plays a key role. The attention mechanism enables the model to pay more attention to the important features related to the emotion target value. For example, in the current commercial block scene, for the facial images of passers-by, the attention mechanism may pay more attention to the key parts of the facial expressions such as eyes and eyebrows that can reflect the emotional state, as well as the key information related to the current scene in the semantic features, such as whether the physical discomfort is mentioned, whether anxiety is shown, etc. When the server loads the emotion feature set into the target value recognition model, the model uses its optimized updated model parameters and attention mechanism to conduct in-depth analysis and processing of the input emotion feature set. It assigns different attention weights according to the importance of each feature, and focuses more on those features that play a key role in the recognition of the emotion target value. After processing and analysis by the target value recognition model, the server obtains the target emotion target value of the current face image generated by the model. In the commercial block scene, the target value recognition model comprehensively considers various standardized data in the emotion feature set and the key features focused on by the attention mechanism. For example, the model noticed that the normalized value of facial expression of the current face image is high, indicating that the facial expression shows a strong emotion; the normalized value of semantic feature correlation also shows a high correlation with semantic features related to certain negative emotions. Based on these analyses, the target value recognition model generates a target emotion target value for the current face image. Assuming that the generated target emotion target value is 8.5 (out of 10), this value indicates that the passerby is currently in a state of high tension or anxiety. After the server obtains this target emotion target value, it can determine the emotion type corresponding to the current face image according to the pre-set emotion type classification rules. For example, if 7-10 is set as a highly nervous and uneasy type, then the server can determine that the current emotion type of the passerby is highly nervous and uneasy. Based on this judgment result, the server can take corresponding emergency measures. For example, through the IP network video intercom emergency equipment, a notification is sent to nearby security personnel or medical personnel to inform them that a passerby is in a highly nervous and uneasy state and needs to go to check and provide help in time. At the same time, the server can also send some comforting messages to passersby through the emergency equipment to inform them that relevant personnel are coming and ask them to stay calm as much as possible. In this way, the server uses the emotion feature association matrix to convert emotion element features, loads them into the target value recognition model, and obtains the target emotion target value. It can accurately identify the emotional state of the current facial image and provide strong support for timely and effective emergency response.
[0079] In the embodiment of the present invention, the training step of the target value recognition model can be implemented through the following examples.
[0080] Acquire a first training enhanced image feature, a first training emotion identifier, and a first training emotion target value corresponding to the first training face image, and a second training enhanced image feature, a second training emotion identifier, and a second training emotion target value corresponding to the second training face image;
[0081] Performing feature synchronization operations on the first training enhanced image features and the second training enhanced image features according to the pre-trained initial feature synchronization network, respectively, to obtain a first training image identification set and a second training image identification set, and establishing a training emotion feature association matrix according to the first training image identification set, the first training emotion identification and the first training emotion target value, and the second training image identification set and the second training emotion identification;
[0082] Convert the training emotion feature association matrix into emotion element features to obtain the training emotion feature set;
[0083] Adding the attention mechanism to the multimodal large model to add the updated model parameters to the multimodal large model through the attention mechanism;
[0084] The basic model parameters of the multimodal large model are maintained unchanged, and the updated model parameters of the multimodal large model are optimized according to the training emotion feature set and the second training emotion target value to obtain the target value recognition model.
[0085] In an embodiment of the present invention, by way of example, in a large medical data center, a server is responsible for managing and processing a large amount of face image data, which will be used to train a target value recognition model. The server first obtains a first training face image from the stored massive data. Suppose this first training face image is from a patient suffering from a cold and is captured and recorded during the patient's visit to the hospital. The server analyzes the image using advanced image processing algorithms to extract the first training enhanced image features. For example, it is observed from the image that the patient's face is slightly pale and the eyes are a bit swollen. Through quantitative analysis, the server quantifies the degree of facial paleness as 7 (out of 10 full marks) and the degree of eye swelling as 6 (out of 10 full marks). These quantified values form part of the first training enhanced image features. At the same time, based on the patient's medical record information and current manifestations, the server determines that the first training emotion label of this patient is "physically uncomfortable and slightly anxious". This is because when the patient is in a sick state, not only does the patient feel physically uncomfortable, but may also be anxious due to concerns about the illness. In addition, after evaluation and annotation by professional doctors, the emotion target value of this patient at this time is set as 8 (out of 10 full marks) as the first training emotion target value, and this value reflects the degree of tension and uneasiness of the patient's current emotion. Then, the server obtains a second training face image. Suppose this is an image of a working professional who has encountered a major stress event at work. The server also processes this image to extract the second training enhanced image features. For example, it is observed that the working professional has a frowning brow and tight lips. Through quantitative analysis, the degree of brow frowning is 8 (out of 10 full marks) and the degree of lip tightness is 7 (out of 10 full marks). Based on their work status and behavioral manifestations, the server determines that their second training emotion label is "anxiety caused by work stress". And after professional psychological evaluation, the emotion target value of this working professional at this time is marked as 9 (out of 10 full marks) as the second training emotion target value. After the server obtains the first training enhanced image features and the second training enhanced image features, it performs a feature synchronization operation on them using a pre-trained initial feature synchronization network. For the first training enhanced image features, the initial feature synchronization network maps them to a pre-set semantic feature space. For example, the feature with a facial paleness degree of 7 is mapped to the "physically weak" label in the semantic feature space, and the feature with an eye swelling degree of 6 is mapped to the "possible inflammation" label, thus forming a first training image label set, such as {"physically weak", "possible inflammation"}. Similarly, for the second training enhanced image features, after the feature synchronization operation, the feature with a brow frowning degree of 8 is mapped to the "nervous" label, and the feature with a lip tightness degree of 7 is mapped to the "emotionally anxious" label, obtaining a second training image label set, such as {"nervous", "emotionally anxious"}.Next, the server establishes a training emotion feature association matrix based on the first training image identifier set, the first training emotion identifier and the first training emotion target value, as well as the second training image identifier set and the second training emotion identifier. In this matrix, the first row corresponds to the relevant information of the first training face image, including the first training image identifier set {"weak body", "may have inflammation"}, the first training emotion identifier "physical discomfort and slightly anxious" and the first training emotion target value 8. The second row corresponds to the information of the second training face image, including the second training image identifier set {"mental tension", "emotional anxiety"}, the second training emotion identifier "anxiety caused by work pressure" and the second training emotion target value 9. The server performs emotion element feature conversion on the established training emotion feature association matrix, with the purpose of converting various eigenvalues in the matrix into standardized data that can be processed by the target value recognition model, thereby obtaining a training emotion feature set. For example, for the quantified value of facial pallor 7 of the first training face image, the server converts it into a standardized value of 0.6 through a specific standardization algorithm (such as subtracting the mean and dividing by the standard deviation). For the first training emotion marker "feeling unwell and slightly anxious", the server performs semantic encoding and converts it into a standardized representation in the form of a vector, such as [0.8, 0.2], where 0.8 represents the degree of "feeling unwell" and 0.2 represents the degree of "feeling slightly anxious". Similarly, similar conversion operations are performed on each feature of the second training face image. After comprehensive emotional element feature conversion, the server organizes the converted data into a training emotion feature set. The data in this set has a unified format and range, which is convenient for subsequent model training and processing. After preparing the training emotion feature set, the server adds the attention mechanism to the preset multimodal large model. The attention mechanism enables the model to pay more attention to important feature information when processing data. In this scenario, by adding the attention mechanism to the multimodal large model, the server enables the model to focus more on features that are closely related to the emotional target value when processing the training emotion feature set. For example, for the features of the first training face image, the attention mechanism may pay more attention to features related to physical conditions such as "weakness" and "possible inflammation", as well as emotional markers such as "feeling unwell and slightly anxious", because these features are important for determining the emotional target value. Through the attention mechanism, the multimodal large model assigns higher weights to these important features, thereby adding updated model parameters related to these important features to the model. After adding the attention mechanism and updating the model parameters, the server keeps the basic model parameters of the multimodal large model unchanged. This is because the basic model parameters have learned some common knowledge and feature representations on large-scale data, and keeping them unchanged can retain this useful information. Then, the server optimizes the updated model parameters of the multimodal large model based on the training emotional feature set and the second training emotional target value.During the optimization process, the server inputs the training emotion feature set as input data into the multimodal large model. After the model performs weighted processing on the features according to the attention mechanism, it generates a predicted emotion target value. For example, for the second training face image, the predicted emotion target value generated by the model based on its training emotion feature set may be initially 8.5. The server compares this predicted value with the actual second training emotion target value of 9 and calculates the difference between the two (such as mean square error, etc.). Based on this difference, the server uses an optimization algorithm (such as gradient descent) to adjust and update the model parameters so that the predicted emotion target value generated by the model gradually approaches the actual target value. After multiple iterations of optimization, when the difference between the model's prediction result and the actual target value meets certain convergence conditions, the server obtains the optimized target value recognition model. Through such training steps, the server successfully trained a target value recognition model, which can more accurately identify the emotion target values of different face images, providing strong support for subsequent emergency emotion recognition and processing.
[0086] In the embodiment of the present invention, before the feature synchronization operation is performed on the enhanced image features according to the pre-trained initial feature synchronization network, the following implementation manner is also provided.
[0087] Obtaining training face images and training emotional semantic information corresponding to the training face images;
[0088] Performing feature extraction on the training face image to obtain training face emotion features, and loading the training face emotion features into an original untrained network, so that the original untrained network aligns the training face emotion features to an input semantic feature domain of the multimodal large model;
[0089] Acquire the training semantic features corresponding to the synchronization request, and load the training semantic features and the target training image semantic features generated by the feature synchronization network into the multimodal large model, wherein the model parameters of the multimodal large model remain unchanged;
[0090] The training predicted emotional semantic information generated by the multimodal large model is obtained, and the original untrained network is trained according to the training emotional semantic information and the training predicted emotional semantic information to obtain the initial feature synchronization network.
[0091] In an embodiment of the present invention, illustratively, the server first obtains training face images from its storage system. Assume that these training face images come from a psychology research project, and researchers have taken a large number of participants' facial photos in different scenarios. For example, one of the training face images shows the expression of a participant when facing a stress test, and the face in the image shows a frown and focused eyes. At the same time, the server also obtains the training emotional semantic information corresponding to this training face image. This information is annotated by professional psychology researchers based on the actual emotional state of the participants when taking the photos. For the above-mentioned face image, the corresponding training emotional semantic information is annotated as "focus and tension when facing stress." This detailed semantic information provides clear goals and guidance for subsequent model training. After obtaining the training face images, the server uses advanced image processing and feature extraction algorithms to analyze them. For the above-mentioned face image showing the participant in the stress test, the server extracts the training face emotional features by analyzing the morphology and changes of various parts of the face, such as eyes, eyebrows, mouth, etc. For example, the server detects that the participant in the picture has a high degree of frowning, which is quantified as 8 (out of 10); the pupil of the eye is slightly enlarged, which is quantified as 6 (out of 10); the lips are tightly closed, which is quantified as 7 (out of 10). These quantized feature values together constitute the training face emotion features. Next, the server loads these training face emotion features into the original untrained network. The original untrained network is a neural network model that has not been fully trained. Its main task is to learn how to effectively process and transform the input feature data. In this process, the original untrained network will try to align the training face emotion features to the input semantic feature domain of the multimodal large model. The input semantic feature domain of the multimodal large model is a pre-defined feature space that can process and understand multiple types of semantic information. The original untrained network adjusts its own network structure and parameters so that the input training face emotion features can be correctly represented and understood in the input semantic feature domain of the multimodal large model. For example, the original untrained network may learn to match the feature of frowning degree 8 with the semantic identifier of "nervous emotion" in the input semantic feature domain, and match the feature of pupil dilation degree 6 with the semantic identifier of "focused attention", thereby achieving feature alignment. After completing feature alignment, the server will receive a synchronization request. This synchronization request triggers the server to obtain the corresponding training semantic features. The training semantic features are professionally annotated semantic information related to the training face image, which can more accurately describe the emotional state of the person in the image. For the above face image, the corresponding training semantic features may include "nervous emotion under stressful situations" and "psychological state of focusing on coping with challenges". At the same time, the feature synchronization network will generate target training image semantic features based on the training face image.The feature synchronization network is a network model specifically used to convert image features into semantic features. It can generate semantic features corresponding to image features by learning the correspondence between a large amount of image and semantic data. Assume that the feature synchronization network generates the target training image semantic features as "facial expression shows high tension, body posture suggests that it is concentrating on coping with difficulties" based on the facial expressions and postures of the participants in the picture. The server loads the acquired training semantic features and the target training image semantic features generated by the feature synchronization network into the multimodal large model. In this process, the model parameters of the multimodal large model remain unchanged, in order to ensure that the basic structure and existing knowledge of the model are not changed, and only use its powerful feature processing and semantic understanding capabilities to process the input feature data. After receiving the training semantic features and the target training image semantic features, the multimodal large model will process and analyze them according to its own model structure and parameters to generate training prediction emotional semantic information. For the face picture in the above example, the training prediction emotional semantic information that the multimodal large model may generate is "under stressful environment, showing a high degree of tension and concentration". After the server obtains the training predicted emotional semantic information generated by the multimodal large model, it will compare and analyze it with the previously obtained training emotional semantic information. By comparing the differences between the two, the server can evaluate the accuracy and effect of the original untrained network in feature alignment and semantic conversion. For example, if the training emotional semantic information is labeled "concentration and tension when facing pressure", and the training predicted emotional semantic information is "under pressure, showing a high degree of tension and concentration", there is a certain semantic difference between the two. According to this difference, the server will use a suitable training algorithm (such as back propagation algorithm) to adjust and optimize the original untrained network. During the training process, the server will adjust the network structure and parameters of the original untrained network according to the size and direction of the difference, so that the original untrained network can more accurately align the training face emotional features to the input semantic feature domain of the multimodal large model, thereby generating training predicted emotional semantic information that is closer to the training emotional semantic information. After multiple iterations of training, when the difference between the training predicted emotional semantic information and the training emotional semantic information meets certain convergence conditions, the server considers that the original untrained network has been trained and the initial feature synchronization network is obtained. This initial feature synchronization network can more accurately convert image features into semantic features, providing an important foundation for the subsequent establishment of the emotion feature association matrix and the training of the target value recognition model. Through the above steps, the server completes the training process of the initial feature synchronization network, laying a solid foundation for building a more accurate and efficient emotion recognition model.
[0092] In an embodiment of the present invention, the updating model parameters of the multimodal large model are optimized according to the training emotion feature set and the second training emotion target value to obtain the target value recognition model, which can be implemented through the following examples.
[0093] Loading the training emotion feature set into the multimodal large model to which the attention mechanism is added, and obtaining the predicted training target value generated by the multimodal large model;
[0094] According to the difference between the predicted training target value and the second training emotion target value, optimizing the updated model parameters of the multimodal large model to obtain the target value recognition model;
[0095] The method further comprises:
[0096] According to the difference between the predicted training target value and the second training emotion target value, the network parameters of the initial feature synchronization network are optimized.
[0097] In an embodiment of the present invention, exemplarily, in an intelligent security monitoring center, the server is responsible for processing a large amount of face image data from various monitoring areas to train a target value recognition model for emergency emotion recognition. At this point, the server has completed the construction of the training emotion feature set and has added an attention mechanism to the multimodal large model. Assume that the training emotion feature set is obtained by a series of operations such as feature extraction and conversion of face image data in multiple different scenes. These data contain rich information, such as facial expression features (such as the degree of frowning of eyebrows, changes in the size of eyes, etc.), semantic features (such as "anxiety" and "fatigue" and other emotional markers), etc., and have been standardized for model processing and analysis. The server loads this training emotion feature set into the multimodal large model with the attention mechanism added. The role of the attention mechanism is to enable the model to pay more attention to key features related to the emotional target value. For example, when processing a face image showing someone in an emergency, the attention mechanism will pay more attention to the features of facial expressions that reflect nervous emotions, such as frowning eyebrows, wide eyes, etc., as well as semantic features related to emergency situations. After receiving the training emotion feature set, the multimodal large model will process and analyze it according to its internal complex neural network structure and algorithm. Each neuron in the model will perform weighted calculations, activation function processing and other operations on the input features, and finally generate a predicted training target value. For example, for a specific training sample, its training emotion feature set includes facial expression features (the standardized value of frowning degree is 0.8, and the standardized value of wide eyes is 0.7) and semantic features (the semantic identifier standardized vector of "high tension" is represented as [0.9, 0.1]). After a series of calculations and processing, the multimodal large model generates a predicted training target value of 8.5 (out of 10), indicating that the model believes that the emotional tension corresponding to the sample is 8.5. After obtaining the predicted training target value generated by the multimodal large model, the server will compare it with the second training emotion target value to evaluate the prediction accuracy of the model and optimize the updated model parameters based on the difference between the two. Assume that for the above-mentioned specific training sample, its second training emotion target value is marked by professionals according to the actual situation, which is 9 (out of 10). The server calculates the difference between the predicted training target value (8.5) and the second training emotion target value (9). For example, the mean square error (MSE) is used as the evaluation indicator, and the calculated mean square error is (8.5-9)^2=0.25. This difference indicates that there is a certain deviation between the model's prediction result and the actual labeled value, and the model's updated model parameters need to be adjusted and optimized. The server will use an optimization algorithm (such as a gradient descent algorithm) to update the model parameters. Specifically, the server will adjust the value of the updated model parameters according to a certain learning rate based on the gradient information of the mean square error to the updated model parameters, so that the model's prediction results can be closer to the actual second training emotion target value.For example, suppose the initial value of a certain update model parameter is 0.5, and the gradient of the parameter with respect to the mean square error is calculated to be -0.1, and the learning rate is set to 0.01. Then, according to the gradient descent algorithm, the update value of the update model parameter is 0.5-0.01*(-0.1)=0.501. The server will make similar adjustments to all update model parameters. After multiple iterations of optimization, the update model parameters are continuously adjusted until the difference between the predicted training target value and the second training emotion target value meets certain convergence conditions (such as the mean square error is less than a preset threshold). When the convergence conditions are met, the server obtains the optimized target value recognition model. This target value recognition model can more accurately predict the emotion target value based on the input emotion feature set, thereby improving the accuracy of emergency emotion recognition. In addition to optimizing the update model parameters of the multimodal large model, the server will also optimize the network parameters of the initial feature synchronization network based on the difference between the predicted training target value and the second training emotion target value. The function of the initial feature synchronization network is to convert image features into semantic features, and the accuracy of its conversion is crucial to the entire emotion recognition process. If there is a large difference between the predicted training target value and the second training emotion target value, it may mean that there are certain problems in the feature conversion process of the initial feature synchronization network, and its network parameters need to be adjusted. Continuing with the above training sample as an example, since there is a difference between the predicted training target value (8.5) and the second training emotion target value (9), the server will analyze which parts of the initial feature synchronization network may cause this difference. For example, when converting facial expression features into semantic features, the weight distribution of some features may be unreasonable, resulting in inaccurate representation of semantic features. Based on this difference, the server will use methods such as back propagation algorithms to adjust the network parameters of the initial feature synchronization network. Specifically, the server will start from the output end of the multimodal large model (i.e., the predicted training target value), reversely calculate the gradient of the error to the network parameters of each layer of the initial feature synchronization network, and then update the network parameters based on the gradient information and the set learning rate. For example, assuming that the initial value of a network parameter of a layer in the initial feature synchronization network is 0.4, the gradient of the parameter to the difference between the predicted training target value and the second training emotion target value is 0.05 after back propagation calculation, and the learning rate is set to 0.02. Then, the updated value of the network parameter is 0.4+0.02*0.05=0.401. By continuously optimizing the network parameters of the initial feature synchronization network according to the difference between the predicted training target value and the second training emotion target value, the feature conversion accuracy of the initial feature synchronization network can be improved, so that the converted semantic features can more accurately reflect the emotional state of the face image, further improving the performance and accuracy of the entire emotion recognition system.In summary, the server loads the training emotion feature set into the multimodal large model with an added attention mechanism, and optimizes the updated model parameters of the multimodal large model and the network parameters of the initial feature synchronization network according to the difference between the predicted training target value and the second training emotion target value. Finally, a more accurate and reliable target value recognition model is obtained, which provides strong support for emergency emotion recognition.
[0098] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned artificial intelligence-based IP network video intercom emergency emotion recognition method. Figure 2 As shown, Figure 2 The computer device 100 is a block diagram of a structure of a computer device 100 provided in an embodiment of the present invention. The computer device 100 includes a memory 111, a processor 112 and a communication unit 113. To achieve data transmission or interaction, the memory 111, the processor 112 and the communication unit 113 are directly or indirectly electrically connected to each other.
[0099] For illustrative purposes, the foregoing description is made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise form disclosed. Numerous modifications and variations are possible in accordance with the above teachings. These embodiments are selected and described in order to best illustrate the principles of the present disclosure and its practical application, so that those skilled in the art can best utilize the present disclosure and utilize various embodiments with different modifications to suit the intended specific application.
Claims
1. An artificial intelligence-based IP network video intercom emergency emotion recognition method, characterized in that: include: In response to an alarm instruction of the IP network video intercom emergency device triggered by the current user, starting a camera of the IP network video intercom emergency device; Collecting current video data through the camera, and determining the current face image from the video data; An emotion feature association matrix of the current face image is constructed, and the emotion type corresponding to the current face image is determined according to the emotion feature association matrix.
2. The method according to claim 1, characterized in that: The method further comprises: When the emotion type belongs to the first-level emotion type, the intercom function of the IP network visual intercom emergency device is turned on to realize manual real-time reply interaction; When the emotion type belongs to the secondary emotion type, current speech data from the current user is loaded into the first generative dialogue model of the integrated model to obtain the basic goal of the current user; the urgency of the secondary emotion type is greater than that of the primary emotion type in an emergency scenario; Loading the past speech contents of all users within a preset period range and the basic goal into the second generative dialogue model of the integrated model to obtain the final goal of the current user; The final goal is adapted to a plurality of preset emergency strategies in a preset emergency strategy database to obtain a target emergency strategy adapted to the current user, and the target emergency strategy is provided to the current user.
3. The method according to claim 1, characterized in that: The step of constructing an emotion feature association matrix of the current face image and determining the emotion type corresponding to the current face image according to the emotion feature association matrix comprises: acquiring the current face image, extracting at least two images from the current face image, and dividing the at least two images into at least two image sets; Selecting a target image from the at least two image sets, and performing enhanced image feature extraction on the target image to obtain a first image feature; Acquire a first emotion marker corresponding to the current facial image, and perform semantic feature extraction on the first emotion marker to obtain a first semantic feature; Obtaining the first multidimensional feature according to the first image feature and the first semantic feature, wherein the first multidimensional feature includes a plurality of dimensional features corresponding to a plurality of dimensions respectively; Matching at least two reference facial images close to the current facial image according to the first multi-dimensional feature, each of the reference facial images being preset with an emotional target value; Establishing the emotion feature association matrix according to the second multidimensional features corresponding to each of the reference facial images, the preset emotion target value of each of the reference facial images, and the first multidimensional features of the current facial image; Performing emotional element feature conversion on the emotional feature association matrix to obtain an emotional feature set, wherein the emotional element feature is standardized data corresponding to a target value recognition model; The emotion feature set is loaded into the target value recognition model, wherein the target value recognition model is obtained by maintaining the basic model parameters of the preset multimodal large model unchanged and optimizing the updated model parameters of the multimodal large model according to the training emotion feature set, wherein the updated model parameters are related to the attention mechanism added to the multimodal large model; Obtaining a target emotion target value of the current face image generated by the target value recognition model; The emotion type corresponding to the current facial image is determined according to the recognition result of the emotion target value.
4. The method according to claim 3, characterized in that The first multidimensional feature includes a first image feature, the second multidimensional feature includes a second image feature, and the number of the first image features includes at least two; the emotional feature association matrix is established according to the second multidimensional feature corresponding to each of the reference facial images, the emotional target value preset for each of the reference facial images, and the first multidimensional feature of the current facial image, including: Performing a feature integration operation on at least two of the first image features to obtain a target enhanced image feature; The target enhanced image features are synchronized to a preset semantic feature domain through a pre-trained feature synchronization network to obtain the first image semantic identification set, and the second image features corresponding to each of the reference face images are feature mapped to obtain the second image semantic identification set corresponding to each of the reference face images; Obtaining the location where the current face image was collected and the voice content recognized by the voice data of the current face image; Generate emergency scene information of the current face image according to the background image of the current face image; Generate the first emotion identifier according to the collection location, the voice content and the emergency scenario information; Establishing emotion feature association parameters according to a first semantic identifier for distinguishing different information in the same facial image, and the second image semantic identifier set corresponding to the same reference facial image, the second emotion identifier, and the emotion target value, so as to obtain at least two emotion feature association parameters corresponding to at least two reference facial images; Merging the at least two emotion feature association parameters to obtain an initial emotion association matrix, wherein the at least two emotion feature association parameters in the initial emotion association matrix are separated by a second semantic identifier for distinguishing information of different facial images; Establishing an emotion classification identifier according to the first semantic identifier, the first image semantic identifier set, and the first emotion identifier; The emotion feature association matrix is generated according to the initial emotion association matrix and the emotion classification identifier.
5. The method according to claim 3, characterized in that: The first multidimensional feature includes a target enhanced image feature and a first semantic feature; Matching at least two reference face images close to the current face image according to the first multi-dimensional feature includes: Obtaining a preset enhanced image feature matching pool and a semantic feature matching pool, wherein the enhanced image feature matching pool includes a correspondence between each pending face image and the enhanced image feature corresponding to the pending face image, and the semantic feature matching pool includes a correspondence between each pending face image and the semantic feature corresponding to the pending face image; According to the target enhanced image feature, the enhanced image feature matching pool and the semantic feature matching pool are matched respectively, and according to the first semantic feature, the enhanced image feature matching pool and the semantic feature matching pool are matched respectively, to obtain at least two target pending facial images close to the current facial image; The reference facial image is determined from the at least two target facial images to be determined according to the image matching coefficients between the at least two target facial images to be determined and the current facial image.
6. The method according to claim 5, characterized in that The at least two target facial images to be determined include at least two target facial images to be determined obtained by matching each matching strategy; and determining the reference facial image from the at least two target facial images to be determined according to the image matching coefficients between the at least two target facial images to be determined and the current facial image, comprises: For each matching strategy, the at least two target facial images to be determined are matched, and a matching coefficient between the current facial image and each target facial image to be determined is determined; For each target undetermined face image, determining the average value of the matching coefficient between the target undetermined face image and the current face image in each matching strategy; The reference face image is determined from the at least two target face images according to the matching coefficient mean.
7. The method according to claim 3, characterized in that The training step of the target value recognition model includes: Acquire a first training enhanced image feature, a first training emotion identifier, and a first training emotion target value corresponding to the first training face image, and a second training enhanced image feature, a second training emotion identifier, and a second training emotion target value corresponding to the second training face image; Performing feature synchronization operations on the first training enhanced image features and the second training enhanced image features according to the pre-trained initial feature synchronization network, respectively, to obtain a first training image identification set and a second training image identification set, and establishing a training emotion feature association matrix according to the first training image identification set, the first training emotion identification and the first training emotion target value, and the second training image identification set and the second training emotion identification; Convert the training emotion feature association matrix into emotion element features to obtain the training emotion feature set; Adding the attention mechanism to the multimodal large model to add the updated model parameters to the multimodal large model through the attention mechanism; The basic model parameters of the multimodal large model are maintained unchanged, and the updated model parameters of the multimodal large model are optimized according to the training emotion feature set and the second training emotion target value to obtain the target value recognition model.
8. The method according to claim 7, characterized in that Before performing the feature synchronization operation on the enhanced image features according to the pre-trained initial feature synchronization network, the method further includes: Obtaining training face images and training emotional semantic information corresponding to the training face images; Performing feature extraction on the training face image to obtain training face emotion features, and loading the training face emotion features into an original untrained network, so that the original untrained network aligns the training face emotion features to an input semantic feature domain of the multimodal large model; Acquire the training semantic features corresponding to the synchronization request, and load the training semantic features and the target training image semantic features generated by the feature synchronization network into the multimodal large model, wherein the model parameters of the multimodal large model remain unchanged; The training predicted emotional semantic information generated by the multimodal large model is obtained, and the original untrained network is trained according to the training emotional semantic information and the training predicted emotional semantic information to obtain the initial feature synchronization network.
9. The method according to claim 7, characterized in that: The step of optimizing the updated model parameters of the multimodal large model according to the training emotion feature set and the second training emotion target value to obtain the target value recognition model includes: Loading the training emotion feature set into the multimodal large model to which the attention mechanism is added, and obtaining the predicted training target value generated by the multimodal large model; According to the difference between the predicted training target value and the second training emotion target value, optimizing the updated model parameters of the multimodal large model to obtain the target value recognition model; The method further comprises: According to the difference between the predicted training target value and the second training emotion target value, the network parameters of the initial feature synchronization network are optimized.
10. A server system, characterized in that: The method comprises a server, wherein the server is used to execute the method described in any one of claims 1 to 9.
Citation Information
Patent Citations
A method for emotion recognition of interactive language
CN109308466A
Alarm receiving processing method and device, machine readable storage medium and processor
CN112542180A
Emotion counseling system and method for mobile ward-round scene
CN112842337A
Multi-modal emotion emergency decision-making system based on multi-stage long short-term memory network
CN115393927A
Multi-view face feature and audio feature fused emotion recognition method and system
CN117312992A
Cited By
Notification system with emotion-inciting image generator
US12518439B2
Notification system with emotion-inciting image generator
US20250218057A1