Emotion and behavior anomaly detection and early warning method and device, equipment and storage medium

Through wearable detection devices and multi-modal emotion perception modules, the lack of emotional and behavior monitoring in industries such as housekeeping services is solved, real-time early warning and data recording are achieved, and the security and management efficiency of the service process are improved.

CN120374916AInactive Publication Date: 2025-07-25SOUTH CHINA UNIV OF TECH +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510561160.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In industries such as housekeeping services and hotel management, there is a lack of real-time emotional perception and behavior monitoring methods and equipment during the service process, making it difficult to accurately capture expression changes and voice information, resulting in increased risks in the service process and it is difficult to restore the scene after emergencies.

Method used

Wearable detection device is adopted to record the expressions, behaviors and voices of the service provider and the serviced person through the camera and microphone, and use the multimodal collaborative emotion perception module to perform emotion analysis, and trigger early warning when abnormal emotions are detected to record data to support subsequent services and event restoration.

Benefits of technology

It realizes accurate identification of real-time emotions and behaviors, timely warning, ensures the safety and stability of the service process, and provides traceable process data support to improve service quality and management level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374916A_ABST
    Figure CN120374916A_ABST
Patent Text Reader

Abstract

The invention discloses an emotion and behavior anomaly detection and early warning method and device, equipment and a storage medium, and the method employs a wearable detection device which can detect the speech and behavior of a server, the speech, expression and posture of a serviced person, and the environment where the serviced person is located. And the emotion and the holding behavior of the server and the serviced person are comprehensively analyzed, so that early warning is carried out on the severe emotion change and the violent behavior. Through a camera and a microphone on the multi-mode cooperative emotion sensing module, emotion changes and voice changes of a server and a serviced person can be clearly recorded, and the emotions and actions of the server and the serviced person are analyzed through the multi-mode cooperative emotion sensing module. On one hand, emotion and behavior changes of a server and a serviced person are recorded to provide data support for subsequent better services, and on the other hand, events can be better restored through the recorded data after accidents occur. The invention relates to a character and environment detection technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technology of human and environment detection, and particularly to a method, device, equipment and storage medium for detecting and warning of emotional and behavioral abnormalities. Background Art

[0002] In industries such as home service and hotel management, including diverse service scenarios such as home care, infant care, postoperative companionship, and customer complaint handling, there are typical requirements for emotion and behavior recognition. However, during the service process, due to the fact that service providers and service recipients may be in a high-pressure and easily fatigued working or living state, combined with physical discomfort, psychological burden, and emotional interference in a complex environment, problems such as psychological stress, emotional fluctuations, behavioral intensification, and even violent behaviors are extremely likely to be induced. Currently, in most service scenarios, there is still a lack of real-time emotion perception and behavior monitoring methods and equipment, and it is difficult to timely capture and intervene in the emotional changes and behavioral abnormalities of service recipients and service providers during the service process. This not only increases the risks during the service process but also makes it difficult to effectively restore the scene after a service dispute occurs, lacking objective evidence to support the responsibility division and handling decision.

[0003] Especially in a multi-person interaction environment, the recognition of individual emotions and behaviors is easily affected by visual occlusion and environmental noise. Although some places are equipped with fixed indoor cameras with a certain ability to record behaviors, limited by the perspective and clarity, it is difficult to accurately capture facial expression changes and voice information, and it is even more difficult to effectively record verbal conflicts and abnormal behaviors in a noisy environment.

[0004] Therefore, there is an urgent need for a portable device with the capabilities of automatically recognizing, continuously monitoring, intelligently warning, and recording the process of emotional and behavioral changes. On the one hand, this device can track and accurately recognize the emotional states and behavioral dynamics of service providers and service recipients in real time during the service process, and issue risk warnings in a timely manner, thus effectively protecting the physical and mental health of service recipients and the safety and stability of the service process. On the other hand, after an unexpected event occurs, the device can provide traceable process data support, which helps to restore the event process and clarify the responsibility attribution, further improving the service quality and overall management level. Summary of the Invention

[0005] To solve at least one of the technical problems existing in the prior art to a certain extent, the purpose of the present invention is to provide a method, device, equipment and storage medium for detecting and warning of emotional and behavioral abnormalities,

[0006] The first technical solution adopted by the present invention is: A method for detecting and warning of emotional and behavioral abnormalities, including:

[0007] Wear the detection device on the service provider;

[0008] Turn on the detection device, start the camera on the detection device, and the camera captures the face image of the person being served;

[0009] Perform face recognition on the face image captured by the camera through the face recognition system to determine the identity of the person being served;

[0010] The camera periodically records the behavior of the service provider and the expression changes of the person being served. At the same time, the microphone on the detection device is activated, and the microphone continuously records the ambient sound and the voices of the service provider and the person being served;

[0011] The images recorded by the camera and the voice data recorded by the microphone are both collected by the multi-modal collaborative emotion perception module. The multi-modal collaborative emotion perception module performs emotion analysis on the voice of the service provider, the expression, behavior, and voice of the person being served, and records them in the database.

[0012] According to some embodiments of the present application, the images recorded by the camera and the voice data recorded by the microphone are both uploaded to the multi-modal data acquisition module. The multi-modal data acquisition module sorts out the data and then collects it to the multi-modal collaborative emotion perception module.

[0013] According to some embodiments of the present application, the multi-modal data acquisition module performs preprocessing such as grayscale conversion or image compression on the images recorded by the camera, and performs noise reduction, voiceprint feature extraction, and voiceprint enhancement processing on the voice data recorded by the microphone.

[0014] According to some embodiments of the present application, the multi-modal data acquisition module separates the ambient sound and human voices, extracts voiceprint features and performs voiceprint comparison on different human voices among them, and then completes the determination of voiceprint identity.

[0015] According to some embodiments of the present application, the multi-modal collaborative emotion perception module performs comprehensive analysis on the images recorded by the camera and the voice data recorded by the microphone. If abnormal emotions are detected, it can trigger an early warning mechanism and alarm or remind the service provider.

[0016] According to some embodiments of the present application, the analysis results of the multi-modal collaborative emotion perception module can be uploaded to a mobile device, and service providers, managers, and authorized personnel can all understand the emotional and behavioral changes of the service provider and the person being served through the mobile device.

[0017] The second technical solution adopted by the present invention is: an emotion and behavior abnormality detection and early warning device, including:

[0018] A detection device, which is used to be worn on the service provider, and the detection device is provided with a camera and a microphone;

[0019] A face recognition system for performing face recognition on the images captured by the camera;

[0020] A multi-modal data acquisition module for collating the images captured by the camera and the voice data collected by the microphone;

[0021] A multi-modal collaborative emotion perception module for analyzing the images captured by the camera and the voice data collected by the microphone;

[0022] A mobile device for reading the analysis results of the multi-modal collaborative emotion perception module.

[0023] According to some embodiments of the present application, the detection device is in the shape of glasses and is worn on the face of the service provider.

[0024] The third technical solution adopted by the present invention is:

[0025] An electronic device, the electronic device includes a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the emotion and behavior anomaly detection and warning method as described above.

[0026] The fourth technical solution adopted by the present invention is:

[0027] A computer-readable storage medium, and at least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the emotion and behavior anomaly detection and warning method as described above.

[0028] The beneficial effects of the present invention are: The detection device can be worn on the service provider. Through the camera and microphone thereon, it can clearly record the emotional changes and voice changes of the service provider and the service recipient, and analyze the emotions and words and deeds of both the service provider and the service recipient through the multi-modal collaborative emotion perception module. On the one hand, it records the emotional and behavioral changes of the service provider and the service recipient to provide data support for subsequent better-quality services. On the other hand, after an accident occurs, it can better restore the event through the recorded data. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the accompanying drawings of the related technical solutions in the embodiments of the present invention or the prior art. It should be understood that the accompanying drawings in the following introduction are only for conveniently and clearly presenting some embodiments of the technical solutions in the present invention. For those skilled in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0030] Figure 1 It is the working flowchart of the wearing state sensing module in the method for detecting and warning emotional and behavioral abnormalities according to the first aspect embodiment of the present invention;

[0031] Figure 2 It is the working flowchart of the environment detection and face detection in the method for detecting and warning emotional and behavioral abnormalities according to the first aspect embodiment of the present invention;

[0032] Figure 3 It is the working flowchart of the service management module in the method for detecting and warning emotional and behavioral abnormalities according to the first aspect embodiment of the present invention;

[0033] Figure 4 It is the working flowchart of the multi-modal data acquisition module in the method for detecting and warning emotional and behavioral abnormalities according to the first aspect embodiment of the present invention;

[0034] Figure 5 It is the multi-voiceprint recognition flowchart in the multi-person environment in the method for detecting and warning emotional and behavioral abnormalities according to the first aspect embodiment of the present invention;

[0035] Figure 6 It is the working flowchart of the multi-modal collaborative emotion perception module in the method for detecting and warning emotional and behavioral abnormalities according to the first aspect embodiment of the present invention;

[0036] Figure 7 It is the alarm flowchart for the abnormal emotions of the service recipient in the method for detecting and warning emotional and behavioral abnormalities according to the first aspect embodiment of the present invention;

[0037] Figure 8 It is the three-dimensional view of the device for detecting and warning emotional and behavioral abnormalities according to the second aspect embodiment of the present invention;

[0038] Figure 9 It is the cross-sectional view of the device for detecting and warning emotional and behavioral abnormalities according to the second aspect embodiment of the present invention.

[0039] Reference numerals: 100 - detection device, 200 - camera, 300 - microphone, 400 - speaker, 500 - power indicator light, 600 - charging interface, 700 - Bluetooth device, 800 - microprocessor, 900 - capacitive touch sensor, 1000 - battery. Detailed implementation manners

[0040] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention. For the step numbers in the following embodiments, they are only set for the convenience of description and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0041] In the description of the present invention, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc. is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.

[0042] In the description of the present invention, the meaning of "several" is one or more, the meaning of "multiple" is two or more, "greater than", "less than", "exceeding", etc. are understood as not including the recited number, and "above", "below", "within", etc. are understood as including the recited number. If there is a description of "first" and "second", it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.

[0043] In the description of the present invention, unless otherwise clearly defined, words such as "set", "installed", "connected", etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.

[0044] In industries such as domestic service and hotel management, including diverse service scenarios such as home care, infant care, postoperative escort, and customer complaint handling, there are typical requirements for emotion and behavior recognition. However, during the service process, since service providers and service recipients may be in a high-pressure and easily fatigued working or living state, combined with physical discomfort, psychological burden, and emotional interference in a complex environment, it is extremely easy to induce problems such as psychological stress, emotional fluctuations, behavioral intensification, and even violent behaviors. Currently, most service scenarios still lack real-time emotion perception and behavior monitoring methods and devices, and it is difficult to timely capture and intervene in the emotional changes and behavioral abnormalities of service recipients and service providers during the service process. This not only increases the risks during the service process but also makes it difficult to effectively restore the scene after a service dispute occurs, lacking objective evidence to support the responsibility division and handling decisions.

[0045] Especially in a multi-person interaction environment, the recognition of individual emotions and behaviors is vulnerable to visual occlusion and environmental noise interference. Although some places are equipped with fixed indoor cameras with certain behavior recording capabilities, limited by the viewing angle and clarity, it is difficult to accurately capture facial expression changes and voice information, and it is even more difficult to effectively record verbal conflicts and abnormal behaviors in a noisy environment.

[0046] Therefore, there is an urgent need for a portable device that can automatically recognize, continuously monitor, intelligently warn, and record the process of emotional and behavioral changes. On the one hand, this device can track and accurately identify the emotional states and behavioral dynamics of service providers and service recipients in real time during the service process, and issue risk warnings in a timely manner, thus effectively protecting the physical and mental health of service recipients and the safety and stability of the service process. On the other hand, after an unexpected event occurs, the device can provide traceable process data support, which helps to restore the event process, clarify the responsibility attribution, and further improve the service quality and overall management level.

[0047] In response to this, this application proposes an emotion and behavior anomaly detection and warning method, device, equipment, and storage medium. The detection device can be worn on the service provider. Through the camera and microphone on it, it can clearly record the emotional changes and voice changes of the service provider and the service recipient, and analyze the emotions and words and deeds of both the service provider and the service recipient through a multi-modal collaborative emotion perception module. On the one hand, it records the emotional and behavioral changes of the service provider and the service recipient to provide data support for subsequent better services. On the other hand, after an accident occurs, it can better restore the event through the recorded data.

[0048] Embodiment 1

[0049] This embodiment provides an emotion and behavior anomaly detection and warning method, including the following steps:

[0050] S100. Wear the detection device 100 on the service provider;

[0051] S200. Turn on the detection device 100, start the camera 200 on the detection device 100, and the camera 200 captures the face image of the service recipient;

[0052] S300. Perform face recognition on the face image captured by the camera 200 through the face recognition system to determine the identity of the service recipient, read the previous detection data of the service recipient, and enter the newly detected data into the database of this service recipient;

[0053] S400. The camera 200 periodically records the behaviors of the service provider and the facial expressions of the service recipient. Meanwhile, the microphone 300 on the detection device 100 is activated, and the microphone 300 continuously records the ambient sound and the voices of the service provider and the service recipient. Subsequently, the emotional changes of the service recipient can be comprehensively analyzed based on the image and sound information, improving the accuracy of the analysis;

[0054] S500. The images recorded by the camera 200 and the voice data recorded by the microphone 300 are both aggregated to the multi-modal collaborative emotion perception module. The multi-modal collaborative emotion perception module performs emotion analysis on the voice of the service provider, the expressions, behaviors, and voices of the service recipient, and records them in the database.

[0055] For the above process, specifically, in order to ensure that the detection device 100 can be powered on in time for recording after being normally worn on the service provider, a wearing state sensing module is provided on the detection device 100. Refer to Figure 1 , the wearing state sensing module is provided with a capacitance sensor. When a human body approaches the capacitance sensor, the distribution of the capacitance will change. After this change is detected by the capacitance sensor, the wearing state can be recognized. Correspondingly, when the service provider removes the detection device 100, the capacitance sensor can also accurately sense the physiological signal of the service provider, thereby triggering a corresponding system response.

[0056] Refer to Figure 2 , after the camera 200 is powered on, on the one hand, it can capture the behaviors of the service provider and the face of the service recipient, and at the same time, it can also capture the surrounding environment, convert the optical signal according to the ambient light intensity of the scene, and change the aperture or exposure, so that the portrait image captured by the camera 200 is clearer.

[0057] During the recording process of the camera 200 and the microphone 300, the service provider can also control the detection device 100 through the capacitance sensor on the detection device 100. Refer to Figure 3 , for example, the service provider can input commands by double-clicking, swiping, or long-pressing the capacitance sensor to control the detection device 100 to perform corresponding operations. Moreover, the service provider can also customize the gestures through the mobile APP on the mobile device, so as to have more freedom in operation.

[0058] Furthermore, the images recorded by the camera 200 and the voice data recorded by the microphone 300 are both uploaded to the multi-modal data acquisition module, and the multi-modal data acquisition module sorts out the data and then aggregates it to the multi-modal collaborative emotion perception module. Refer to Figure 4, the multimodal data acquisition module is used to preprocess the materials collected by the camera 200 and the microphone 300. For example, it grayscales or compresses the images recorded by the camera 200, and performs noise reduction, voiceprint feature extraction, and voiceprint enhancement processing on the voice data recorded by the microphone 300. This improves the accuracy of feature recognition and reduces the impact of impurity signals on the analysis results. Additionally, the multimodal data acquisition module is also used to collect gesture data from the capacitive sensor, and after preprocessing, enter the gesture data into the memory for operations such as task management, personnel switching, and exception reporting.

[0059] Furthermore, for an environment with multi-person interaction, referring to Figure 5 , the multimodal data acquisition module can separate environmental sounds and human voices, extract voiceprint features and compare voiceprints of different human voices, and then complete the determination of voiceprint identities. Specifically, in order to better separate the voice of the person being served from environmental sounds, the microphone 300 can adopt an array layout and use spatial spectrum estimation technology to enhance the voice of each speaker while suppressing background noise. The design of the microphone array enables the system to enhance the sound from a specific direction through sound source localization, especially when there are multiple people talking, and it can identify sounds from different directions. Different people can also pre-enter 3 to 5 segments of speech (each segment is 5 to 10 seconds), extract acoustic features such as MFCC, generate and encrypt and store the user's voiceprint template in a secure database, and subsequently be able to distinguish different speakers based on the voiceprints saved in the database.

[0060] In addition, for noise reduction operations, a Deep Beamforming Network can be used to achieve adaptive beamforming through a Convolutional Neural Network (CNN), enhance the signal of the target speaker, and filter out noise from other directions at the same time. Based on the Short-Time Fourier Transform (STFT) and the Complex Wavelet Transform, the system can analyze signals in the time-frequency domain and predict and eliminate noise through a Deep Convolutional Neural Network (CNN) or a Long Short-Term Memory Network (LSTM).

[0061] For voiceprint feature extraction and voiceprint enhancement, traditional voiceprint recognition methods usually rely on features such as MFCC (Mel Frequency Cepstral Coefficients) or i-vector. However, these features have limited robustness to complex background noise or multi-speaker environments. Therefore, it is necessary to introduce multi-dimensional time-frequency feature extraction methods. Time-frequency domain decomposition and deep feature learning: Time-frequency feature learning is adopted, that is, after the signal is transformed into the time-frequency domain through STFT, a deep convolutional neural network (CNN) is used to extract features from the time-frequency map. CNN can automatically capture local patterns in the time-frequency map, further enhancing the signal's identification ability. Time-frequency enhancement based on Self-Attention: The self-attention mechanism is introduced to process the signals in the time-frequency domain, thereby effectively enhancing the speech of the target speaker and reducing interference when multiple speakers overlap. In a multi-speaker environment, the system separates the overlapping speech signals through a speaker separation network, identifies each independent speaker, and processes the speech signals of each speaker independently. Self-supervised learning methods are used to train the voiceprint features of the speaker using unlabeled data, enhancing the recognition ability for unknown speakers. By constructing positive and negative sample pairs, the network's discrimination of voiceprints is improved. Through voiceprint recognition and matching, each speaker is compared with a known identity database to accurately identify the speaker.

[0062] Refer to Figure 6 , the multi-modal collaborative emotion perception module integrates the image and video materials of the camera 200 and the voice materials collected by the microphone 300 for comprehensive analysis and judgment, thereby analyzing the emotions of the service provider and the service recipient. The multi-modal collaborative emotion perception module specifically includes a posture emotion proxy module, an expression emotion proxy module, and a voice emotion proxy module, which can respectively judge the emotions of the service recipient's posture, expression, and tone of voice. In addition, the multi-modal collaborative emotion perception module also includes an ambient sound anomaly detection module to detect whether there is unexpected noise in the environment; and a video emotion tendency discrimination module to continuously analyze the behavior of the service recipient based on the video materials, thereby making a comprehensive judgment on their behavior.

[0063] In the multi-modal collaborative emotion perception module, all emotion analysis results, including the emotion analysis results of images, voices, and videos, will be transmitted to the storage management module for storage. This module stores all multi-modal emotion data in a database for subsequent analysis, retrospective review, and model optimization. By recording the process of each emotion analysis, the system can continuously learn and optimize the accuracy of emotion recognition, and provide data support for subsequent task evaluation.

[0064] Furthermore, the multi-modal collaborative emotion perception module comprehensively analyzes the images recorded by the camera 200 and the voice data recorded by the microphone 300. If abnormal emotions are detected, it can trigger an early warning mechanism and alarm or remind the service provider. Refer to Figure 7, once the multi-modal collaborative emotion perception module detects that the serviced person has abnormal emotions, or detects a specific sound pattern in the surrounding environment (such as screams, smashing sounds or calls for help), it can judge that there is potential physical violence. At the same time, when the serviced person makes abnormal movements (such as suddenly falling, violent defensive or aggressive movements), the effective speech segment is judged by combining the environmental noise level and semantic content, and abnormal sound capture is performed. The trigger sampling system for non-steady-state sounds such as screams / sobs can identify these movements as possible signals of violent behavior through the model.

[0065] When the system detects emotional abnormalities, it automatically triggers an early warning mechanism and sends a voice warning reminder to the service provider to remind them to take appropriate intervention measures.

[0066] Furthermore, the analysis results of the multi-modal collaborative emotion perception module can be uploaded to the mobile device, and service providers, managers and authorized personnel can all understand the emotional and behavioral changes of the service provider and the serviced person through the mobile device. In this application, it refers to the personnel who directly provide services to customers, such as domestic service personnel, nursing staff, escort personnel, etc. Service providers can view the basic information of their service objects according to their duties and cooperate with the system to record emotional and behavioral data, but do not have system management permissions. Managers refer to the internal personnel of organizations or enterprises responsible for service process management, personnel scheduling and service quality supervision, such as project supervisors, dispatchers, institutional leaders, etc. Managers have the permissions to view, review and conduct preliminary analysis of service process data, and can also access abnormal behavior records within a specific time period for service improvement and risk prevention. Authorized personnel refer to authorized technical and management personnel who have the right to access the system background, operate sensitive data or perform technical maintenance, as well as service users or subscribers, specifically including system administrators, security auditors, data privacy protection personnel, etc. The permissions of such personnel are restricted by institutional regulations and need to be strictly authorized and filed before they can access or operate relevant data content.

[0067] The mobile device can be a mobile phone, a tablet computer, a computer or other electronic devices with human-computer interaction capabilities, which will not be elaborated here.

[0068] Embodiment 2

[0069] Refer to Figure 8 , this embodiment provides an emotion and behavior abnormality detection and early warning device, which is used to execute the above-mentioned emotion and behavior abnormality detection and early warning method, including a detection device 100, a face recognition system, a multi-modal data acquisition module, a multi-modal collaborative emotion perception module and a mobile device, where:

[0070] The detection device 100 is used to be worn on the service provider, and is provided with a camera 200 and a microphone 300 thereon. The camera 200 is used to capture the expressions and behaviors of the service recipient, the behaviors of the service provider, and the surrounding environment, and the microphone 300 is used to collect the sounds in the environment. By combining the image materials of the camera 200 and the sound materials collected by the microphone 300, the emotional changes of the service provider and the service recipient can be comprehensively analyzed.

[0071] The face recognition system is used to perform face recognition on the images captured by the camera 200, so as to determine the identity of the service recipient, and avoid data confusion caused by the collected data being entered into different service recipients.

[0072] The multimodal data acquisition module is used to organize the images captured by the camera 200 and the voice data collected by the microphone 300.

[0073] The multimodal collaborative emotion perception module is used to analyze the images captured by the camera and the voice data collected by the microphone.

[0074] The mobile device is used to read the analysis results of the multimodal collaborative emotion perception module.

[0075] Specifically, for the specific structure of the detection device 100, it is in the shape of glasses and is worn on the face of the service provider. Refer to Figure 9 , on the detection device 100, there are a camera 200, a microphone 300, a speaker 400, a power indicator 500, a charging interface 600, a Bluetooth device 700, a microprocessor 800, a capacitive touch sensor 900, and a battery camera 1000. Among them, the camera 200 is installed at the front of the detection device 100 and can capture the service recipient and the surrounding environment from the first-person perspective of the service provider. The microphone 300 adopts an array design for collecting voices, supports identifying the voice characteristics of the service object and the domestic helper, and helps analyze the conversation content, emotional changes, and environmental sound information. The speaker 400 is used to play voice feedback or notifications, supports the voice interaction function, and makes the communication between the device and the user more convenient. The power indicator 500 displays the power status of the device in real time, helping the user to understand the remaining battery situation in time, so as to reasonably arrange charging. The charging interface 600 adopts a magnetic attraction design for connecting a charging device to provide power support. The Bluetooth device 700 supports wireless connection to other intelligent devices, such as mobile phones or home intelligent assistants, facilitating data transmission and remote control. The microprocessor 800 is responsible for core operations and data processing, ensuring the coordination and efficient operation of the device's various functions. The capacitive touch sensor 900 supports the touch control function, and through gestures or touch operations, realizes the intuitive control of the device. The battery 1000 provides the power support required by the detection device 100 to ensure that the device can operate stably for a long time.

[0076] Embodiment 3

[0077] An embodiment of the present invention further provides an electronic device, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the above-mentioned emotion and behavior anomaly detection and warning method.

[0078] It can be understood that the memory may include a Random Access Memory (RAM), and may also include a Read-Only Memory. Optionally, the memory includes a non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, codes, code sets or instruction sets. The memory may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store data created according to the use of the server, etc.

[0079] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect various parts within the entire server, and by running or executing instructions, programs, code sets or instruction sets stored in the memory, and by calling data stored in the memory, it executes various functions of the server and processes data. Optionally, the processor can be implemented in at least one of the following hardware forms: Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processor can integrate one or several combinations of a Central Processing Unit (CPU) and a modem, etc. Among them, the CPU mainly processes the operating system and application programs, etc.; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor and can be implemented separately by a single chip.

[0080] Since this electronic device is the electronic device corresponding to the closed-loop transcranial electrical stimulation method of the embodiment of the present invention, and the principle of the electronic device to solve problems is similar to that of this method, the implementation of this electronic device can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.

[0081] Embodiment 4

[0082] An embodiment of the present invention further provides a computer-readable storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the above-mentioned emotion and behavior anomaly detection and warning method.

[0083] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically-erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium capable of carrying or storing data.

[0084] Since this storage medium is the storage medium corresponding to a closed-loop transcranial electrical stimulation method of an embodiment of the present invention, and the principle of solving problems by this storage medium is similar to that of this method, the implementation of this storage medium can refer to the implementation process of the above method embodiment, and the repeated parts will not be elaborated.

[0085] Embodiment 5

[0086] In some possible embodiments, aspects of the method of the embodiments of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps of an emotion and behavior anomaly detection and warning method according to various exemplary embodiments described above in this specification. Among them, the executable computer program code or "code" for executing each embodiment can be written in high-level programming languages such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.

[0087] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0088] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0089] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable those of ordinary skill in the art to understand the content of the present invention and implement it accordingly, and cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the content of the present invention should be covered within the protection scope of the present invention.

Claims

1. A method for detecting and warning of abnormal emotions and behaviors, characterized in that, It includes the following steps: Wear the detection device on the service provider; Turn on the detection device, start the camera on the detection device, and the camera captures the face image of the person being served; Perform face recognition on the face image captured by the camera through the face recognition system to determine the identity of the person being served; The camera periodically records the behavior of the service provider and the expression changes of the person being served. Meanwhile, the microphone on the detection device is started, and the microphone continuously records the ambient sound and the voices of the service provider and the person being served; The images recorded by the camera and the voice data recorded by the microphone are both collected into the multi-modal collaborative emotion perception module. The multi-modal collaborative emotion perception module performs emotion analysis on the voice of the service provider and the expressions, behaviors, and voices of the person being served, and records them in the database.

2. The emotional and behavioral abnormality detection and warning method according to claim 1, wherein: The images recorded by the camera and the voice data recorded by the microphone are both uploaded to the multi-modal data acquisition module. The multi-modal data acquisition module sorts out the data and then collects it into the multi-modal collaborative emotion perception module.

3. The emotional and behavioral anomaly detection and warning method according to claim 2, wherein: The multi-modal data acquisition module performs preprocessing such as grayscale conversion or image compression on the images recorded by the camera, and performs noise reduction, voiceprint feature extraction, and voiceprint enhancement processing on the voice data recorded by the microphone.

4. The emotional and behavioral anomaly detection and warning method according to claim 3, characterized in that: The multi-modal data acquisition module separates the ambient sound and human voices, extracts voiceprint features and performs voiceprint comparison on different human voices among them, and then completes the determination of voiceprint identity.

5. The emotional and behavioral anomaly detection and warning method according to claim 1, wherein: The multi-modal collaborative emotion perception module comprehensively analyzes the images recorded by the camera and the voice data recorded by the microphone. If abnormal emotions are detected, it can trigger an early warning mechanism and alarm or remind the service provider.

6. The emotional and behavioral anomaly detection and early warning method according to claim 1, characterized in that: The analysis results of the multi-modal collaborative emotion perception module can be uploaded to the mobile device. The service provider, manager, and authorized personnel can all understand the emotional and behavioral changes of the service provider and the person being served through the mobile device.

7. An emotion and behavior abnormality detection and warning device for implementing the emotion and behavior abnormality detection and warning method according to any one of claims 1 to 6, characterized in that, It includes: A detection device, which is used to be worn on the service provider, and the detection device is provided with a camera and a microphone; A face recognition system, which is used to perform face recognition on the images captured by the camera; A multi-modal data acquisition module, which is used to sort out the images captured by the camera and the voice data collected by the microphone; A multi-modal collaborative emotion perception module, which is used to analyze the images captured by the camera and the voice data collected by the microphone; A mobile device, which is used to read the analysis results of the multi-modal collaborative emotion perception module.

8. The emotional and behavioral anomaly detection and warning device according to claim 7, wherein: The detection device is in the shape of glasses and is worn on the face of the service provider.

9. An electronic device, characterized in that, The electronic device includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the emotion and behavior anomaly detection and early warning method as described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the emotion and behavior anomaly detection and warning method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Emotion recognition method based on large model and related device

    CN119904901A

  • Automotive ambient temperature sensor

    KR1020220023563A

  • Alarm system and method for using AI technology to monitor emotion of caregivers to achieve the purpose of proactive warning and real-time prediction of violent incidents through a variety of image recognition and analysis

    TW202435176A

  • KR20240103865A