Object demand processing method and device, equipment, storage medium and program product
By conducting emotional analysis of pet audio and video data, combined with environmental information and context, the accuracy of pet needs identification is solved, and a comprehensive personalized analysis and timely satisfaction of pet needs is achieved.
Patent Information
- Application Number
- CN202510629377.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-26
AI Technical Summary
It is difficult for humans to accurately understand pet emotions and needs, and the existing technology cannot effectively use information such as pet behavior, voice and posture to identify demands.
Through sentiment analysis based on audio data and video data, combined with environmental monitoring and context information, the pet's demand analysis results are determined, and personalized analysis is achieved using the sentiment analysis module, environmental monitoring module and demand analysis module.
It realizes comprehensive and accurate identification and personalized analysis of pet needs, can meet pet needs in a timely manner, and improves the accuracy of pet emotional understanding.
Smart Images

Figure CN120541769A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to an object demand processing method, device, electronic device, storage medium, and program product. Background Art
[0002] As people's material living standards improve and their spiritual lives become more enriched, more and more people are keeping pets for companionship. However, due to language barriers between humans and animals, humans often cannot accurately understand their pets' emotions and needs. While pets cannot directly express their thoughts and feelings in human language, their behavior, sounds, and gestures contain a wealth of information. Therefore, how to rationally and fully utilize this information to accurately identify pets' needs has become a pressing technical challenge. Summary of the Invention
[0003] The embodiments of the present disclosure provide an object demand processing method, an object demand processing device, an electronic device, a computer-readable storage medium, and a computer program product.
[0004] In the first aspect, an embodiment of the present disclosure proposes a method for processing object needs, including: performing emotion analysis on the target object based on audio data and video data associated with the target object to obtain a first emotion classification result; monitoring the current environment of the target object to obtain environmental information associated with the target object; and determining the demand analysis result of the target object based on the first emotion classification result, the environmental information, and the contextual information associated with the target object.
[0005] In a second aspect, embodiments of the present disclosure provide an object demand processing device, comprising: a sentiment analysis module, an environment monitoring module, and a demand analysis module. The sentiment analysis module is configured to perform sentiment analysis on a target object based on audio and video data associated with the target object, obtaining a first sentiment classification result; the environment monitoring module is configured to monitor the target object's current environment, obtaining environmental information associated with the target object; and the demand analysis module is configured to determine a target object demand analysis result based on the first sentiment classification result, the environmental information, and contextual information associated with the target object.
[0006] In a third aspect, an embodiment of the present disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can implement the object demand processing method described in any implementation method in the first aspect when executing the instructions.
[0007] In a fourth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, which are used to enable a computer to implement the object demand processing method described in any implementation method of the first aspect when executed.
[0008] In a fifth aspect, an embodiment of the present disclosure provides a computer program product comprising a computer program, which, when executed by a processor, can implement the object demand processing method as described in any implementation manner in the first aspect.
[0009] The object demand processing solution provided by the embodiments of the present disclosure can not only perform real-time emotional analysis on the target object based on the audio and video data associated with the target object, but also monitor the environmental information corresponding to the target object's current environment in real time. Furthermore, it is necessary to fully combine the first emotion classification result of the target object obtained and its associated environmental information and contextual information to comprehensively analyze the target object's needs. In this way, by fully considering the target object's emotional state, the state of the environment in which it is located, and the contextual information, it is possible to more comprehensively and accurately determine the target object's needs, and to achieve personalized analysis of the target object so as to truly understand the target object's needs.
[0010] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Other features, objects and advantages of the present disclosure will become more apparent from a reading of the detailed description of non-limiting embodiments made with reference to the following drawings:
[0012] Figure 1 is an exemplary system architecture in which the present disclosure may be applied;
[0013] Figure 2 A flowchart of a method for processing object requirements provided by an embodiment of the present disclosure;
[0014] Figure 3 A flowchart of another object demand processing method provided by an embodiment of the present disclosure;
[0015] Figure 4 A schematic diagram of the overall process of an object demand processing method in an application scenario provided by an embodiment of the present disclosure;
[0016] Figure 5 The above embodiments of the present disclosure provide Figure 4 A schematic diagram of the processing process of the sentiment analysis module in FIG;
[0017] Figure 6 The above embodiments of the present disclosure provide Figure 4 A schematic block diagram of a processing process of the feature extractor in ;
[0018] Figure 7 The above embodiments of the present disclosure provide Figure 4 A schematic diagram of the interaction between the encoder and the adapter;
[0019] Figure 8 The above embodiments of the present disclosure provide Figure 4 A structural block diagram of the fusion network in ;
[0020] Figure 9 The above embodiments of the present disclosure provide Figure 4 A schematic block diagram of a processing process of the environmental analysis module in FIG.
[0021] Figure 10 The above embodiments of the present disclosure provide Figure 4 A schematic block diagram of the context information update process in the corresponding embodiment;
[0022] Figure 11 The above embodiments of the present disclosure provide Figure 4 A schematic block diagram of the processing process of the demand analysis module in FIG;
[0023] Figure 12 An example block diagram of an object demand processing device provided by an embodiment of the present disclosure;
[0024] Figure 13 A schematic diagram of the structure of an electronic device suitable for executing an object demand processing method provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0025] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description. It should be noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other unless there is a conflict.
[0026] It should be pointed out that in the technical solutions disclosed herein, the collection, acquisition, storage, processing, transmission, provision, disclosure and application of user personal information (such as gender, age, images, etc.) are all carried out with the user's knowledge and explicit authorization, and are in compliance with relevant laws and regulations and do not violate public order and good morals.
[0027] Figure 1 An exemplary system architecture 100 is shown to which embodiments of the object demand processing method, apparatus, electronic device, and computer-readable storage medium disclosed herein can be applied.
[0028] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. Network 104 is a medium for providing communication links between terminal devices 101, 102, 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0029] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications for enabling information communication between the terminal devices 101, 102, and 103 and server 105 can be installed, such as object demand analysis applications, object demand processing applications, and instant messaging applications.
[0030] Terminal devices 101, 102, 103 and server 105 can be either hardware or software. When terminal devices 101, 102, 103 are hardware, they can be various electronic devices with display screens, including but not limited to smartphones, tablet computers, laptop computers, and desktop computers. When terminal devices 101, 102, 103 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software or software modules, or as a single software or software module, and are not specifically limited here. When server 105 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules, or as a single software or software module, and are not specifically limited here.
[0031] The server 105 can provide various services through various built-in applications. For example, the server 105 can provide an object demand processing application based on audio and video data analysis. When running the object demand processing application, the server 105 can achieve the following effects:
[0032] It should be noted that, in addition to being obtained from the terminal devices 101, 102, and 103 via the network 104, the audio and video data associated with the object can also be pre-stored locally on the server 105 in various ways. Therefore, when the server 105 detects that such data is already stored locally (for example, when it begins processing previously stored audio and video data associated with the object), it can choose to directly obtain such data locally. In this case, the exemplary system architecture 100 may also not include the terminal devices 101, 102, and 103 and the network 104.
[0033] Because demand processing based on object-associated audio and video data requires significant computing resources and significant computational power, the object demand processing methods provided in the subsequent embodiments of this disclosure are generally executed by a server 105 possessing significant computing power and resources. Accordingly, the object demand processing apparatus is generally also located within the server 105. However, it should also be noted that, if terminal devices 101, 102, and 103 also possess sufficient computing power and resources, the terminal devices 101, 102, and 103 may also utilize the object demand processing application installed thereon to perform the aforementioned operations delegated to the server 105, thereby outputting the same results as the server 105. In particular, in the presence of multiple terminal devices with varying computing capabilities, if the object demand processing application determines that the terminal device it is in possession of possesses significant computing power and resources, the terminal device may be instructed to perform the aforementioned operations, thereby appropriately alleviating the computational burden on the server 105. Accordingly, the object demand processing apparatus may also be located within the terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may also not include the server 105 and the network 104 .
[0034] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0035] Please refer to Figure 2 , Figure 2 This is a flowchart of a method for processing object requirements provided by an embodiment of the present disclosure, wherein process 200 includes the following steps:
[0036] Step 201: performing emotion analysis on the target object based on audio data and video data associated with the target object to obtain a first emotion classification result.
[0037] This step is intended to be performed by the subject of the object demand processing method (e.g. Figure 1 When the server 105 or the terminal devices 101, 102, 103 shown in the figure obtains audio data and video data associated with a target object, it can perform emotion analysis on the target object based on the audio and video data to obtain a corresponding emotion classification result. The target object includes but is not limited to pets, users with special needs, etc.
[0038] In some optional implementations of the disclosed embodiments, simple recording and video equipment can be used to collect audio and video data without the need for wearable sensors. This is particularly true for pets, as it can avoid the discomfort or tension caused by wearable sensors to the pets. It can also avoid the problems of wearable sensors being easily dropped or devices being easily powered off, thereby ensuring the stability and accuracy of the acquisition or collection of audio and video data for the target object. The audio and video data associated with the target object can be used to obtain information of multiple modalities of the target object, including but not limited to the target object's voice information, facial expression information, posture and movement information, etc., for accurate emotional analysis of the target object.
[0039] Step 202: Monitor the current environment of the target object to obtain environmental information associated with the target object.
[0040] In the embodiment of the present disclosure, in addition to the need to perform emotional analysis on the target object, it is also necessary to monitor the environment in which the target object is currently located to obtain environmental information associated with the target object. In some optional implementations of the embodiment of the present disclosure, the environmental information associated with the target object may include but is not limited to at least one of the following: location information of the target object in the environment and information on the presence of items in the environment, wherein the information on the presence of items in the environment may refer to information on the presence of items in the environment that are closely related to the life or survival of the target object, for example, whether the target object currently has food, water, toys, vomit, essential medicines, etc. in the environment. Furthermore, in some optional implementations of the embodiment of the present disclosure, the above-mentioned environmental information may also include but is not limited to at least one of the following: noise information, temperature information, humidity information, and lighting information in the environment.
[0041] In some optional implementations of the disclosed embodiments, real-time monitoring of the aforementioned environmental information can be achieved through a multi-target tracking algorithm. Thus, the target object and related items in the environment can be monitored simultaneously to obtain the aforementioned environmental information in real time.
[0042] Step 203: Determine a demand analysis result of the target object based on the first emotion classification result, the environmental information, and the context information associated with the target object.
[0043] In some optional implementations of the disclosed embodiments, the contextual information associated with the target object includes, but is not limited to, at least one of the following: attribute information of the target object, medical information of the target object, information about the target object's activities in the environment, information about the target object's sleep status, and information about visits by objects other than the target object to the environment in which the target object is currently located. The attribute information of the target object includes, but is not limited to, at least one of the following: age, gender, breed or type; the activity information of the target object in the environment includes, but is not limited to, at least one of the following: activity frequency, activity area, whether or not the target object has not eaten for a long time; and the sleep information of the target object includes, but is not limited to, at least one of the following: duration of continuous sleep, duration of deep sleep, etc.
[0044] The object demand processing method provided by the embodiment of the present disclosure can not only perform real-time emotion analysis on the target object based on the audio data and video data associated with the target object, but also monitor the environmental information corresponding to the target object's current environment in real time. Furthermore, it is necessary to fully combine the first emotion classification result of the target object obtained and its associated environmental information and context information to comprehensively analyze the target object's needs. In this way, by fully considering the target object's emotional state, the state of the environment in which it is located, and the context information, it is possible to more comprehensively and accurately determine the target object's needs, and to achieve personalized analysis of the target object so as to truly understand the target object's needs.
[0045] Please refer to Figure 3 , Figure 3 This is a flowchart of another object demand processing method provided by an embodiment of the present disclosure, wherein process 300 includes the following steps:
[0046] Step 301: Perform emotion analysis on the target object based on audio data and video data associated with the target object to obtain a first emotion classification result.
[0047] Step 302: Monitor the current environment of the target object to obtain environmental information associated with the target object.
[0048] Step 303: Determine a demand analysis result of the target object based on the first emotion classification result, the environmental information, and the context information associated with the target object.
[0049] The above steps 301-303 are similar to the following Figure 2 Steps 201-203 shown are consistent. For the same contents, please refer to the corresponding parts of the previous embodiment and will not be repeated here.
[0050] Step 304: Based on the demand analysis result, a prompt message and / or an execution command is sent to the smart device.
[0051] This step is intended to be performed by the subject of the object demand processing method (e.g. Figure 1 The server 105 shown) sends prompt information and / or execution commands to the smart device in real time based on the demand analysis results of the target object accurately determined by fully and comprehensively considering the emotional state of the target object, the environmental state and context information, so as to meet the needs of the target object in a timely manner. In some optional implementations of the embodiments of the present disclosure, the prompt information can be used to indicate the demand analysis results and / or care suggestions for the target object to the user, and the execution command is used to control the smart device to perform the operation corresponding to the execution command. In some optional implementations of the embodiments of the present disclosure, the smart device may include but is not limited to: smart phones, smart home appliances (such as smart feeders, waterers, water dispensers, etc.).
[0052] In some optional implementations of the embodiments of the present disclosure, based on the above-mentioned demand analysis results, prompt information and / or execution commands can be sent to the smart device through an AI agent.
[0053] Step 305: Receive feedback from the smart device regarding the prompt information and / or execution command. In this embodiment, further feedback from the smart device regarding the prompt information and / or execution command may be received to indicate whether the prompt information was received in a timely manner and / or whether the execution command was executed in a timely and effective manner.
[0054] Step 306: re-analyze the emotion of the target object based on the new audio data and new video data associated with the target object to obtain a second emotion classification result.
[0055] In this embodiment, in order to verify whether the needs of the target object accurately determined based on full and comprehensive consideration of the target object's emotional state, environmental conditions, and contextual information are accurate and satisfied, after sending a prompt message and / or executing a command to the smart device based on the needs analysis result, the emotional state of the target object can be re-analyzed, and the accuracy of the previously obtained target object needs analysis result can be evaluated based on the obtained new emotion classification result (i.e., the second emotion classification result). The method for obtaining the new audio data and new video data associated with the target object is the same as the method for obtaining the audio data and video data in step 201 or 301 above, and will not be repeated here.
[0056] In some optional implementations of the embodiments of the present disclosure, the above step 306 may be executed when the above execution command is executed, or may be executed within a fixed time after the above prompt information and / or execution command is issued.
[0057] It should be noted that the execution of step 305 is optional, and the above does not constitute a limitation on the execution order of step 305 and step 306, and no specific limitation is made here.
[0058] Step 307: In response to determining that the needs analysis result is incorrect based on the second emotion classification result, a new needs analysis result for the target object is determined based on the second emotion classification result. In this embodiment, if the previously obtained needs analysis result, which serves as the basis for determining the prompt information and / or execution command, is determined to be incorrect based on the new emotion classification result, a new needs analysis result is required for the target object based on the new emotion classification result, i.e., the second emotion classification result, to obtain a new needs analysis result.
[0059] In some optional implementations of the embodiments of the present disclosure, the process of determining the new demand analysis result of the target object based on the second emotion analysis result can be specifically performed as follows: monitoring the environment in which the target object is located to obtain new environmental information associated with the target object; based on the second emotion analysis result, the new environmental information and the new context information associated with the target object, determining the new demand analysis result of the target object. The types or categories of the content contained in the new environmental information and the new context information can refer to the above Figure 2 The corresponding contents in the corresponding embodiments will not be repeated here.
[0060] According to the object demand processing method provided by the embodiment of the present disclosure, prompt information and / or execution commands can be sent to relevant smart devices in a timely manner based on the demand analysis results of the determined target object to meet the needs of the target object. Furthermore, the emotional changes of the target object can be used to verify the accuracy of the demand analysis results of the determined target object to verify whether the needs of the target object are met. In the event that it is confirmed that its needs are not met, the demand analysis of the target object can be performed again in a timely manner based on the new emotional analysis results of the target object to obtain a new demand analysis result of the target object, thereby ensuring the high accuracy of the demand analysis of the target object.
[0061] In some optional implementations of the disclosed embodiments, the method for processing target needs may further include: in response to the second emotion classification result being consistent with the first emotion classification result, determining that the target need analysis result is incorrect. In this embodiment, if the target target's emotional state does not change before and after a prompt message is sent to the smart device and / or a command is executed based on the determined need analysis result, it can be efficiently and accurately determined that the previously obtained need analysis result is inaccurate, requiring a new need analysis of the target target to promptly meet its needs.
[0062] In the above Figure 2 or Figure 3 On the basis of the corresponding embodiments, in some optional implementation methods of the embodiments of the present disclosure, the above-mentioned step 201 or 301 can be specifically executed as follows: based on the audio data and video data associated with the target object, determine the audio spectrum graph and video key frames of the target object; input the audio spectrum graph and video key frames into a pre-trained emotion analysis model to obtain a first emotion classification result.
[0063] In this embodiment, the audio data and video data associated with the target object can be first processed into audio spectrograms and video key frames suitable for input into a pre-trained emotion analysis model, and then the audio spectrograms and video key frames of the target object can be input into the emotion analysis model for analysis and processing, so as to accurately and efficiently obtain the analysis results of the emotional state of the target object, that is, the first emotion classification result.
[0064] Furthermore, in some optional implementations of the disclosed embodiments, the step of determining the target object's audio spectrogram and video keyframes based on the audio and video data associated with the target object can be specifically performed as follows: preprocessing the audio and video data associated with the target object, respectively; separating the audio spectrogram belonging to the target object from the preprocessed audio data; and extracting the video keyframes belonging to the target object from the preprocessed video data. This can reduce the amount of data subsequently processed, thereby improving the accuracy of the target object's emotion analysis.
[0065] In some optional implementations of the embodiments of the present disclosure, the above-mentioned method of preprocessing the audio data associated with the target object includes but is not limited to performing noise reduction, spectrum analysis and spectrum graph separation in sequence; and the above-mentioned method of preprocessing the video data associated with the target object includes but is not limited to performing denoising and key frame extraction in sequence.
[0066] Furthermore, in some optional implementations of the embodiments of the present disclosure, the above-mentioned step of inputting the audio spectrogram and video key frames into a pre-trained emotion analysis model to obtain a first emotion classification result can be specifically performed as follows: inputting the audio spectrogram and video key frames into the audio feature extractor and visual feature extractor of the emotion analysis model respectively to obtain the audio features and visual features of the target object; and obtaining the first emotion classification result based on the audio features and visual features.
[0067] In this embodiment, in the process of processing the input audio spectrogram and video key frames through a pre-trained emotion analysis model to obtain the result of emotion analysis of the target object, the audio feature extractor of the emotion analysis model can be used to extract features of the audio spectrogram to obtain the audio features of the target object, and the visual feature extractor of the emotion analysis model can be used to extract features of the video key frames to obtain the visual features of the target object, and then accurate analysis can be performed based on the obtained audio features and visual features to obtain the emotion classification result of the target object.
[0068] Furthermore, in some optional implementations of the embodiments of the present disclosure, the above-mentioned step of obtaining the first emotion classification result based on audio features and visual features can be specifically executed as follows: the encoder and adapter based on the emotion analysis model process the audio features and visual features to obtain the audio knowledge representation and visual knowledge representation of the target object; the first fully connected layer based on the emotion analysis model processes the audio knowledge representation to obtain the audio fusion features of the target object, and the second fully connected layer based on the emotion analysis model processes the visual knowledge representation to obtain the visual fusion features of the target object; the fusion network based on the emotion analysis model performs multimodal fusion on the audio fusion features and the visual fusion features to obtain multimodal fusion features; and based on the multimodal fusion features, the first emotion classification result is obtained.
[0069] In this embodiment, in order to ensure the accuracy of the emotion analysis of the target object, the audio data and video data of the target object are processed by the audio feature extractor and the visual feature extractor in the emotion analysis model respectively to obtain the corresponding audio features and visual features. Further, the audio knowledge representation and visual knowledge representation of the target object can be obtained through the encoder and the adapter in the emotion analysis model. Specifically, the audio features and visual features of the target object can be input into the encoder in the emotion analysis model respectively, and then the outputs of different layers of the encoder can be used as inputs of different modules of the adapter. Further, the audio knowledge representation and visual knowledge representation of the target object can be obtained through the processing of the encoder and the adapter. Further, the audio knowledge representation and visual knowledge representation can be internally fused respectively through the first fully connected layer and the second fully connected layer of the emotion analysis model, and then the audio fusion features and visual fusion features output by the two fully connected layers are multimodally fused through the fusion network of the emotion analysis model. Further, based on the multimodal fusion features output by the fusion network, an accurate emotion analysis result of the target object can be obtained.
[0070] Furthermore, in some optional implementations of the embodiments of the present disclosure, the step of obtaining a first emotion classification result based on the multimodal fusion features can be specifically performed as follows: performing emotion label classification based on the multimodal fusion features to obtain the first emotion classification result. In this way, by performing a multi-label classification task after obtaining the multimodal fusion features, an accurate emotion classification result can be ensured. The labels or emotion labels may include, but are not limited to, excitement, anger, coquettishness, fear, depression, normal, etc.
[0071] Furthermore, in some optional implementations of the embodiments of the present disclosure, the audio feature extractor and the visual feature extractor are constructed based on the Vision Transformer (ViT) model framework; and the encoder, adapter, fully connected layer, and fusion network are constructed based on the ConKI (Contrastive Knowledge Injection) model framework. The sentiment analysis model of the embodiments of the present disclosure combines some structural modules of the ViT model that can perform powerful feature extraction with some structural modules of the ConKI model for multimodal sentiment analysis. Through this sentiment analysis model, the emotional state of the target object can be accurately determined.
[0072] In the above Figure 2 or Figure 3 On the basis of the corresponding embodiments, in some optional implementation methods of the embodiments of the present disclosure, the above-mentioned step 203 or 303 can be specifically executed as follows: based on the first emotion classification result, environmental information and contextual information associated with the target object, paragraph information is constructed; the encoded information generated based on the paragraph information and the demand problem information for the target object is input into the pre-trained demand analysis model to obtain the demand analysis result.
[0073] In this embodiment, accurate demand analysis of the target object can be achieved through a pre-trained demand analysis model. Specifically, the emotion classification results, environmental information and context information of the target object can be pieced together into paragraph information, and then corresponding coding information can be generated based on the paragraph information and the demand problem information for the target object. By inputting the coded information into the pre-trained demand analysis model for processing, an accurate demand analysis result for the target object can be obtained.
[0074] It should be noted that the above-mentioned emotion analysis model in the embodiment of the present disclosure can be obtained by training a training set constructed based on sample audio data and sample video data related to the target object as input and the corresponding real emotion classification results as the expected output, and the above-mentioned demand analysis model can be obtained by training a training set constructed based on sample emotion classification results, sample environmental information and sample contextual information related to the target object as input and the corresponding real demand analysis results as the expected output.
[0075] To deepen understanding, this disclosure also provides a specific implementation scheme in combination with a specific application scenario, which involves performing a demand analysis on pets (corresponding to the target objects in the above embodiments) to meet the needs of pets. Figure 4 In this specific implementation scheme, it mainly involves a voice processing module 401, a video processing module 402, a sentiment analysis module 403 (also called an emotion analysis module), an environment analysis module 404, a demand analysis module 405 and an AI agent 406.
[0076] In this specific implementation, a pet needs analysis method based on emotion analysis and goal tracking is provided. This method analyzes multiple pet characteristics (such as voice, expression, movement, and posture) to determine the pet's emotional state. It also performs goal tracking on the pet's surrounding environment. Combining the results of the pet's emotion analysis, the goal tracking results of the surrounding environment, and the pet's contextual information, a needs analysis can be performed on the pet. Furthermore, an AI agent can process the needs analysis results and transmit instructions to the device (a smart appliance or the pet owner's terminal device, corresponding to the execution command sent to the smart device in the above embodiment) to execute the corresponding operation to meet the pet's needs. At this time, the environmental information will change, and the pet's emotions will be analyzed again to verify whether the pet's needs have been met. Specifically, when the AI agent receives feedback that the device has executed the corresponding instruction, it schedules the emotion analysis module to perform emotion analysis again and compares the results of this analysis with the previous one. If the emotion analysis results change, the needs analysis results receive positive feedback; if the emotion analysis results remain unchanged, the needs analysis is performed again to predict the pet's needs.
[0077] This specific implementation solution can be used for devices such as smart pet companion robots and home mobile robots, and can communicate and interact with Internet of Things (IoT) devices and clients to provide real-time pet demand analysis functions.
[0078] In this specific implementation, the voice processing module 401 can be used to pre-process the required pet audio data collected by the recording device to obtain the required audio data. Specifically, considering that during the audio data collection process, due to the interference of ambient sound or other sounds, real-time noise reduction processing is required to improve the audio quality; further, the noise-reduced audio data is converted into a digital signal through analog-to-digital signal conversion, and the audio data is converted to the frequency domain for analysis using spectrum analysis technology (such as short-time Fourier transform (STFT)). The audio data is then distinguished from the pet's audio data through the spectrogram. Specifically, the spectrogram can be processed using blind source separation technology, such as independent component analysis (ICA), such as filtering out ambient noise or other irrelevant information to separate the spectrogram containing only the pet's voice (corresponding to the audio spectrogram in the above embodiment).
[0079] In this specific implementation, the video processing module 402 can be used to pre-process the video data of the pet captured by the camera. Specifically, the collected video data can be denoised to improve video quality and clarity. Furthermore, a frame difference method can be used to extract key frames (corresponding to the video key frames in the above embodiment) by calculating the pixel differences between adjacent frames, thereby obtaining continuous frame data to reduce the amount of data required for subsequent processing.
[0080] In this specific implementation, the sentiment analysis module 403 can be implemented by a sentiment analysis model, specifically by two feature extractors (such as Figure 5 The audio feature extractor and video feature extractor shown in FIG) obtain audio feature I a and visual features I v , and obtain knowledge representation through encoder and adapter, such as Figure 5 As shown, by inputting the output of each feature extractor into the corresponding encoder, the corresponding audio knowledge representation can be output. a and video knowledge representation O v , and the corresponding audio knowledge representation A can be obtained by taking the output of each encoder as the input of the corresponding adapter a and video knowledge representation A v Then, the audio knowledge representation and video knowledge representation output by the encoder and adapter can be respectively subjected to multimodal fusion within each modality through the corresponding fully connected (FC) layer, as shown in the following example: Figure 5As shown, the output of the audio encoder O a and output A of adapter 1 a Internal fusion is performed through the corresponding fully connected layer, and the output O of the visual encoder is converted to v and output A of adapter 1 v Internal fusion is performed through the corresponding fully connected layer. Furthermore, the output of the fully connected layer is multimodally fused in the fusion network, and then a multi-label classification task is performed to obtain the final sentiment analysis result (corresponding to the sentiment classification result in the above embodiment). In this specific implementation scheme, the sentiment analysis model can draw on some structural modules of the ConKI model and the ViT model, focusing on the importance of video and audio modalities, avoiding the defects of pet language expression, and in the analysis, the self-attention mechanism of the ViT model can be mainly used for feature extraction.
[0081] The following is a description of the training of the sentiment analysis model in this specific implementation scheme, which may specifically involve the following contents:
[0082] During the data collection phase, audio and video files of different pets can be collected, and 50% of the data can be used as external data to train the adapter, and the other 50% of the data can be used as the target dataset. This is used to train the overall model. Considering the possibility of a small dataset, K-Fold Cross-Validation can be used to split the data multiple times for training, thereby ensuring model performance while avoiding overfitting. Furthermore, the dataset needs to be labeled, with each sample data in the dataset being labeled with a corresponding emotion (or mood) label y, such as excitement, anger, coquettishness, fear, depression, normal, etc.
[0083] The model training process can include two different training tasks. The first is the Multimodal Sentiment Analysis (MSA) task, which is the main task of the model and is essentially a classification task. The second is the contrastive knowledge subtask, which expects the outputs of the adapter and encoder to capture different knowledge representations. By integrating this hierarchical contrast information, the model can capture the comprehensive dynamic changes between knowledge representations, making the MSA results more accurate and reliable.
[0084] During specific training, the adapter is first trained to obtain the corresponding weights and parameters, and then the weights and parameters of the adapter are frozen before training the encoder: (1) Input external data ε and its label y. For each batch of data, the audio features and video features I are first extracted by the feature extractor. a and Iv , and then obtain the output O through the encoder and adapter respectively a With A a , O v With A v (2) Concatenate the specific knowledge representation and general knowledge representation of each modality to obtain the concatenated vector [ a ; A a ] and [O v ; A v ], and first perform an internal multimodal fusion through the fully connected layer. (3) Perform multimodal fusion within the fusion network, calculate the prediction results of the model through the multi-layer perceptron, and at the same time obtain the loss of the multimodal sentiment analysis task (4) During the training process of the adapter, the parameters of the two encoders are frozen, and only the weights and parameters of the adapter are updated. Finally, the adapter parameters related to the optimal validation set are saved, and the encoder is trained after completing its pre-training.
[0085] Further fine-tune the downstream tasks and input the target dataset and its label y, and the same process as the adapter pre-training, the same audio feature and video feature extraction, encoder and adapter output, splicing, multimodal fusion are performed, and the global loss can be obtained at the same time Among them, L con This is the loss function for multimodal hierarchical contrastive learning. During this training process, only the adapter weights and parameters are frozen. After training, the multi-label classification task is performed. The test data is input, and the model sorts the probabilities of each label in reverse order, and selects the top two labels as the final predicted labels.
[0086] The processing of each feature extractor mentioned above can be found in Figure 6 ,include:
[0087] (1) Image segmentation: The audio spectrogram or video keyframe is segmented into, for example, N 16*16 patch sequences, which are then sequentially reduced in dimensionality and expanded through a linear projection layer to obtain a one-dimensional vector, where the one-dimensional vector carries information related to the patch.
[0088] (2) Position encoding: Use one-dimensional absolute position encoding to embed the obtained position encoding information into the above one-dimensional vector to obtain an N*D two-dimensional vector.
[0089] (3) Transformer encoder processing: The above two-dimensional vector is sent to the Transformer encoder for processing. The Transformer encoder can include an embedding layer (Embedded), layer normalization (Norm), multi-head attention (Multi-Head Attention) and residual connection (plus sign) and a multilayer perceptron (Multilayer Perceptron, MLP). Specifically, the temporal and spatial relationship between images can be captured through the multi-head self-attention mechanism (Multi-Head Attention) to obtain a feature vector. The self-attention mechanism in this process is only performed between all image blocks in the same frame and corresponding positions in different frames, which can reduce time complexity.
[0090] like Figure 7 The adapter shown can be composed of two modules with the same structure, each of which is mainly composed of two FC layers and two Transformer layers. Specifically, these two modules can be inserted between the 1st-2nd and 3rd-4th layers of the Transformer in each modality (audio or visual) encoder, that is, the output of the 1st layer of each modality encoder can be used as the input of the first module of the corresponding adapter, and the output of the 3rd layer of each modality encoder and the output of the first module of the adapter can be used as the input of the second module of the adapter.
[0091] like Figure 8 The fusion network shown is essentially an MLP, consisting primarily of an FC layer, a Batch Normalization layer, and activation functions (ReLU and Sigmoid). Feature vectors entering the fusion network are first concatenated to combine features from different modalities. They then enter the fully connected layer, where the Batch Normalization layer constrains the input to a standard normal distribution to prevent distribution shift in the subsequent ReLU activation function. After passing through the fully connected layer, the classified output is generated.
[0092] In this specific implementation, the environment analysis module 404 can achieve real-time monitoring of environmental information through a multi-target tracking algorithm, and obtain environmental information related to the pet (such as the pet's location information, whether there is food and water, whether there are toys, etc.), such as Figure 9As shown in Figure 2, this process can use a deep learning-based multi-target tracking algorithm DeepSort (Deep Simple Online and Realtime Tracking) algorithm (a deep learning-based target tracking algorithm that achieves continuous and stable tracking of multiple targets by extracting and matching target features). First, the target object in each frame is detected using the target detection algorithm (for example, Figure 9 The position and bounding box of the pet, feeder, toy in the frame are then filtered by a Kalman filter based on the motion variables of the previous frame (which can be obtained by Figure 9 The tracking box (Tracks) shown in the figure is used to predict the motion features of the next frame, which can be specifically represented by the Mahalanobis distance. The appearance features are then obtained using the person re-identification (ReID) algorithm (also known as the cross-camera tracking algorithm), which can be specifically represented by the cosine distance. Finally, cascade matching and intersection over union (IOU) matching are used to match the data and calculate the object classification result. The classification results of the target tracking are mapped into the parameters required by the demand analysis module 405, such as indicating whether there is food in the feeder.
[0093] In this specific implementation, the demand analysis module 405 combines the emotion classification results output by the sentiment analysis module 403, the external environment information obtained by the environmental information detection module, the environmental information obtained by the sensor, and the context information stored in the cloud, and combines the natural language processing (NLP) pre-training model, such as Bidirectional Encoder Representations from Transformers (BERT), XLNet (eXtreme Language Network, a pre-trained language model based on Transformer), RoBERTa (Robustly Optimized BERT Pretraining Approach), T5 (Text-to-Text Transfer Transformer), etc., to infer the needs of pets, and then provide personalized care suggestions based on this. Among them:
[0094] (1) Emotional information (corresponding to the above-mentioned emotion classification results): Based on the emotion analysis results, the emotional information of the pet (excited, angry, coquettish, afraid, depressed, normal, etc.) is obtained.
[0095] (2) Environmental information: The environmental analysis module 404 determines whether there is food, water, toys, vomit, or a litter box. The sensor obtains noise, temperature, and other data, and pre-processes the data based on the threshold of appropriate conditions to obtain information such as whether the environment is too noisy.
[0096] (3) Contextual information: information about the pet’s age and breed, recent medical history, activity frequency, duration of sleep, whether it has not eaten for a long time, whether other animals have visited, etc. Contextual information comes from two sources, such as Figure 9 As shown, based on the physical classification results of the target tracking in the above-mentioned environmental analysis module 404, data filtering is performed and the object categories appearing in the space (in addition to environmental information) are stored. This can obtain information such as the presence of other animals. Object category data within a fixed time period is retained, and object category data outside the fixed time period is deleted. The updated data is stored in the object category library, and the upstream context information library is then updated. A recurrent neural network (RNN) or a long short-term memory network (LSTM) is used to process video information to obtain pet behavior or action information. The current pet behavior information is stored. Behavior data within a fixed time period is retained, and behavior data outside the fixed time period is deleted. The updated data is stored in the behavior information library. Based on the pet information within a fixed (or specified) time period, information such as the pet's sleep duration and activity frequency is analyzed and updated in the context information library. For example, the pet's activity information within a 24-hour or 48-hour period is updated based on the above information. For example, if the pet's sleep time is too long, if more than 95% of the behavior monitoring results within a certain period are sleep, the pet can be determined to have not eaten for a long time.
[0097] The model loss function corresponding to the demand analysis module 405 can be calculated based on the loss of the sentiment analysis model. (i.e. the above ), the loss of the target tracking model The final loss function is It can be used to evaluate and optimize the model, thereby improving the model's prediction accuracy.
[0098] During the demand analysis process, various information about pets is combined and a NLP pre-trained model similar to the BERT model is used to perform question and answer (QA) tasks and output the pet's demand results. Figure 11 As shown, including:
[0099] (1) Data preparation: The above information is pieced together to form the model's paragraph data. The questions for this model can be fixed, such as "What are the current needs of pets?" During training, relevant questions, paragraph information, answer information, etc. are collected to adjust the model parameters.
[0100] (2) Data coding stage: First, the input question information, such as Figure 11 The "What" shown in the figure is marked with [CLS] to help understand the question; in the paragraph information, the [SEP] token is added to the end of each sentence to separate the sentences, such as Figure 11 "Corgi", "toys", etc. shown in.
[0101] (3) Model analysis phase: The encoded text is input into the BERT model, and the model learns the relationship prediction results between texts through the Transformer encoder. The model output sequence is the probability prediction of the answer at each position in the context information. For each predicted position i, the model outputs the probability of the answer starting position and the probability of the end position Finally, the softmax function is used to obtain the probability distribution, and the position with the highest probability is selected as the final demand result.
[0102] Furthermore, AI Agents establish connections with related applications to analyze demand results based on the model, send information to corresponding IoT devices, or schedule connected smart devices to take corresponding measures. In this specific implementation, a smart pet translation method based on a large language model is proposed. It analyzes various types of pet biometric information, analyzes pets' emotional intentions and demand expressions in real time, and uses emotional changes to verify demand analysis to increase the accuracy of pet demand analysis.
[0103] Demand analysis model: Combines the emotion classification results output by the emotion analysis module, the environmental information obtained by the environmental analysis module 404, the pet context information and the environmental sensor data, performs a demand analysis on the pet through the NLP pre-training model, and passes the results of the demand analysis to the AI Agent for processing.
[0104] The following is an example description of the application scenarios of the above specific implementation scheme:
[0105] Example 1: Integrate the module into a home monitoring system for pet health monitoring and management
[0106] (1) Health warning: By analyzing the pet's emotional changes, the system can predict the pet's possible health problems, such as pain, discomfort, or signs of illness, thereby reminding the owner to take medical measures in time. For example, the above-mentioned sentiment analysis module analyzes the pet's unhappy or depressed mood, and the above-mentioned environment analysis module analyzes that there is food, water, toys, and vomit in the room. The context information indicates that other animals have been there, the pet has not eaten within eight hours, and the sleep time exceeds four hours. The environmental information measured by the sensor is analyzed to obtain the appropriate temperature, appropriate humidity, moderate light, and normal ambient noise. Furthermore, the above-mentioned demand analysis module can be used to determine the pet's needs: possible health problems may occur, and it is recommended to seek medical attention in time. Specifically, the AI Agent can send a reminder to the pet's mobile phone.
[0107] (2) Identification of behavioral problems: For bad behaviors of pets, such as excessive barking and furniture destruction, emotion analysis can be used to help owners identify the emotional reasons behind these behaviors and provide solutions. For example, the above-mentioned emotion analysis module can be used to analyze the emotional analysis results of the pet being bored or anxious. The above-mentioned environmental analysis module can be used to analyze that there is food in the feeder, water in the water dispenser, and no pet toys are detected. The context information indicates that there are no abnormalities, the pet has eaten within two hours, and the activity frequency is normal. The environmental information measured by the sensor can be analyzed to obtain the appropriate temperature, appropriate humidity, moderate light, and normal ambient noise. Furthermore, the above-mentioned demand analysis module can be used to obtain the pet's needs: the pet feels bored and lacks stimulation, and needs to be provided with toys or companionship and entertainment. Specifically, the AI Agent can send information to the corresponding IOT device or dispatch the connected smart device to take corresponding measures. For example, the emotional expression of the pet's language and actions can be translated to the owner in combination with the current video call application to improve the quality of communication with the pet.
[0108] At the same time, in the process of demand analysis in the above two situations, the model not only meets the needs of the pet, but also ensures the correctness of the demand analysis results through model verification. After the first demand analysis, through the first adjustment, for example, the pet has obtained a toy, the emotion analysis module can be called again within a fixed time after the AI Agent successfully dispatches the device or receives feedback from the director. If the emotion analysis result is still boredom or anxiety, and the verification analysis confirms that the emotion analysis result is the same as the last time, then the demand analysis result is wrong and a second demand analysis is required; if the emotion analysis result is comfortable or happy, and the verification analysis confirms that the emotion analysis result is different from the last time, then the demand analysis result is correct.
[0109] Example 2: Apply this module to pet smart appliances. Specifically, you can combine it with camera devices to control the operation of the appliances based on the analysis of pet needs.
[0110] The sentiment analysis module analyzes the pet's excitement or coquettishness, and the environmental analysis module analyzes the presence of food in the feeder, the inoperative water dispenser, and the presence of toys in the activity space. Contextual information indicates normal sleep time, no recent medical history, and normal activity frequency. Furthermore, environmental information measured by sensors indicates suitable temperature, humidity, lighting, and ambient noise levels. Furthermore, the demand analysis module determines the pet's needs: the pet is being coquettish, and the water dispenser may be out of water and needs to be turned on. If the pet needs water, the dispenser can be turned on to meet the pet's needs.
[0111] The sentiment analysis module analyzes the pet's excitement or coquettishness, and the environmental analysis module analyzes that there is no food in the feeder, the water dispenser is functioning normally, and toys are detected in the activity space. Contextual information indicates normal sleep time, no recent medical records, and normal activity frequency. Furthermore, environmental information measured by sensors indicates suitable temperature, humidity, lighting, and ambient noise levels. Furthermore, the demand analysis module analyzes the pet's needs, such as being coquettish and possibly out of food in the feeder, requiring it to be opened. If the pet is detected to need food, the feeder can be activated to meet the pet's needs.
[0112] In summary, this specific implementation comprehensively analyzes the pet's current movements, sounds, and expressions. Through a multimodal sentiment analysis process, the pet's status is accurately analyzed, making the sentiment analysis process more precise. A real-time multi-target tracking algorithm is used to obtain environmental information related to the pet. A neural network is constructed to detect the pet's current behavior, and the pet's contextual information is updated based on this behavior. Combining the results of pet sentiment analysis with external environmental information and the pet's contextual information to analyze pet needs can effectively improve the accuracy of pet needs analysis and provide personalized care recommendations based on the pet's emotional state and needs. Specifically, it can be deployed on devices such as televisions and mobile phones, allowing users to monitor their pet's status in real time and make video calls with their pet. It can also be used to recommend personalized pet products, assist in pet training, and recommend smart pet toys. It can also be used in care systems for user groups with special needs.
[0113] Further references Figure 12 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an object demand processing device, which is similar to Figure 2Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0114] like Figure 2 As shown, the object demand processing device 1200 of this embodiment may include: an emotion analysis module 1201, an environment monitoring module 1202, and a demand analysis module 1203. The emotion analysis module 1201 is configured to perform emotion analysis on the target object based on the audio data and video data associated with the target object to obtain a first emotion classification result; the environment monitoring module 1202 is configured to monitor the environment in which the target object is currently located to obtain environmental information associated with the target object; and the demand analysis module 1203 is configured to determine a demand analysis result of the target object based on the first emotion classification result, the environmental information, and the context information associated with the target object.
[0115] In this embodiment, in the object demand processing device 1200, the specific processing of the emotion analysis module 1201, the environment monitoring module 1202 and the demand analysis module 1203 and the technical effects thereof can be referred to respectively. Figure 2 The relevant descriptions of steps 201-203 in the corresponding embodiment are not repeated here.
[0116] In some optional implementations of this embodiment, the object demand processing apparatus 1200 further includes a sending module and a receiving module (not shown). The sending module is configured to send prompt information and / or execution commands to the smart device based on the demand analysis results; and the receiving module is configured to receive feedback from the smart device regarding the prompt information and / or execution commands.
[0117] In some optional implementations of this embodiment, the above-mentioned emotion analysis module 1201 is further configured to: re-perform emotion analysis on the target object based on new audio data and new video data associated with the target object to obtain a second emotion classification result; the above-mentioned demand analysis module 1203 is further configured to: in response to determining that the demand analysis result is wrong based on the second emotion classification result, determine a new demand analysis result for the target object based on the second emotion classification result.
[0118] In some optional implementations of this embodiment, the demand analysis module 1203 is further configured to: in response to the second emotion classification result being consistent with the first emotion classification result, determine that the demand analysis result is wrong.
[0119] In some optional implementations of this embodiment, the above-mentioned emotion analysis module 1201 is further specifically configured to: determine the audio spectrum and video key frames of the target object based on the audio data and video data associated with the target object; input the audio spectrum and video key frames into a pre-trained emotion analysis model to obtain a first emotion classification result.
[0120] In some optional implementations of this embodiment, the above-mentioned emotion analysis module 1201 is further specifically configured to: input the audio spectrogram and video key frames into the audio feature extractor and visual feature extractor of the emotion analysis model respectively to obtain the audio features and visual features of the target object; based on the audio features and visual features, obtain the first emotion classification result.
[0121] In some optional implementations of this embodiment, the above-mentioned emotion analysis module 1201 is further specifically configured as follows: based on the encoder and adapter of the emotion analysis model, the audio features and visual features are processed to obtain the audio knowledge representation and visual knowledge representation of the target object; based on the first fully connected layer of the emotion analysis model, the audio knowledge representation is processed to obtain the audio fusion features of the target object, and based on the second fully connected layer of the emotion analysis model, the visual knowledge representation is processed to obtain the visual fusion features of the target object; based on the fusion network of the emotion analysis model, the audio fusion features and the visual fusion features are multimodally fused to obtain the multimodal fusion features; based on the multimodal fusion features, a first emotion classification result is obtained.
[0122] In some optional implementations of this embodiment, the above-mentioned audio feature extractor and visual feature extractor are constructed based on the visual transformation ViT model framework; and the above-mentioned encoder, adapter, fully connected layer and fusion network are constructed based on the contrastive knowledge injection ConKI model framework.
[0123] In some optional implementations of this embodiment, the above-mentioned demand analysis module 1203 is further specifically configured to: construct paragraph information based on the first emotion classification result, environmental information and contextual information associated with the target object; input the encoded information generated based on the paragraph information and the demand problem information for the target object into a pre-trained demand analysis model to obtain the demand analysis result.
[0124] In some optional implementations of this embodiment, the above-mentioned environmental information includes at least one of the following: location information of the target object in the environment, information on the presence of items in the environment; the context information associated with the above-mentioned target object includes at least one of the following: attribute information of the target object, medical information of the target object, activity information of the target object in the environment, sleep information of the target object, and visit information of other objects other than the target object to the environment.
[0125] This embodiment exists as a device embodiment corresponding to the above-mentioned method embodiment. The object demand processing device 1200 provided in this embodiment can not only perform real-time emotion analysis on the target object based on the audio data and video data associated with the target object, but also monitor the environmental information corresponding to the target object's current environment in real time. Furthermore, it is necessary to fully combine the first emotion classification result of the target object obtained and its associated environmental information and contextual information to comprehensively analyze the target object's needs. In this way, by fully considering the emotional state of the object, the state of the environment in which it is located, and the contextual information, the needs of the target object can be determined more comprehensively and accurately, and personalized analysis of the target object can be achieved to truly understand the needs of the target object.
[0126] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the object demand processing method described in any of the above embodiments can be implemented when the at least one processor executes them.
[0127] According to an embodiment of the present disclosure, the present disclosure further provides a readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to implement the object demand processing method described in any of the above embodiments when executed.
[0128] According to an embodiment of the present disclosure, the present disclosure further provides a computer program product, which, when executed by a processor, can implement the object demand processing method described in any of the above embodiments.
[0129] Figure 13 A schematic block diagram of an example electronic device 1300 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0130] like Figure 13As shown, electronic device 1300 includes a computing unit 1301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1302 or a computer program loaded from a storage unit 1308 into a random access memory (RAM) 1303. Various programs and data required for the operation of electronic device 1300 can also be stored in RAM 1303. Computing unit 1301, ROM 1302, and RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface 1305 is also connected to bus 1304.
[0131] Multiple components in electronic device 1300 are connected to I / O interface 1305, including: an input unit 1306, such as a keyboard, mouse, etc.; an output unit 1307, such as various types of displays, speakers, etc.; a storage unit 1308, such as a magnetic disk, optical disk, etc.; and a communication unit 1309, such as a network card, modem, wireless communication transceiver, etc. The communication unit 1309 allows electronic device 1300 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0132] The computing unit 1301 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 1301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1301 performs the various methods and processes described above, such as the object demand processing method. For example, in some embodiments, the object demand processing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1308. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 1300 via the ROM 1302 and / or the communication unit 1309. When the computer program is loaded into the RAM 1303 and executed by the computing unit 1301, one or more steps of the object demand processing method described above can be performed. Alternatively, in other embodiments, the computing unit 1301 may be configured to execute the object demand processing method in any other appropriate manner (for example, by means of firmware).
[0133] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0134] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0135] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0136] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0137] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0138] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The client-server relationship arises through computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within a cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and virtual private server (VPS) services.
[0139] According to the technical solution of the embodiment of the present disclosure, not only can the target object's emotions be analyzed in real time based on the audio and video data associated with the target object, but the environmental information corresponding to the target object's current environment can also be monitored in real time. Furthermore, it is necessary to fully combine the first emotion classification result of the target object and its associated environmental information and contextual information to comprehensively analyze the target object's needs. In this way, by fully considering the target object's emotional state, the state of the environment in which it is located, and the contextual information, the target object's needs can be determined more comprehensively and accurately, and personalized analysis of the target object can be achieved to truly understand the target object's needs.
[0140] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0141] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for processing object requirements, comprising: Performing emotion analysis on the target object based on audio data and video data associated with the target object to obtain a first emotion classification result; Monitor the environment in which the target object is currently located to obtain environmental information associated with the target object; Determine a demand analysis result of the target object based on the first emotion classification result, the environmental information, and context information associated with the target object.
2. The method according to claim 1, further comprising: Based on the demand analysis results, sending prompt information and / or executing commands to the smart device; Receive feedback from the smart device regarding the prompt information and / or the execution command.
3. The method according to claim 2, further comprising: re-performing emotion analysis on the target object based on new audio data and new video data associated with the target object to obtain a second emotion classification result; In response to determining that the demand analysis result is wrong based on the second emotion classification result, a new demand analysis result of the target object is determined based on the second emotion classification result.
4. The method according to claim 3, further comprising: In response to the second emotion classification result being consistent with the first emotion classification result, it is determined that the demand analysis result is incorrect.
5. The method according to claim 1, wherein The performing emotion analysis on the target object based on the audio data and video data associated with the target object to obtain a first emotion classification result includes: Determining an audio spectrogram and video keyframes of the target object based on audio data and video data associated with the target object; The audio spectrogram and the video key frame are input into a pre-trained emotion analysis model to obtain the first emotion classification result.
6. The method according to claim 5, wherein: Inputting the audio spectrogram and the video key frame into a pre-trained emotion analysis model to obtain the first emotion classification result includes: Inputting the audio spectrogram and the video key frame into the audio feature extractor and the visual feature extractor of the emotion analysis model respectively to obtain the audio features and visual features of the target object; The first emotion classification result is obtained based on the audio features and the visual features.
7. The method according to claim 6, wherein: Obtaining the first emotion classification result based on the audio feature and the visual feature includes: Processing the audio features and the visual features based on the encoder and the adapter of the emotion analysis model to obtain the audio knowledge representation and the visual knowledge representation of the target object; The audio knowledge representation is processed based on the first fully connected layer of the emotion analysis model to obtain an audio fusion feature of the target object, and the visual knowledge representation is processed based on the second fully connected layer of the emotion analysis model to obtain a visual fusion feature of the target object; Performing multimodal fusion on the audio fusion feature and the visual fusion feature based on the fusion network of the emotion analysis model to obtain a multimodal fusion feature; Based on the multimodal fusion features, the first emotion classification result is obtained.
8. The method according to claim 7, wherein: The audio feature extractor and the visual feature extractor are constructed based on a visual transformation (ViT) model framework; and The encoder, the adapter, the fully connected layer and the fusion network are constructed based on the Contrastive Knowledge Injection ConKI model framework.
9. The method according to claim 1, wherein: The determining of a demand analysis result of the target object based on the first emotion classification result, the environmental information, and context information associated with the target object includes: constructing paragraph information based on the first emotion classification result, the environmental information, and context information associated with the target object; The coded information generated based on the paragraph information and the demand problem information for the target object is input into a pre-trained demand analysis model to obtain the demand analysis result.
10. The method according to claim 1, wherein The environmental information includes at least one of the following: location information of the target object in the environment, and information about the presence of items in the environment; The context information associated with the target object includes at least one of the following: attribute information of the target object, medical information of the target object, activity information of the target object in the environment, sleep information of the target object, and visit information of other objects other than the target object to the environment.
11. An object demand processing device, comprising: an emotion analysis module configured to perform emotion analysis on the target object based on audio data and video data associated with the target object to obtain a first emotion classification result; An environment monitoring module is configured to monitor the environment in which the target object is currently located and obtain environmental information associated with the target object; The demand analysis module is configured to determine a demand analysis result of the target object based on the first emotion classification result, the environmental information, and context information associated with the target object.
12. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the object demand processing method according to any one of claims 1 to 10.
13. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute the object demand processing method according to any one of claims 1 to 10.
14. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the steps of the object requirement processing method according to any one of claims 1 to 10.