Pet emotional soothing method, apparatus, computing device, and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN PEIMI SMART HOME TECHNOLOGY CO LTD
- Filing Date
- 2026-06-16
- Publication Date
- 2026-08-04
AI Technical Summary
[0004]本申请实施例提供一种宠物情绪安抚方法、装置及介质,可以解决目前的宠物情绪安抚方案无法实现宠物情绪的识别与智能差异化安抚,难以满足宠物的陪护需求的问题
[0015]This application provides a pet emotion soothing method, device, computing device, and medium. Responding to a sound acquisition command, it detects sound data in a target environment. When the sound data contains pet sound information of a target pet, it extracts the feature vector corresponding to the pet sound information. Then, based on the feature vector and a pre-trained reference feature vector set, it determines the pet's emotion. Finally, based on the pet's emotion and a preset soothing strategy, it soothes the target pet. In the pet emotion soothing solution provided in this application, a pre-trained reference feature vector set is used to complete feature comparison and determine the pet's current emotional state. Based on this, a preset soothing strategy is matched according to the identified pet emotion, and corresponding devices are linked to execute targeted soothing actions. This breaks away from the single audio playback or fixed feeding mode, achieving intelligent and differentiated emotional guidance and care, effectively improving the companionship effect and meeting the pet's companionship needs.
Smart Images

Figure CN122498460A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of human-computer interaction technology, specifically to a method, device, computing device, and medium for calming a pet's emotions. Background Technology
[0002] As the number of pets living alone in cities continues to increase, pets are prone to loneliness, anxiety, hunger, and other negative emotions when their owners are away, leading to a growing market demand for smart pet companion devices. Currently, mainstream pet companion products mainly integrate basic functions such as video monitoring, timed feeding, and fixed sound effects, offering only passive viewing and mechanical care with a low level of intelligence.
[0003] While existing devices with sound acquisition capabilities can capture pet vocalizations, they struggle to determine a pet's true emotional state. They can only provide simple sound alarms and cannot proactively sense emotions. Furthermore, traditional devices offer limited soothing modes, with various peripherals operating independently. This makes it impossible to tailor solutions to the pet's real-time emotions. Relying solely on fixed feeding or audio playback is insufficient to effectively alleviate negative emotions. Therefore, current pet emotional soothing solutions fail to recognize and intelligently differentiate pet emotions, thus failing to meet the companionship needs of pets. Summary of the Invention
[0004] This application provides a method, device, and medium for soothing pet emotions, which can solve the problem that current pet emotion soothing solutions cannot recognize pet emotions and provide intelligent differentiated soothing, thus failing to meet the pet's companionship needs.
[0005] In a first aspect, embodiments of this application provide a method for calming a pet's emotions, including: In response to a sound acquisition command, detect sound data in the target environment; When the sound data contains the pet sound information of the target pet, the feature vector corresponding to the pet sound information is extracted; The pet's emotions are determined based on the feature vectors and the pre-trained reference feature vector set. Based on the pet's emotions and preset soothing strategies, the target pet is emotionally comforted.
[0006] Optionally, in some embodiments of this application, determining the pet emotion of the target pet based on the feature vector and a pre-trained reference feature vector set includes: Obtain the reference feature vector set obtained through pre-training; Calculate the similarity between the feature vector and each reference feature vector in the reference feature vector set; The pet's emotional state is determined based on the calculation results.
[0007] Optionally, in some embodiments of this application, determining the pet's mood based on the calculation results includes: Based on the calculation results, the reference feature vector with the highest similarity to the feature vector is determined as the target feature vector; Obtain the emotion label corresponding to the target feature vector, and determine the emotion label corresponding to the target feature vector as the pet emotion of the target pet.
[0008] Optionally, in some embodiments of this application, the step of calming the target pet based on the pet's emotions and a preset calming strategy includes: Obtain the target soothing strategy corresponding to the pet's emotion from the preset soothing strategy; Determine the target device invoked by the target appeasement strategy; The target device is controlled according to the target soothing strategy to soothe the target pet's emotions.
[0009] Optionally, in some embodiments of this application, controlling the invoked target device according to the target soothing strategy to soothe the target pet includes: If the target soothing strategy is the first soothing strategy, then the audio device to be invoked will be controlled to play a preset voice message, and the interactive device to be invoked will be controlled to interact with the target pet. If the target soothing strategy is the second soothing strategy, then the feeding device invoked by the control will perform the feeding operation.
[0010] Optionally, in some embodiments of this application, it further includes: Obtain sample sound data of the target pet; In response to the classification operation on the sample sound data, an emotion label corresponding to the sample sound data is added, and a reference feature vector corresponding to the sample sound data after the label is added is extracted.
[0011] Optionally, in some embodiments of this application, the step of adding an emotion label corresponding to the sample sound data in response to a classification operation on the sample sound data, and extracting a reference feature vector corresponding to the labeled sample sound data, includes: A sound classification interface is displayed, which includes multiple sound tag controls; In response to a tag selection operation on the sound classification interface, the target emotion tag corresponding to the selected sound tag control is determined; Bind the target emotion tag to the sample sound data; The labeled sample sound data is segmented according to a preset time interval to obtain the sound source slice sequence corresponding to the labeled sample sound data. Generate the Mel spectrogram corresponding to the sound source slice sequence, and parse the Mel spectrogram to obtain the reference feature vector corresponding to the sample sound data after the label.
[0012] Secondly, embodiments of this application provide a pet emotional comforting device, comprising: The detection module is used to detect sound data in the target environment in response to sound acquisition commands; The extraction module is used to extract the feature vector corresponding to the pet sound information when the sound data contains the pet sound information of the target pet. The determination module is used to determine the pet emotion of the target pet based on the feature vector and a pre-trained set of reference feature vectors. The soothing module is used to soothe the target pet's emotions based on the pet's emotions and preset soothing strategies.
[0013] Thirdly, embodiments of this application also provide a computing device, including: At least one processor; and A memory communicatively connected to the at least one processor; wherein: The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the pet emotional soothing steps as described above.
[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the pet emotional soothing method described in the first aspect.
[0015] This application provides a pet emotion soothing method, device, computing device, and medium. Responding to a sound acquisition command, it detects sound data in a target environment. When the sound data contains pet sound information of a target pet, it extracts the feature vector corresponding to the pet sound information. Then, based on the feature vector and a pre-trained reference feature vector set, it determines the pet's emotion. Finally, based on the pet's emotion and a preset soothing strategy, it soothes the target pet. In the pet emotion soothing solution provided in this application, a pre-trained reference feature vector set is used to complete feature comparison and determine the pet's current emotional state. Based on this, a preset soothing strategy is matched according to the identified pet emotion, and corresponding devices are linked to execute targeted soothing actions. This breaks away from the single audio playback or fixed feeding mode, achieving intelligent and differentiated emotional guidance and care, effectively improving the companionship effect and meeting the pet's companionship needs. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of a scenario illustrating the pet emotional soothing method provided in the embodiments of this application; Figure 2 This is a flowchart illustrating the pet emotional soothing method provided in the embodiments of this application; Figure 3 This is a schematic diagram of the pet emotional calming device provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the lawnmower robot provided in the embodiments of this application. Detailed Implementation
[0018] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of systems and methods consistent with those detailed in the appended claims or with some aspects of this application.
[0019] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover descriptions such as non-exclusive inclusion, so that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment.
[0020] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0021] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.
[0022] To address the aforementioned technical problems and overcome the shortcomings of existing technologies, this application provides a method, device, computing device, and medium for calming pet emotions, which can improve the accuracy of the lawnmower robot's judgment of the working environment and its level of intelligent control.
[0023] In the embodiments of the pet emotion-soothing method in this application, a pet emotion-soothing device is used as the executing entity. For simplicity and ease of description, this executing entity will be omitted in subsequent method embodiments. The pet emotion-soothing device is applied to a computing device equipped with a camera, such as a mobile phone, tablet, PC, smart TV, vehicle terminal, network camera, high-definition camera, or other terminal device. Optionally, in some embodiments of this application, please refer to... Figure 1 , Figure 1 This diagram illustrates the application environment of the pet emotional soothing method provided in this embodiment. The pet emotional soothing method can be applied to a pet emotional soothing system. This system may include a terminal 110 and a server 120. The terminal 110 and server 120 are connected via a network. The terminal 110 can be a desktop terminal or a mobile terminal; the mobile terminal can be at least one of a mobile phone, tablet, or laptop. The server 120 can be implemented using a standalone server or a server cluster consisting of multiple servers. Specifically, the server 120 can be used for: In response to a sound acquisition command, the system detects sound data in the target environment. When the sound data contains the pet's voice information, the system extracts the feature vector corresponding to the pet's voice information. Based on the feature vector and a pre-trained set of reference feature vectors, the system determines the target pet's emotional state. Based on the pet's emotional state and a preset soothing strategy, the system soothes the target pet.
[0024] The pet emotion-soothing solution provided in this application uses a pre-trained reference feature vector set to complete feature comparison and determine the pet's current emotional state. Based on this, it matches the identified pet emotion with a corresponding preset soothing strategy and links the corresponding device to perform targeted soothing actions. This breaks away from the single audio playback or fixed feeding mode, realizing intelligent and differentiated emotional guidance and care, effectively improving the companionship effect and meeting the pet's companionship needs.
[0025] Please see Figure 2 , Figure 2 This is a flowchart illustrating the pet emotional soothing method provided in this application embodiment. The pet emotional soothing method provided in this application embodiment may specifically include the following steps: S1. In response to the sound acquisition command, detect sound data in the target environment.
[0026] The sound acquisition command is a control command used to trigger the device to start its sound recording function. It can be automatically generated after the device is powered on, triggered by a scheduled task, or manually issued by the user. It is the trigger signal to start the sound acquisition process. The target environment refers to the indoor environment of the home where the target pet usually lives, which is also the deployment area of this smart companion device. It is generally the living room, bedroom, or other spaces where the pet frequently resides. The sound data is the raw digital audio data stream picked up by the device's sound receiving device in the target environment. It includes all audio content such as pet barks, environmental noise, and other indoor sounds. It is the raw audio signal that has not been filtered or processed.
[0027] For example, specifically, when a sound acquisition command is received, the built-in microphone and other audio-receiving hardware are activated, putting the audio-receiving module into operation. The activated microphone continuously acquires various audio signals within the target environment. Then, the acquired analog audio signals are converted into standard digital audio formats to generate complete sound data. For instance, if a pet owner is worried about their pet while away, they can open the accompanying mobile app and click the real-time audio recording function; the app sends a sound acquisition command to the indoor companion device, and the device immediately activates its microphone upon receiving the command; the microphone picks up the audio in the target environment (pet's bedroom); the audio signal undergoes analog-to-digital conversion to generate the corresponding sound data.
[0028] S2. When the sound data contains the pet sound information of the target pet, extract the feature vector corresponding to the pet sound information.
[0029] Pet sound information refers to the specific vocal segments of a target pet separated from mixed raw sound data after noise reduction, sound source screening, and noise removal. It contains only the various sounds made by the pet and excludes environmental noises such as air conditioner operation, outdoor traffic noise, appliance noise, and human voices. For example, in mixed audio collected by the device in a living room, after removing fan noise and outside noise, sound segments such as dog barking, cat whimpering, and pet whimpering are separated out; this pure vocal content is the pet sound information. The feature vector is a digital acoustic representation vector obtained after standardizing the audio processing of the pet sound information. In this embodiment, the feature vector includes two core dimensions: frequency and energy, used to quantitatively describe the acoustic characteristics of the pet's vocalizations. For example, after processing the sound of a dog barking continuously, a set of digital vectors representing the frequency and energy of the vocalizations is generated; this set of digital vectors is the feature vector. For example, noise reduction is performed on the original sound data to filter out environmental interference and separate the independent and complete pet sound information. Then, the separated pet sound information is segmented to generate multiple standardized sound source slices. Next, for each sound source slice, a corresponding Mel spectrogram is generated to simulate the characteristics of human hearing and serve as the input data for the model. Then, the ResNet model is called to analyze the Mel spectrograms one by one, extracting multiple sets of vectors, which are combined to form a feature vector sequence. Finally, the average value of the above feature vector sequence is calculated, and the result is determined as the feature vector corresponding to that segment of pet sound information. Specifically, in a home setting, noise such as air conditioning and outdoor traffic can be filtered out first to separate the dog's barking. Then, the barking audio is cut into 4-second segments to obtain multiple short audio slices. Each slice is then used to generate a Mel spectrogram, which is then fed into a ResNet model to output a set of feature vector sequences. Finally, the average value of the vector sequences is calculated to obtain the feature vector corresponding to the barking segment.
[0030] S3. Determine the target pet's emotional state based on the feature vector and the pre-trained reference feature vector set.
[0031] The reference feature vector set is a benchmark dataset constructed and persistently stored during the offline modeling phase. It consists of multiple reference feature vectors bound to emotion labels. Each reference feature vector is obtained from sample sound data of the target pet under different emotions, after slicing, generating Mel spectrograms, extracting using a ResNet model, and averaging the sequences. It can be stored locally on the device or in a cloud database. For example, this reference feature vector set contains multiple sets of vectors, each bound to emotion labels such as loneliness, anxiety, and hunger. Each set of vectors corresponds to the acoustic features of the pet's vocalizations for a specific emotion. Pet emotions are categorized into emotion and state types based on the pet's vocalization characteristics. In some embodiments of this application, it is possible to identify various negative emotions and feeding needs of pets when they are alone.
[0032] For example, a pre-trained set of reference feature vectors is retrieved from the device's local storage or a cloud database. Then, each reference feature vector and its associated emotion tag are read sequentially from the set. Next, a cosine similarity algorithm is used to calculate the similarity value between the current real-time feature vector and the reference feature vector. The similarity value and the corresponding emotion tag are temporarily associated and saved. Then, all similarity values are compared, the set of data with the highest similarity is selected, and the emotion tag corresponding to the highest similarity result is determined as the target pet's current emotion.
[0033] Optionally, in some embodiments of this application, the step of "determining the pet's emotion based on the feature vector and a pre-trained set of reference feature vectors" may specifically include: Obtain the reference feature vector set obtained through pre-training; Calculate the similarity between the feature vector and each reference feature vector in the reference feature vector set; The pet's emotional state is determined based on the calculation results.
[0034] For example, a pre-trained set of reference feature vectors is retrieved from the device's local storage or a cloud database. Then, individual reference feature vectors are read sequentially from the set, along with the emotion tag associated with each vector. Next, a cosine similarity algorithm is used to calculate the similarity between the currently extracted feature vector and the reference feature vector, and the calculated similarity value and corresponding emotion tag are recorded simultaneously. Finally, the set of results with the highest similarity is selected, and the emotion tag corresponding to the highest similarity result is determined as the target pet's current emotion.
[0035] Optionally, in some embodiments of this application, the step "determining the pet's emotions based on the calculation results" may specifically include: Based on the calculation results, the reference feature vector with the highest similarity to the feature vector is determined as the target feature vector; Obtain the emotion label corresponding to the target feature vector, and determine the emotion label corresponding to the target feature vector as the pet emotion of the target pet.
[0036] For example, specifically, after selecting the reference feature vector (i.e. the target feature vector) with the highest similarity, the emotion tag that was pre-bound to the target feature vector in the offline modeling stage is retrieved, and then the read emotion tag is determined as the current pet emotion of the target pet.
[0037] S4. Based on the pet's emotions and preset soothing strategies, soothe the target pet's emotions.
[0038] The soothing strategy is a pre-configured solution for handling different pet emotions. Optionally, in some embodiments of this application, the soothing strategy may include a first soothing strategy and a second soothing strategy. The first soothing strategy corresponds to negative emotions such as loneliness or anxiety in the pet, controls the audio device to play preset voice messages, and drives the interactive device to play and interact with the pet. The second soothing strategy corresponds to the pet's hunger state, controls the feeding device to perform automatic feeding operations.
[0039] For example, based on the currently identified pet emotion, a preset strategy library is searched to retrieve the corresponding soothing strategy. Then, based on the retrieved soothing strategy type, the hardware device to be linked is identified: if it is the first soothing strategy, the target device is determined to be an audio device or an interactive device; if it is the second soothing strategy, the target device is determined to be a feeding device. Next, a communication connection is established with the target device to ensure that control commands can be issued normally. If the first soothing strategy is executed, the audio device is controlled to play preset owner interaction voice messages, and a start command is issued to the interactive device to control its operation, allowing it to play with the pet and alleviate negative emotions. If the second soothing strategy is executed, a feeding command is issued to the feeding device to control it to complete the quantitative feeding operation.
[0040] Optionally, in some embodiments of this application, the step of "soothing the target pet's emotions based on the pet's emotions and preset soothing strategies" may specifically include: Obtain the target soothing strategy corresponding to the pet's emotions from the preset soothing strategies; Identify the target device to be invoked in the target appeasement strategy; The target device is controlled according to the target soothing strategy in order to soothe the target pet's emotions.
[0041] For example, the system retrieves the preset soothing strategy library within the device, combines it with the currently determined pet emotion, extracts the target soothing strategy corresponding to that pet emotion, then queries the preset device association relationships for the target soothing strategy, and then establishes a communication link with the identified target device via home WiFi. According to the execution rules of the target soothing strategy, it sends control commands to the target device to execute the emotion soothing operation. If it is the first soothing strategy, it controls the audio device to play a preset voice and simultaneously drives the interactive device to operate and interact with the target pet; if it is the second soothing strategy, it controls the feeding device to perform a quantitative feeding operation.
[0042] Optionally, in some embodiments of this application, the step "controlling the invoked target device according to the target soothing strategy to soothe the target pet's emotions" may specifically include: If the target soothing strategy is the first soothing strategy, then control the invoked audio device to play the preset voice and control the invoked interactive device to interact with the target pet; If the target soothing strategy is the second soothing strategy, then the feeding device invoked by the control will perform the feeding operation.
[0043] For example, if the target soothing strategy is the first soothing strategy, a control command is sent to the activated audio device to start the device and play a pre-stored preset soothing voice. At the same time, a start command is sent to the activated interactive device to drive the interactive device to operate and interact with the target pet through actions, sounds, and lights. After the preset voice playback duration and device interaction duration are reached, the audio device and interactive device are turned off respectively, and both devices switch to standby mode.
[0044] If the target soothing strategy is the second soothing strategy, a feeding command is sent to the activated feeding device, controlling the device to complete the food dispensing operation according to preset parameters. Once the feeding action is completed, the feeding device automatically stops working and returns to standby mode.
[0045] Specifically, if the system determines that the pet is agitated, the first soothing strategy is activated. The main control device instructs the audio device to play a series of soothing voice messages recorded by the owner, while simultaneously starting an electric pet toy to move back and forth. This combination of voice companionship and fun interaction helps alleviate the pet's agitation. After the set time, both devices automatically shut down. If the system determines that the pet is hungry, the second soothing strategy is activated. The main control device sends a command to the smart feeder, which dispenses a preset amount of pet food. Once feeding is complete, the device immediately stops operating.
[0046] Optionally, in some embodiments of this application, it may further include: Obtain sample sound data of the target pet; In response to the classification operation on the sample audio data, an emotion label corresponding to the sample audio data is added, and the reference feature vector corresponding to the sample audio data after the label is added is extracted.
[0047] The sample audio data consists of raw audio data of the target pet in different emotions and states, collected to build the emotion recognition benchmark model. This data is recorded using devices such as microphones, smart collars, voice recorders, and mobile phones. The sample audio data includes various vocalizations of the target pet; for example, audio files of the pet whimpering alone, barking restlessly, and begging for food were recorded using a mobile phone and a pet smart collar. This batch of audio files constitutes the sample audio data. Emotion tags are manually defined classification labels used to categorize and bind the sample audio data, achieving a mapping between audio and emotion. For example, audio of a pet whimpering alone is labeled "loneliness," audio of frequent barking is labeled "anxiety," and short begging sounds are labeled "hunger." These labels are all emotion tags. The reference feature vector is a digitized acoustic vector obtained from the sample audio data with bound emotion tags after standardized slicing, generating a Mel spectrogram, extracting features using a ResNet model, and calculating the sequence mean. This reference feature vector carries the acoustic features of the corresponding audio and is bound to an emotion tag. These are then aggregated to form a reference feature vector set, serving as a comparison benchmark for real-time emotion recognition.
[0048] For example, a smartphone, voice recorder, or smart collar with recording function can be used to record the vocalizations of a target pet under different emotions in different scenarios and time periods, collecting multiple raw audio segments. The collected sample sound data is then exported to the terminal device, and a corresponding sound classification interface is opened. Next, based on the pet's actual state corresponding to each audio segment, a matching emotion tag is selected and bound to the current sample sound data. Then, the tagged sample sound data is segmented at 4-second intervals to generate multiple standardized sound source slice sequences. For each sound source slice, a corresponding Mel spectrogram is generated, converting the audio signal into image features that the model can recognize. Subsequently, a ResNet model is used to analyze each Mel spectrogram, extracting multiple sets of feature vectors, which are combined to form a feature vector sequence. Finally, the average value of the feature vector sequence is calculated, and the calculated average value is used as the reference feature vector corresponding to that tagged sample sound data segment.
[0049] Specifically, the owner used a mobile phone and a smart collar to record the sounds of their pet dog whimpering, barking, and begging for food when alone for several consecutive days, compiling multiple audio clips (sample sound data). Then, the owner opened the accompanying app, categorized the whimpering audio and added a "lonely" tag, the barking audio a "restless" tag, and the begging sounds a "hungry" tag. The mobile phone then sliced each tagged audio clip into 4-second units, generated a Mel spectrogram, and fed it into a ResNet model for processing. The mean of the output feature vector sequence was calculated to obtain multiple reference feature vectors bound to the "lonely," "restless," and "hungry" tags, respectively.
[0050] Optionally, in some embodiments of this application, the step "in response to a classification operation on the sample sound data, adding an emotion label corresponding to the sample sound data, and extracting a reference feature vector corresponding to the labeled sample sound data" may specifically include: The interface displays a sound categorization feature, which includes multiple sound tag controls. In response to a tag selection operation on the sound classification interface, determine the target emotion tag corresponding to the selected sound tag control; Bind the target emotion label to the sample sound data; The labeled sample audio data is segmented according to a preset time interval to obtain the sound source slice sequence corresponding to the labeled sample audio data. Generate Mel spectrograms corresponding to the sound source slice sequence, and parse the Mel spectrograms to obtain the reference feature vectors corresponding to the labeled sample sound data.
[0051] A sound source slice sequence is an ordered set of multiple independent short audio segments obtained by equally dividing a whole segment of sample sound data with attached emotion tags according to a preset time interval. Each short audio segment is called a sound source slice, and the sequential combination of all slices forms the sound source slice sequence. A Mel-spectrum image is an audio feature map that simulates the characteristics of human hearing. The time-domain format sound source slice audio signal is algorithmically converted and mapped to a Mel frequency scale, transforming the audio's time, frequency, and sound energy into two-dimensional image data.
[0052] For example, specifically, the operating terminal loads and presents a sound classification interface, which deploys multiple sound tag controls, each corresponding to different emotion markers such as loneliness, anxiety, and hunger, for the operator to select and use. When the operator clicks on a sound tag control, the system retrieves the preset association information of that control to determine the currently selected target emotion tag. The target emotion tag determined in the previous step is uniquely associated with the currently processed sample sound data, completing the tag binding. Further, the system segments the tagged sample sound data into segments at preset 4-second intervals; after segmentation, all independent short audio segments are arranged sequentially to form a corresponding sound source slice sequence. Next, the system traverses the sound source slice sequence, performing signal conversion on each sound source slice to generate a corresponding Mel spectrogram; and calls the ResNet model to analyze the Mel spectrograms one by one, extracting multiple sets of feature vectors, which are combined to form a feature vector sequence; finally, the average value of the feature vector sequence is calculated, and the result is determined as the reference feature vector corresponding to the tagged sample sound data segment.
[0053] In summary, the pet emotion-soothing method provided in this embodiment responds to a sound acquisition command, detects sound data in the target environment, and when the sound data contains the pet's sound information, extracts the feature vector corresponding to the pet's sound information. Then, based on the feature vector and a pre-trained reference feature vector set, the pet's emotion is determined. Finally, based on the pet's emotion and a preset soothing strategy, the target pet is soothed. In the pet emotion-soothing solution provided in this application, a pre-trained reference feature vector set is used to complete feature comparison and determine the pet's current emotional state. Based on this, the preset soothing strategy is matched according to the identified pet emotion, and corresponding devices are linked to execute targeted soothing actions. This breaks away from the single audio playback or fixed feeding mode, realizing intelligent and differentiated emotional guidance and care, effectively improving the companionship effect and meeting the pet's companionship needs.
[0054] It should be understood that, although Figure 2 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 2 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0055] To facilitate better implementation of the pet emotional soothing method of this application embodiment, this application embodiment also provides a pet emotional soothing device based on the above-described pet emotional soothing method. The meanings of the terms used are the same as in the above-described pet emotional soothing method, and specific implementation details can be found in the description of the method embodiment.
[0056] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of a pet emotion-soothing device provided in an embodiment of this application. Specifically, the pet emotion-soothing device may include a parameter detection module 201, an extraction module 202, a determination module 203, and an adjustment module 204, as follows: Detection module 201 is used to detect sound data in the target environment in response to a sound acquisition command; Extraction module 202 is used to extract the feature vector corresponding to the pet sound information when the sound data contains the pet sound information of the target pet; The determination module 203 is used to determine the pet's emotions based on the feature vectors and a pre-trained set of reference feature vectors. The soothing module 204 is used to soothe the target pet's emotions based on the pet's mood and preset soothing strategies.
[0057] Optionally, in some embodiments of this application, the determining module 203 may specifically be used for: Obtain the reference feature vector set obtained through pre-training; Calculate the similarity between the feature vector and each reference feature vector in the reference feature vector set; The pet's emotional state is determined based on the calculation results.
[0058] Optionally, in some embodiments of this application, the determining module 203 may specifically be used for: Based on the calculation results, the reference feature vector with the highest similarity to the feature vector is determined as the target feature vector; Obtain the emotion label corresponding to the target feature vector, and determine the emotion label corresponding to the target feature vector as the pet emotion of the target pet.
[0059] Optionally, in some embodiments of this application, the soothing module 204 may specifically be used for: Obtain the target soothing strategy corresponding to the pet's emotions from the preset soothing strategies; Identify the target device to be invoked in the target appeasement strategy; The target device is controlled according to the target soothing strategy in order to soothe the target pet's emotions.
[0060] Optionally, in some embodiments of this application, the soothing module 204 may specifically be used for: If the target soothing strategy is the first soothing strategy, then the audio device to be called will play a preset voice message, and the interactive device to be called will interact with the target pet. If the target soothing strategy is the second soothing strategy, then the feeding device invoked by the control will perform the feeding operation.
[0061] The pet emotion-soothing device provided in this embodiment includes a detection module 201 that responds to a sound acquisition command to detect sound data in the target environment; an extraction module 202 that extracts the feature vector corresponding to the pet's sound information when the sound data contains such information; a determination module 203 that determines the pet's emotion based on the feature vector and a pre-trained reference feature vector set; and finally, a soothing module 204 that soothes the pet based on its emotion and a preset soothing strategy. In this pet emotion-soothing solution, a pre-trained reference feature vector set is used to perform feature comparison and determine the pet's current emotional state. Based on this, the preset soothing strategy is matched to the identified pet emotion, and corresponding devices are linked to perform targeted soothing actions. This breaks away from the single audio playback or fixed feeding mode, achieving intelligent and differentiated emotional guidance and care, effectively improving companionship and meeting the pet's companionship needs.
[0062] Furthermore, embodiments of this application also provide a computing device, such as... Figure 4 As shown, it illustrates a schematic diagram of the computing device involved in the embodiments of this application, specifically: The computing device may include: a processor 402, a communications interface 404, a memory 406, and a communications bus 408.
[0063] The processor 402, communication interface 404, and memory 406 communicate with each other via communication bus 408. Communication interface 404 is used to communicate with other network elements such as clients or other servers. The processor 402 executes program 410, specifically performing the relevant steps in the above-described embodiment of the bandwidth acquisition method for computing devices.
[0064] Specifically, program 410 may include program code that includes computer operation instructions.
[0065] Processor 402 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The computing device includes one or more processors, which may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.
[0066] Memory 406 is used to store program 410. Memory 406 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0067] Specifically, program 410 can be used to cause processor 402 to execute the bandwidth acquisition method in any of the above method embodiments. The specific implementation of each step in program 410 can be found in the corresponding descriptions of the steps and units in the above bandwidth acquisition embodiments, and will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments, and will not be repeated here.
[0068] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The required structure for constructing such systems is apparent from the above description. Furthermore, the embodiments of this application are not directed to any particular programming language. It should be understood that the contents of the embodiments of this application described herein can be implemented using various programming languages, and the above description of specific languages is for the purpose of disclosing the best implementation of the embodiments of this application.
[0069] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.
[0070] Furthermore, those skilled in the art will understand that although some embodiments described herein include certain features but not others included in other embodiments, combinations of features from different embodiments are meant to be within the scope of the embodiments of this application and form different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.
[0071] The various component embodiments of this application can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components according to the embodiments of this application. The embodiments of this application can also be implemented as device or apparatus programs (e.g., computer programs and computer program products) for performing part or all of the methods described herein. Such programs implementing the embodiments of this application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form: Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the pet emotional soothing methods provided in embodiments of this application. For example, the instructions can execute the following steps: In response to a sound acquisition command, sound data in the target environment is detected. When the sound data contains the pet sound information of the target pet, the feature vector corresponding to the pet sound information is extracted. Based on the feature vector and a pre-trained reference feature vector set, the pet emotion of the target pet is determined. Based on the pet emotion and a preset soothing strategy, the target pet is emotionally soothed.
[0072] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0073] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0074] Since the instructions stored in the computer-readable storage medium can execute the steps of any of the pet emotional soothing methods provided in the embodiments of this application, the beneficial effects that any of the pet emotional soothing methods provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0075] The above provides a detailed description of a pet emotional soothing method, apparatus, computing device, and computer-readable storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method of pet emotional pacification, characterized by, include: In response to a sound acquisition command, it detects sound data in the target environment; When the sound data contains the pet sound information of the target pet, the feature vector corresponding to the pet sound information is extracted; The pet's emotions are determined based on the feature vectors and the pre-trained reference feature vector set. Based on the pet's emotions and preset soothing strategies, the target pet is emotionally comforted.
2. The pet emotional soothing method according to claim 1, characterized in that, Determining the pet's emotion based on the feature vector and a pre-trained reference feature vector set includes: Obtain the reference feature vector set obtained through pre-training; Calculate the similarity between the feature vector and each reference feature vector in the reference feature vector set; The pet's emotional state is determined based on the calculation results.
3. The pet emotional soothing method according to claim 2, characterized in that, The step of determining the target pet's emotional state based on the calculation results includes: Based on the calculation results, the reference feature vector with the highest similarity to the feature vector is determined as the target feature vector; Obtain the emotion label corresponding to the target feature vector, and determine the emotion label corresponding to the target feature vector as the pet emotion of the target pet.
4. The pet emotional soothing method according to claim 1, characterized in that, The process of emotionally soothing the target pet based on its emotions and a preset soothing strategy includes: Obtain the target soothing strategy corresponding to the pet's emotion from the preset soothing strategy; Determine the target device invoked by the target appeasement strategy; The target device is controlled according to the target soothing strategy to soothe the target pet's emotions.
5. The pet emotional soothing method according to claim 4, characterized in that, The step of controlling the target device according to the target soothing strategy to soothe the target pet includes: If the target soothing strategy is the first soothing strategy, then the audio device to be invoked will be controlled to play a preset voice message, and the interactive device to be invoked will be controlled to interact with the target pet. If the target soothing strategy is the second soothing strategy, then the feeding device invoked by the control will perform the feeding operation.
6. The pet emotional soothing method according to any one of claims 1 to 5, characterized in that, Also includes: Obtain sample sound data of the target pet; In response to the classification operation on the sample sound data, an emotion label corresponding to the sample sound data is added, and a reference feature vector corresponding to the sample sound data after the label is added is extracted.
7. The pet emotional soothing method according to claim 6, characterized in that, The step of responding to the classification operation on the sample audio data, adding an emotion label corresponding to the sample audio data, and extracting a reference feature vector corresponding to the labeled sample audio data includes: A sound classification interface is displayed, which includes multiple sound tag controls; In response to a tag selection operation on the sound classification interface, the target emotion tag corresponding to the selected sound tag control is determined; Bind the target emotion tag to the sample sound data; The labeled sample sound data is segmented according to a preset time interval to obtain the sound source slice sequence corresponding to the labeled sample sound data. Generate the Mel spectrogram corresponding to the sound source slice sequence, and parse the Mel spectrogram to obtain the reference feature vector corresponding to the sample sound data after the label.
8. A pet emotional comforting device, characterized in that, include: The detection module is used to detect sound data in the target environment in response to sound acquisition commands; The extraction module is used to extract the feature vector corresponding to the pet sound information when the sound data contains the pet sound information of the target pet. The determination module is used to determine the pet emotion of the target pet based on the feature vector and a pre-trained set of reference feature vectors. The soothing module is used to soothe the target pet's emotions based on the pet's emotions and preset soothing strategies.
9. A computing device, characterized in that, include: At least one processor; and A memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the steps of the pet emotional soothing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the pet emotional soothing method as described in any one of claims 1 to 7.