Multi-mode-based dog training method and system for intelligently correcting pet behaviors
Through multimodal data fusion and personalized training feedback mechanism, the problems of weak modal information fusion capability and insufficient personalized response in pet behavior recognition in existing technologies are solved, efficient and safe pet behavior correction is achieved, and training effects and pet compliance are improved.
Patent Information
- Application Number
- CN202510759324.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-06-09
AI Technical Summary
Existing technologies have weak modal information fusion capabilities in pet behavior and emotion recognition, making it difficult to extract unified behavioral features from multi-source data. Behavior classification relies on artificial rules or static models, and has insufficient generalization and real-time performance. It lacks a personalized response mechanism and cannot dynamically optimize training strategies based on pet personality differences.
A multimodal tensor structure is constructed to fuse acoustic signals, motion signals, and trajectory data into a model. Combined with pre-trained end-to-end behavior recognition and emotion perception models, a flexible and continuously optimized behavior intervention strategy is implemented through personalized labels and training feedback mechanisms. A "stimulation-comfort" dual-channel feedback system is used to dynamically adjust the intervention strategy to suit the individual needs of pets.
It significantly improves the accuracy and response timeliness of pet behavior detection, avoids blind stimulation under negative emotions, improves training effects and pet compliance through personalized training strategies, and ensures the effectiveness and safety of long-term intervention.
Smart Images

Figure CN120632629A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pet products, and in particular to a dog training method and system for intelligently correcting pet behavior based on multimodality. Background Art
[0002] As the proportion of pet owners in cities continues to rise, pet behavior management and training have become a major concern for many pet-owning families. Traditional dog training methods, which rely heavily on observation and empirical judgment, lack scientific and systematic principles. These methods can easily lead to pet rebelliousness or behavioral deterioration due to delayed responses and inappropriate training methods.
[0003] In recent years, the development of artificial intelligence and the Internet of Things (IoT) has provided new insights into pet behavior recognition and correction. Some studies have attempted to identify and analyze pet emotions and behaviors using single-modality data, such as image recognition and speech analysis. For example, prior art 1 (CN201910350198.9) discloses a method for training a dog emotion recognition model. This method primarily determines a canine's emotional state by independently extracting and recognizing image and sound data. However, this method fails to effectively integrate the dynamic relationships between different modalities and lacks the ability to continuously model the evolution of behavior and emotions.
[0004] Furthermore, prior art 2 (CN110175526B) provides a knowledge graph-based method for intelligently recognizing canine posture and behavior in surveillance videos. This method relies on a pre-built canine behavior rule graph to infer behavioral categories. While somewhat scalable, this method relies heavily on manual knowledge definition, making it difficult to cover complex or emergent behavior scenarios and lacking end-to-end adaptability.
[0005] In summary, existing technologies for pet behavior and emotion recognition have weak modal information fusion capabilities and difficulty in extracting unified behavioral features from multi-source data. At the same time, behavior classification relies on artificial rules or static models, with insufficient generalization and real-time performance. There is a lack of personalized response mechanisms, and it is impossible to dynamically optimize training strategies based on differences in pet personalities.
[0006] To this end, this paper proposes a dog training method and system for intelligently correcting pet behavior based on multimodal data. This system constructs a unified multimodal tensor structure, integrating and modeling signals such as sound, motion, and trajectory. This system, combined with pre-trained end-to-end behavior recognition and emotion perception models, improves the accuracy of recognizing complex behaviors and underlying emotional states. Furthermore, the system incorporates personalized labeling and training feedback mechanisms to achieve a flexible and continuously optimized behavioral intervention strategy, effectively overcoming the limitations of existing technologies. Summary of the Invention
[0007] The purpose of the present invention is to provide a dog training method and system for intelligently correcting pet behavior based on multimodality, so as to solve the problem in the prior art that the pet behavior cannot be automatically detected and corrected when manual training of the pet dog is not required.
[0008] To achieve the above object, the present invention provides the following technical solutions:
[0009] A multimodal intelligent dog training method for correcting pet behavior comprises the following steps:
[0010] Step S1: Acquire multimodal perception data of the pet, wherein the multimodal perception data includes acoustic signals, motion signals, and trajectory data;
[0011] Step S2: Time alignment, feature extraction, and tensor modeling are performed on the multimodal perception data to construct an acoustic feature tensor, a motion feature tensor, and a position feature tensor, and the three are spliced to form a unified coupling tensor;
[0012] Step S3: input the coupling tensor into a pre-trained behavior recognition model, identify the behavior type, and output a behavior label;
[0013] Step S4: Acquire the pet's physiological characteristics, extract emotion-related parameters from the acoustic signal and motion signal, and input them into the emotion recognition module, so that the emotion recognition module outputs an emotion label;
[0014] Step S5: Statistical modeling is performed based on the pet's historical behavior data, training response data, and emotional response data to form a pet personality label, and the personality label is updated using a sliding window method. The personality label includes behavioral stability, emotional susceptibility, stimulus sensitivity, stimulus response delay, and reward acceptance preference indicators;
[0015] Step S6: inputting the behavior label, emotion label, and personality label into the intervention strategy generation model and outputting an adapted intervention strategy, wherein the intervention strategy includes a soothing strategy and a stimulating strategy;
[0016] Step S7: Adjust the feedback intensity parameters of the intervention strategy based on historical feedback effectiveness records, calculate the feedback convergence index and emotional recovery index, and dynamically modify the feedback level and delay strategy of the intervention strategy;
[0017] Step S8: After the intervention strategy is executed, the pet's emotional and behavioral response status is monitored, and a feedback response score is calculated. If the score does not meet expectations, the feedback type is automatically switched or the intervention intensity is increased.
[0018] Preferably, in step S2, the step of constructing the acoustic feature tensor includes: performing frame division and frequency domain analysis on the original sound, the spectrum analysis unit is used to perform frequency domain transformation processing on each frame, extracting multiple acoustic feature vectors based on Mel-frequency cepstral coefficients, short-time energy, and spectral centroid, and combining the evolution of multiple acoustic features over time frames to form an acoustic feature tensor;
[0019] The steps of constructing a motion feature tensor include: obtaining linear acceleration information of the pet, using an angular velocity measurement unit to obtain angular velocity information of the pet in various directions, aligning and fusing the acceleration information and angular velocity information according to a unified time window, and constructing the fused motion feature sequence into a motion feature tensor;
[0020] The steps of constructing the position feature tensor include: collecting the pet's current coordinates in real time, analyzing the pet's movement trajectory, deviation pattern and spatial boundary state in combination with the timestamp, and constructing the trajectory sequence into a position feature tensor.
[0021] Preferably, in step S3, the training method of the behavior recognition model is: using a model with a kernel nesting mechanism, training a set of disturbance response kernels, inputting a tensor field, mapping it to a high-dimensional embedding space, using a soft geodesic kernel classifier to construct behavior probabilities, using cross entropy loss combined with a tensor regularization term for training, and using it for online inference after training is completed.
[0022] Preferably, in step S4, the step of obtaining the emotion label by the emotion recognition module includes: performing time series difference and frequency domain jitter rate analysis on the acoustic signal to extract frequency fluctuation rate, peak offset amplitude and high-frequency burst ratio; identifying continuous barking segments and calculating the average duration and maximum duration of each barking; extracting the pet's linear acceleration direction change and rhythm fluctuation from the motion signal to calculate the pace rhythm consistency index; analyzing the high-frequency vibration components of the angular velocity within a small range to determine whether there is fear or stress response; combining the emotion parameters extracted from the acoustic signal and the motion signal into a set of emotion feature vectors, and outputting the corresponding emotion label according to the emotion feature vector.
[0023] Preferably, in step S5, the steps of extracting feedback response data based on the pet's long-term behavioral history and extracting its personality characteristics to generate a personality label include: collecting performance records of each pet in different behavioral situations and marking the behavioral characteristics exhibited; constructing a personality characteristic vector based on indicators such as barking intensity, aggressive tendency, feedback response delay, and reward acceptance frequency in the long-term behavioral records; and mapping the personality characteristic vector with the current emotional characteristic vector to generate a personality label.
[0024] Preferably, in step S6, the intervention strategy includes at least: sound prompts, voice praise, vibration stimulation, electrical stimulation, feeding rewards, playing owner's voice, and soothing light; each intervention strategy supporting parameters include: intensity level, duration, delay time window, feedback mode; for each combination of behavior label and emotion label, an initial intervention strategy candidate combination is generated based on preset rules; the intervention strategy generation model is a behavior-emotion joint scoring matrix, which is used to jointly judge the current behavior recognition results and emotion recognition results.
[0025] Preferably, in step S7, the behavior convergence index and the emotion recovery index of the intervention strategy are dynamically calculated according to the behavior convergence speed and the emotion fluctuation trend, so as to modify the feedback level and delay strategy of the intervention strategy.
[0026] Preferably, in step S8, feedback data is recorded after the intervention strategy is executed, and the feedback data includes strategy type, response time, emotion change rate, and behavior improvement extent; a feedback response scoring model is constructed based on the feedback data to evaluate the effectiveness of the intervention strategy; if the score is lower than the preset threshold, the intervention strategy is automatically adjusted.
[0027] Preferably, in step S8, when the intervention strategy is executed after behavior recognition, the current behavior label, emotion label, selected intervention strategy and behavior and emotion changes in a short time after the intervention are recorded; the recorded information is formed into a training sample quadruple and stored in a database; the performance dimensions of the pet under different types of behavior and emotional states are extracted from the database, and the performance dimensions include performance stability, soothing acceptance, and stimulation tolerance; each performance dimension is normalized by statistical quantities to form a local personality vector; the local personality vector is embedded into the behavior recognition model, emotion recognition model and intervention strategy generation model as an auxiliary input.
[0028] A multimodal intelligent dog training system for correcting pet behavior is based on the aforementioned multimodal intelligent dog training method for correcting pet behavior, and includes the following modules:
[0029] The sensing and acquisition module is used to collect pet acoustic signals, motion signals, and position signals;
[0030] Edge recognition module, including behavior recognition model and emotion recognition model;
[0031] The intervention decision module includes a behavior-emotion-feedback mapping unit and an intensity adjustment unit;
[0032] an executive control module for executing intervention strategies on pets;
[0033] The data feedback module is used to record data and continuously optimize behavior recognition, emotion recognition, and intervention strategies based on the recorded data.
[0034] Compared to existing technologies, the present invention offers the following advantages: By utilizing multimodal sensing (including acoustic signals, motion signals, and trajectory data) for tensor modeling, combined with a deep learning-based behavior recognition model, it can identify specific pet behaviors such as excessive barking, running around, and crossing boundaries in real time, significantly improving the accuracy and timeliness of behavior detection. It also comprehensively assesses a pet's emotional state, such as anxiety, fear, and excitement, thereby avoiding blindly applying stimulation to individuals with negative emotions and preventing stress reactions. By statistically modeling a pet's historical behavioral data and feedback response records, it dynamically generates a "personality profile" encompassing multiple dimensions such as emotional susceptibility, feedback latency, and reward preference. This allows each pet to be matched with the most appropriate intervention plan based on a personalized profile. Unlike traditional single-stimulation training methods, the present invention utilizes a dual-channel "stimulation-comfort" feedback system. Based on the combined behavioral and emotional profile and a feedback strategy matrix, it automatically selects methods such as sound composure, feeding rewards, vibration, and electrical stimulation, prioritizing non-invasive training and improving training effectiveness and pet compliance. By recording the behavioral improvements and emotional recovery trends after each feedback, calculating the feedback response score, and dynamically adjusting the intervention strategy, a closed-loop control logic of "behavior-feedback-response-readjustment" is formed to ensure the effectiveness and safety of long-term intervention. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is a flow chart of a dog training method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0036] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0037] See also Figure 1 The present invention provides a dog training method for correcting pet behavior based on multimodal intelligence, comprising the following steps:
[0038] Step S1: Acquire multimodal perception data of the pet, wherein the multimodal perception data includes acoustic signals, motion signals, and trajectory data;
[0039] Step S2: Time alignment, feature extraction, and tensor modeling are performed on the multimodal perception data to construct an acoustic feature tensor, a motion feature tensor, and a position feature tensor, and the three are spliced to form a unified coupling tensor;
[0040] Step S3: input the coupling tensor into a pre-trained behavior recognition model, identify the behavior type, and output a behavior label;
[0041] Step S4: Acquire the pet's physiological characteristics, extract emotion-related parameters from the acoustic signal and motion signal, and input them into the emotion recognition module, so that the emotion recognition module outputs an emotion label;
[0042] Step S5: Statistical modeling is performed based on the pet's historical behavior data, training response data, and emotional response data to form a pet personality label, and the personality label is updated using a sliding window method. The personality label includes behavioral stability, emotional susceptibility, stimulus sensitivity, stimulus response delay, and reward acceptance preference indicators;
[0043] Step S6: inputting the behavior label, emotion label, and personality label into the intervention strategy generation model and outputting an adapted intervention strategy, wherein the intervention strategy includes a soothing strategy and a stimulating strategy;
[0044] Step S7: Adjust the feedback intensity parameters of the intervention strategy based on historical feedback effectiveness records, calculate the feedback convergence index and emotional recovery index, and dynamically modify the feedback level and delay strategy of the intervention strategy;
[0045] Step S8: After the intervention strategy is executed, the pet's emotional and behavioral response status is monitored, and a feedback response score is calculated. If the score does not meet expectations, the feedback type is automatically switched or the intervention intensity is increased.
[0046] Preferably, in step S2, the step of constructing an acoustic feature tensor includes: performing frame division and frequency domain analysis on the original sound, a spectrum analysis unit is used to perform frequency domain transformation processing on each frame, extracting multiple acoustic feature vectors based on Mel-frequency cepstral coefficients, short-time energy, and spectral centroid, and combining the evolution of multiple acoustic features over time frames to form an acoustic feature tensor; expressed as: ,in is the acoustic feature dimension, is the number of time frames;
[0047] The steps of constructing the motion feature tensor include: obtaining the pet's linear acceleration information, using the angular velocity measurement unit to obtain the pet's angular velocity information in various directions, aligning and fusing the acceleration information and angular velocity information according to a unified time window, and constructing the fused motion feature sequence into a motion feature tensor; expressed as: ,in is the motion feature dimension, is the time step length.
[0048] The steps of constructing the position feature tensor include: collecting the pet's current coordinates in real time, analyzing the pet's movement trajectory, deviation pattern, and spatial boundary status in combination with the timestamp, and constructing the trajectory sequence into a position feature tensor. The position tensor generation unit constructs the trajectory sequence into a position feature tensor by analyzing the pet's movement trajectory, deviation pattern, and spatial boundary status in combination with the timestamp, and expresses it as: ,in is the position feature dimension, is the length of the sampling period.
[0049] Concatenate the three tensors into a coupled tensor field:
[0050]
[0051] It is used to comprehensively reflect the sound behavior characteristics, movement behavior characteristics and position behavior characteristics within the current time window, and serve as the input of the behavior recognition model.
[0052] After the three feature tensors are spliced and constructed into a coupled tensor field, the second processor ensures that each modal feature has an effective temporal correspondence and behavioral correlation in the recognition model by maintaining the feature channel dimension and time synchronization processing, thereby avoiding recognition failure caused by information misalignment between modalities. After the coupled tensor field is input into the pre-trained behavior recognition model, the global spatiotemporal feature representation is first extracted through the encoding network, and then the classification network is used to distinguish the behavioral state. The recognition model adopts a multi-layer nonlinear mapping structure to learn the implicit coupling relationship between cross-modal features and has the ability to quantitatively evaluate the intensity of behavioral disturbances. The output of the recognition model includes two result branches: one is the behavior type prediction result, which indicates whether there is a predefined bad behavior type; the other is the disturbance amplitude score result, which reflects the degree of deviation of the current behavior from the normal state. The second processor selects the corresponding feedback unit in the activation warning module based on the relationship between the recognition output result and the set multi-level response threshold to form a hierarchical corrective feedback mechanism. The multi-level feedback mechanism includes at least three output states: when the disturbance amplitude score is lower than the first threshold, only the sound prompt is triggered; when the score is between the first and second thresholds, the vibration prompt is superimposed and triggered; when the score is higher than the second threshold, the electric shock stimulation module is triggered, realizing an upgraded behavioral intervention strategy from mild to forced.
[0053] In this embodiment, the tensor perturbation response function family The response of the trimodal feature tensor to external stimuli under different behavior modes is simulated as follows:
[0054]
[0055] in, is the perturbation kernel function, representing the typical behavior pattern The response characteristics can be obtained by offline training based on empirical behavioral data, and its structure can be a Gaussian function family, Laplace kernel, radial basis function (RBF) or tensor attention kernel; is the spatial support of the tensor field, which represents the global weighted modeling of the information of all spatiotemporal positions in the current observation window; the convolution symbol It represents that kernel convolution or feature correlation operation is performed between the kernel function and the specific modal channel or combination channel of the coupled tensor field, which can realize the extraction of disturbance-sensitive features in the local area; Represents a predicted response to a behavior.
[0056] The family of tensor disturbance response functions Based on multi-modal coupled tensor field Based on the disturbance distribution characteristics of The pet's combined response to external stimuli, including sound, movement, and position, within a specific time window is captured. This family of functions is constructed using a kernel response mapping approach, essentially a kernel convolution-type functional mapping mechanism that captures the nonlinear interference characteristics and spatiotemporal evolution of different modes in a tensor field.
[0057] In this embodiment, the second processor couples the tensor field in a non-Euclidean space structure. To classify behaviors, the domain is defined as a Riemann manifold The classification function is established on it, which is expressed as:
[0058]
[0059] This classification function no longer uses traditional linear distance metrics for discrimination, but instead performs geodesic distance evaluation in the manifold structure to achieve higher fitting ability and discrimination accuracy for complex behavioral tensor distribution characteristics.
[0060] Expressed as:
[0061]
[0062] in, Indicates the The tensor core center of a typical behavior sample can be obtained through prior sample clustering or learning mechanism to represent the typical characteristic state of a specific behavior category (such as "quiet", "violent running", "crossing the boundary"); is a tensor and behavioral center samples In the manifold The geodesic distance on , which can be calculated by the shortest path length under the Riemannian metric, avoids the projection distortion problem caused by the Euclidean distance in high-dimensional curved space; is the weight coefficient corresponding to each sample center, indicating its representativeness and confidence in the recognition judgment, which can be determined by the training data sample frequency or supervised learning strategy; is the Sigmoid function, which is used to map the weighted exponential distance to the probability of behavior occurrence between 0 and 1.
[0063] Preferably, in step S3, the training method of the behavior recognition model is: using a model with a kernel nesting mechanism, training a set of disturbance response kernels, inputting a tensor field, mapping it to a high-dimensional embedding space, using a soft geodesic kernel classifier to construct behavior probabilities, using cross entropy loss combined with a tensor regularization term for training, and using it for online inference after training is completed.
[0064] In this example, a model with a kernel nesting mechanism is used. , train a set of perturbation response kernels { }, expressed as:
[0065]
[0066] Different behavior modes are assigned different perturbation kernel response models.
[0067] In this embodiment, the input tensor field , mapped to a high-dimensional embedding space, expressed as:
[0068]
[0069] The behavior probability is constructed using the soft geodesic kernel classifier and is expressed as:
[0070] .
[0071] In this embodiment, cross entropy loss is used in combination with a tensor regularization term for training, which is expressed as:
[0072] Among them: the first term is the behavior classification loss; the second term is the perturbation energy regularization term, which is used to penalize non-stationary tensor perturbations; is the Frobenius norm.
[0073] In step S4, the emotion recognition module obtains the emotion label by performing time series difference and frequency domain jitter rate analysis on the acoustic signal to extract frequency fluctuation rate, peak offset amplitude and high-frequency burst ratio; identifying continuous barking segments and calculating the average duration and maximum duration of each bark; extracting the pet's linear acceleration direction change and rhythm fluctuation from the motion signal to calculate the pace rhythm consistency index; analyzing the high-frequency vibration components of the angular velocity within a small range to determine whether there is a fear or stress response; combining the emotion parameters extracted from the acoustic signal and the motion signal into a set of emotion feature vectors, and outputting the corresponding emotion label according to the emotion feature vectors.
[0074] Features such as dominant frequency, frequency fluctuation amplitude, barking duration, interval density, and high-frequency spectral energy ratio are extracted from acoustic signals. Acceleration fluctuation frequency, gait balance (based on left-right acceleration symmetry), tail swing amplitude, and tremor rhythm are extracted from motion signals. Inter-loiter path, acceleration rate, and duration of dwell in a region are extracted from position signals. All perceptual features are uniformly mapped into a standard emotion feature space, combining them to form a temporal emotion feature sequence. A standard emotion classification system is established, encompassing at least five categories: excitement, anxiety, fear, calmness, and aggression. A labeled training sample library is constructed using supervised learning. Different emotion response mapping templates are defined for different dog breeds and personalities to enhance cross-individual transferability. A Transformer temporal modeling neural network architecture is used to model the temporal evolution of features. The input is a multimodal emotion feature sequence, and the output is the probability distribution of the emotion at the current moment. Training is performed using a soft-label cross-entropy loss function combined with a temporal smoothing regularization term. If the current emotion recognition output indicates a negative emotion (e.g., anxiety / fear), the electrical stimulation intervention signal is automatically suppressed and replaced with soothing speech or shock-absorbing feedback.
[0075] In step S5, the steps of extracting feedback response data based on the pet's long-term behavioral history and extracting its personality traits to generate a personality label include: collecting performance records for each pet in different behavioral situations and labeling the behavioral traits exhibited; constructing a personality trait vector based on indicators such as barking intensity, aggressive tendencies, feedback response delay, and reward acceptance frequency in the long-term behavioral records; and mapping the personality trait vector with the current emotion trait vector to generate a personality label. Each pet's performance records in different behavioral situations are collected and labeled for whether it exhibits behavioral traits such as fear, dependence, resistance, or playfulness; constructing a personality vector based on indicators such as barking intensity, aggressive tendencies, feedback response delay, and reward acceptance frequency in the long-term behavioral records; and mapping the personality vector with the current emotion vector to generate a personality label (including, for example, preference for sound soothing over vibration feedback or sensitivity to electrical stimulation).
[0076] In step S6, intervention strategies include at least: sound prompts, voice praise, vibration stimulation, electrical stimulation, feeding rewards, owner voice playback, and soothing light. Each intervention strategy's supporting parameters include intensity level, duration, delay window, and feedback modality. For each behavior label and emotion label combination, initial intervention strategy candidate combinations are generated based on pre-set rules. The intervention strategy generation model is a behavior-emotion joint scoring matrix, which is used to jointly judge the current behavior recognition results and emotion recognition results. A behavior-emotion-feedback comparison matrix is established based on the behavior classification results (excessive barking, aimless running, chasing behavior, etc.) and the emotion labels (anxiety, excitement, fear, indifference, etc.). In this comparison matrix, different combinations are mapped to corresponding feedback types: ("excitement + out-of-bounds" is mapped to low-intensity vibration, and "anxiety + non-aggressive behavior" is mapped to voice soothing or delayed treatment). The treatment principle is non-invasive priority, that is, when the intervention effectiveness is equal, sound prompts and vibration prompts are used first, and electrical stimulation has a lower priority than sound prompts and vibration prompts.
[0077] In step S7, the intervention strategy's behavioral convergence index and emotional recovery index are dynamically calculated based on the behavioral convergence speed and emotional fluctuation trends, thereby modifying the intervention strategy's feedback level and delay strategy. The stimulation intensity is dynamically adjusted based on the behavioral convergence speed and emotional fluctuation trends. The behavioral convergence index and emotional disturbance recovery index for each feedback type are recorded to set the feedback intensity correction factor. When the pet shows rapid adaptation to a certain low-intensity stimulus, the stimulation level of that type is automatically reduced; otherwise, the intensity is moderately increased until it is effective. If the emotional state is identified as "excessively frightened" or "continuously depressed," the system selects the sound soothing mode and plays the owner's voice clips or soothing background sounds from a preset audio database. The system connects to other smart devices in the home (automatic feeders, smart toys) to collaboratively trigger physical soothing feedback such as stroking and releasing snacks.
[0078] In step S8, feedback data is recorded after the intervention strategy is executed. The feedback data includes strategy type, response time, emotion change rate, and behavior improvement degree. A feedback response scoring model is constructed based on the feedback data to evaluate the effectiveness of the intervention strategy. If the score is lower than the preset threshold, the intervention strategy is automatically adjusted.
[0079] In step S8, when implementing an intervention strategy after behavior recognition, the current behavior label, emotion label, selected intervention strategy, and behavioral and emotional changes shortly after the intervention are recorded. This recorded information is then combined into a training sample quadruple and stored in a database. From the database, the pet's performance dimensions for different behavioral and emotional states are extracted, including performance stability, soothing acceptance, and stimulus tolerance. Each performance dimension is statistically normalized to form a local personality vector. This local personality vector serves as an auxiliary input and is embedded in the behavior recognition model, emotion recognition model, and intervention strategy generation model. Long-term behavioral observation data (such as behavior frequency, duration, and emotional response curve) is used to extract pet features. The extracted metrics include, but are not limited to, the standard deviation of barking frequency (reflecting sensitivity to emotional fluctuations), behavioral response latency (measuring training speed), and behavioral avoidance rate after specific stimuli (assessing stimulus sensitivity). These features are combined into a pet personality feature vector and stored in the individual profile.
[0080] In specific implementation, for example, when dealing with excessive barking and anxious behavior caused by the ringing of the doorbell, if the owner has installed the system of the present invention at home and has established Doudou's personality label through preliminary data collection (the personality trait vector shows that he is emotionally sensitive, has poor stimulation tolerance, and is highly dependent on the owner's voice).
[0081] When the doorbell rang, Doudou suddenly barked and scurried across the living room. The system's acoustic sensors recorded its continuous, high-intensity barking, the inertial measurement unit recorded its dramatic acceleration changes, and the positioning module identified its high-frequency switching between the door and the living room.
[0082] The system converts acoustic data into Mel-Frequency Cepstral Coefficient (MFCC) tensors to extract barking frequency burst indicators and intensity fluctuations. Motion signals are converted into acceleration change sequences and angular velocity data to extract gait rhythm and high-frequency tremor signals. Position data is used to generate trajectory offset tensors. These three types of tensors are time-aligned and concatenated into a coupled tensor.
[0083] The system uses the pre-trained behavior recognition model to identify it as "excessive barking + aimless running" and outputs the behavior label as "behavior type B3".
[0084] Emotional characteristic parameters were extracted (the average duration of barking was longer than normal, the rhythm fluctuated greatly, the tail swing amplitude was small but the tremor value was high), and it was determined to be "anxiety state E2".
[0085] Calling Doudou's personality label shows that he is slow to respond to electrical stimulation but recovers slowly from emotional fluctuations, and prefers to be comforted by his owner's voice.
[0086] Based on behavior "B3" + emotion "E2" + personality label, the system selects a feedback plan in the behavior-emotion-feedback matrix: "play owner's voice clip + light vibration prompt + delayed snack release", forming a soothing multimodal intervention combination.
[0087] Considering that Doudou has shown avoidance behavior towards high-intensity vibrations, the system selects the vibration level as L1 (the lowest). After playing the owner's voice soothing message for 10 seconds, if the barking stops, a snack will be released through the feeder within 5 seconds.
[0088] After the intervention was implemented, the system monitored that the sound signal dropped by 90%, the acceleration weakened, and the spatial residence time increased. The emotion recognition module determined that the emotion returned to "calm E5"; the system calculated the feedback response score as 92 points (out of 100), recorded it as a positive response case and wrote it into the "feedback-response-personality" history model.
[0089] This embodiment effectively suppresses excessive barking of pets due to external stimuli by integrating behavior recognition, emotion perception and personality adaptation, uses non-stimulating means to reduce emotional fluctuations, and improves the sustainability of training and the compliance of pets.
[0090] The present invention also provides an embodiment of a multimodal intelligent dog training system for correcting pet behavior, based on the aforementioned multimodal intelligent dog training method for correcting pet behavior, comprising the following modules:
[0091] The sensing and acquisition module is used to collect pet acoustic signals, motion signals, and position signals;
[0092] Edge recognition module, including behavior recognition model and emotion recognition model;
[0093] The intervention decision module includes a behavior-emotion-feedback mapping unit and an intensity adjustment unit;
[0094] an executive control module for executing intervention strategies on pets;
[0095] The data feedback module is used to record data and continuously optimize behavior recognition, emotion recognition, and intervention strategies based on the recorded data.
[0096] In this embodiment, the training process of the behavior recognition model includes: collecting sound modal data, motion modal data and position modal data and unifying them into a time frame, constructing a three-modal tensor, using a model with a kernel nesting mechanism, training a set of disturbance response kernels, inputting a tensor field, mapping it to a high-dimensional embedding space, using a soft geodesic kernel classifier to construct behavior probabilities, using cross-entropy loss combined with a tensor regularization term for training, and using it for online inference after training is completed.
[0097] Tensor modeling of multimodal perception (including acoustic signals, motion signals, and trajectory data) combined with a deep learning-based behavior recognition model enables real-time identification of specific pet behaviors such as excessive barking, running around, and crossing boundaries, significantly improving the accuracy and timeliness of behavior detection. A comprehensive assessment of a pet's emotional state, such as anxiety, fear, and excitement, helps prevent blind stimulation during negative emotions and prevent stress reactions. By statistically modeling a pet's historical behavioral data and feedback response records, a "personality signature" is dynamically generated, encompassing multiple dimensions such as emotional susceptibility, feedback latency, and reward preference. This allows each pet to be matched with the most appropriate intervention plan based on a personalized profile. Unlike traditional single-stimulation training methods, this invention utilizes a dual-channel "stimulation-comfort" feedback system. Based on the combined behavioral and emotional signature and a feedback strategy matrix, it automatically selects methods such as sound composure, feeding rewards, vibration, and electrical stimulation, prioritizing non-invasive training and improving training effectiveness and pet compliance. By recording the behavioral improvements and emotional recovery trends after each feedback, calculating the feedback response score, and dynamically adjusting the intervention strategy, a closed-loop control logic of "behavior-feedback-response-readjustment" is formed to ensure the effectiveness and safety of long-term intervention.
[0098] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A dog training method for correcting pet behavior based on multimodal intelligence, characterized in that: The following steps are involved: Step S1: Acquire multimodal perception data of the pet, wherein the multimodal perception data includes acoustic signals, motion signals, and trajectory data; Step S2: Time alignment, feature extraction, and tensor modeling are performed on the multimodal perception data to construct an acoustic feature tensor, a motion feature tensor, and a position feature tensor, and the three are spliced to form a unified coupling tensor; Step S3: input the coupling tensor into a pre-trained behavior recognition model, identify the behavior type, and output a behavior label; Step S4: Acquire physiological characteristic parameters of the pet, extract emotion-related parameters from acoustic signals and motion signals, and input them into the emotion recognition module, so that the emotion recognition module outputs an emotion label; Step S5: Statistical modeling is performed based on the pet's historical behavior data, training response data, and emotional response data to form a pet personality label, and the personality label is updated using a sliding window method; Step S6: inputting the behavior label, emotion label, and personality label into the intervention strategy generation model and outputting an adapted intervention strategy, wherein the intervention strategy includes a soothing strategy and a stimulating strategy; Step S7: Adjust the feedback intensity parameters of the intervention strategy based on historical feedback effectiveness records, calculate the feedback convergence index and emotional recovery index, and dynamically modify the feedback level and delay strategy of the intervention strategy; Step S8: After the intervention strategy is executed, the pet's emotional and behavioral response status is monitored, and a feedback response score is calculated. If the score does not meet expectations, the feedback type is automatically switched or the intervention intensity is increased.
2. The multimodal intelligent dog training method for correcting pet behavior according to claim 1, characterized in that: In step S2, the step of constructing an acoustic feature tensor includes: performing frame division and frequency domain analysis on the original sound, a spectrum analysis unit is used to perform frequency domain transformation processing on each frame, extracting multiple acoustic feature vectors based on Mel-frequency cepstral coefficients, short-time energy, and spectral centroid, and combining the multiple acoustic features as they evolve over time frames to form an acoustic feature tensor; The steps of constructing a motion feature tensor include: obtaining linear acceleration information of the pet, using an angular velocity measurement unit to obtain angular velocity information of the pet in various directions, aligning and fusing the acceleration information and angular velocity information according to a unified time window, and constructing the fused motion feature sequence into a motion feature tensor; The steps of constructing the position feature tensor include: collecting the pet's current coordinates in real time, analyzing the pet's movement trajectory, deviation pattern and spatial boundary state in combination with the timestamp, and constructing the trajectory sequence into a position feature tensor.
3. The multimodal intelligent dog training method for correcting pet behavior according to claim 2, characterized in that: In step S3, the training method of the behavior recognition model is as follows: using a model with a kernel nesting mechanism, training a set of disturbance response kernels, inputting a tensor field, mapping it to a high-dimensional embedding space, using a soft geodesic kernel classifier to construct behavior probabilities, using cross entropy loss combined with a tensor regularization term for training, and using it for online inference after training.
4. The multimodal intelligent dog training method for correcting pet behavior according to claim 2, characterized in that: In step S4, the emotion recognition module obtains the emotion label by performing time series difference and frequency domain jitter rate analysis on the acoustic signal to extract frequency fluctuation rate, peak offset amplitude and high-frequency burst ratio; identifying continuous barking segments and calculating the average duration and maximum duration of each bark; extracting the pet's linear acceleration direction change and rhythm fluctuation from the motion signal to calculate the pace rhythm consistency index; analyzing the high-frequency vibration components of the angular velocity within a small range to determine whether there is a fear or stress response; combining the emotion parameters extracted from the acoustic signal and the motion signal into a set of emotion feature vectors, and outputting the corresponding emotion label according to the emotion feature vectors.
5. The multimodal intelligent dog training method for correcting pet behavior according to claim 2, characterized in that: In step S5, the steps of extracting feedback response data based on the pet's long-term behavioral history and extracting its personality characteristics to generate a personality label include: collecting performance records of each pet in different behavioral situations and marking the behavioral characteristics it exhibits; constructing a personality characteristic vector based on indicators such as barking intensity, aggressive tendency, feedback response delay, and reward acceptance frequency in the long-term behavioral records; and mapping the personality characteristic vector with the current emotional characteristic vector to generate a personality label.
6. The multimodal intelligent dog training method for correcting pet behavior according to claim 1, characterized in that: In step S6, the intervention strategies include at least: sound prompts, voice praise, vibration stimulation, electrical stimulation, feeding rewards, playing owner's voice, and soothing light; the supporting parameters of each intervention strategy include: intensity level, duration, delay time window, and feedback mode; for each combination of behavior label and emotion label, an initial intervention strategy candidate combination is generated based on preset rules; the intervention strategy generation model is a behavior-emotion joint scoring matrix, which is used to jointly judge the current behavior recognition results and emotion recognition results.
7. The multimodal intelligent dog training method for correcting pet behavior according to claim 6, characterized in that: In step S7, the behavior convergence index and the emotion recovery index of the intervention strategy are dynamically calculated according to the behavior convergence speed and the emotion fluctuation trend, so as to modify the feedback level and delay strategy of the intervention strategy.
8. The multimodal intelligent dog training method for correcting pet behavior according to claim 3, characterized in that: In step S8, feedback data is recorded after the intervention strategy is executed. The feedback data includes strategy type, response time, emotion change rate, and behavior improvement degree. A feedback response scoring model is constructed based on the feedback data to evaluate the effectiveness of the intervention strategy. If the score is lower than the preset threshold, the intervention strategy is automatically adjusted.
9. The multimodal intelligent dog training method for correcting pet behavior according to claim 3, characterized in that: In step S8, when the intervention strategy is executed after behavior recognition, the current behavior label, emotion label, selected intervention strategy, and behavior and emotion changes in the short period after the intervention are recorded; the recorded information is formed into a training sample quadruple and stored in a database; the performance dimensions of the pet under different types of behavior and emotional states are extracted from the database, and the performance dimensions include performance stability, soothing acceptance, and stimulation tolerance; each performance dimension is normalized by statistical quantities to form a local personality vector; the local personality vector is embedded as an auxiliary input into the behavior recognition model, emotion recognition model, and intervention strategy generation model.
10. A multimodal intelligent dog training system for correcting pet behavior, based on the multimodal intelligent dog training method for correcting pet behavior according to any one of claims 1 to 9, characterized in that: Includes the following modules: The sensing and acquisition module is used to collect pet acoustic signals, motion signals, and position signals; Edge recognition module, including behavior recognition model and emotion recognition model; The intervention decision module includes a behavior-emotion-feedback mapping unit and an intensity adjustment unit; an executive control module for executing intervention strategies on pets; The data feedback module is used to record data and continuously optimize behavior recognition, emotion recognition, and intervention strategies based on the recorded data.
Citation Information
Patent Citations
Dog emotion recognition model training method and device, computer equipment and storage medium
CN110175526A
Dog emotion recognition model training method, device, computer equipment and storage medium
CN110175526B
AI-based pet emotion recognition system
CN119049086A
Pet pacifying method
CN119646600A
Cited By
Pet behavior training system based on AI vision
CN120982435A
Pet behavior training system based on AI vision
CN120982435B
Intelligent human-pet interaction method and system, storage medium and program product
CN121242593A
Pet behavior recognition and emotion detection method
CN121415134A
Group behavior regulation and control method and system based on multi-dog cooperative training
CN121605938A