A pet state intelligent translation method and system based on multi-modal data

By analyzing multimodal data, we can generate pet stress level predictions and avoidance paths, which solves the problem of insufficient accuracy in pet state translation in existing technologies. This enables personalized environmental responses and emotional support, and improves the safety and comfort of pets during outdoor activities.

CN120612649BActive Publication Date: 2025-11-07BEIJING CHONGYOUDAO TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510706731.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-11-07
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Existing technologies have limitations in identifying pet stress responses, including limited coverage, misjudgments and omissions, and neglect of intrinsic physiological indicators, resulting in insufficient accuracy and reliability of intelligent translation of pet status.

Method used

By acquiring multimodal data, including noise levels, crowd density, and 3D spatial data, an obstacle distribution map is generated and jointly encoded with the multimodal data. Combined with pet limb movement time-series data, abnormal behavioral fragments are identified, stress levels are predicted, and translated instructions containing avoidance paths are generated.

Benefits of technology

It improves the accuracy and reliability of intelligent translation of pet status, provides personalized natural language descriptions and emotional comfort suggestions, and enhances the safety and comfort of pets in outdoor environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612649B_ABST
    Figure CN120612649B_ABST
Patent Text Reader

Abstract

The application provides a pet state intelligent translation method and system based on multi-modal data. Wherein, by collecting multi-modal data of the outdoor environment where the target pet is located, an obstacle distribution atlas is generated to quantify the environmental obstacle density and spatial occlusion relationship, and a feature vector reflecting environmental stress is generated by joint coding combined with multi-modal data. The limb action time sequence data of the pet is synchronously acquired, and mode matching is performed with the preset pet body language database to extract behavior abnormality segments associated with the spatio-temporal changes of the outdoor environment. By analyzing the mapping relationship between the environmental stress feature vector and the behavior abnormality segments, the stress level of the pet is predicted, and finally a translation instruction containing a high-density obstacle area avoidance path prompt is generated, realizing the intelligent translation of the linkage between the pet behavior state and the environmental risk. The technical scheme provided by the application can improve the efficiency and accuracy of pet state intelligent translation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of pet state intelligent translation, and in particular to a pet state intelligent translation method and system based on multi-modal data. BACKGROUND

[0002] In outdoor complex environments, the stress response of pets is crucial for their health and safety. With the acceleration of urbanization, pets have more opportunities to participate in outdoor activities with their owners. Changes in the environment, such as noise, unfamiliar creatures or crowds, can all trigger stress responses. Accurate identification and timely response to these states can help prevent potential health problems and improve communication understanding. To achieve this goal, analyzing the pet's posture changes, sound frequencies, and environmental factors can provide in-depth insights into the pet's immediate emotional state, helping owners or trainers to take timely measures to prevent potential risks.

[0003] Currently, a targeted solution to this technical need is to use an intelligent monitoring system installed in outdoor environments to assess pet states. This system integrates high-definition cameras, audio collection devices, and environmental sensors to capture pet visual behavior and auditory signals comprehensively, and combines surrounding environmental parameters for comprehensive analysis. By identifying features such as pet behavior trajectories, posture changes, and call frequencies, timely measures can be taken. This method provides a new approach to understanding and managing pet stress responses in complex outdoor environments.

[0004] Although the above solution meets the demand for intelligent translation of pet states to some extent, it still has some significant shortcomings. First, due to the reliance on fixed-position monitoring devices, its coverage is limited, and for pets with a wide range of activities, there may be monitoring blind spots. Second, it has limitations in dealing with complex background noise and differences between different pet individuals, which may lead to misjudgment or missed cases. Third, this solution mainly focuses on the analysis of external behavior and sound, ignoring changes in internal physiological indicators such as heart rate and body temperature, which are also crucial for a comprehensive understanding of the pet's true state. Therefore, how to overcome these shortcomings and improve the accuracy and reliability of pet state intelligent translation remains an important direction for future research. SUMMARY

[0005] The present application provides a pet state intelligent translation method and system based on multi-modal data to solve the problem of low efficiency and poor accuracy of pet state intelligent translation in the prior art.

[0006] In a first aspect, the present application provides a pet state intelligent translation method based on multi-modal data, comprising:

[0007] Obtain multi-modal data of an outdoor complex environment where a target pet is located, the multi-modal data including noise decibel, crowd density, and three-dimensional space data captured by a laser radar;

[0008] Generate an obstacle distribution map based on the three-dimensional space data, and jointly encode the multi-modal data and the obstacle distribution map to generate an environmental stress feature vector;

[0009] Synchronously collect limb action time series data of the target pet, and perform pattern matching on the limb action time series data and a preset pet body language database to extract a behavior abnormality segment that has a spatio-temporal correlation with the outdoor complex environment;

[0010] Determine a stress level prediction result of the target pet in the outdoor complex environment through association mapping of the environmental stress feature vector and the behavior abnormality segment;

[0011] Generate a translation instruction according to the stress level prediction result, the translation instruction including an evasive path prompt for a high-density area in the obstacle distribution map.

[0012] Optionally, after the generating of the translation instruction according to the stress level prediction result, the method further includes:

[0013] Construct a base large model based on multi-modal data, the base large model including cross-modal analysis results of human voice, environmental sound, and animal call;

[0014] Jointly label the cross-modal analysis results of the animal call according to the environmental stress feature vector and the behavior abnormality segment to obtain a labeling result, the labeling result including a mapping relationship among dog breed category, emotion type, and environmental stress;

[0015] Extract a mel-frequency spectrum feature of the labeling result by using the base large model to obtain a spectrum vector, perform feature alignment on the spectrum vector and the environmental stress feature vector, adjust parameters of the base large model based on an alignment result to obtain an adjusted base large model;

[0016] Construct a dog language reasoning logic based on the adjusted base large model, and adjust a language generation strategy of the translation instruction by using the dog language reasoning logic to generate a natural language description that matches personality characteristics of the target pet;

[0017] Identify an emotion state of the target pet based on the natural language description, the limb action time series data, and the noise decibel, and add an emotional pacification suggestion corresponding to the emotion state to the translation instruction;

[0018] Push a translation instruction with the emotional pacification suggestion to the user terminal, and dynamically update an avoidance path prompt in the translation instruction according to changes in the high-density area of the obstacle distribution map.

[0019] Optionally, the obstacle distribution map is constructed by quantifying the mobile obstacle density and the space occlusion coefficient in the target pet activity range.

[0020] The generating of the obstacle distribution map based on the three-dimensional space data comprises:

[0021] Based on the three-dimensional space data, a plurality of space grid units in the target pet activity range are divided, the moving track of the obstacle in each space grid unit and its appearance frequency in a preset time period are extracted, the change in the line-of-sight occlusion degree between adjacent space grid units is combined, and the mobile obstacle density and the space occlusion coefficient of each space grid unit are calculated.

[0022] The mobile obstacle density and the space occlusion coefficient of each space grid unit are mapped into a three-dimensional heat map layer, and the obstacle distribution map is generated by superposition.

[0023] Optionally, the joint encoding of the multi-modal data and the obstacle distribution map to generate an environmental stress feature vector comprises:

[0024] According to the change trend of the noise decibel with time and the distribution change of the crowd density in space, a parameter set in the multi-modal data that overlaps with the high-density area in the obstacle distribution map is extracted, parameters in the parameter set whose fluctuation amplitude exceeds a preset threshold are time-series accumulated, and a parameter sequence reflecting an environmental sudden interference event is generated.

[0025] The mobile obstacle density, the space occlusion coefficient, and the parameter sequence are associated and matched to determine the diffusion path of the environmental sudden interference event in the obstacle distribution map, and the mobile obstacle density and the space occlusion coefficient of each space grid unit are adjusted according to the diffusion path.

[0026] Based on the time-series accumulation results of the adjusted mobile obstacle density, the space occlusion coefficient, and the parameter sequence, the time-series fluctuation features of the multi-modal data and the spatial topological features of the obstacle distribution map are tensor-spliced, and the splicing results are encoded by a convolutional neural network to generate an environmental stress feature vector.

[0027] Optionally, the limb action time-series data comprises a continuous change sequence of limb swing frequency, trunk contraction amplitude, and head deflection angle.

[0028] The limb movement time sequence data is matched with a preset pet limb language database in a mode, and behavior abnormality segments with space-time correlation in the outdoor complex environment are extracted, including:

[0029] The limb movement time sequence data is divided into analysis windows with variable lengths, and a sliding interval of the analysis window is dynamically adjusted according to a spatial density change of an obstacle distribution map in the outdoor complex environment;

[0030] In each analysis window, a matching degree of a limb movement parameter in the current analysis window to a standard behavior mode in the pet limb language database is calculated based on a movement change feature of adjacent time points in the limb movement time sequence data, a dynamic matching condition is generated in combination with a parameter related to visual interference in an environmental stress feature vector;

[0031] Analysis windows with a matching degree lower than a preset reference are screened according to the dynamic matching condition, a periodic difference between a limb swing frequency and a trunk contraction amplitude in the analysis window is extracted, and a behavior abnormality segment with position correlation in the outdoor complex environment is marked in combination with position information of a high-density region in the obstacle distribution map;

[0032] The behavior abnormality segment is compared and matched with a spatial distribution change of a crowd density in the multi-modal data, non-stress action data generated by a pet autonomous behavior is excluded, and a set of the space-time correlated behavior abnormality segments is generated.

[0033] Optionally, the emotion state of the target pet is identified based on the natural language description, in combination with the limb movement time sequence data and the noise decibel, and a sentiment pacification suggestion corresponding to the emotion state is attached to the translation instruction, including:

[0034] The natural language description is processed by word segmentation, context correlation features are extracted, and a sentiment feature vector is generated, the sentiment feature vector including a word segmentation weight distribution and a semantic dependency strength;

[0035] A two-dimensional feature space with limb movement amplitude and movement frequency as axes is established, a dynamic trajectory point set is constructed based on the sentiment feature vector and the limb movement time sequence data of the target pet in a set time window, and an emotion intensity level of the emotion state is divided by a distribution density of the dynamic trajectory point set in the two-dimensional feature space;

[0036] A voiceprint pulse interval with a duration longer than a threshold value in the noise decibel sequence is extracted, an energy integral value of each voiceprint pulse interval is coupled with the movement frequency of the dynamic trajectory point set by weighting, and an emotion state judgment coefficient is generated;

[0037] construct a bidirectional mapping table of emotion labels and body action patterns based on harmonic components in a frequency domain of the voiceprint pulse interval, and when the emotion state judgment coefficient exceeds a dynamic baseline, use the context-related features to correct a confidence level of an emotion label corresponding to a current body action pattern;

[0038] When the corrected emotion label confidence level exceeds a preset threshold, match an interactive instruction template positively correlated with the emotion intensity level from an emotional pacification strategy library, and perform semantic alignment of the interactive instruction template and the context of the translation instruction to output a translation instruction with an additional emotional pacification suggestion.

[0039] Optionally, the parameters in the parameter set with fluctuation amplitudes exceeding a preset threshold are time-series accumulated to generate a parameter sequence reflecting an environmental sudden interference event, including:

[0040] A dynamic fluctuation threshold is established in time slices, and based on historical change trends of noise decibels in the parameter set over time and historical distribution changes of crowd density in space, a range exceeding twice the standard deviation of the median of the historical change trends or the historical distribution changes is set as an initial trigger boundary;

[0041] The parameter value fluctuation differences of adjacent time slices are calculated, and when the difference directions of three consecutive time slices are consistent and the fluctuation amplitudes exceed the initial trigger boundary, the current time slice is locked to an interference event end marker;

[0042] According to the locked time slice, a frequency energy sudden increase interval of noise decibels and a mutation direction of a crowd density spatial gradient are extracted, and an overlapping region in the time dimension is divided into independent interference segments, and interference segments with time continuity or spatial coverage overlap are accumulated;

[0043] The accumulated interference segments are parameter-weighted, and the duration of the frequency energy sudden increase interval and the spatial diffusion speed of the crowd density spatial gradient are used as weight factors to generate a parameter sequence reflecting an environmental sudden interference event.

[0044] Optionally, based on the action change features of adjacent time points in the body action time series data, the matching degree of the body action parameters in the current analysis window with the standard behavior patterns in the pet body language database is calculated, including:

[0045] The continuous change sequence of the body swing frequency, the trunk contraction amplitude, and the head deflection angle in the body action time series data is decomposed into multi-joint action trajectories, and the action change features of each joint point at adjacent time points are extracted;

[0046] Based on the gradient direction of the spatial density change in the obstacle distribution map, the joint point weight distribution of the decomposed action trajectories is dynamically adjusted, and the offset of the joint points of the body parts corresponding to the high-density area is calculated.

[0047] According to the offset calculation result, the cosine of the included angle of the motion change feature corresponding to the standard behavior mode in the analysis window is analyzed, and the cosine of the included angle is dynamically thresholded in combination with the time distribution characteristics of the visual interference parameter in the environmental stress feature vector;

[0048] The truncated cosine of the included angle is weighted and summed according to the joint node weight to generate a local matching degree, and the local matching degree is phase compensated based on the difference in periodicity between the limb swing frequency in the analysis window and the standard behavior mode in the pet body language database, to obtain the matching degree of the current analysis window.

[0049] Optionally, the determination of the stress level prediction result of the target pet in the outdoor complex environment through the association mapping of the environmental stress feature vector and the behavior abnormal segment comprises:

[0050] The spatio-temporal alignment relationship between the environmental stress feature vector and the behavior abnormal segment is constructed, and a joint feature set of the environmental stress feature vector and the behavior abnormal segment in the overlapping time window is extracted;

[0051] The joint feature set is subjected to nonlinear correlation analysis, the coupling relationship between the environmental stress feature vector and the behavior abnormal segment is captured through a multi-layer feature interaction network, and a coupling strength coefficient representing the association strength between the environment and the behavior is generated;

[0052] Based on the numerical distribution range of the coupling strength coefficient, the stress level threshold interval of the target pet in the outdoor complex environment is divided;

[0053] According to the falling point position of the coupling strength coefficient in the stress level threshold interval, the confidence weight of the stress level is dynamically allocated, and the confidence weight is corrected in combination with the same environmental stress in the historical stress event, to obtain the stress level prediction result.

[0054] In a second aspect, the present application provides a pet state intelligent translation system based on multi-modal data, comprising:

[0055] An acquisition module acquires multi-modal data of a target pet in an outdoor complex environment, the multi-modal data including noise decibels, crowd density, and three-dimensional space data captured by a laser radar;

[0056] An encoding module generates an obstacle distribution atlas based on the three-dimensional space data, and jointly encodes the multi-modal data and the obstacle distribution atlas to generate an environmental stress feature vector;

[0057] extracting a behavior abnormality segment with a space-time correlation with the outdoor complex environment;

[0058] mapping the environment stress feature vector and the behavior abnormality segment to determine a stress level prediction result of the target pet in the outdoor complex environment;

[0059] generating a translation instruction according to the stress level prediction result, the translation instruction including an evasive path prompt for a high-density area in the obstacle distribution atlas.

[0060] In a third aspect, an embodiment of the present application provides a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the pet state intelligent translation method based on multi-modal data as described in the first aspect.

[0061] In a fourth aspect, an embodiment of the present application provides a computer storage medium storing a computer program, when the computer program is executed by a computer, a pet state intelligent translation method based on multi-modal data as described in the first aspect is implemented.

[0062] In the embodiment of the present application, multi-modal data of an outdoor complex environment where a target pet is located is acquired, the multi-modal data including noise decibel, crowd density, and three-dimensional space data captured by a laser radar; an obstacle distribution atlas is generated based on the three-dimensional space data, and the multi-modal data and the obstacle distribution atlas are jointly encoded to generate an environment stress feature vector; limb action time series data of the target pet is synchronously collected, and the limb action time series data is matched with a preset pet body language database to extract a behavior abnormality segment with a space-time correlation with the outdoor complex environment; the environment stress feature vector and the behavior abnormality segment are associated and mapped to determine a stress level prediction result of the target pet in the outdoor complex environment; and a translation instruction is generated according to the stress level prediction result, the translation instruction including an evasive path prompt for a high-density area in the obstacle distribution atlas.

[0063] The technical scheme of the present application has the following beneficial effects:

[0064] The present application can comprehensively understand the characteristics of the environment where the pet is located by collecting noise decibels, crowd density and three-dimensional space data captured by the laser radar, and provide basic data support for subsequent analysis. The multi-modal data is combined with the obstacle distribution map to obtain an environmental stress feature vector, which can effectively identify potential risk factors in the environment and quantify the stress level caused by these factors on the pet, which helps to understand the stressors that the pet may face. By extracting the abnormal behavior segments of the pet, the monitoring and analysis of the pet's behavior can accurately find the behavior abnormalities of the pet in a specific environment, and provide a direct basis for evaluating the stress state of the pet. The stress level of the pet is evaluated in combination with the environmental information and the abnormal behavior of the pet, so that the judgment of the stress state is more scientific and reasonable. Finally, the avoidance path prompt in the high-density area of the obstacle distribution map is obtained, which aims to give corresponding action guidance according to the stress level of the pet, to help the pet avoid high-risk areas and improve safety.

[0065] Further, after generating the translation instruction, the language generation strategy is optimized according to the historical behavior data and individualized behavior mode of the target pet to obtain a natural language description conforming to its personality characteristics, and the emotional state of the pet is recognized by combining the body movement time sequence data and noise decibel changes, and emotional pacification suggestions are added to the translation instruction. Finally, the instruction fused with emotional pacification and avoidance path is pushed to the user terminal, and the path prompt content is dynamically adjusted according to the change of the obstacle distribution map, which not only considers the individual differences of the pet, but also enhances the ability of the pet to cope with environmental changes through dynamic adjustment strategy, provides emotional support, improves user experience, and improves the safety and comfort of the pet.

[0066] These aspects or other aspects of the present application will be more apparent in the following description of the embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0068] Figure 1 A flowchart of a pet state intelligent translation method based on multi-modal data provided by the present application is shown;

[0069] Figure 2 A structural schematic diagram of a pet state intelligent translation system based on multi-modal data provided by the present application is shown;

[0070] Figure 3A structural schematic diagram of a computing device provided by the present application is shown. DETAILED DESCRIPTION

[0071] To make the personnel in the technical field better understand the present application scheme, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0072] In some processes described in the specification and claims of the present application and the above description, a plurality of operations appearing in a specific order are included, but it should be clearly understood that these operations can be executed or in parallel without the order in which they appear in this text. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and these operations can be executed in sequence or in parallel. It should be noted that the "first", "second", etc. in this text are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence. Also, "first" and "second" are not of different types.

[0073] The technical solutions of the present application are suitable for stress response translation scenarios in outdoor complex environments. By integrating multi-modal data, including noise decibels, crowd density, and three-dimensional spatial data captured by laser radar, a comprehensive perception system of the environment where the target pet is located is constructed, and the timing data of the target pet's body movements are synchronously collected and analyzed. Finally, according to the evaluation results of the pet stress level, translation instructions containing avoidance path prompts are generated, aiming to help pets avoid high-risk areas in outdoor complex environments and improve the safety and comfort of pet outdoor activities. This process embodies a systematic solution from environmental monitoring to pet behavior analysis to personalized guidance suggestion generation.

[0074] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0075] Figure 1 A flowchart of a pet state intelligent translation method based on multi-modal data is provided for the embodiments of the present application, as shown in Figure 1 The method comprises:

[0076] 101, obtaining multi-modal data of the outdoor complex environment where the target pet is located, the multi-modal data including noise decibels, crowd density, and three-dimensional spatial data captured by laser radar;

[0077] In this step, multi-modal data refers to the simultaneous acquisition of acoustic, visual, and spatial perception data by multiple sensors, specifically including noise decibel, crowd density, and laser radar three-dimensional spatial data.

[0078] Noise decibel is an environmental sound pressure level quantification index collected by a microphone array, used to reflect the instantaneous or sustained intensity of acoustic interference.

[0079] Crowd density is the number of human targets per unit area calculated by visual cameras or infrared thermal imaging technology, used to represent the distribution density of social stress sources in the environment.

[0080] Three-dimensional spatial data is the depth point cloud data generated by laser radar by emitting laser pulses and receiving reflected signals, used to reconstruct the three-dimensional topology of the environment, including obstacle location, shape, and moving trajectory.

[0081] In the embodiments of the present application, laser radar is used to scan the target area in all directions, the distance of the object is calculated by the reflection time of the laser beam, the point cloud data containing obstacle coordinates is generated, and the obstacle contour is extracted by using a density-based spatial clustering algorithm. A distributed microphone array is deployed to collect environmental soundprint signals, the frequency energy integral is calculated after Fourier transform of the sound wave signal, and the noise decibel value of different time slices is obtained. The visual camera captures the picture, and the pre-trained detection convolutional neural network model is used to identify and count the number of pet targets in the picture. Finally, the laser radar, microphone, and camera data are correlated under the unified time sequence reference through timestamp alignment technology, ensuring the spatiotemporal consistency of multi-modal data.

[0082] In the park path scene, laser radar scans the trees, benches, and moving pedestrians on both sides of the path, generating point cloud data containing three-dimensional coordinates of obstacles. The distributed microphone array detects a sudden increase in the sound pressure level of the children's laughter in the distance to 85 decibels, with continuous fluctuations. The visual camera captures the picture of the crowd gathering at the entrance of the path, and the pedestrian detection model counts the crowd density in this area as 3 people per square meter. The three are synchronized through timestamp alignment to form a synchronized multi-modal data set.

[0083] 102、Based on the three-dimensional spatial data, an obstacle distribution map is generated, and the multi-modal data is jointly encoded with the obstacle distribution map to generate an environmental stress feature vector;

[0084] In this step, the obstacle distribution map is a spatial topology model generated by analyzing three-dimensional spatial data, quantifying the density gradient, line-of-sight obstruction coefficient, and moving obstacle trajectory of obstacles in the activity area.

[0085] The environmental stress feature vector is a composite feature vector generated by fusing noise decibel, crowd density, and obstacle distribution map, used to represent the comprehensive interference intensity of the environment on pets.

[0086] In the embodiments of the present application, the three-dimensional space data generated by the laser radar is divided into uniform grid cells, the frequency of obstacles in each cell is counted, and the line-of-sight blocking coefficient (i.e. the probability that the line-of-sight is blocked by obstacles) between adjacent cells is calculated. The trajectory of a moving obstacle (such as a pedestrian or a vehicle) is tracked, and its direction and speed of movement are predicted by a Kalman filter algorithm. The noise decibel sequence is mapped to the corresponding grid cell according to the time slice, and the mean and variance of the sound pressure level of each cell are calculated. The crowd density distribution is superimposed on the grid cell, and the coverage range of the high-density area is calculated. The noise, crowd and obstacle data are weighted and fused by an attention mechanism, where the weight of the sudden high-decibel area and the high-density crowd area is increased, and an environmental stress feature vector containing spatial, acoustic and visual joint features is generated.

[0087] Continuing the above case, the laser radar point cloud data is divided into 0.5m x 0.5m grid cells, the obstacle density at the entrance of the sidewalk is calculated to be 0.8 (the highest is 1), the line-of-sight blocking coefficient in the bench area is 0.9, and the moving obstacle tracking module predicts that a runner is approaching the pet at a speed of 2 meters per second. The noise decibel sequence is mapped to the corresponding grid cell, the mean sound pressure level at the entrance is 80 decibels, and the crowd density distribution shows that the entrance coverage rate reaches 70%. After weighted fusion by the attention mechanism, the environmental stress feature vector strengthens the composite features of the entrance "high decibel + high density + moving obstacle approaching".

[0088] 103、Synchronously collecting the time sequence data of the target pet's limb movements, and performing pattern matching on the time sequence data of the limb movements and a preset pet body language database to extract a behavior abnormality segment that has a spatiotemporal correlation with the outdoor complex environment;

[0089] In this step, the time sequence data of the limb movements is continuously recorded by the inertial sensor in the pet wearable device, including the limb swing frequency, the trunk contraction amplitude, and the head deflection angle.

[0090] The pet body language database is a pre-defined standardized behavior pattern library, including joint motion feature templates of common stress actions (such as rigidity and escape) and non-stress actions (such as walking and sniffing).

[0091] The behavior abnormality segment refers to a motion sequence that deviates significantly from the preset standard behavior pattern exhibited by the pet within a specific time period. This segment includes abnormal fluctuations in the joint motion parameters (such as limb swing frequency, trunk contraction amplitude, and head deflection angle) in the time sequence data of the limb movements, and needs to have a strong correlation with the spatiotemporal tags of the environmental stress feature vector (such as a sudden increase in noise decibels and a dense obstacle area).

[0092] In this embodiment, a nine-axis inertial measurement unit (accelerometer, gyroscope, magnetometer) is embedded in a pet collar or vest to collect triaxial acceleration, angular velocity, and orientation data of the limb joints. The raw data is segmented into sliding windows, and parameters such as swing frequency and contraction amplitude within each window are extracted. A dynamic time warping algorithm is used to align the action sequence of the current window with standard behavior templates in the database, calculating a similarity score for the joint trajectories. Combined with spatiotemporal labels from environmental pressure feature vectors (such as high-noise periods or areas with dense obstacles), time windows with similarity below a preset threshold are selected and marked as abnormal behavioral segments spatiotemporally associated with the outdoor environment.

[0093] The nine-axis inertial sensor in the pet wearable device detected a sudden increase in trunk contraction amplitude from 5 cm to 15 cm within 2 seconds (the average normal walking amplitude is 8 cm), and the head deflection angle continuously turned to the right by more than 30 degrees. The dynamic time warping algorithm calculated that its similarity to the "normal exploration" mode was only 40%. Combined with the high-risk label at the entrance in the current environmental pressure feature vector, this period (10:05:23-10:05:25) was determined to be a segment of abnormal behavior.

[0094] 104. By mapping the environmental stress feature vector with the abnormal behavior segments, determine the predicted stress level of the target pet in the complex outdoor environment;

[0095] In this step, the association mapping establishes a non-linear relationship between environmental stress feature vectors and abnormal behavioral segments through a machine learning model, quantifying the impact of environmental disturbances on pet behavior.

[0096] The stress level prediction results are dynamic classification results based on historical data clustering, including three levels: low stress, moderate stress, and high stress.

[0097] In this embodiment, environmental stress feature vectors (such as noise decibel gradient and obstacle density) and motion parameters of abnormal behavior segments (such as sway frequency shift and contraction amplitude difference) are input into a multilayer perceptron model. The model learns the nonlinear relationship between the environment and behavior through fully connected layers and outputs a stress intensity score. Based on a Gaussian mixture model, historical stress scores are clustered to dynamically divide threshold intervals for low, medium, and high stress levels. For example, when the product of noise decibels and obstacle density in the environmental stress feature vector exceeds a threshold, and the motion parameter shift of the abnormal behavior segment reaches a critical value, it is determined to be a high stress level.

[0098] The environmental pressure characteristics at the entrance (noise decibel gradient + 15 dB / s, obstacle density 0.8) and the abnormal behavior segment parameters (contraction amplitude offset + 87.5%) are input into the multi-layer perception model, and the output is a stress intensity score of 92 points (full score 100). Gaussian mixture model clustering analysis shows that the historical high stress threshold is 85 points, and the current score triggers the "high stress" level determination.

[0099] 105. Generate translation instructions according to the stress level prediction results, which include evasive path prompts for high-density areas in the obstacle distribution map.

[0100] In this step, the translation instructions are composite outputs containing natural language descriptions and path planning instructions to guide users to adjust pet behavior.

[0101] High-density areas refer to spatial units where the frequency of obstacles is significantly higher than the average level or areas where the line-of-sight obstruction coefficient exceeds a preset threshold. This area includes dense distribution of static obstacles (such as fixed buildings, trees) and overlapping areas of dynamic obstacles (such as crowds, moving vehicles).

[0102] Evasive path prompts are dynamic obstacle avoidance routes generated based on the topological structure of high-density areas in the obstacle distribution map, avoiding dense areas of moving obstacles.

[0103] In the embodiments of the present application, a pre-defined instruction template library is matched according to the stress level, for example, "high-risk area detected in front, suggest immediate detour". Combined with the predicted trajectory of moving obstacles in the obstacle distribution map, a path search algorithm is used to calculate the minimum risk path from the current position to the safe area. The path weight is determined by the obstacle density, line-of-sight obstruction coefficient and noise decibel, and the route with low density and high visibility is preferred. The path coordinates are converted into natural language descriptions (such as "turn right along the green belt for 50 meters"), and are fused with emotional soothing suggestions (such as "keep the leash relaxed to reduce pet anxiety"), to generate translation instructions.

[0104] According to the instruction template matched with the high stress level, the obstacle avoidance path is calculated in the path planning module, which is to turn right around the green belt from the current position to avoid the crowd area and bench obstruction area at the entrance (path risk value decreases from 0.78 to 0.21). The translation instructions are generated, such as "crowd gathering and high decibel noise detected in front, suggest immediate right turn along the green belt for 50 meters, keep the leash relaxed to relieve pet anxiety." The user terminal map synchronously highlights the path and labels "high-risk area with dense crowds".

[0105] In summary, steps 101 to 105 perceive the compound risk of crowds, noise, and moving obstacles at the multi-modal data perception portal, combined with abnormal movements such as pet body contraction and head deflection, to accurately determine the high stress state. The generated instructions not only provide a dynamic obstacle avoidance path (around crowds and occluded areas), but also provide calming suggestions, realizing the closed-loop response of environmental risk and pet behavior. Ultimately, it helps users guide the pet to safely escape the high-pressure area, verifying the effectiveness of the entire link from data collection to decision output.

[0106] To further improve the individual adaptation ability and environmental dynamic response accuracy of the pet stress translation instructions, the scheme customizes the language style of the translation instructions by calling the base large model to analyze historical behavior and real-time data, ensuring that the instructions are consistent with the pet's personality and can identify the pet's emotional state, providing appropriate emotional calming suggestions. In addition, these suggestions are pushed to the user, and the path guidance is dynamically updated according to environmental changes to enhance the pet's sense of safety and comfort. In some embodiments, after generating the translation instructions based on the stress level prediction results in step 105, the method further includes: 201, constructing a base large model based on multi-modal data, the base large model including cross-modal analysis results of human speech, environmental sound, and animal vocalization;

[0107] In step 201, the base large model refers to a deep learning model pre-trained based on massive Internet multi-modal audio data. The input includes human speech waveform, environmental sound spectrum, and animal vocalization time domain signal, and the output is cross-modal semantic analysis result. Cross-modal analysis result refers to the intermediate feature representation of the base large model for unified understanding of different types of audio data, which can map emotional tendencies in human speech, danger signals in environmental sound, and emotional fluctuations in animal vocalization to the same semantic space.

[0108] In the embodiments of the present application, a pre-trained general audio large model is loaded as the base, and the model uses a deep convolutional neural network and a self-attention mechanism fusion architecture. First, input the multi-modal data collected in the outdoor environment, and process it through parallel branches. The human speech branch performs speech activity detection and framing processing, the environmental sound branch performs Mel-frequency cepstral coefficient feature extraction, and the animal vocalization branch performs time domain envelope analysis. The output features of the three branches are interactively fused through the cross-modal attention layer to generate cross-modal analysis results containing semantic correlation, which serve as the basis for subsequent steps. This process enables the model to understand the correlation of "people, environment, and animals" sound.

[0109] 202, jointly label the cross-modal analysis results of the animal vocalization according to the environmental stress feature vector and the behavior abnormality segment to obtain a labeled result, the labeled result including the mapping relationship of dog breed, emotional type, and environmental stress;

[0110] In step 202, joint labeling refers to a process of adding multi-dimensional labels to animal call data using environmental stress feature vectors and behavior abnormality segments. The mapping relationship is that in the labeling dimension, the dog breed corresponds to the breed identification of the target pet, the emotion type includes discrete classifications such as fear and excitement, and the environmental stress mapping relationship refers to establishing a quantitative association matrix of call features and specific environmental stress values.

[0111] In the embodiments of the present application, the cross-modal analysis results of the target pet in the complex outdoor environment are collected, and the environmental stress feature vector and the behavior abnormality segment are synchronously associated. The labeling system uses a semi-automatic rule. First, the environmental stress feature vector is divided into high / medium / low risk levels through a clustering algorithm. Second, the emotion label library is matched according to the amplitude of the body movement in the behavior abnormality segment. Finally, the labeling result including the mapping relationship of the dog breed, the emotion type and the environmental stress is generated. This labeling result is used as a supervision signal to guide the model to learn the causal chain of environment and emotion.

[0112] 203, extracting the mel-frequency spectrum features of the labeling result using the base large model to obtain a spectrum vector, and performing feature alignment on the spectrum vector and the environmental stress feature vector, based on the alignment result, adjusting the parameters of the base large model to obtain an adjusted base large model;

[0113] In step 203, the spectrum vector refers to a frequency domain feature sequence converted from the labeled call signal by a mel filter bank, representing the acoustic fingerprint of the call. Feature alignment refers to projecting the spectrum vector and the environmental stress feature vector into the same dimension space to minimize the distribution distance of the two types of features.

[0114] In the embodiments of the present application, the labeling result is input into the call analysis branch of the base large model, the mel-frequency spectrum features are extracted through multiple one-dimensional convolution layers to generate a spectrum vector. At the same time, the environmental stress feature vector is input into the environmental analysis branch of the model. Feature alignment module is used for fusion. Based on the aligned fusion features, the gradient clipping strategy is used to fine-tune the parameters of the base large model. When high-risk environmental stress is detected, the weights of fear-related neurons in the call branch are strengthened. When the behavior abnormality segment shows low stress, the response of irrelevant frequency channels is weakened. Iteration is performed until the environmental-related emotion recognition accuracy of the base large model on the test set is stable above a threshold value, and the adjusted base large model is obtained.

[0115] 204, constructing dog language reasoning logic based on the adjusted base large model, and adjusting the language generation strategy of the translation instruction using the dog language reasoning logic to generate a natural language description matching the personality characteristics of the target pet;

[0116] In step 204, the dog language reasoning logic encapsulates the fine-tuned lightweight inference module of the large model parameters, which can convert the real-time input pet call into a natural language description. The language generation strategy is a command template library based on a natural language processing model, which adjusts the wording style and information density of the natural language description according to the pet personality label, including sensitive, exploratory, stable, etc. The natural language description is a text instruction generated in combination with the pet personality characteristics, which is used to guide the user to take interactive actions that adapt to the behavior characteristics of the pet.

[0117] In the embodiments of the present application, the parameter subset related to animal call analysis in the adjusted base model is derived, and the parameter subset is compressed into dog language reasoning logic through knowledge distillation technology. When the target pet emits a new call, the personalized behavior pattern library is called, and the personalized behavior matching the target pet in the personalized behavior pattern library is input into the pre-trained natural language generation model. The model selects the instruction template matching the personality characteristics through the attention mechanism. For example, for a sensitive pet, a description of "detected dense crowd in front, target pet prone to anxiety, please slowly guide to turn left to bypass" is generated, and for an exploratory pet, a description of "suggest exploring to the left to bypass" is generated. Finally, the natural language instruction containing the personalized description is output.

[0118] 205、based on the natural language description, in combination with the limb action time series data and the noise decibel, the emotional state of the target pet is identified, and the emotional pacification suggestion corresponding to the emotional state is attached to the translation instruction;

[0119] In step 205, the emotional state is the pet instantaneous psychological state classification result determined by multi-modal data fusion, including labels such as fear, anxiety, calm, etc. The limb action time series data is the pet joint motion parameter collected by the inertial sensor in the current time period, such as paw tremor frequency and ear rear angle. The noise decibel is the quantization sequence of the environmental sound pressure level. The emotional pacification suggestion is a set of intervention strategies matching the emotional state, including sound pacification instructions (such as playing specific frequency white noise), physical guidance suggestions (such as light pulling traction rope force), etc. operable instructions.

[0120] In the embodiments of the present application, keywords (such as "anxiety" and "rigidity") are extracted from natural language descriptions, micro-movement features in current limb movement time series data are analyzed through a convolutional neural network, for example, the time series correlation of paw tremor frequency and ear rear attachment angle is calculated. The frequency energy distribution of the noise decibel sequence is synchronously analyzed to identify the time coverage interval of sudden acoustic interference (such as car horn). The micro-movement features, acoustic features, and keyword vectors are input into an emotion classification model to output an emotion label (such as "high fear"). According to the emotion label, the instruction with the highest positive rate of historical interaction feedback is selected from the pacification strategy library, for example, the "call the pet's name softly and pat the back" suggestion matched with the "high fear" label is added to the end of the translation instruction.

[0121] 206. The translation instruction with the emotional pacification suggestion is pushed to the user terminal, and the avoidance path prompt in the translation instruction is dynamically updated according to the changes in the high-density area of the obstacle distribution map.

[0122] In step 206, the dynamic updating mechanism is a path correction logic based on the changes in the high-density area of the obstacle distribution map, which is used to ensure that the avoidance path prompt is synchronized with the current environmental state. The high-density area is a set of spatial units in the obstacle distribution map whose obstacle density or line-of-sight obstruction coefficient exceeds a preset threshold. The avoidance path prompt is a minimum-risk movement route generated according to the topological structure, which avoids the high-density area.

[0123] In the embodiments of the present application, after the translation instruction is pushed to the user terminal, the movement trajectory changes of the high-density area in the obstacle distribution map are continuously monitored, for example, it is detected that the crowd spreads to the northeast direction, causing the original path to be blocked. The conflict probability of the original path and the new obstacle distribution is calculated through a path re-planning algorithm, if the conflict probability exceeds a dynamic threshold, an incremental updating strategy is adopted, the path main section is retained, and the detour direction is locally adjusted (such as changing "turn right to detour" to "slant through the green belt to the right front"). The updated path coordinates and emotional pacification suggestions are re-fused to generate an incremental instruction and push it, for example, "detect crowd movement, please adjust to slant through the green belt to the right front, continue to pat the pet's back".

[0124] The following is a specific example:

[0125] In the park trail scene, the environmental perception system collects real-time crowd density, noise decibel, and laser radar point cloud when the target pet passes through the tourist gathering area. The base large model analyzes the environmental stress feature vector and synchronously identifies the dog's high-frequency whining as "high-frequency whining". Combined with historical behavior annotation, the current call is jointly annotated. The dog language reasoning logic generates a natural language description based on the fine-tuned parameters, and combines real-time body data (tight tail clamping) and noise data to determine high fear. The server generates a pacification suggestion, while the initial avoidance path prompt is "turn right and bypass". When the laser radar detects a sudden turn of the skateboard, the new obstacle map is obtained through the application programming interface, and the instruction is immediately updated to "accelerate straight through", which is pushed to the owner's mobile terminal.

[0126] In summary, steps 201 to 206 build a joint annotation system of environmental stress and call emotion, so that the fine-tuning process of the base large model has clear environmental relevance. The semantic space of multi-modal data is unified by using feature alignment technology, which significantly improves the accuracy of emotion recognition in complex environments. Based on the adjustment of language generation strategy according to individual behavior patterns, it is ensured that the translation result conforms to the pet personality characteristics. Combined with dynamic obstacle perception and real-time path update, a closed-loop decision is formed from emotion recognition to action guidance. Finally, in outdoor scenes, the whole-link intelligent translation of "accurate perception of environmental threats, accurate translation of individual emotions, and dynamic optimization of avoidance paths" is realized.

[0127] In order to improve the recognition accuracy of environmental obstacles in the target pet activity range, the present scheme uses three-dimensional space data to evaluate the obstacle situation in the pet activity area in detail, including obstacle density and spatial occlusion degree, and generates an obstacle distribution map. In some embodiments, the obstacle distribution map in step 102 is constructed by quantifying the moving obstacle density and spatial occlusion coefficient in the target pet activity range;

[0128] The method for generating an obstacle distribution map based on three-dimensional space data comprises:

[0129] 301、Based on the three-dimensional space data, the target pet's corresponding activity range is divided into multiple spatial grid units, the moving trajectory of the obstacle in each spatial grid unit and its appearance frequency in a preset time period are extracted, and the line-of-sight occlusion degree change between adjacent spatial grid units is combined to calculate the moving obstacle density and spatial occlusion coefficient of each spatial grid unit;

[0130] In step 301, a spatial grid cell is a uniform spatial block dividing the target pet's activity range by a fixed size, used to quantify the spatial resolution of obstacle distribution. A movement trajectory is a set of continuous position coordinates of an obstacle within a preset time period, reflecting changes in its direction and speed. Occurrence frequency is the percentage of times an obstacle is detected within a single spatial grid cell, used to characterize the density of static obstacle distribution. Line-of-sight occlusion is the proportion of reduced visibility between adjacent spatial grid cells due to the presence of obstacles, calculated by the probability of the line-of-sight path being blocked. Moving obstacle density is a weighted value of the number of trajectories of dynamic obstacles (such as pedestrians and vehicles) within a spatial grid cell and their occurrence frequency. The spatial occlusion coefficient is a cell-level oppressive index generated by combining line-of-sight occlusion and static obstacle density.

[0131] In this embodiment, the park trail is divided into square grid cells with sides of one meter using LiDAR point cloud data. The frequency of static obstacles such as benches and trees is counted for each cell, and the number of trajectories is calculated by tracking pedestrian movement. A line-of-sight tracking algorithm is used to calculate the probability of line-of-sight occlusion between adjacent cells; for example, the line-of-sight occlusion coefficient between the cell containing the bench and the cell to the east is 0.9. The density of moving obstacles is calculated by combining the frequency of occurrence and the number of trajectories; for example, the density at the trail entrance cell is 0.7 (maximum 1).

[0132] 302. Map the moving obstacle density and spatial occlusion coefficient of each spatial grid cell to a three-dimensional thermal layer, and overlay them to generate the obstacle distribution map.

[0133] In step 302, the 3D heatmap is a visualization tool used to show the density of moving obstacles and the intensity of spatial occlusion coefficients within different grid cells. The obstacle distribution map is a map generated by integrating the above information, reflecting the distribution of obstacles throughout the entire activity area.

[0134] In this embodiment, firstly, the moving obstacle density and spatial occlusion coefficient calculated for each spatial grid cell are converted into corresponding color values ​​to generate a three-dimensional thermal layer. Then, the thermal layers of all grid cells are layered together to form a complete obstacle distribution map. This map not only visually displays the location and density of obstacles but also reveals the line-of-sight obstruction caused by the presence of obstacles, providing fundamental data support for subsequent environmental stress assessments.

[0135] Here is a specific example:

[0136] In the park path scene, the system divides the path and surrounding area into a one-meter-square three-dimensional grid. In a certain grid cell at the entrance, the pedestrian trajectory density is detected to be 0.7, the frequency of benches is 0.9, the degree of line-of-sight obstruction is 0.8, the mobile obstacle density is calculated to be 0.78, and the space obstruction coefficient is 0.62. The adjacent green belt unit has no dynamic obstacles and the line of sight is transparent, with a density of 0.1 and an obstruction coefficient of 0.08. Through three-dimensional heat map layer mapping, the entrance unit is displayed as deep red superimposed on deep blue (high density and high obstruction), and the green belt unit is displayed as light red superimposed on light blue.

[0137] In summary, steps 301 to 302 achieve accurate depiction of the distribution of environmental obstacles by quantitatively analyzing the mobile obstacle density and space obstruction coefficient in the target pet's activity range. By monitoring and analyzing the distribution of obstacles, the safety and comfort of pets and other visitors in public places are improved.

[0138] To improve the accuracy of identifying the environmental stress characteristics of the target pet in a complex outdoor environment, in some embodiments, the joint encoding of the multi-modal data and the obstacle distribution atlas in step 102 to generate an environmental stress feature vector includes:

[0139] 401. According to the change trend of noise decibels over time and the distribution change of crowd density in space, extract the parameter set in the multi-modal data that overlaps with the high-density area in the obstacle distribution atlas, and perform time series accumulation on the parameters in the parameter set whose fluctuation amplitude exceeds the preset threshold to generate a parameter sequence reflecting environmental sudden disturbance events;

[0140] In step 401, the parameter set is a set of noise decibel sequences and crowd density distribution data in the multi-modal data that overlap with the high-density area in space. The fluctuation amplitude is the deviation of the parameter value from the historical mean, which is measured by the standard deviation multiple. Time series accumulation is the energy integration and spatial superposition calculation of parameters exceeding the fluctuation threshold within a time window. Environmental sudden disturbance events are complex interference scenarios triggered by abnormal fluctuations in multi-modal parameters, such as crowds gathering with high-decibel noise. The parameter sequence is a time-sequential feature set that records the start time, intensity peak, and spatial range of the sudden disturbance event.

[0141] In this embodiment, noise decibel sequences and crowd density distribution data corresponding to high-density areas in the obstacle distribution map are extracted, and the fluctuation amplitude of each parameter is calculated. An interval where the fluctuation amplitude exceeds a preset threshold is detected using a sliding time window. Energy integration (e.g., frequency domain energy accumulation of noise decibels) and spatial coverage calculation (e.g., the number of grids for crowd density diffusion) are performed on the parameter values ​​within each window. Windows that meet the conditions are marked as candidate interference events. Adjacent or overlapping windows are merged in chronological order to generate a parameter sequence, and its start time, peak intensity, and influence range are recorded.

[0142] 402. The moving obstacle density, spatial occlusion coefficient and parameter sequence are correlated and matched to determine the diffusion path of the sudden environmental interference event in the obstacle distribution map, and the moving obstacle density and spatial occlusion coefficient of each spatial grid cell are adjusted according to the diffusion path.

[0143] In step 402, the diffusion path is the direction and coverage of the sudden environmental disturbance event in the obstacle distribution map, which is determined through spatial correlation analysis of the parameter sequence. Correlation matching involves aligning the density of moving obstacles, spatial occlusion coefficient, and parameter sequence in time and space to calculate the probability model of event diffusion.

[0144] In this embodiment, the spatial range of the interference events recorded in the parameter sequence is spatiotemporally aligned with the density distribution of moving obstacles in the obstacle distribution map, and the event propagation direction is calculated using a probabilistic graphical model. For example, based on the consistency between the crowd density propagation trend and the trajectory direction of moving obstacles, the propagation of the interference event along a specific path is predicted. The density of moving obstacles in the grid cells along the propagation path is updated (e.g., by increasing the weight of the trajectory number) and the spatial occlusion coefficient.

[0145] 403. Based on the adjusted moving obstacle density, spatial occlusion coefficient, and the temporal cumulative results of the parameter sequence, the temporal fluctuation characteristics of the multimodal data and the spatial topological characteristics of the obstacle distribution map are tensor-concatenated, and the concatenation result is encoded by a convolutional neural network to generate an environmental pressure feature vector.

[0146] In step 403, the temporal accumulation result is a weighted aggregation of the duration, peak intensity, and spatial coverage of sudden disturbance events in the parameter sequence. Tensor concatenation is a technique for combining data from different dimensions. Convolutional neural networks are a type of deep learning model that excels at processing data with spatial topological structures. The environmental stress feature vector is a composite feature generated by integrating and adjusting the density of moving obstacles, spatial occlusion coefficient, and temporal accumulation results, used to characterize the overall stress intensity of the environment on pets.

[0147] In the embodiments of the present application, first, according to the adjusted mobile obstacle density, the spatial occlusion coefficient and the time sequence accumulation result of the parameter sequence, the tensor splicing method is used to combine the multi-modal data with the time sequence fluctuation characteristics and the spatial topological characteristics of the obstacle distribution atlas. Then, the convolutional neural network is used to encode the spliced data, and finally a feature vector that can fully reflect the environmental stress characteristics is generated.

[0148] The following is a specific example:

[0149] In the park path scene, the system divides the activity range into one-meter-square grid cells. The crowd density in the entrance cell is detected to surge with a sudden whistle sound, and a parameter sequence labeled as "entrance sudden interference event" is generated. It is predicted that the event spreads northeast along the path, increasing the obstacle density and occlusion coefficient of the grid along the way. Finally, the environmental stress feature vector is generated by fusing the adjusted parameters, which is used for subsequent stress level prediction.

[0150] In summary, steps 401 to 403 realize the fine quantization of environmental stress sources by spatial grid modeling and dynamic association with multi-modal parameters. Based on the diffusion path prediction and parameter adjustment mechanism of the sudden interference event, the obstacle distribution atlas is improved. The environmental stress feature vector effectively fuses the spatial occlusion, mobile obstacle density and sudden interference intensity, providing high-discrimination input for stress level prediction. In the park path scene, the system accurately identifies the compound interference of the crowd and noise at the entrance, dynamically adjusts the grid parameters along the path, and makes the environmental stress evaluation results highly consistent with the real risk scene, laying a data foundation for generating accurate obstacle avoidance paths.

[0151] In order to further improve the spatio-temporal correlation and environmental adaptability of pet behavior abnormal segment extraction, in some embodiments, the limb movement time sequence data in step 103 includes continuous change sequences of limb swing frequency, trunk contraction amplitude and head deflection angle;

[0152] The mode matching of the limb movement time sequence data with the preset pet body language database to extract the behavior abnormal segment with spatio-temporal correlation in the outdoor complex environment includes:

[0153] 501, divide the limb movement time sequence data into analysis windows with variable length, and dynamically adjust the sliding interval of the analysis window according to the spatial density change of the obstacle distribution atlas in the outdoor complex environment;

[0154] In step 501, the analysis window is a variable-length interval that divides the limb movement timing data in the time dimension, and the sliding interval is dynamically adjusted according to the gradient direction of the spatial density change in the obstacle distribution map. The spatial density change is the rate of increase or decrease of the number of obstacles or the number of movement trajectories in a unit area in the obstacle distribution map, which is calculated by the difference between the grid cell statistics of adjacent time slices. The dynamic sliding interval refers to the time span between the start point and the end point of the analysis window, which is automatically adjusted according to the fluctuation of environmental pressure. For example, when the high-density area is expanded, the interval is shortened to improve the detection sensitivity.

[0155] In the embodiments of the present application, the spatial density change rate of each grid cell in the obstacle distribution map is monitored, and the density gradient direction (such as diffusion from southeast to northwest) is calculated. The sliding interval of the analysis window is adjusted according to the gradient direction. If the density change rate exceeds the threshold value, the interval is shortened to increase the sampling frequency; if the change is gentle, the interval is lengthened to reduce the computational load. The limb movement timing data (such as swing frequency, contraction amplitude) in each window is cut and cached according to the time stamp.

[0156] 502、In each analysis window, based on the action change characteristics of adjacent time in the limb movement timing data, the matching degree of the limb movement parameters in the current analysis window and the standard behavior mode in the pet limb language database is calculated, and the dynamic matching condition is generated by combining the parameters related to visual interference in the environmental pressure feature vector;

[0157] In step 502, the action change characteristics are a set of difference values of the limb movement parameters at adjacent time in the analysis window, such as the increment of swing frequency and the change rate of contraction amplitude. The matching degree is a similarity quantification index of the current window action sequence and the standard behavior mode in the pet limb language database, which is calculated by the trajectory alignment algorithm. The dynamic matching condition is a similarity judgment threshold generated by combining the visual interference parameters (such as obstacle density and line-of-sight blocking coefficient) in the environmental pressure feature vector, which is used to dynamically adjust the matching reference.

[0158] In the embodiments of the present application, for the action sequence in each analysis window, the dynamic time warping algorithm is used to calculate the trajectory similarity with the templates in the standard behavior mode library. According to the visual interference parameters (such as high blocking coefficient) in the environmental pressure feature vector, the matching threshold is dynamically adjusted. When the visual interference is enhanced, the similarity threshold is lowered to capture more subtle action abnormalities. For example, if the blocking coefficient exceeds 0.7, the matching threshold is lowered from 0.65 to 0.55.

[0159] 503、According to the dynamic matching condition, the analysis window with a matching degree lower than a preset reference is screened, a periodic difference of limb swing frequency and trunk contraction amplitude in the analysis window is extracted, position information of a high-density area in the obstacle distribution map is combined, and a behavior abnormality segment with a position correlation to the outdoor complex environment is marked;

[0160] In step 503, the periodic difference is a non-autonomous rhythmic fluctuation feature of the limb swing frequency and the trunk contraction amplitude in the analysis window, for example, high-frequency irregular jitter. The position information is a set of spatial coordinates of the high-density area in the obstacle distribution map, used to associate the occurrence position of the behavior abnormality segment. The behavior abnormality segment is a motion sequence that simultaneously satisfies low matching degree, significant periodic difference, and spatial overlap with the high-density area.

[0161] In the embodiment of the application, the analysis window with a matching degree lower than the dynamic threshold is screened, and the time sequence fluctuation feature of the limb swing frequency and the trunk contraction amplitude is extracted. The periodic component is detected by Fourier transform, and if the fundamental frequency deviates from the normal range and the harmonic energy is dispersed, it is determined that the periodic difference is significant. The timestamp of the abnormal window is matched with the spatial coordinates of the high-density area in the obstacle distribution map, and the spatiotemporally overlapping segment is screened and marked as a position-correlated abnormality.

[0162] 504、The behavior abnormality segment is associated and compared with the spatial distribution change of the crowd density in the multi-modal data, non-stress action data generated by pet autonomous behavior is excluded, and a set of spatiotemporally associated behavior abnormality segments is generated.

[0163] In step 504, the non-stress action data is a motion sequence generated by pet autonomous behavior (such as chasing flying insects and sniffing the ground), and the spatial distribution change of the crowd density has no spatiotemporal overlap with the behavior segment. The set of spatiotemporally associated behavior abnormality segments is the remaining abnormality segment set after excluding the non-stress data, which satisfies the strong correlation between the behavior abnormality and the crowd density mutation in time synchronization and spatial coverage.

[0164] In the embodiment of the application, the marked position-correlated abnormality segment is spatiotemporally compared with the crowd density distribution in the multi-modal data. If the crowd density in the grid unit where the abnormal segment occurs has no significant change during the period, it is determined as a pet autonomous action and is excluded. Finally, the abnormal segment that is spatiotemporally overlapped with the crowd density surge area is retained, and a set of spatiotemporally associated behavior abnormality segments is generated.

[0165] The following is a specific example:

[0166] In the park path scene, the system detects a sudden increase in obstacle density at the entrance, shortens the analysis window interval to 3 seconds. Within a certain window, the pet's torso contraction amplitude fluctuates irregularly, similar to the "normal walking" mode with a similarity of 0.55, which is lower than the dynamic threshold of 0.6. Extract the sudden increase in swing frequency feature in this window, and combine it with the high-density coordinate label to mark the abnormal segment. After comparing with the crowd density data, exclude similar segments in the area far from the crowd, and generate the final abnormal set.

[0167] In summary, steps 501 to 504 achieve synchronous behavior anomaly detection and environmental stress fluctuation by dynamically adjusting the analysis window and matching threshold, reducing false negatives or false positives caused by fixed parameters. Combined with the correlation screening of high-density area spatial coordinates, it effectively distinguishes between stress behaviors and autonomous actions. In the park path scene, the system accurately captures the abnormal contraction and shaking of pets when the crowd is dense at the entrance, excludes the interference of running in the empty area, provides high-confidence input for subsequent stress level prediction, and significantly improves the spatio-temporal relevance and environmental adaptability of behavior analysis.

[0168] To improve the accuracy of pet emotional state recognition and the adaptability of soothing strategies, in some embodiments, step 202 includes identifying the emotional state of the target pet based on the natural language description, combined with the limb action time series data and the noise decibel, and adding emotional soothing suggestions corresponding to the emotional state in the translation instruction, including:

[0169] 601. Perform word segmentation processing on the natural language description, extract context-related features, and generate an emotional feature vector containing word segmentation weight distribution and semantic dependency strength.

[0170] In step 601, the natural language description refers to information obtained through voice or text form. Word segmentation processing is the division of continuous natural language text into individual words. Context-related features refer to the semantic relationship between words. The emotional feature vector is a mathematical representation containing emotion-related information extracted from the natural language description, such as word segmentation weight distribution and semantic dependency strength.

[0171] In the embodiments of the present application, natural language processing techniques are first applied to perform word segmentation processing on the natural language description, and word frequency and inverse document frequency methods are used to calculate the importance weight of each word. Then, by analyzing the sentence structure and the relationship between words, an emotional feature vector is generated, which not only reflects the emotional tendency in the text, but also contains the semantic dependency degree between words.

[0172] 602、establish a two-dimensional feature space with limb movement amplitude and movement frequency as axes, construct a dynamic trajectory point set based on the emotional feature vector and the limb movement time sequence data of the target pet within a set time window, and divide the emotional intensity level of the emotional state through the distribution density of the dynamic trajectory point set in the two-dimensional feature space;

[0173] In step 602, the two-dimensional feature space is a quantitative analysis space constructed with limb movement amplitude and movement frequency as coordinate axes, the limb movement amplitude is the maximum displacement of joint movement, the movement frequency is the number of repeated limb periodic movements per unit time, the dynamic trajectory point set is a set of time sequence data points of the movement amplitude and frequency of the target pet within a set time window, the distribution density is the number of points in a unit area of the trajectory point set in the two-dimensional space, and the emotional intensity level is a pet emotion quantitative indicator divided by the distribution density.

[0174] In the embodiments of the present application, the inertial sensor worn on the pet's torso is used to collect limb movement time sequence data, the torso contraction amplitude and limb swing frequency at each time point are extracted, the emotional feature vector, the torso contraction amplitude and the limb swing frequency are fused, a dynamic trajectory point set is constructed, and the dynamic trajectory point set is mapped to a two-dimensional coordinate system to form a dynamic trajectory point set. A density-based spatial clustering algorithm is used to calculate the distribution density of the trajectory points in different regions, the density threshold is adjusted in combination with the part-of-speech weight distribution of the emotional feature vector, and low-density, medium-density and high-density regions are divided. The density threshold is set according to historical data, and the emotional intensity level of the emotional state is divided.

[0175] 603、extract the voiceprint pulse interval with a duration exceeding a threshold value in the noise decibel sequence, weight and couple the energy integral value of each voiceprint pulse interval with the movement frequency of the dynamic trajectory point set to generate an emotional state judgment coefficient;

[0176] In step 603, the voiceprint pulse interval is a continuous high-energy period in the noise decibel sequence with a duration exceeding a preset threshold, the energy integral value is the total energy of the sound pressure level accumulated over time in the voiceprint pulse interval, the weight coupling is a calculation process of fusing the voiceprint energy integral value with the movement frequency of the dynamic trajectory point set according to a weight coefficient, and the emotional state judgment coefficient is a comprehensive index quantifying the correlation between acoustic interference and movement features.

[0177] In the embodiments of the present application, the noise decibel sequence is detected by a sliding window, a period exceeding the threshold decibel continuously is identified, the voiceprint signal of the period is extracted and subjected to fast Fourier transform, and the frequency energy integral value is calculated. The average movement frequency of the dynamic trajectory point set in the same period is extracted synchronously, the voiceprint energy integral value and the movement frequency are linearly weighted according to a preset weight coefficient, and the emotional state judgment coefficient is generated.

[0178] 604、construct a bidirectional mapping table of emotional labels and body action patterns based on the frequency domain harmonic components of the voiceprint pulse interval, when the emotional state determination coefficient exceeds the dynamic baseline, the context-related features are used to modify the confidence of the emotional label corresponding to the current body action pattern;

[0179] In step 604, the frequency domain harmonic component is the energy distribution characteristic of the integer multiple fundamental frequency in the voiceprint pulse frequency spectrum extracted by Fourier transform, the bidirectional mapping table is a lookup table recording the correspondence between harmonic components and pet body action patterns, the emotional label is the classification result of the pet emotional state, and the confidence is the credibility of the current emotional label.

[0180] In the embodiments of the present application, the frequency spectrum of the voiceprint pulse interval is analyzed by harmonic analysis, and the three harmonic components with the highest energy proportion and their frequency values are extracted. The pre-constructed bidirectional mapping table is queried to match the harmonic components with the body action patterns recorded in the historical data (such as high-frequency harmonic associated with ear posterior sticking action). When the emotional state determination coefficient exceeds the dynamic baseline, the confidence of the emotional label is adjusted according to the context-related features extracted from the natural language description and the harmonic energy distribution ratio, for example, when the harmonic energy proportion exceeds half, the confidence is increased by 0.1.

[0181] 605、when the modified emotional label confidence exceeds the preset threshold, match the interactive instruction template positively correlated with the emotional intensity level from the emotional pacification strategy library, and perform semantic alignment between the interactive instruction template and the context of the translation instruction to output the translation instruction with additional emotional pacification suggestions.

[0182] In step 605, the emotional pacification suggestion is a set of intervention strategies matched with the modified emotional label, and the additional process is the logical integration of the suggestion text and the path prompt in the translation instruction. The emotional pacification strategy library is a set of pre-defined interactive instruction templates designed to help alleviate the negative emotions of pets. Context alignment refers to matching the content of the interactive instruction template with the background information of the translation instruction to ensure the relevance and applicability of the suggestions.

[0183] In the embodiments of the present application, once the modified emotional label confidence reaches the preset threshold, the system selects the interactive instruction template with the highest positive rate of historical interaction feedback from the emotional pacification strategy library, and performs semantic alignment between it and the current context information to output the suggestion text with additional emotional pacification suggestions. The suggestion text is inserted into the specified position of the translation instruction, for example, the pacification action description is appended after the path prompt to ensure semantic coherence. For example, the "fear" label matches the "call the pet's name softly and keep the leash relaxed" suggestion, which is appended to the end of the instruction.

[0184] The following is a specific example:

[0185] In the park trail scene, the system detects that the pet's torso contraction amplitude in the densely populated area jumps from 5 cm to 15 cm, and the motion frequency decreases from 5 times per second to 2 times per second, generating a high-density trajectory point set. The energy integral value of the continuous whistle soundprint pulse is extracted synchronously, and the weighted calculation determines the coefficient to be 0.75. Soundprint spectral analysis shows that the five-hundred-hertz harmonic is dominant, and the bidirectional mapping table is associated with the "fear" label and the confidence is corrected from 0.7 to 0.8. Finally, the translation instruction is generated, such as "detecting a crowd ahead, suggesting a left turn to bypass, please calm down and relax the leash".

[0186] In summary, steps 601 to 605 quantize the limb movement features in a two-dimensional feature space, and combine the weighted coupling analysis of soundprint pulse energy and motion frequency to achieve fine division of emotion intensity levels. Based on the dynamic confidence correction mechanism of frequency domain harmonic components and bidirectional mapping table, the reliability of emotion labels is improved. The system accurately identifies the fear state of pets caused by complex interference, generates instructions that integrate obstacle avoidance paths and personalized calming strategies, significantly enhances the nature and effectiveness of human-pet interaction, and ultimately reduces the risk of pet stress and loss of control and improves outdoor activity safety.

[0187] To solve the problems of accuracy and timing relevance of environmental sudden disturbance event detection, in some embodiments, the step 401 of time-series accumulation of parameters in the parameter set whose fluctuation amplitude exceeds the preset threshold to generate a parameter sequence reflecting the environmental sudden disturbance event includes:

[0188] 701. Establish a dynamic fluctuation threshold in time slices, based on the historical change trend of noise decibels in the parameter set over time and the historical distribution change of crowd density in space, set the range exceeding twice the standard deviation of the median of the historical change trend or historical distribution change as the initial trigger boundary;

[0189] In step 701, the time slice is a fixed-length analysis unit that divides the continuous time axis for discrete processing of time-series data. The dynamic fluctuation threshold is a parameter anomaly judgment boundary that changes with time slices, dynamically adjusted according to historical data to adapt to environmental changes. The historical change trend of noise decibels over time refers to the statistical distribution range in the historical time period, including statistical quantities such as maximum, minimum, and median. Twice the standard deviation of the median is a statistical fluctuation range calculated based on historical data to exclude occasional noise interference. The initial trigger boundary is the parameter value range exceeding twice the standard deviation of the median, which is the starting condition for disturbance event detection. The historical distribution change of crowd density in space refers to the change of the number of people in a certain area in the past.

[0190] In the embodiments of the present application, first, the noise level and crowd density data in the environment within a period of time are collected. Then, the historical change trend and distribution of these parameters are calculated using statistical methods, especially the median and standard deviation. Next, a dynamic fluctuation threshold is set in time slices, which is based on the above calculation results, and the range exceeding twice the standard deviation of the median of the historical change trend or historical distribution change is set as the initial trigger boundary. When the monitored data exceeds this boundary, it is considered that a significant fluctuation has occurred.

[0191] 702、Calculate the parameter value fluctuation difference of adjacent time slices, and when the difference direction of three consecutive time slices is consistent and the fluctuation amplitude exceeds the initial trigger boundary, lock the current time slice to the end marker of the interference event;

[0192] In step 702, the parameter value fluctuation difference is the numerical increase and decrease of the same parameter in adjacent time slices, reflecting the instantaneous intensity of parameter change. The consistency of the difference direction means that the parameter change direction of multiple consecutive time slices is the same, such as continuous rise or fall. The end marker of the interference event is the time point when the parameter value falls below the trigger boundary, which is used to define the time span of the event.

[0193] In the embodiments of the present application, the parameter difference values of adjacent time slices are calculated in time slice order, and whether the difference direction of three consecutive time slices is consistent and the amplitude exceeds the initial trigger boundary is detected. If the conditions are met, the event start time slice is locked, and continuous monitoring is performed until the parameter value falls below the boundary, and the end time slice is recorded to determine the event period. For example, if the noise decibel values of three consecutive time slices (each 5 seconds) are 75 decibels, 85 decibels, and 95 decibels, the difference direction is positive and exceeds the trigger boundary (80 decibels), it is determined that the conditions are met.

[0194] 703、According to the locked time slice, extract the frequency energy sudden increase interval of noise decibels and the mutation direction of crowd density spatial gradient, and divide the overlapping area in the time dimension into independent interference segments, and accumulate the interference segments with time continuity or spatial coverage overlap;

[0195] In step 703, the frequency energy sudden increase interval is a set of frequency bands in the acoustic signal whose energy is significantly higher than the fundamental frequency, reflecting a specific type of acoustic interference. The crowd density spatial gradient is the rate and direction of change of crowd density per unit distance, representing the diffusion characteristics of crowd movement. The time dimension overlapping area is the interval where the frequency energy sudden increase and the crowd density mutation completely or partially coincide in the time axis. The independent interference segment is the minimum continuous time period divided from the overlapping area, which has the characteristics of time continuity or spatial coverage overlap.

[0196] In the embodiments of the present application, the time slice of the lock is subjected to fast Fourier transform, and the frequency band and time period of the sudden increase of frequency domain energy are extracted. The gradient change direction and rate of the crowd density in space are calculated synchronously, the acoustic sudden increase time period and the crowd diffusion time period are aligned on the time axis, and the interval completely overlapping the two is intercepted as an independent interference segment. For example, if the acoustic sudden increase time period is 10:10-10:30 and the crowd diffusion time period is 10:15-10:40, the overlapping interval of 10:15-10:30 is intercepted as an independent interference segment (15 minutes long).

[0197] 704, the accumulated interference segment is subjected to parameter weighting, wherein the duration of the frequency domain energy sudden increase interval and the spatial diffusion speed of the crowd density spatial gradient are taken as weight factors, and a parameter sequence reflecting the environmental sudden interference event is generated.

[0198] In step 704, the duration weight factor is the importance coefficient of the total time span of the frequency domain energy sudden increase interval in the weighting calculation, reflecting the influence strength in the time dimension. The spatial diffusion speed weight factor is the importance coefficient of the change rate in the direction of the crowd density gradient in the weighting calculation, reflecting the influence range in the space dimension. The parameter sequence is a time-sequenced feature set recording the starting time, peak intensity and spatial range of the interference event, used to quantify the comprehensive intensity of the event.

[0199] In the embodiments of the present application, the weight coefficients of the duration and the spatial diffusion speed are assigned according to the preset rules, for example, the duration weight is 0.6 and the diffusion speed weight is 0.4. The duration and the spatial diffusion speed of the independent interference segment are calculated to generate a weighted value 9.2. The event starting time, the weighted peak intensity (9.2) and the spatial influence range (eastward diffusion 50 meters) are encoded into a structured parameter sequence, for example, “starting time 10:15, peak intensity 9.2, influence range eastward 50 meters”. The parameter sequence is stored in association with the environmental stress feature vector as the input of the subsequent stress level prediction model.

[0200] The following is a specific example:

[0201] In the park path scene, the noise decibel trigger boundary is set to 40 decibels to 80 decibels based on historical data, and the crowd density boundary is set to 1 person per square meter to 3 persons per square meter. When it is detected that the obstacle density in a certain area increases from 0.3 to 0.8 within 10 seconds, the time slice is shortened to 5 seconds to improve the detection sensitivity. The overlapping interval of the acoustic sudden increase time period and the crowd diffusion time period is intercepted as an independent interference segment. The weighted value is calculated according to the duration weight 0.6 and the diffusion speed weight 0.4, and the parameter sequence “peak intensity 9.2, influence range eastward 50 meters” is generated, which is used for subsequent path planning and stress level prediction.

[0202] In summary, steps 701 to 704 accurately capture real interference events and filter incidental noise through dynamic threshold and continuous triggering mechanism. Based on frequency domain and spatial multi-dimensional feature segmentation and weighted aggregation, the intensity quantification and spatio-temporal representation of complex interference events are realized. By correlating the acoustic surge and crowd diffusion, a high-discrimination parameter sequence is generated, providing reliable input for environmental stress feature fusion, significantly improving the accuracy of stress level prediction and the performance of path planning, and ultimately enhancing the overall performance of pet behavior management in complex environments.

[0203] To solve the problem of matching accuracy decline caused by dynamic environmental interference in the pet behavior recognition system, in some embodiments, the step 502 calculates the matching degree of the limb movement parameters in the current analysis window with the standard behavior mode in the pet limb language database based on the movement change features of adjacent time in the limb movement time series data, including:

[0204] 801. Decompose the continuous change sequence of limb swing frequency, trunk contraction amplitude, and head deflection angle in the limb movement time series data into movement trajectories of multiple joint nodes, and extract the movement change features of each joint node at adjacent time;

[0205] In step 801, the movement trajectory refers to converting continuous limb movement data into displacement sequences of multiple joint nodes. The limb swing frequency is the number of periodic limb movements per unit time, the trunk contraction amplitude represents the range of length change of the trunk, and the head deflection angle describes the three-axis rotation angle of the head coordinate system relative to the trunk coordinate system. The joint node refers to the motion connection point with degrees of freedom in the pet three-dimensional skeletal model. The continuous change sequence is a sequence composed of continuous data points of limb movement parameters arranged in chronological order, representing the trend of parameter change over time.

[0206] In the embodiments of the present application, a pet three-dimensional motion model is first established through skeletal binding technology, and the original data collected by the accelerometer and gyroscope are mapped to twelve main joint nodes. The sliding window method is used for time domain segmentation of the continuous change sequence, the first-order difference algorithm is used to calculate the displacement vector of each joint node at adjacent time, and the coordinate transformation is used to eliminate the error caused by the installation position difference of the inertial measurement unit. Finally, the movement change features of each joint node at adjacent time are output.

[0207] 802. Based on the gradient direction of the spatial density change in the obstacle distribution map, dynamically adjust the weight distribution of the joint nodes decomposed by the movement trajectory, and calculate the offset of the joint nodes of the limb parts corresponding to the high-density area;

[0208] In step 802, the obstacle distribution map is a two-dimensional gridded environment feature map constructed by a laser radar, and the spatial density variation gradient represents the change rate of the number of obstacles per unit distance. The joint weight distribution refers to adjusting the calculation coefficient according to the influence degree of the environment region on the limb movement. The offset reflects the coefficient of the degree to which the joint movement trajectory is affected by the environment. The gradient direction refers to the change direction of the obstacle density in space, which is the vector direction from the high-density area to the low-density area.

[0209] In the embodiments of the present application, the simultaneous localization and mapping technology is used to update the obstacle distribution map, and the obstacle density gradient vector of the region corresponding to the current window is calculated. According to the included angle between the gradient direction and the pet movement direction, the weight coefficient of the limb (such as the front leg joint) in the front side of the movement direction is dynamically adjusted, and the offset of the joint of the limb part is calculated. When a high-density obstacle area is detected in front, the weight of the front limb joint is increased to 2 times of the reference value, and the weight of the rear limb joint is reduced to 0.7 times of the reference value. The environmental adaptability is enhanced through the weighted attention mechanism.

[0210] 803. According to the calculation result of the offset, the included angle cosine value of the motion change characteristic and the reference trajectory corresponding to the standard behavior mode in the analysis window is calculated, and the included angle cosine value is dynamically threshold truncated in combination with the time distribution characteristics of the visual interference parameter in the environmental pressure feature vector;

[0211] In step 803, the reference trajectory is the theoretical movement path of the standard behavior mode in three-dimensional space, and the included angle cosine value represents the direction similarity between the actual movement trajectory and the standard trajectory. The visual interference parameter includes the light intensity change rate, the motion blur degree and other environmental factors that affect the visual feature extraction. The analysis window is a fixed length data slice for local feature extraction, and the length is determined by the minimum period of the behavior mode. The time distribution characteristics refer to the change law of the visual interference parameter in the time dimension.

[0212] In the embodiments of the present application, the time axes of the actual trajectory and the reference trajectory are aligned through the dynamic time warping algorithm, and the movement direction vector of each joint in each frame of data is calculated. The cosine similarity algorithm is used to obtain the cosine value of the included angle between the vectors, and the current light sensor data and the image definition score of the visual interference parameter are combined to construct a threshold function that changes with time. When the light mutation exceeds the threshold, the abnormal data caused by overexposure or shadow is eliminated through adaptive filtering.

[0213] 804. The truncated included angle cosine value is weighted and summed according to the joint weight to generate a local matching degree, and the local matching degree is phase compensated based on the periodicity difference between the limb swing frequency in the analysis window and the standard behavior mode in the pet limb language database to obtain the matching degree of the current analysis window.

[0214] In step 804, the phase compensation refers to the time domain correction of the matching degree according to the motion frequency difference, and the periodic difference is obtained by calculating the fundamental frequency component extracted by Fourier transform. Weighted summation refers to the operation mode of linear superposition of local similarity according to the joint importance coefficient. The standard behavior mode refers to the typical pet behavior three-dimensional motion trajectory and frequency feature template pre-stored in the database.

[0215] In the embodiments of the present application, the weighted summation formula is used to aggregate the corrected similarity of 12 joint nodes, and the weight coefficient comes from the dynamic allocation result of step 802. The limb swing frequency of the current window is subjected to fast Fourier transform, and the phase shift is obtained by cross-correlation operation with the standard frequency of the database. The local matching degree is aligned in phase using the linear interpolation method, and the standardized matching degree in the range of 0 to 1 is finally output.

[0216] The following is a specific example:

[0217] In the park path, the pet encounters the alternation of stone road and lawn terrain when running in the morning. When entering the high-density obstacle area of the stone road, the system detects that the front limb joint weight is automatically increased. At this time, the pet's front leg appears to avoid the high action of the stone, and the cosine value of the front limb trajectory of the "alert walking" mode in the database reaches 0.7, but the visual interference parameter trigger threshold is down-regulated to 0.65 due to the influence of tree shadow shaking. After phase compensation correction, the final matching degree reaches 0.85, and the alert behavior is accurately recognized. The swing frequency is detected on the lawn, and the phase difference of the "happy running" mode is 0.3, and the matching degree is improved to 0.92 after compensation.

[0218] In summary, steps 801 to 804 significantly improve the behavior recognition accuracy in complex environments through multi-dimensional feature fusion and dynamic parameter adjustment. First, the joint trajectory decomposition combined with the weight distribution mechanism of environment perception effectively highlights the key motion features affected by the terrain. Second, the phase compensation technique eliminates the system error caused by the difference in motion rhythm, ensuring the consistency of recognition under different motion speeds. At the same time, the weight distribution and threshold adjustment mechanism supports the feature adaptation of pets of different sizes and motion habits, significantly improves the generalization ability, and meets the interaction needs in complex dynamic scenes.

[0219] In order to further improve the prediction accuracy of the stress level of pets in complex outdoor environments, in some embodiments, the association mapping of the environmental stress feature vector and the behavior abnormal segment in step 104 to determine the stress level prediction result of the target pet in the complex outdoor environment includes:

[0220] 901, construct the spatio-temporal alignment relationship between the environmental stress feature vector and the behavior abnormal segment, and extract the joint feature set of the environmental stress feature vector and the behavior abnormal segment in the overlapping time window;

[0221] In step 901, the environmental stress feature vector contains various physical and chemical parameters that may trigger stress in the target pet, such as temperature, humidity, noise level, air quality, etc. The behavior abnormality segment refers to a time period that is significantly different from normal behavior identified by analyzing the pet's behavior patterns. The spatio-temporal alignment relationship refers to matching these environmental data and behavior abnormality segments in chronological order. The overlapping time window refers to selecting time periods where both environmental changes and behavior abnormalities exist to perform more accurate analysis.

[0222] In the embodiments of the present application, first, data from different sensors are collected to construct the environmental stress feature vector, and video analysis techniques are used to identify and extract behavior abnormality segments of the pet. Then, the overlapping parts of the two in the same time window are determined by an algorithm, and a joint feature set is extracted from them, which may include but is not limited to the rate of change of environmental factors and the response speed of pet behavior, etc. The final result is to obtain a feature set that can reflect the interaction between the environment and the pet behavior in a specific time period, for example, the sudden change of air temperature in a specific period and the pet's restless movements.

[0223] 902, performing nonlinear correlation analysis on the joint feature set to capture the coupling relationship between the environmental stress feature vector and the behavior abnormality segment through a multi-layer feature interaction network, and generating a coupling strength coefficient representing the correlation strength between the environment and the behavior;

[0224] In step 902, the joint feature set is composed of data obtained in the previous step, revealing the potential relationship between environmental factors and pet behavior. Nonlinear correlation analysis is a statistical method used to explore the complex interaction between variables. Multi-layer feature interaction network is a deep learning architecture specially designed to handle deep interactions between high-dimensional data. The coupling strength coefficient is a quantitative indicator generated in this process, which represents the closeness between environmental factors and pet behavior.

[0225] In the embodiments of the present application, a pre-trained multi-layer feature interaction network is used to process the joint feature set. The network adjusts its internal parameters through learning from a large amount of historical data, so that it can accurately calculate the coupling strength coefficient. This process involves adjusting the weights of the network to minimize the prediction error, and evaluating the impact of environmental changes on pet behavior according to the output coupling strength coefficient.

[0226] 903, based on the numerical distribution range of the coupling strength coefficient, dividing the stress level threshold interval of the target pet in the outdoor complex environment;

[0227] In step 903, the numerical distribution range of the coupling strength coefficient represents the value interval of all calculated coupling strength coefficients. The stress level threshold interval is set based on this numerical distribution, aiming to distinguish different degrees of stress state according to different coupling strength coefficients.

[0228] In the embodiments of the present application, first, statistical analysis is performed on all collected coupling strength coefficients to understand their distribution characteristics. By applying a clustering algorithm, the data is divided into different groups according to the numerical value of the coupling strength coefficient. Each group represents the possibility of a different degree of stress reaction. Then, according to the results of these groupings, specific stress level threshold intervals are set. For example, lower coupling strength coefficients may correspond to low stress levels, while higher coefficients indicate high stress levels.

[0229] 904、According to the falling position of the coupling strength coefficient in the stress level threshold interval, the confidence weight of the stress level is dynamically allocated, and the confidence weight is corrected in combination with the same environmental pressure in the historical stress event, to obtain a stress level prediction result.

[0230] In step 904, the confidence weight is a parameter that measures the credibility of the current coupling strength coefficient falling within a certain stress level interval. The same environmental pressure in the historical stress event means that the pet stress cases that occurred in similar environments in the past are referred to adjust the current confidence weight.

[0231] In the embodiments of the present application, first, according to the stress level threshold interval obtained in the previous step, it is judged which interval the current coupling strength coefficient falls into, and a confidence weight is preliminarily allocated accordingly. Then, the system searches the historical database for records of stress events that have occurred in similar environments in the past. By analyzing these records, the similarity between the current situation and the historical cases is evaluated, and the initially allocated confidence weight is adjusted accordingly. Finally, the weighted average value obtained after considering the above factors is the final stress level prediction result.

[0232] The following is a specific example:

[0233] In the park trail scene, the environmental pressure feature vector shows that the noise decibel gradient is 25 decibels per second, the obstacle density change rate is 0.5 per second, the limb swing frequency offset in the behavior abnormality segment is -3 hertz, and the trunk contraction amplitude difference is 6 centimeters in the 10:10-10:25 period. The coupling strength coefficient generated by the multi-layer feature interaction network is 0.65. According to the historical threshold interval, it is divided into the medium stress level (0.3-0.7). In combination with the historical co-occurrence frequency of 80%, the confidence weight is corrected to 0.78, and the output is "medium stress level, confidence 0.78".

[0234] In summary, steps 901 to 904 accurately quantify the coupling relationship between environmental stress and behavioral abnormalities through spatio-temporal alignment and nonlinear correlation analysis, avoiding the limitations of single parameter analysis. Dynamic threshold division combined with confidence correction mechanism optimizes the reliability of prediction results by combining historical data. The system effectively identifies the joint influence of noise and obstacle density on pet contraction behavior, outputs stress level and high confidence weight, and provides high-precision input for subsequent generation of personalized path prompts and soothing strategies.

[0235] Figure 2 A structure diagram of a pet state intelligent translation system based on multi-modal data is provided for the embodiments of the present application, as shown in Figure 2 The system comprises:

[0236] The acquisition module 21 acquires multi-modal data of the target pet in an outdoor complex environment, wherein the multi-modal data includes noise decibel, crowd density, and three-dimensional space data captured by a laser radar.

[0237] The encoding module 22 generates an obstacle distribution atlas based on the three-dimensional space data, and jointly encodes the multi-modal data and the obstacle distribution atlas to generate an environmental stress feature vector.

[0238] The extraction module 23 synchronously acquires limb action time series data of the target pet, and performs pattern matching on the limb action time series data and a preset pet body language database to extract a behavioral abnormality segment that has a spatio-temporal correlation with the outdoor complex environment.

[0239] The mapping module 24 determines a stress level prediction result of the target pet in the outdoor complex environment through the association mapping of the environmental stress feature vector and the behavioral abnormality segment.

[0240] The generation module 25 generates a translation instruction according to the stress level prediction result, wherein the translation instruction includes an evasive path prompt for a high-density area in the obstacle distribution atlas.

[0241] Figure 2 The pet state intelligent translation system based on multi-modal data can perform Figure 1 The pet state intelligent translation method based on multi-modal data described in the embodiments shown in the above has the same implementation principles and technical effects, which will not be described in detail. The specific operation of each module and unit of the pet state intelligent translation system based on multi-modal data in the above embodiments has been described in detail in the embodiments related to the method, which will not be described in detail.

[0242] In one possible design, Figure 2The pet state intelligent translation system based on multi-modal data of the illustrated embodiment can be implemented as a computing device, such as Figure 3 As shown, the computing device can include a storage component 31 and a processing component 32.

[0243] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32.

[0244] The processing component 32 is configured to perform the above Figure 1 The pet state intelligent translation method based on multi-modal data of the embodiment.

[0245] The processing component 32 can include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component can also be one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components, for executing the above method.

[0246] The storage component 31 is configured to store various types of data to support the operation of the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0247] Of course, the computing device can also include other components, such as input / output interfaces, display components, communication components, etc.

[0248] The input / output interface provides an interface between the processing component and the peripheral interface module, which can be an output device, an input device, etc.

[0249] The communication component is configured to facilitate wired or wireless communication between the computing device and other devices, etc.

[0250] The computing device can be a physical device or an elastic computing host provided by a cloud computing platform, and the computing device can be a cloud server, and the processing component, the storage component, etc. can be basic server resources rented or purchased from the cloud computing platform.

[0251] The computer storage medium provided by the embodiment of the present application stores a computer program, and the computer program can implement the aboveFigure 1 An intelligent pet state translation method based on multi-modal data of the illustrated embodiment.

[0252] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.

[0253] The device embodiments described above are only schematic, and the units illustrated as separate components can or can not be physically separated, and the components illustrated as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement it without creative labor.

[0254] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course, they can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software products, which can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including a number of instructions to make a computer device (which can be a personal computer, server, or network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0255] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A pet state intelligent translation method based on multi-modal data, characterized in that, The method comprises the following steps: acquiring multi-modal data of an outdoor complex environment where a target pet is located, the multi-modal data including noise decibel, crowd density, and three-dimensional space data captured by a laser radar; generating an obstacle distribution map based on the three-dimensional space data, and jointly encoding the multi-modal data and the obstacle distribution map to generate an environmental stress feature vector; synchronously collecting limb action time series data of the target pet, and performing pattern matching on the limb action time series data and a preset pet body language database to extract a behavior abnormality segment having a space-time correlation with the outdoor complex environment; determining a stress level prediction result of the target pet in the outdoor complex environment through associated mapping of the environmental stress feature vector and the behavior abnormality segment; generating a translation instruction according to the stress level prediction result, the translation instruction including an evasive path prompt for a high-density area in the obstacle distribution map; wherein the generation of the obstacle distribution map based on the three-dimensional space data comprises: dividing a plurality of spatial grid units in a corresponding activity range of the target pet based on the three-dimensional space data, extracting a moving track of an obstacle in each spatial grid unit and an occurrence frequency of the obstacle in a preset time period, and calculating a moving obstacle density and a space occlusion coefficient of each spatial grid unit in combination with a line-of-sight occlusion degree change between adjacent spatial grid units; mapping the moving obstacle density and the space occlusion coefficient of each spatial grid unit into a three-dimensional heat map layer to generate the obstacle distribution map by superposition; wherein the joint encoding of the multi-modal data and the obstacle distribution map to generate the environmental stress feature vector comprises: extracting a parameter set in the multi-modal data that overlaps with a high-density area in the obstacle distribution map according to a change trend of noise decibel over time and a distribution change of crowd density in space, performing time series accumulation on parameters in the parameter set whose fluctuation amplitude exceeds a preset threshold to generate a parameter sequence reflecting an environmental sudden disturbance event; performing associated matching on the moving obstacle density, the space occlusion coefficient, and the parameter sequence to determine a diffusion path of the environmental sudden disturbance event in the obstacle distribution map, and adjusting the moving obstacle density and the space occlusion coefficient of each spatial grid unit according to the diffusion path; based on the time series accumulation results of the adjusted moving obstacle density, the space occlusion coefficient, and the parameter sequence, tensor splicing is performed on the time series fluctuation features of the multi-modal data and the spatial topology features of the obstacle distribution map, and a convolutional neural network is used to encode the splicing result to generate the environmental stress feature vector. 2.The pet status intelligent translation method based on multi-modal data according to claim 1, characterized in that, After the generation of the translation instruction according to the stress level prediction result, the method further comprises: constructing a base large model based on the multi-modal data, the base large model including cross-modal analysis results of human voice, environmental sound, and animal sound; jointly labeling the cross-modal analysis results of the animal sound according to the environmental stress feature vector and the behavior abnormality segment to obtain a labeling result, the labeling result including a mapping relationship among dog breed category, emotion type, and environmental stress; Mel frequency spectrum features of the labeled results are extracted from the base large model to obtain spectrum vectors, and the spectrum vectors are aligned with the environmental stress feature vectors, parameters of the base large model are adjusted based on the alignment results, and an adjusted base large model is obtained; Dog language reasoning logic is constructed based on the adjusted base large model, and the language generation strategy of the translation instruction is adjusted using the dog language reasoning logic to generate a natural language description matching the personality characteristics of the target pet; Based on the natural language description, the emotional state of the target pet is identified in combination with the limb movement time series data and the noise decibel, and an emotional pacification suggestion corresponding to the emotional state is added to the translation instruction; The translation instruction with the emotional pacification suggestion is pushed to the user terminal, and the evasive path prompt in the translation instruction is dynamically updated according to the change of the high-density area in the obstacle distribution map.

3. The method of claim 1, wherein, The limb movement time series data includes continuous change sequences of limb swing frequency, trunk contraction amplitude, and head deflection angle; The limb movement time series data is matched with the preset pet body language database to extract behavior abnormal segments with spatio-temporal correlation in the outdoor complex environment, including: The limb movement time series data is divided into variable-length analysis windows, and the sliding interval of the analysis window is dynamically adjusted according to the spatial density change of the obstacle distribution map in the outdoor complex environment; In each analysis window, based on the action change characteristics of adjacent time points in the limb movement time series data, the matching degree of the limb movement parameters in the current analysis window with the standard behavior mode in the pet body language database is calculated, and the dynamic matching condition is generated in combination with the parameters related to visual interference in the environmental stress feature vector; According to the dynamic matching condition, the analysis window with a matching degree lower than a preset reference is filtered, the periodic difference between the limb swing frequency and the trunk contraction amplitude in the analysis window is extracted, and the position-related behavior abnormal segment in the outdoor complex environment is marked in combination with the position information of the high-density area in the obstacle distribution map; The behavior abnormal segment is associated and compared with the spatial distribution change of the crowd density in the multi-modal data to exclude non-stress action data generated by pet autonomous behavior, and a set of spatio-temporal correlated behavior abnormal segments is generated.

4. The method of claim 2, wherein, Based on the natural language description, in combination with the limb movement time series data and the noise decibel, the emotional state of the target pet is identified, and an emotional pacification suggestion corresponding to the emotional state is added to the translation instruction, including: The natural language description is processed by word segmentation to extract context-related features and generate an emotional feature vector, the emotional feature vector includes word segmentation weight distribution and semantic dependency strength; A two-dimensional feature space with limb movement amplitude and movement frequency as axes is established, a dynamic trajectory point set is constructed based on the emotional feature vector and the limb movement time series data of the target pet within a set time window, and the emotional intensity level of the emotional state is divided by the distribution density of the dynamic trajectory point set in the two-dimensional feature space; extracting a voiceprint pulse interval with a duration exceeding a threshold value in a noise decibel sequence, coupling an energy integral value of each voiceprint pulse interval with a motion frequency of the dynamic trajectory point set to generate an emotional state judgment coefficient; constructing a bidirectional mapping table of emotional labels and body action patterns based on a frequency domain harmonic component of the voiceprint pulse interval, and when the emotional state judgment coefficient exceeds a dynamic baseline, using the context-related features to correct a confidence level of an emotional label corresponding to a current body action pattern; when the corrected emotional label confidence level exceeds a preset threshold value, matching an interactive instruction template positively correlated with the emotional intensity level from an emotional pacification strategy library, and performing semantic alignment between the interactive instruction template and a context of a translation instruction to output a translation instruction with an additional emotional pacification suggestion.

5. The method of claim 1, wherein, The parameters in the parameter set that fluctuate by more than a preset threshold value are time-series accumulated to generate a parameter sequence reflecting environmental sudden disturbance events, including: establishing a dynamic fluctuation threshold in time slices, setting a range exceeding twice the standard deviation of the median of the historical change trend of noise decibels over time and the historical distribution change of crowd density in space as the initial trigger boundary based on the parameter set; calculating the fluctuation difference of parameter values of adjacent time slices, and when the difference directions of three consecutive time slices are consistent and the fluctuation amplitude exceeds the initial trigger boundary, locking the current time slice to the end marker of the disturbance event; extracting the frequency energy sudden increase interval of noise decibels and the mutation direction of crowd density spatial gradient according to the locked time slice, and dividing the overlapping region in the time dimension into independent disturbance segments, and accumulating disturbance segments with time continuity or spatial coverage overlap; parameter weighting is performed on the accumulated disturbance segments, wherein the duration of the frequency energy sudden increase interval and the spatial diffusion speed of the crowd density spatial gradient are used as weight factors to generate a parameter sequence reflecting environmental sudden disturbance events.

6. The method of claim 3, wherein, Based on the action change characteristics of adjacent time points in the body action time series data, the matching degree of the body action parameters in the current analysis window with the standard behavior patterns in the pet body language database is calculated, including: The continuous change sequence of the body swing frequency, trunk contraction amplitude, and head deflection angle in the body action time series data is decomposed into multi-joint action trajectories, and the action change characteristics of each joint at adjacent time points are extracted; Based on the gradient direction of the spatial density change in the obstacle distribution map, the joint weight distribution of the action trajectory decomposition is dynamically adjusted, and the offset of the joint of the body part corresponding to the high-density area is calculated; According to the offset calculation result, the analysis window is used to calculate the cosine value of the included angle between the action change characteristics and the reference trajectory corresponding to the standard behavior pattern, and the cosine value is dynamically threshold truncated in combination with the time distribution characteristics of the visual interference parameter in the environmental stress feature vector. The truncated included angle cosine value is weighted and summed according to the weight of the node to generate a local matching degree, and the local matching degree is phase compensated based on the difference in periodicity between the limb swing frequency in the analysis window and the standard behavior pattern in the pet limb language database, to obtain a matching degree of the current analysis window.

7. The method of claim 1, wherein, The stress level prediction result of the target pet in the outdoor complex environment is determined through the association mapping of the environmental stress feature vector and the behavior abnormal segment, including: The spatio-temporal alignment relationship between the environmental stress feature vector and the behavior abnormal segment is constructed, and a joint feature set of the environmental stress feature vector and the behavior abnormal segment in the overlapping time window is extracted; Nonlinear correlation analysis is performed on the joint feature set, and the coupling relationship between the environmental stress feature vector and the behavior abnormal segment is captured through a multi-layer feature interaction network to generate a coupling strength coefficient representing the association strength between the environment and the behavior; Based on the numerical distribution range of the coupling strength coefficient, the stress level threshold interval of the target pet in the outdoor complex environment is divided; According to the falling point position of the coupling strength coefficient in the stress level threshold interval, the confidence weight of the stress level is dynamically allocated, and the confidence weight is corrected in combination with the same environmental stress in the historical stress event to obtain the stress level prediction result.

8. A pet state intelligent translation system based on multi-modal data, configured to perform the pet state intelligent translation method based on multi-modal data according to any one of claims 1-7. It includes: An acquisition module acquires multi-modal data of an outdoor complex environment where a target pet is located, and the multi-modal data includes noise decibels, crowd density, and three-dimensional space data captured by a laser radar; An encoding module generates an obstacle distribution atlas based on the three-dimensional space data, and jointly encodes the multi-modal data and the obstacle distribution atlas to generate an environmental stress feature vector; An extraction module synchronously collects limb action time series data of the target pet, and performs pattern matching on the limb action time series data and a preset pet limb language database to extract a behavior abnormal segment that has a spatio-temporal association with the outdoor complex environment; A mapping module determines a stress level prediction result of the target pet in the outdoor complex environment through the association mapping of the environmental stress feature vector and the behavior abnormal segment; A generation module generates a translation instruction according to the stress level prediction result, and the translation instruction includes an evasive path prompt for a high-density area in the obstacle distribution atlas.

Citation Information

Patent Citations

  • Campus inspection robot navigation method based on large model fusion environment and biological multi-modal information

    CN118329044A

  • AI-based pet emotion recognition system

    CN119049086A