Pet health analysis method and device, electronic equipment and storage medium

By combining artificial intelligence technology and large language models with multimodal data analysis, the problem of traditional pet health monitoring being labor-intensive and inaccurate has been solved, achieving high efficiency and accuracy in pet health analysis.

CN120526457BActive Publication Date: 2026-05-05SHENZHEN LIBRO TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN LIBRO TECH CO LTD
Filing Date
2025-07-17
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Traditional pet health monitoring relies on pet owners' observation, which is time-consuming and inaccurate. Existing smart devices have limited functions and cannot comprehensively analyze pet health.

Method used

Using artificial intelligence technology, the system combines multimodal data with pet management clients and servers, employs a pre-trained large language model for pet health analysis, and integrates image, text, and behavioral data collected by smart devices for comprehensive analysis.

Benefits of technology

It improves the accuracy and reliability of pet health analysis, enabling timely detection of potential health risks and reducing reliance on the subjective judgment of pet owners.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120526457B_ABST
    Figure CN120526457B_ABST
Patent Text Reader

Abstract

This application provides a pet health analysis method, apparatus, electronic device, and storage medium. The method includes receiving a pet health analysis instruction, which includes a target pet image and target health analysis text; performing pet recognition on the target pet image to obtain the target pet; acquiring target health record information and target historical behavior data of the target pet within a preset time period; extracting image features from the target pet image to obtain target image features; and performing health analysis based on the target image features, target health analysis text, target health record information, and target historical behavior data using a pre-trained large language model to obtain the target pet's target health analysis result. This method can improve the accuracy of pet health analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, specifically to a pet health analysis method and device, electronic device and storage medium. Background Technology

[0002] In recent years, with the continuous development of technology and the accelerating pace of life, people's life pressure has been constantly increasing. The companionship of pets can greatly alleviate people's mental stress and bring them physical and mental pleasure, which has led to the rapid development of the pet supplies market in recent years.

[0003] Pet health is always a major concern for pet owners, and traditional pet health monitoring mainly relies on the daily observation of pet owners. However, this method not only consumes a lot of the pet owner's energy, but also has certain limitations due to the pet owner's limited health knowledge, resulting in relatively low accuracy in the pet's health analysis. Summary of the Invention

[0004] The main objective of this application is to provide a pet health analysis method, device, electronic device, and storage medium, aiming to improve the accuracy of pet health analysis.

[0005] To achieve the above objectives, a first aspect of this application provides a pet health analysis method, the method comprising:

[0006] Receive pet health analysis instructions, which include a target pet image and target health analysis text;

[0007] Perform pet recognition on the target pet image to obtain the target pet;

[0008] Obtain the target pet's health record information and the target pet's historical behavior data within a preset time period;

[0009] Image features are extracted from the target pet image to obtain target image features. Based on the target image features, target health analysis text, target health record information, and target historical behavior data, a pre-trained large language model is called to perform health analysis and obtain the target pet's target health analysis results.

[0010] To achieve the above objectives, a second aspect of this application provides a pet health analysis device, the device comprising:

[0011] The instruction receiving unit is used to receive pet health analysis instructions, which include a target pet image and target health analysis text.

[0012] The pet recognition unit is used to identify the target pet from the target pet image;

[0013] The data acquisition unit is used to acquire the target pet's target health record information and the target pet's target historical behavior data within a preset time period;

[0014] The health analysis unit is used to extract image features from the target pet image to obtain target image features, and then call a pre-trained large language model to perform health analysis based on the target image features, target health analysis text, target health record information, and target historical behavior data to obtain the target pet's target health analysis results.

[0015] Optionally, in some embodiments, the health analysis unit includes:

[0016] The image segmentation subunit is used to call a preset visual encoder to segment the target pet image and obtain multiple image blocks;

[0017] The first coding subunit is used to perform linear mapping on each image block to obtain an image vector, and generate a position code based on the position of the image block in the target pet image.

[0018] The feature construction subunit is used to construct the target image features based on the image vectors and positional encodings corresponding to multiple image blocks.

[0019] Optionally, in some embodiments, the health analysis unit further includes:

[0020] The second encoding subunit is used to call a preset text encoder to encode the target health analysis text, target health record information and target historical behavior data to obtain target text features;

[0021] The feature fusion subunit is used to construct target fusion features based on target image features and target text features;

[0022] The first health analysis subunit is used to call the pre-trained large language model to determine the attention weights of the target fusion features, and perform health analysis based on the attention weights and target fusion features to obtain the target pet's target health analysis results.

[0023] Optionally, in some embodiments, the pet health analysis device further includes a model training unit for training a large language model, including:

[0024] The training data acquisition subunit is used to acquire multimodal data of pets and the corresponding sample health analysis results. The multimodal data includes sample health analysis text of pets, sample pet images, sample health record information and sample historical behavior data.

[0025] The third encoding subunit is used to call the visual encoder to determine the sample image features of the sample pet image, and to call the text encoder to determine the sample text features of the sample health analysis text, sample health record information and sample historical behavior data.

[0026] The second health analysis subunit is used to construct sample fusion features based on sample image features and sample text features, and call a pre-trained visual language model to perform health analysis based on the sample fusion features to obtain the predicted health analysis results.

[0027] The fine-tuning training subunit is used to fine-tune the visual language model based on the predicted health analysis results and the sample health analysis results, and to determine the fine-tuned visual language model as the large language model.

[0028] Optionally, in some embodiments, the second health analysis subunit includes:

[0029] The first feature fusion module is used to determine the initial image weights of the sample image features and the initial text weights of the sample text features, and to construct sample fusion features based on the initial image weights, sample image features, initial text weights and sample text features.

[0030] Fine-tuning the training sub-units includes:

[0031] The fine-tuning training module is used to update the network parameters of the visual language model based on the predicted health analysis results and the sample health analysis results, and to update the initial image weights and initial text weights. The updated visual language model is determined as the large language model, the updated initial image weights are determined as the target image weights, and the updated initial text weights are determined as the target text weights.

[0032] Feature fusion subunit, including:

[0033] The second feature fusion module is used to construct target fusion features based on target image weights, target text weights, target image features, and target text features.

[0034] Optionally, in some embodiments, the data acquisition unit includes:

[0035] The feeding data acquisition subunit is used to acquire historical pet videos and historical feeding data within a preset time period from associated smart devices;

[0036] The data filtering subunit is used to select key frames that are related to the target pet image from historical pet videos, and to filter historical feeding data according to the key time period corresponding to the key frame;

[0037] The data construction subunit is used to construct target historical behavior data based on keyframes and filtered historical eating data.

[0038] Optionally, in some embodiments, the health analysis unit further includes:

[0039] The abnormal data identification subunit is used to generate statistical features based on the target's historical behavior data, and to extract abnormal data from the target's historical behavior data based on the statistical features.

[0040] The health knowledge acquisition subunit is used to acquire health knowledge from a pre-built pet health knowledge base based on abnormal data and target health analysis text.

[0041] The third health analysis subunit is used to call a large language model to perform health analysis based on health knowledge, abnormal data, target image features, target health analysis text, and target health record information, and obtain the target pet's target health analysis results.

[0042] Optionally, in some embodiments, the pet recognition unit includes:

[0043] The pet comparison subunit is used to detect pets in the target pet image and compare the detected pets with the pet profile information of registered pets;

[0044] The pet identification unit is used to identify the target pet from the detected pets based on the comparison results.

[0045] To achieve the above objectives, a third aspect of this application provides an electronic device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the pet health analysis method described in the first aspect.

[0046] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the pet health analysis method described in the first aspect.

[0047] To achieve the above objectives, a fifth aspect of this application provides a computer program product comprising a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the pet health analysis method described in the first aspect.

[0048] The pet health analysis method provided in this application embodiment receives a pet health analysis instruction, which includes a target pet image and a target health analysis text; performs pet recognition on the target pet image to obtain the target pet; extracts image features from the target pet image and obtains the target pet's target health record information and target historical behavior data of the target pet within a preset time period; and performs health analysis by calling a preset large language model based on the target health analysis text, image features, target health record information, and target historical behavior data to obtain the target pet's health analysis result.

[0049] Therefore, the pet health analysis method provided in this application embodiment can call a pre-trained large language model to perform health analysis on the target pet based on multimodal data such as target health analysis text, target image features, target health record information, and target historical behavior data, thereby improving the accuracy of pet health analysis. Attached Figure Description

[0050] The accompanying drawings are provided to further understand the technical solutions of this disclosure and constitute a part of the specification. They are used together with the embodiments of this disclosure to explain the technical solutions of this disclosure and do not constitute a limitation on the technical solutions of this disclosure.

[0051] Figure 1 A schematic diagram of the system used in the pet health analysis method provided in this application;

[0052] Figure 2 A flowchart illustrating the pet health analysis method provided in this application;

[0053] Figure 3 A schematic diagram of the model network for the pet health analysis method provided in this application;

[0054] Figure 4 A schematic diagram illustrating the pet health analysis method provided in this application;

[0055] Figure 5 A schematic diagram of the pet health analysis device provided in this application;

[0056] Figure 6 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this disclosure.

[0058] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0059] Before providing a further detailed description of the embodiments of this application, the nouns and terms used in the embodiments of this application are explained, and the nouns and terms used in the embodiments of this application shall be interpreted as follows:

[0060] Large Language Models (LLMs) are deep learning-based AI systems that analyze massive amounts of text data to grasp language patterns and generate and understand human language. The core technology of LLMs relies on the transformer architecture, utilizing self-attention mechanisms to dynamically capture semantic relationships in long texts. They are adapted to different tasks through pre-training (such as GPT's autoregressive prediction and BERT's masked language modeling) and fine-tuning. With a large number of parameters, these models can generate coherent text (such as articles and code), answer questions, or perform comprehension tasks such as summarizing and translating.

[0061] In related technologies, pet health monitoring mainly relies on pet owners' daily observations. However, this method not only consumes a lot of the pet owner's energy, but also has certain limitations due to the pet owner's limited health knowledge, and the accuracy of the pet owner's analysis of the pet's health is also relatively low.

[0062] With the development of artificial intelligence technology, intelligent devices have emerged that can analyze pet health. However, most of these devices are single-function, only able to interact with pet owners as chatbots. That is, these devices can only answer pet owners' questions about pet health, but cannot perform health analysis based on the pet's actual condition, resulting in low accuracy in pet health analysis.

[0063] Therefore, to address the issue of low accuracy in pet health analysis in related technologies, this application provides a pet health analysis method, device, electronic device, and storage medium. This method uses artificial intelligence technology combined with acquired multimodal data to perform pet health analysis. The pet health analysis method provided in this application embodiment will be described below.

[0064] Reference Figure 1This is a schematic diagram of the system used in the pet health analysis method provided in this application. As shown, the system includes at least one smart device 110, at least one pet management client 130 corresponding to the smart device, and a pet device management server 120. The smart device 110 can be a smart litter box, a smart feeder, or a smart water dispenser, etc. The pet management client 130 is a terminal that runs a pet management application. This terminal can be a mobile terminal, such as a smartphone, or a tablet, personal computer, head-mounted device, or vehicle-mounted terminal, etc. The pet management application can control smart devices such as smart litter boxes, smart feeders, and smart water dispensers. The pet device management server 120 is a server that provides management, data processing, and interaction services to multiple pet management clients. The pet device management server 120 can be a single high-performance computer, a cluster of multiple high-performance computers, a portion of a single high-performance computer (e.g., a virtual machine), or a combination of portions of multiple high-performance computers (e.g., virtual machines) in a network platform. The pet device management server 120 can be a physical server or a cloud server.

[0065] In this embodiment, the pet device management server 120 can receive pet health analysis instructions sent by the pet management client 130. These instructions may include a target pet image and target health analysis text. The pet device management server 120 can perform pet recognition on the target pet image to obtain the target pet. Furthermore, the pet device management server 120 can also obtain the target pet's health record and retrieve the target pet's historical behavior data within a preset time period from the smart device 110. The pet device management server 120 can extract image features from the target pet image to obtain target image features, and based on the target image features, target health analysis text, target health record information, and target historical behavior data, it can call a pre-trained large language model to perform health analysis, obtaining the target pet's target health analysis result.

[0066] Reference Figure 2 This is a flowchart illustrating the pet health analysis method provided in this application.

[0067] In some embodiments, the pet health analysis method provided in this application includes, but is not limited to, steps S201 to S204.

[0068] Step S201: Receive pet health analysis command;

[0069] Step S202: Perform pet recognition on the target pet image to obtain the target pet;

[0070] Step S203: Obtain the target pet's target health record information and the target pet's target historical behavior data within a preset time period;

[0071] Step S204: Extract image features from the target pet image to obtain target image features, and use a pre-trained large language model to perform health analysis based on the target image features, target health analysis text, target health record information, and target historical behavior data to obtain the target pet's target health analysis results.

[0072] In steps S201 to S204 of this embodiment, the pet management client can respond to the pet owner's relevant operations and generate corresponding pet health analysis instructions. The pet health analysis instructions include a target pet image and target health analysis text. For example, the pet management client can display a pet health analysis control in the interface corresponding to the pet management application. The pet management client can respond to the pet owner's touch operation on the pet health analysis control and generate a pet health analysis instruction. At this time, the target pet image corresponding to the pet health analysis instruction can be an image automatically obtained by the pet management application (e.g., it can be automatically obtained from images pre-uploaded by the pet owner, thus reducing the need for the pet owner to repeatedly upload images). The target health analysis text can be a template text pre-set by the pet management application (e.g., "Please perform a health analysis on the pet in the picture"), or text pre-set by the pet owner in the corresponding interface.

[0073] Alternatively, pet management applications can have a pet health Q&A function. The pet management client can respond to the target pet image and target health analysis text uploaded by the pet owner on the pet health Q&A interface, and generate corresponding pet health analysis instructions. In this case, the target health analysis text can include specific health inquiries, such as "Pet A is X months old, has recently experienced decreased appetite and reduced urine output. Please analyze whether pet A has any health problems." After receiving the pet health analysis instructions sent by the pet management client, the pet device management server can perform pet recognition on the target pet image corresponding to the instructions to identify the pet to be analyzed (i.e., the target pet). In addition, the pet device management server can also obtain the target pet's target health record information, as well as the target's historical behavioral data collected by the smart device within a preset time period. The target health record information can include the target pet's basic information (such as breed, age, sex, and weight), past medical history, vaccination records, and previous physical examination and diagnostic results. This data can be provided by the pet owner during the initial pet registration or veterinary treatment process, and the target health record information reflects the pet's overall health background. For example, pet owners can upload their pet's veterinary check-up reports to a pet management application. The application can then process these reports and generate a health record for the pet. Alternatively, the health record can be stored in other applications (such as the veterinary clinic's application), and the pet device management server can access this information via the application programming interface (API) of those other applications. Historical behavioral data can refer to data collected by smart devices within a preset time period. For instance, when the smart device is a smart feeder, historical behavioral data could refer to the pet's food intake within that period. When the smart device is a smart litter box, historical behavioral data could refer to the pet's littering data within that period. When the smart device is a camera module, historical behavioral data could refer to video data of the pet captured by the camera module within a preset time period. It is understood that there can be multiple smart devices, and these devices can be of different types.

[0074] Historical behavioral data of the target pet is stored in a backend database in time-series format, objectively reflecting the daily physiological and behavioral characteristics of the target pet. For example, for smart water fountains, historical behavioral data can refer to daily water intake and frequency. Abnormal drinking patterns (such as sudden increases or decreases in water intake) are often early signs of health problems. For example, persistent excessive water intake may indicate problems such as kidney disease or diabetes. Conversely, a significant decrease in water intake may be associated with factors such as fever, oral pain, or depression. By incorporating drinking data into the analysis, large language models can more sensitively capture abnormal changes in pet water intake. For smart feeders, historical behavioral data can refer to changes in the pet's food intake, feeding time, and frequency. Changes in eating habits often directly reflect the pet's health status: if a pet has a poor appetite or decreased food intake, it may indicate digestive system diseases, oral diseases, or other health discomforts; if a pet overeats or eats too quickly, it may also suggest endocrine disorders or changes in stress. By combining feeder data, large language models can quantify a pet's appetite and nutrient intake, thereby identifying abnormal eating behaviors early and linking them to potential diseases. For smart litter boxes, target historical behavioral data can refer to indicators such as the frequency of urination, excrement volume, and pet weight. Changes in excretion patterns are often important clues to many diseases; for example, increased urination accompanied by increased water intake strongly suggests urinary system or endocrine problems; chronic constipation or diarrhea corresponds to digestive system disorders. Data from smart litter boxes allows large language models to quantify these phenomena and cross-validate them with data from other smart devices. For example, when a significant increase in water intake recorded by the water fountain and an increase in urination recorded by the litter box are detected, along with a decrease in weight recorded by the feeder, the large language model correlates these cross-device information and infers with high confidence that there may be problems such as diabetes or kidney disease. The reliability of this conclusion is far higher than that of a judgment based on a single indication.

[0075] Secondly, image features can be extracted from the target pet image to obtain target image features. Target image features refer to pet-related features that can be obtained from the target pet image.

[0076] In this way, the pet device management server can use the acquired multimodal data (i.e., target health analysis text, target image features, target health record information, and target historical behavior data) to call a pre-trained large language model to perform health analysis on the target pet, thereby obtaining the target pet's health analysis results. The pet device management server can then distribute the health analysis results to the pet management client, allowing the client to display the results on a corresponding interface (such as a pet health Q&A interface).

[0077] The pet health analysis method provided in this application can utilize a pre-trained large language model to perform health analysis on a target pet based on multimodal data such as target health analysis text, target image features, target health record information, and target historical behavior data. This improves the accuracy of pet health analysis. Specifically, as the above analysis shows, since the target pet image is uploaded by the pet owner, it usually implicitly contains the pet owner's attention bias or a direct representation of the pet's current health status. Other multimodal data (i.e., target health analysis text, target health record information, and target historical behavior data) can complement the target pet image for analysis. Related technologies for pet health monitoring often limit themselves to data from a single source or simple rule-based judgments, such as relying solely on symptoms subjectively described by the pet owner or monitoring results from a single device. Therefore, they are easily affected by occasional factors and struggle to detect potential health problems in a timely and comprehensive manner. In contrast, the embodiments of this application collect objective data (i.e., target health record information, target historical behavior data, etc.) from multiple aspects of the pet's daily life and perform deep correlation analysis using a large language model, providing richer context and evidence for health analysis. Thus, the embodiments of this application can improve the accuracy and reliability of health analysis.

[0078] In some embodiments, performing pet recognition on a target pet image to obtain the target pet includes:

[0079] Perform pet detection on the target pet image and compare the detected pets with the pet profile information of registered pets;

[0080] The target pet is determined from the detected pets based on the comparison results.

[0081] In this embodiment, the pet device management server can perform health analysis on registered pets. Therefore, the pet device management server needs to identify the target pet image. Specifically, the pet device management server can perform pet detection on the target pet image corresponding to the pet health analysis command to determine whether the target pet image contains a pet. When the pet detection result indicates that a pet has been detected, the pet device management server can compare the detected pet with the pet profile information of registered pets to determine which pet in the target pet image belongs to the pet owner. Registered pets can refer to pets whose profile information is stored in the pet management application. The pet profile information includes information such as the pet's birthday, weight, breed, and feature reference pool. The feature reference pool can include multiple images of the corresponding pet. For example, pet owners can pre-register their pets or resident pets by manually (or automatically) entering pet profile information in the pet management application.

[0082] In this way, the pet device management server can determine which of the detected pets is a registered pet based on the comparison results, and select the registered pet in the target pet image as the target pet for health analysis.

[0083] Understandably, in some embodiments, when the comparison result indicates that the target pet image contains multiple registered pets, the pet device management server can also determine the target pet based on the target health analysis text. For example, if the pet device management server detects that the target pet image contains three pets, A, B, and C, it can determine that pets A and B are registered pets by comparing them with the registered pet profile information. In this case, the pet device management server can analyze the target health analysis text. If the target health analysis text is "Perform a health analysis on pet A," then the target pet is pet A. Alternatively, when it is impossible to determine from the target health analysis text which pet to perform the health analysis (e.g., the target health analysis text is "Please perform a health analysis on the pet in the picture"), the pet device management server can send a corresponding instruction to the pet management client, causing the pet management client to display a question on the corresponding interface (e.g., a pet health Q&A interface): "Please perform a health analysis on pet A or pet B?" Thus, based on the pet owner's response to the question (e.g., "Perform a health analysis on pet A"), pet A can be determined as the target pet. It is understood that the above embodiments are illustrated using the health analysis of a single pet as an example, but the health analysis of multiple pets can also be performed simultaneously as needed, such as simultaneously performing health analysis on all registered pets in the target pet image. This application embodiment does not limit this.

[0084] In this embodiment of the application, the method of determining the target pet through pet profile information can reduce the need to perform health analysis on temporary pets (such as guests' pets) in the target pet image, and can achieve automatic identification of the target pet, reducing the need for pet owners to manually specify the pet.

[0085] In some embodiments, acquiring target historical behavior data of a target pet within a preset time period includes:

[0086] Retrieve historical pet videos and feeding data for a preset time period from associated smart devices;

[0087] Select key frames from historical pet videos that are related to the target pet image, and filter historical feeding data according to the key time periods corresponding to the key frames;

[0088] The target's historical behavior data is constructed based on keyframes and filtered historical eating data.

[0089] In this embodiment, the pet device management server can obtain historical pet videos and historical feeding data of the target pet within a preset time period from associated smart devices. Associated smart devices can refer to devices capable of collecting information related to the target pet. For example, in a home environment, associated smart devices can include all smart devices in the home environment that can interact with the pet device management server and collect information related to the target pet.

[0090] Thus, after identifying the associated smart devices, the pet device management server can obtain historical pet videos within a preset time period from smart devices with camera modules (such as smart litter boxes), and historical feeding data (such as historical feeding data and historical drinking data) from smart feeding devices (such as smart feeders and smart water fountains). The preset time period can refer to a historical period such as last week or last month; this application embodiment does not limit the duration of the preset time period. After obtaining the historical pet videos, the pet device management server can select keyframes from the historical pet videos. These keyframes can refer to important video frames in the historical pet videos. It is understood that when pet owners wish to perform health analysis on their pets, the uploaded target pet image is usually an image of the current or recent target pet, or an image that can reflect the pet's health status. However, historical pet videos may contain video frames unrelated to the current health analysis (such as video frames from a long time ago, which have no reference value for health analysis). Therefore, keyframes can be selected based on the target pet image. For example, video frames in historical pet videos with a similarity greater than a preset threshold to the target pet image can be used as the first keyframe. Furthermore, adjacent frames to the first keyframe can also be used as keyframes, etc., which is not limited in this embodiment. In this way, video frames in historical pet videos can be filtered, thereby improving the accuracy and efficiency of health analysis based on keyframes. It is understood that the specific value of the preset threshold can be adaptively set according to actual conditions, such as setting it to 80% or other values, and is not specifically limited thereto.

[0091] Thus, after selecting keyframes, the pet device management server can filter historical feeding data based on the time periods (i.e., key time periods) corresponding to multiple keyframes. For example, it can filter historical feeding data indicating that the feeding time falls within a key time period, or it can filter historical feeding data where the time difference between the feeding time and the key time period is within a preset range. In this way, historical feeding data related to keyframes (i.e., related to the image of the target pet) can be filtered from a large amount of historical feeding data, thereby improving the accuracy and efficiency of subsequent health analysis based on the filtered historical feeding data.

[0092] In some embodiments, image feature extraction is performed on the target pet image to obtain target image features, including:

[0093] The preset visual encoder is invoked to perform image segmentation on the target pet image, resulting in multiple image blocks;

[0094] For each image block, a linear mapping is performed on the image block to obtain an image vector, and a location code is generated based on the position of the image block in the target pet image;

[0095] The target image features are constructed by using the image vectors and positional encodings corresponding to multiple image blocks.

[0096] In the embodiments of this application, such as Figure 3 As shown, image feature extraction can be performed based on a visual encoder (using a visual transformer (ViT) as the visual encoder). Specifically, the visual encoder can segment an image (such as a target pet image) into fixed-size blocks, thus obtaining multiple image blocks. Then, each image block is flattened into a vector and linearly mapped to a fixed spatial dimension to obtain an image vector. Furthermore, to preserve spatial relationships, a positional encoding can be added to each image block (e.g., the positional encoding of the top-left image block is [0.1, 0.9], and the positional encoding of the bottom-right image block is [0.9, 0.1]). Thus, a visual token (i.e., image features, such as the target image features corresponding to the target pet image) can be constructed based on the image vectors corresponding to all image blocks and the vectors corresponding to the positional encodings.

[0097] The method for determining visual tokens based on a visual encoder in this application can transform an image into a sequence of feature vectors that preserves spatial relationships, including both local details (such as a pet's eyes and ears) and implicit global structures (such as body posture). Thus, the visual token can implicitly contain both the pet's facial expression features (determined based on local details) and posture features (determined based on global structure). These facial expression and posture features typically reflect the pet's health status. For example, when the facial expression features indicate that the target pet is sticking out its tongue and its tongue is cyanotic, it suggests that the target pet may have respiratory problems. When the posture features indicate that the target pet has an arched back, it suggests that the target pet may have digestive problems. Therefore, the method for determining visual tokens based on a visual encoder in this application can provide data support for health analysis.

[0098] In some embodiments, a pre-trained large language model is invoked to perform health analysis based on target image features, target health analysis text, target health record information, and target historical behavior data to obtain the target pet's target health analysis results, including:

[0099] The preset text encoder is invoked to encode the target health analysis text, target health record information, and target historical behavior data to obtain target text features;

[0100] Construct target fusion features based on target image features and target text features;

[0101] The pre-trained large language model is invoked to determine the attention weights of the target fusion features, and health analysis is performed based on the attention weights and target fusion features to obtain the target pet's target health analysis results.

[0102] In the embodiments of this application, such as Figure 3 As shown, the acquired text data (such as target health analysis text, target health record information, and target historical behavior data) can be converted into structured vector sequences (i.e., text tokens, or text features, such as target text features) using a text encoder. In this way, discrete text can be converted into continuous vector sequences for subsequent semantic association by a large language model. By concatenating target image features and target text features, a complete contextual prompt (i.e., target fusion feature) can be constructed. After obtaining the target fusion feature, it is input into a pre-defined large language model for analysis. In this way, the originally heterogeneous multimodal data is uniformly represented in a sequence that the large language model can understand.

[0103] The large language model can compute the association weights (i.e., determine the attention weights) among all tokens contained in the target fusion features based on a self-attention mechanism, and perform health analysis based on the association weights and all tokens to obtain the target pet's health analysis result. Understandably, the large language model can contain parallel transformer branches for different tasks (such as understanding tasks and result generation tasks). Within each transformer layer, a shared multimodal self-attention operation is performed on all input tokens. This design enables the simultaneous capture of semantic information and pixel-level details of images within a unified network, without the need to build separate sub-modules.

[0104] Understandably, compared to related technologies that use keypoint detection for pose and expression detection, this application's embodiments achieve end-to-end feature extraction and fusion through a unified transformer architecture. Specifically, the extracted visual tokens directly enter a multi-layer transformer, eliminating the need for pre-localization of facial keypoints or body joints using specialized algorithms. The transformer implicitly learns expression and pose patterns through a self-attention mechanism, without relying on manually defined intermediate features. This end-to-end design simplifies the process and improves robustness, reducing potential error accumulation in related technologies, while enabling more flexible capture of complex facial expressions and pose features.

[0105] Please see Figure 4In some embodiments, a pre-trained large language model is invoked to perform health analysis based on target image features, target health analysis text, target health record information, and target historical behavior data to obtain the target pet's target health analysis results, including:

[0106] Statistical features are generated based on the target's historical behavior data, and abnormal data is extracted from the target's historical behavior data based on the statistical features.

[0107] Extract health knowledge from a pre-built pet health knowledge base based on abnormal data and target health analysis text;

[0108] Based on health knowledge, abnormal data, target image features, target health analysis text, and target health record information, a large language model is invoked to perform health analysis and obtain the target pet's target health analysis results.

[0109] In this embodiment, the method for determining the target health analysis result based on the large language model is similar to that in the aforementioned embodiments. The difference is that, in this embodiment, the acquired text data includes not only the target health analysis text, target health record information, and target historical behavior data, but also health knowledge. Furthermore, in this embodiment, the target historical behavior data can be filtered, and the filtered abnormal data can be used as input data for the large language model.

[0110] Specifically, the pet device management server can statistically analyze the historical behavior data of targets corresponding to the same smart device to obtain the statistical characteristics of each smart device. The statistical characteristics can be the mean, variance, etc., which are not limited in this embodiment. The pet device management server can compare the statistical characteristics with the historical behavior data of the target to extract data that deviates significantly from the statistical characteristics (i.e., abnormal data). It is understood that under normal circumstances, pet behavior data follows certain patterns or ranges. Therefore, when abnormal data appears, it indicates that the target pet may have health abnormalities. Based on this, health knowledge related to abnormal data, target image features, and target health analysis text can be retrieved from a pre-built pet health knowledge base. Thus, health knowledge can be combined with multimodal data (i.e., abnormal data, target image features, target health analysis text, and target health record information), and based on the combination result, a large language model can be called to perform health analysis on the target pet to obtain the target health analysis result. Specifically, as... Figure 3As shown, text data (including abnormal data, target health analysis text, target health record information, and health knowledge) can be converted into text tokens based on a text encoder. Then, a target fusion feature can be constructed based on this text token and target image features (i.e., visual token). A pre-trained large language model can then be used to perform health analysis based on this target fusion feature to obtain the target health analysis result.

[0111] The pet health knowledge base can be constructed based on authoritative pet knowledge sources, authoritative veterinary papers, and other data. Specifically, this application embodiment can utilize Retrieval Augmentation (RAG) technology to integrate the authoritative pet health knowledge base into the analysis process, thereby further improving the accuracy and reliability of the health analysis results. Specifically, before calling the large language model for health analysis, relevant background knowledge (such as authoritative information on possible causes of corresponding symptoms and treatment suggestions) can be retrieved from the pre-constructed pet health knowledge base based on the pet's current symptom characteristics and abnormal indicators. This retrieved knowledge is added to the input context of the large language model as supplementary text. Through this mechanism, the large language model, when generating analysis conclusions, not only relies on its own training parameters but also refers to the latest professional knowledge, thereby effectively reducing the risk of the large language model generating false information (i.e., "hallucinations"). By allowing the large language model to answer based on real and credible data, the RAG method has been shown to improve answer accuracy and reduce hallucination phenomena. In this embodiment, the veterinary literature and guidelines provided by the pet health knowledge base provide a solid foundation for the analysis of the large language model. For example, when the large language model needs to determine the possible causes of a combination of symptoms, the retrieved authoritative information will provide strong evidence for that cause. Therefore, the fusion of multimodal features combined with the retrieved knowledge allows the reasoning of the large language model to encompass both individual pet data and universal medical knowledge, achieving an organic combination of data-driven and knowledge-driven approaches. This fusion processing ultimately produces accurate health analysis results tailored to specific pets, offering significant advantages over methods that rely solely on a single modality or general model memory.

[0112] The target pet's health analysis results can include whether the target pet has health abnormalities, the reasons for these abnormalities, and methods to alleviate them. Furthermore, because the large language model incorporates a pet health knowledge base for health analysis, it can also cite sources for this knowledge in the analysis results, reducing the potential for misleading interpretations.

[0113] In the embodiments of this application, the method of filtering out abnormal data from target historical behavior data and performing health analysis based on abnormal data can improve the efficiency of health analysis by large language models.

[0114] In some embodiments, the training method for a large language model includes:

[0115] Acquire multimodal data of pets and corresponding sample health analysis results. The multimodal data includes sample health analysis text, sample pet images, sample health record information, and sample historical behavior data.

[0116] The visual encoder is invoked to determine the sample image features of the sample pet images, and the text encoder is invoked to determine the sample text features of the sample health analysis text, sample health record information, and sample historical behavior data.

[0117] Based on the features of the sample images and the features of the sample text, a sample fusion feature is constructed, and a pre-trained visual language model is called to perform health analysis based on the sample fusion feature to obtain the predicted health analysis result;

[0118] The visual language model was fine-tuned based on the results of the predicted health analysis and the sample health analysis, and the fine-tuned visual language model was determined as the large language model.

[0119] In this embodiment, the large language model can be based on an open-source Visual Language Model (VLM). The VLM model is pre-trained using massive image and text corpora, thus possessing broad general visual understanding and language generation capabilities. Based on this, the VLM model can be fine-tuned for the pet health analysis scenario applied in this embodiment, making it more suitable for the application requirements of this embodiment. It is understood that fine-tuning is essentially the process of further training the pre-trained model on a domain-specific dataset. Through fine-tuning, the model's network parameters can be adjusted, allowing for a better understanding of data patterns and terminology in the pet domain, thereby improving the model's analytical performance in pet health analysis scenarios.

[0120] Specifically, training samples can be constructed based on multimodal data and corresponding sample health analysis results to fine-tune the VLM model. Similar to the multimodal data corresponding to the target pets introduced earlier, the multimodal data used for fine-tuning training here includes sample health analysis text, sample pet images, sample health record information, and sample historical behavior data. Furthermore, the processing method for this multimodal data is the same as described earlier: visual tokens are obtained based on the visual encoder, text tokens are obtained based on the text encoder, and sample fusion features are constructed based on the visual tokens and text tokens. The model performs self-attention processing on the sample fusion features and determines the final health analysis result (i.e., the predicted health analysis result), which will not be elaborated further. Sample health analysis results can refer to high-quality result data obtained by professional individuals through diagnostic annotation based on the corresponding multimodal data, or by reviewing and correcting the model's preliminary analysis results. Sample health analysis results are used to guide the model to learn more accurate predictive patterns.

[0121] In this embodiment, supervised learning can be used to repeatedly fine-tune the VLM model, so that the predicted health analysis results output by the VLM model based on the training samples approximate the corresponding sample health analysis results. It is understood that the open-source VLM model itself possesses cross-modal processing capabilities, and pre-training enables the VLM model to accept input signals mixed with images and text. In this embodiment, the VLM model has already encountered similar multimodal contexts (i.e., contextual cues) during training, thus it can freely associate information from different sources during inference. That is, this embodiment, by designing contextual cues, integrates data from multiple sources for processing by a unified large language model, thereby fully utilizing the large language model's ability to integrate and reason about complex information.

[0122] Internally, the analysis process relies on the powerful self-attention mechanism and deep semantic representation capabilities of the large language model. When multimodal data is input, the large language model generates vector representations for each part of the sequence (whether it's a text fragment, feature descriptions extracted by a visual encoder or text encoder, or textual representations of sensor values), and captures the correlations between different parts of the information through attention calculations in a multi-layer transformer network. Simply put, when reading the entire context, the large language model automatically focuses on important associations between features of different modalities, thereby achieving semantic fusion and comprehensive analysis of the data. For example, for the numerical feature "significantly increased historical water intake" provided in the input, the large language model can combine the information "advanced age / previous history of kidney disease" from the health record information, and the description "excessive water intake in pets may be a sign of kidney problems" retrieved from the knowledge base, comprehensively considering these data to infer a more accurate health conclusion. For example, the large language model might correlate information such as "pet is thin" extracted from pet images with records of "reduced food intake" in historical feeding data, while also referencing health record information indicating "no recent changes in diet or environment." This would lead to the inference that the pet may have underlying health issues. Through a self-attention mechanism, the large language model automatically completes the association reasoning of multimodal features without requiring manually defined rules. The information from each modality forms a unified semantic representation within the health record information, supporting the final decision generation.

[0123] Through the aforementioned fine-tuning training, the large language model effectively grasps the important feature patterns in pet health analysis scenarios and their correlation with health status, thereby improving its adaptability and accuracy in pet health analysis. For example, the fine-tuned large language model can "learn" to associate a significant increase in water intake with the possibility of kidney disease or diabetes, and to link reduced food intake accompanied by weight loss with indigestion or oral health problems, thus providing more realistic judgments in health analysis. Furthermore, in practical applications, a continuous learning and optimization mechanism can be established, which involves regularly collecting new device data, pet owner feedback, and labeled cases confirmed by professionals to periodically and incrementally fine-tune the large language model. Such continuous optimization ensures that the performance of the large language model continuously improves with data accumulation, gradually reducing misjudgments and omissions, and enabling the large language model to maintain a good adaptability to the latest pet health trends and individual differences.

[0124] In some embodiments, sample fusion features are constructed based on sample image features and sample text features, including:

[0125] Determine the initial image weights of the sample image features and the initial text weights of the sample text features, and construct sample fusion features based on the initial image weights, sample image features, initial text weights, and sample text features.

[0126] The visual language model was fine-tuned based on the predicted health analysis results and the sample health analysis results. The fine-tuned visual language model was then determined as the large language model, including:

[0127] Based on the predicted health analysis results and the sample health analysis results, the network parameters of the visual language model are updated, and the initial image weights and initial text weights are also updated. The updated visual language model is determined as the large language model, the updated initial image weights are determined as the target image weights, and the updated initial text weights are determined as the target text weights.

[0128] Target fusion features are constructed based on target image features and target text features, including:

[0129] Target fusion features are constructed based on target image weights, target text weights, target image features, and target text features.

[0130] In this embodiment, target image weights corresponding to target image features and target text weights corresponding to target text features can be set to filter or enhance specific modalities before inputting data into a large language model. The specific values ​​of the target image weights and target text weights can be determined during the training of the large language model. That is, the weight values ​​are continuously updated during the training process. After the large language model is trained, the final determined weight values ​​are saved, and during actual inference, the corresponding weight values ​​are directly used to perform weighted fusion of the data from each modality to obtain the target fused features. This method of fixing weights during actual inference can reduce weight drift caused by temporary fluctuations in data quality during inference. In particular, in a home setting, the pets for which health analysis is performed are usually fixed, so the distribution of the data corresponding to that pet is relatively stable (e.g., the data collection method remains relatively fixed, and the statistical characteristics of multimodal data usually do not fluctuate significantly). Therefore, the fixed weight method can improve the accuracy of health analysis in this scenario.

[0131] In this embodiment, the method for determining the specific values ​​of the target image weights and target text weights includes two stages: the first stage is the initial value setting stage (i.e., determining the initial image weights and initial text weights), and the second stage is the training stage (i.e., updating the values ​​of the initial image weights and initial text weights to determine the final target image weights and target text weights). It is understood that before training the large language model, a validation set and a training set can be pre-constructed based on multimodal data and sample health analysis results. The training set is used to train the large language model, and the validation set can be used for the experiments in the first stage.

[0132] Specifically, in the first stage, the value ranges of image weights (e.g., w1) and text weights (e.g., w1 ∈ [0.5, 0.95], w2 = 1 - w1 are first set (the value range can be adaptively set according to actual needs and is not specifically limited), and a grid search is performed based on the value range to determine all possible weight combinations. For example, the step size can be set to 0.05, thus searching for weight combinations such as (w1 = 0.5, w2 = 0.5), (w1 = 0.55, w2 = 0.45), etc. Then, for each weight combination, the same network is trained based on the training set (i.e., ...). Figure 3 The network structure shown is used for several epochs (i.e., training rounds), and the health analysis accuracy is calculated based on the validation set after each epoch. For example, the similarity between the health analysis results output by the large language model and the corresponding expert annotations (i.e., sample health analysis results) in the text encoder projection space can be calculated, and this similarity can be used as the accuracy. Finally, the weight combination with the highest similarity is selected as the initial weight (i.e., as the initial image weight and initial text weight). It is understandable that the first stage is executed only once to provide the large language model with a reasonable and stable training starting point, ensuring that the subsequent large language model learning process converges quickly, and reducing gradient instability caused by extreme weights.

[0133] In the second stage, the initial image weights and initial text weights are treated as learnable parameters and jointly optimized with the large language model. Specifically, a weighted fusion of the initial image weights, sample image features, initial text weights, and sample text features is first performed to obtain sample fusion features. Then, these sample fusion features are used as input data to the pre-trained VLM model to obtain the corresponding predicted health analysis results. Thus, the health analysis loss can be determined based on this predicted health analysis result and the corresponding sample health analysis result, and joint optimization is performed based on the health analysis loss. In other words, throughout the entire network training process, the network parameters, initial image weights, and initial text weights of the VLM model are updated simultaneously with backpropagation.

[0134] Thus, after the VLM model converges, the final updated initial image weights are used as the target image weights, and the final updated initial text weights are used as the target text weights. During actual inference, the target image weights, target image features, target text weights, and target text features are weighted and fused to obtain the target fusion features. These target fusion features are then used as input data for the trained large language model, thereby obtaining the target pet's health analysis results.

[0135] Furthermore, to improve the accuracy of the large language model analysis, the pet device management server can also receive feedback data from pet owners regarding health analysis results from the pet management client. The pet device management server can adjust the large language model based on this feedback data, or adjust the weights corresponding to multimodal data. Specifically, the pet device management server can collect feedback data such as likes or dislikes from pet owners regarding health analysis results, and analyze the reasons for positive or negative feedback to determine corresponding optimization strategies. For feedback data indicating negative evaluations, the pet device management server can optimize prompt words, such as adjusting the frequency and depth of professional terminology used in health analysis results for different pet breeds. Simultaneously, it can fine-tune the large language model using reinforcement learning methods, making it more inclined to generate content similar to highly liked results. In addition, the pet device management server can expand the search scope, searching for more relevant information from the pet health knowledge base to improve the professionalism and accuracy of the answers. For example, when a pet owner gives negative feedback on a health analysis result indicating a specific feline disease, the pet device management server can analyze whether the negative feedback is due to unprofessional prompt words or insufficient search knowledge, and then optimize accordingly. The optimization process can also be automatically executed by the large language model, thereby continuously improving the quality of the large language model's responses. Alternatively, the pet device management server can redetermine the specific values ​​of the target image weights and target text weights based on a new batch of training and validation sets to better match the recent state of the target pet (or the latest data distribution pattern).

[0136] Understandably, in some embodiments, the large language model can also combine a device knowledge base to answer pet owners' questions about smart devices. For example, when a smart feeder or smart waterer malfunctions, pet owners can ask questions on the pet health Q&A interface. The large language model can retrieve relevant instructions from the device knowledge base based on the pet owner's question, thereby providing pet owners with targeted troubleshooting methods. The device knowledge base can be constructed based on the smart device's instruction manual and other relevant information.

[0137] In summary, the pet health analysis method provided in this application receives a pet health analysis instruction, which includes a target pet image and a target health analysis text; performs pet recognition on the target pet image to obtain the target pet; acquires the target pet's target health record information and the target pet's target historical behavior data within a preset time period; extracts image features from the target pet image to obtain target image features; and performs health analysis by calling a preset large language model based on the target health analysis text, target image features, target health record information, and target historical behavior data to obtain the target pet's health analysis result.

[0138] Therefore, the pet health analysis method provided in this application embodiment can acquire multimodal data such as the target pet's drinking, eating, and exercise through associated smart devices. Thus, when performing health analysis on the target pet based on this multimodal data and image features extracted from the target pet's image, a more accurate and comprehensive health analysis can be achieved. Furthermore, when using a large language model for health analysis, this application embodiment can set different weights for the multimodal data, and the large language model can also combine health knowledge extracted from a pet health knowledge base for health analysis, thereby improving the accuracy and efficiency of the health analysis. Moreover, the multimodal data provided by various smart devices depict the pet's health status from different perspectives, and there are often potential correlations and complementary relationships between the multimodal data. The large language model in this application embodiment can fuse these scattered clues, perform complex pattern recognition and causal reasoning, greatly reducing the possibility of misjudgment and missed diagnosis. On the one hand, multimodal context provides large language models with a more comprehensive information foundation: data anomalies in one modality can be corroborated or refuted by data from other modalities, reducing the risk of misjudgment due to random factors. On the other hand, advanced reasoning capabilities enable large language models to discover subtle correlations based on rich data, thereby making more refined health assessments. This data-driven and cross-domain knowledge-supported analysis approach significantly improves the accuracy and reliability of pet health analysis results compared to methods relying on experience or single-indicator warnings in related technologies.

[0139] Reference Figure 5 In some embodiments, this application also provides a pet health analysis device 500, which includes:

[0140] The instruction receiving unit 510 is used to receive pet health analysis instructions, which include a target pet image and target health analysis text.

[0141] The pet recognition unit 520 is used to perform pet recognition on the target pet image to obtain the target pet;

[0142] Data acquisition unit 530 is used to acquire target health record information of the target pet and target historical behavior data of the target pet within a preset time period;

[0143] The health analysis unit 540 is used to extract image features from the target pet image to obtain target image features, and to call a pre-trained large language model to perform health analysis based on the target image features, target health analysis text, target health record information and target historical behavior data to obtain the target pet's target health analysis results.

[0144] Optionally, in some embodiments, the health analysis unit 540 includes:

[0145] The image segmentation subunit is used to call a preset visual encoder to segment the target pet image and obtain multiple image blocks;

[0146] The first coding subunit is used to perform linear mapping on each image block to obtain an image vector, and generate a position code based on the position of the image block in the target pet image.

[0147] The feature construction subunit is used to construct the target image features based on the image vectors and positional encodings corresponding to multiple image blocks.

[0148] Optionally, in some embodiments, the health analysis unit 540 further includes:

[0149] The second encoding subunit is used to call a preset text encoder to encode the target health analysis text, target health record information and target historical behavior data to obtain target text features;

[0150] The feature fusion subunit is used to construct target fusion features based on target image features and target text features;

[0151] The first health analysis subunit is used to call the pre-trained large language model to determine the attention weights of the target fusion features, and perform health analysis based on the attention weights and target fusion features to obtain the target pet's target health analysis results.

[0152] Optionally, in some embodiments, the pet health analysis device further includes a model training unit for training a large language model, including:

[0153] The training data acquisition subunit is used to acquire multimodal data of pets and the corresponding sample health analysis results. The multimodal data includes sample health analysis text of pets, sample pet images, sample health record information and sample historical behavior data.

[0154] The third encoding subunit is used to call the visual encoder to determine the sample image features of the sample pet image, and to call the text encoder to determine the sample text features of the sample health analysis text, sample health record information and sample historical behavior data.

[0155] The second health analysis subunit is used to construct sample fusion features based on sample image features and sample text features, and call a pre-trained visual language model to perform health analysis based on the sample fusion features to obtain the predicted health analysis results.

[0156] The fine-tuning training subunit is used to fine-tune the visual language model based on the predicted health analysis results and the sample health analysis results, and to determine the fine-tuned visual language model as the large language model.

[0157] Optionally, in some embodiments, the second health analysis subunit includes:

[0158] The first feature fusion module is used to determine the initial image weights of the sample image features and the initial text weights of the sample text features, and to construct sample fusion features based on the initial image weights, sample image features, initial text weights and sample text features.

[0159] Fine-tuning the training sub-units includes:

[0160] The fine-tuning training module is used to update the network parameters of the visual language model based on the predicted health analysis results and the sample health analysis results, and to update the initial image weights and initial text weights. The updated visual language model is determined as the large language model, the updated initial image weights are determined as the target image weights, and the updated initial text weights are determined as the target text weights.

[0161] Feature fusion subunit, including:

[0162] The second feature fusion module is used to construct target fusion features based on target image weights, target text weights, target image features, and target text features.

[0163] Optionally, in some embodiments, the data acquisition unit includes:

[0164] The feeding data acquisition subunit is used to acquire historical pet videos and historical feeding data within a preset time period from associated smart devices;

[0165] The data filtering subunit is used to select key frames that are related to the target pet image from historical pet videos, and to filter historical feeding data according to the key time period corresponding to the key frame;

[0166] The data construction subunit is used to construct target historical behavior data based on keyframes and filtered historical eating data.

[0167] Optionally, in some embodiments, the health analysis unit 540 further includes:

[0168] The abnormal data identification subunit is used to generate statistical features based on the target's historical behavior data, and to extract abnormal data from the target's historical behavior data based on the statistical features.

[0169] The health knowledge acquisition subunit is used to acquire health knowledge from a pre-built pet health knowledge base based on abnormal data and target health analysis text.

[0170] The third health analysis subunit is used to call a large language model to perform health analysis based on health knowledge, abnormal data, target image features, target health analysis text, and target health record information, and obtain the target pet's target health analysis results.

[0171] Optionally, in some embodiments, the pet recognition unit includes:

[0172] The pet comparison subunit is used to detect pets in the target pet image and compare the detected pets with the pet profile information of registered pets;

[0173] The pet identification unit is used to identify the target pet from the detected pets based on the comparison results.

[0174] Reference Figure 6 , Figure 6 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0175] The processor 601 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0176] The memory 602 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 602 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 602 and is called and executed by the processor 601 using the pet health analysis method of the embodiments of this application.

[0177] The input / output interface 603 is used to implement information input and output;

[0178] The communication interface 604 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0179] Bus 605 transmits information between various components of the device (e.g., processor 601, memory 602, input / output interface 603, and communication interface 604);

[0180] The processor 601, memory 602, input / output interface 603, and communication interface 604 are connected to each other within the device via bus 605.

[0181] This application also provides a storage medium storing a computer program, which, when executed by a processor, implements the pet health analysis method provided in this application.

[0182] This application also provides a computer program product, which includes a computer program. The processor of a computer device reads and executes the computer program, causing the computer device to perform the aforementioned pet health analysis method.

[0183] It should be understood that the various implementation methods provided in this application can be combined arbitrarily to achieve different technical effects.

[0184] The above is a detailed description of the embodiments of this disclosure. However, this disclosure is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this disclosure. All such equivalent modifications or substitutions are included within the scope defined by the claims of this disclosure.

Claims

1. A method for analyzing pet health, characterized in that, The method includes: Receive a pet health analysis instruction, which includes a target pet image and a target health analysis text; Perform pet recognition on the target pet image to obtain the target pet; Obtain the target health record information of the target pet and the target historical behavior data of the target pet within a preset time period; Image feature extraction is performed on the target pet image to obtain target image features. Based on the target image features, the target health analysis text, the target health record information, and the target historical behavior data, a pre-trained large language model is invoked to perform health analysis, resulting in a target health analysis result for the target pet. This includes: invoking a preset text encoder to encode the target health analysis text, the target health record information, and the target historical behavior data to obtain target text features; constructing target fusion features based on target image weights, target text weights, the target image features, and the target text features; invoking a pre-trained large language model to determine the attention weights of the target fusion features; and performing health analysis based on the attention weights and the target fusion features to obtain the target health analysis result for the target pet. The fine-tuned visual language model is defined as the large language model. Training samples are constructed based on the multimodal data of pets and the corresponding sample health analysis results. The visual language model is fine-tuned based on the training samples. During the training process, the network parameters, initial image weights and initial text weights of the visual language model are updated simultaneously. After the visual language model converges, the final updated initial image weights are used as the target image weights, and the final updated initial text weights are used as the target text weights.

2. The method according to claim 1, characterized in that, The step of extracting image features from the target pet image to obtain target image features includes: A preset visual encoder is invoked to perform image segmentation on the target pet image, resulting in multiple image blocks; For each image block, an image vector is obtained by linear mapping of the image block, and a position code is generated based on the position of the image block in the target pet image; The target image features are constructed based on the image vectors corresponding to multiple image blocks and the positional encoding.

3. The method according to claim 1, characterized in that, The training method for the large language model includes: Acquire multimodal data of pets and corresponding sample health analysis results, wherein the multimodal data includes sample health analysis text of pets, sample pet images, sample health record information and sample historical behavior data; The visual encoder is invoked to determine the sample image features of the sample pet image, and the text encoder is invoked to determine the sample text features of the sample health analysis text, the sample health record information, and the sample historical behavior data; Based on the sample image features and the sample text features, a sample fusion feature is constructed, and a pre-trained visual language model is invoked to perform health analysis based on the sample fusion feature to obtain the predicted health analysis result; The visual language model is fine-tuned based on the predicted health analysis results and the sample health analysis results, and the fine-tuned visual language model is determined as the large language model.

4. The method according to claim 3, characterized in that, The step of constructing sample fusion features based on the sample image features and the sample text features includes: Determine the initial image weights of the sample image features and the initial text weights of the sample text features, and construct the sample fusion features based on the initial image weights, the sample image features, the initial text weights, and the sample text features; The step of fine-tuning the visual language model based on the predicted health analysis results and the sample health analysis results, and determining the fine-tuned visual language model as the large language model, includes: The network parameters of the visual language model are updated based on the predicted health analysis results and the sample health analysis results. The initial image weights and the initial text weights are also updated. The updated visual language model is determined as the large language model, the updated initial image weights are determined as the target image weights, and the updated initial text weights are determined as the target text weights.

5. The method according to claim 1, characterized in that, Obtain the target pet's historical behavior data within a preset time period, including: Obtain historical pet videos and historical feeding data within the preset time period from associated smart devices; Select key frames from the historical pet videos that are associated with the target pet image, and filter the historical feeding data according to the key time periods corresponding to the key frames; The target historical behavior data is constructed based on the keyframes and the filtered historical eating data.

6. The method according to claim 1, characterized in that, The process involves using a pre-trained large language model to perform health analysis based on the target image features, the target health analysis text, the target health record information, and the target historical behavior data to obtain the target pet's target health analysis results, including: Statistical features are generated based on the target's historical behavior data, and abnormal data is extracted from the target's historical behavior data based on the statistical features. Based on the abnormal data and the target health analysis text, health knowledge is obtained from a pre-built pet health knowledge base; Based on the health knowledge, the abnormal data, the target image features, the target health analysis text, and the target health record information, the large language model is invoked to perform health analysis, and the target health analysis result of the target pet is obtained.

7. The method according to claim 1, characterized in that, The step of performing pet recognition on the target pet image to obtain the target pet includes: The target pet image is subjected to pet detection, and the detected pets are compared with the pet profile information of registered pets; The target pet is determined from the detected pets based on the comparison results.

8. A pet health analysis device, characterized in that, The device includes: The instruction receiving unit is used to receive pet health analysis instructions, which include a target pet image and target health analysis text. A pet recognition unit is used to perform pet recognition on the target pet image to obtain the target pet; The data acquisition unit is used to acquire the target health record information of the target pet and the target historical behavior data of the target pet within a preset time period; The health analysis unit is used to extract image features from a target pet image to obtain target image features, and to perform health analysis based on the target image features, the target health analysis text, the target health record information, and the target historical behavior data by calling a pre-trained large language model to obtain the target health analysis result of the target pet. This includes: calling a preset text encoder to encode the target health analysis text, the target health record information, and the target historical behavior data to obtain target text features; constructing target fusion features based on target image weights, target text weights, the target image features, and the target text features; calling the pre-trained large language model to determine the attention weights of the target fusion features, and performing health analysis based on the attention weights and the target fusion features to obtain the target health analysis result of the target pet. The fine-tuned visual language model is defined as the large language model. Training samples are constructed based on the multimodal data of pets and the corresponding sample health analysis results. The visual language model is fine-tuned based on the training samples. During the training process, the network parameters, initial image weights and initial text weights of the visual language model are updated simultaneously. After the visual language model converges, the final updated initial image weights are used as the target image weights, and the final updated initial text weights are used as the target text weights.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the pet health analysis method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the pet health analysis method according to any one of claims 1 to 7.

11. A computer program product, characterized in that, The computer program product includes a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the pet health analysis method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Pet health monitoring method and system based on prediction model

    CN119650074A