Pet health analysis method and device, electronic equipment and storage medium
Through pet toilet video analysis and visual language model, the pet health status is automatically identified, which solves the problem of insufficient human judgment in the existing technology, improves the efficiency and accuracy of pet health analysis, and reduces pet medical costs.
Patent Information
- Application Number
- CN202510913173.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-03
AI Technical Summary
The lack of efficient pet health analysis methods in the prior art, and relying on human judgment, leads to failure to detect or deal with pet diseases in a timely manner, increasing medical costs and health risks.
By obtaining pet toilet videos, pet toilet behavior category identification and excrement image frame analysis are carried out, and health analysis is performed in combination with visual language models to automatically identify pet health status.
It has achieved efficient pet health analysis without manual participation, improved analysis efficiency and accuracy, timely discovered health abnormalities and provided disposal solutions, and reduced pet medical costs.
Smart Images

Figure CN120412985A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly relates to a pet health analysis method, device, electronic device, and storage medium. Background Art
[0002] In recent years, with the continuous development of technology and the accelerating pace of life, people's life pressure has also been continuously increasing. The company of pets can, to a large extent, relieve people's mental stress and bring physical and mental pleasure, which has led to the rapid development of the pet supplies market in recent years.
[0003] Among them, during the process of raising pets, pet diseases may be caused by pet diet problems, weather changes, pet living habits, and other unexpected problems. If pet owners fail to detect pet diseases in a timely manner and perform corresponding treatment, it may lead to problems such as the aggravation of the pet's condition.
[0004] However, in the related art, determining whether a pet is sick still depends on manual judgment, lacking an efficient pet health analysis method. Summary of the Invention
[0005] The main purpose of the embodiments of the present application is to propose a pet health analysis method, device, electronic device, and storage medium, aiming to improve the efficiency and accuracy of pet health analysis.
[0006] To achieve the above object, a first aspect of the embodiments of the present application proposes a pet health analysis method, the method including: Obtaining a pet toilet video, where the pet toilet video includes video data captured from multiple perspectives within a pet toilet device; Performing pet toilet behavior category recognition on the pet toilet video to obtain toilet behavior category information; Extracting excrement image frames from the pet toilet video, and performing physiological index analysis on the excrement image frames to obtain a physiological index analysis result; Based on the toilet behavior category information and the physiological index analysis result, calling a vision-language model for health analysis to obtain a pet health analysis result.
[0007] To achieve the above object, a second aspect of the embodiments of the present application proposes a pet health analysis device, the device including: An obtaining unit, configured to obtain a pet toilet video, where the pet toilet video includes video data captured from multiple perspectives within a pet toilet device; An identifying unit, configured to perform pet toilet behavior category recognition on the pet toilet video to obtain toilet behavior category information; An extraction unit for extracting excrement image frames from the pet toileting video and performing physiological index analysis on the excrement image frames to obtain a physiological index analysis result; An analysis unit for performing health analysis by calling a vision-language model based on the toileting behavior category information and the physiological index analysis result to obtain a pet health analysis result.
[0008] Optionally, in some embodiments, the analysis unit includes: A first acquisition subunit for acquiring a pet health analysis prompt text, the pet health analysis prompt text indicating the generation of a pet health analysis result; An input subunit for inputting the pet health analysis prompt text, the toileting behavior category information, and the physiological index analysis result into the vision-language model to obtain the pet health analysis result output by the vision-language model.
[0009] Optionally, in some embodiments, the present application further provides a model training device, including: A pre-training subunit for acquiring a plurality of first pet behavior video samples and pre-training the vision-language model based on the plurality of first pet behavior video samples; A second acquisition subunit for acquiring a plurality of second pet behavior video samples, the second pet behavior video samples including a plurality of pet toileting video samples, pet health analysis prompt text samples corresponding to each pet toileting video sample, and corresponding pet health problem text labels; A first fine-tuning subunit for fine-tuning the vision-language model according to the plurality of second pet behavior video samples to obtain a trained vision-language model.
[0010] Optionally, in some embodiments, the extraction unit includes: An extraction subunit for extracting a sequence of excrement image frames from the pet toileting video; An analysis subunit for performing color development analysis on the sequence of excrement image frames to obtain a color development analysis result; A determination subunit for determining a physiological index analysis result based on the color development analysis result and key image frames extracted from the image frame sequence.
[0011] Optionally, in some embodiments, the pet health analysis device provided by the present application further includes: A third acquisition subunit for acquiring a treatment plan corresponding to the health abnormality when the pet health analysis result indicates that the pet has a health abnormality; A generating subunit, configured to generate a health report based on the health anomaly and the treatment plan, and send the health report to the pet management application client corresponding to the pet.
[0012] Optionally, in some embodiments, the generating subunit includes: An obtaining module, configured to obtain the historical health analysis results of the pet within a preset period; A sending module, configured to generate a health report according to the historical health analysis results, the health anomaly, and the treatment plan, and send the health report to the pet management application client corresponding to the pet.
[0013] Optionally, in some embodiments, the pet health analysis device provided in this application further includes: A receiving subunit, configured to receive the health feedback information for the health anomaly returned by the pet management application client; A second fine-tuning subunit, configured to update the sample data for fine-tuning the vision-language model based on the health feedback information, and fine-tune the vision-language model based on the updated sample data.
[0014] Optionally, in some embodiments, this application further provides a video acquisition device, including: A fourth obtaining subunit, configured to obtain the environmental information around the pet toilet device, and perform pet approach detection based on the environmental information to obtain a detection result; An acquisition subunit, configured to trigger the multiple camera devices in the pet toilet device to perform pet toilet video acquisition when it is determined based on the detection result that the pet enters the pet toilet device.
[0015] Optionally, in some embodiments, the recognition unit includes: A splitting subunit, configured to respectively perform frame splitting on the multi-angle video data included in the pet toilet video to obtain a video frame sequence corresponding to each video data; An alignment subunit, configured to perform time alignment at the frame level on the multiple video frame sequences, and combine the multiple video frame sequences into a target video frame sequence according to the alignment result; A recognition subunit, configured to perform toilet behavior category recognition based on the target video frame sequence to obtain toilet behavior category information.
[0016] To achieve the above object, a third aspect of the embodiments of this application proposes an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, it implements the pet health analysis method described in the first aspect.
[0017] To achieve the above object, a fourth aspect of the embodiments of the present application provides a storage medium storing a computer program, which when executed by a processor implements the pet health analysis method described in the first aspect.
[0018] To achieve the above object, a fifth aspect of the embodiments of the present application provides a computer program product, which includes a computer program that is read and executed by a processor of a computer device, causing the computer device to execute the pet health analysis method described in the first aspect.
[0019] The pet health analysis method proposed in the embodiments of the present application includes obtaining a pet toilet video, which includes video data captured from multiple perspectives inside a pet toilet device; identifying the pet toilet behavior category of the pet toilet video to obtain toilet behavior category information; extracting excrement image frames from the pet toilet video and performing physiological index analysis on the excrement image frames to obtain physiological index analysis results; and calling a vision-language model based on the toilet behavior category information and the physiological index analysis results for health analysis to obtain a pet health analysis result.
[0020] It can be seen that the pet health analysis method provided by the embodiments of the present application collects a toilet video based on a camera device set in the toilet device when the pet uses the toilet, then analyzes the pet's toilet behavior and excrement state according to the toilet video, and then infers the pet's health state based on the vision-language model for the toilet behavior and excrement state. This method does not require manual participation and can automatically perform health analysis on the pet when the pet uses the toilet, which can improve the efficiency of pet health analysis; in addition, this method performs pet health analysis through a trained vision-language model, which can also greatly improve the accuracy of pet health analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The drawings are used to provide a further understanding of the technical solutions of the present application, and constitute a part of the specification. They are used together with the embodiments of the present application to explain the technical solutions of the present application, and do not constitute a limitation to the technical solutions of the present application.
[0022] Figure 1 It is a schematic diagram of the system to which the pet health analysis method provided by the present application is applied; Figure 2 It is a schematic flowchart of the pet health analysis method provided by the present application; Figure 3 It is a schematic structural diagram of the vision-language model provided by the present application; Figure 4 It is a schematic diagram of the data flow during the pet health analysis process in the embodiments of the present application; Figure 5 It is a schematic diagram of an interface of the pet management application provided by the present application; Figure 6 Schematic structural diagram of the pet health analysis device provided for this application; Figure 7 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of this application. Detailed implementation manners
[0023] In order to make the objectives, technical solutions and advantages of this application clearer, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0024] Before further describing the embodiments of this application in detail, the nouns and terms involved in the embodiments of this application are described. The nouns and terms involved in the embodiments of this application are applicable to the following explanations: Vision-Language Models (VLMs): Vision-Language Models combine two modalities, vision and language, and are committed to solving the cross-modal understanding problem between images and texts. Its core task is to understand and generate language information related to visual content through cross-modal representation learning, or generate corresponding visual content from language information. Compared with traditional single-modal models, the advantage of Vision-Language Models is that they can combine visual information and language information for various aspects of reasoning and generation. For example, in the image description task, the Vision-Language Model needs to understand information such as objects, scenes, and actions in the image and generate an accurate description through language. Conversely, in the image question answering task, the model needs to extract relevant information from the image based on a natural language question and give an accurate answer.
[0025] Pet medical expenses are one of the relatively large expenditures in pet-raising expenses. Especially for some novice pet owners, due to the lack of rich pet-raising knowledge, when encountering pet health problems, they don't know how to deal with them and need to go to the pet hospital. The necessary examination items and treatment items in the pet hospital will cost a lot of money, bringing a greater economic pressure to pet owners. In addition, for some novice pet owners, sometimes when the pet has initial health abnormalities, they cannot be detected in time, resulting in the aggravation of the pet's health problems due to the failure to be treated in time. The inventors of this application found in their research that most of the obvious health problems of pets are easily detected in time (such as some external injuries), while the hidden health problems are not easily detected (such as internal medicine diseases), and most of the hidden health problems are related to the pet's gastrointestinal tract. Therefore, the health problems of pets can be accurately analyzed through the analysis of defecation behavior and excrement.
[0026] Based on this, in order to improve the analysis efficiency and accuracy of pet health problems, the present application provides a pet health analysis method. This method sets up a pet toilet device with a camera device to collect video of the toilet process when the pet uses the toilet, and then conducts pet health analysis based on the collected video. Next, the pet health analysis method provided by the embodiments of the present application will be described.
[0027] Referring to Figure 1 , it is a schematic diagram of the system applied to the pet health analysis method provided by the present application. As shown in the figure, the system includes at least one pet toilet device 110, at least one pet management application client 130, and a pet management application server 120. Among them, the pet toilet device 110 can be a device that provides a pet excretion place and collects or processes pet excrement. For example, when the pet is a cat, the pet toilet device 110 can be a litter box. The pet management application client 130 is specifically a terminal loaded with a pet management application. This terminal can be a mobile terminal such as a smart phone, or various forms of terminals such as a tablet, a personal computer, a head-mounted device, or a vehicle-mounted terminal. In addition to managing the pet toilet device, the pet management application can also control other pet devices such as pet feeding devices and pet toy devices. The pet management application server 120 can be a server that provides services such as management, data processing, and interaction for multiple pet management application clients. The pet management application server 120 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part (such as a virtual machine) allocated from a high-performance computer, a combination of parts (such as virtual machines) allocated from multiple high-performance computers, etc. The pet management application server 120 can be a physical server or a cloud server.
[0028] In the embodiments of the present application, the pet owner can initiate a pet health analysis instruction through the pet management application client 130. After receiving the pet health analysis instruction, the pet management application server 120 determines the corresponding pet toilet device 110 according to the pet information in the instruction. Then, the pet toilet device 110 conducts pet approach detection. When it detects that a pet approaches and enters the pet toilet device 110 to use the toilet, it triggers the activation of multiple camera devices in the pet toilet device 110 to collect toilet videos of the pet. The pet toilet device 110 sends the collected pet toilet videos to the pet management application server 120. The pet management application server conducts toilet behavior category recognition based on the pet toilet videos and extracts excrement image frames from the toilet videos for physiological index recognition. Further, it inputs the recognized toilet behavior category information and physiological index information into a vision-language model for health analysis to obtain a pet health analysis result. Then, the pet management application server 120 sends the pet health analysis result to the corresponding pet management application client 130.
[0029] It can be understood that the pet management application server 120 can be connected to multiple pet toilet devices 110 and multiple pet management application clients 130. The three shown in the figure are only for illustration and do not limit the number of pet toilet devices 110 and pet management application clients 130.
[0030] Refer to Figure 2 , which is a schematic flow chart of the pet health analysis method provided by this application.
[0031] In some embodiments, the pet health analysis method provided by the embodiments of this application includes but is not limited to steps S201 to S204.
[0032] Step S201, obtain a pet toilet video.
[0033] Among them, the pet health analysis method provided by this application can be specifically applied to the pet management application server 120. The pet toilet video can be specifically obtained in real time from the pet toilet device 110, or can be obtained from the storage area in the pet management application server 120.
[0034] In the embodiments of the present disclosure, the pet toilet device can not only be internally loaded with multiple camera devices, but also be provided with a variety of sensors such as a gravity sensing device, an infrared sensing device, a sound detection device, and an odor detection device. The multiple camera devices loaded in the pet toilet device can perform video acquisition in real time and upload it to the corresponding storage space in the pet management application server 120 for storage. When health analysis of the pet is required, the video collected by the pet toilet device corresponding to the pet can be retrieved from the storage space, and then the toilet video can be extracted from it for analysis. Or, in some embodiments, the pet toilet device 110 can trigger video acquisition when receiving a video acquisition instruction from the pet management application server 120. For example, when the user sends a health analysis instruction to the pet management application server 120 through the pet management application client 130, the pet management application server 120 sends a video acquisition instruction to the pet toilet device 110 based on the health analysis instruction, and at this time the pet toilet device 110 triggers the camera device to perform video acquisition. In this way, it is possible to avoid video acquisition when health analysis is not required, thus causing waste of the energy of the camera device and the storage space in the server. Or, in still other embodiments, after receiving the health analysis instruction sent by the pet management application client 130, the pet management application server 120 does not directly send a video acquisition instruction to the pet toilet device 110, but performs pet proximity sensing detection based on the sensing data returned by other sensors of the pet toilet device 110. When it is detected that the pet approaches the pet toilet device 110, a video acquisition instruction is sent to the pet toilet device 110 to start collecting the toilet video of the pet. In this way, more accurate video acquisition can be further achieved, avoiding waste of energy and storage resources caused by collecting invalid videos.
[0035] In some embodiments, the pet health analysis instruction may include pet information, such as specifying health analysis for one or several pets. However, for some multi-pet families, there may be a situation where multiple pets share a pet toilet device. In this case, after the pet toilet video is collected, the pets in the video collected by the pet toilet device can be further subjected to pet identity detection, and then according to the identity detection result, the pet toilet video corresponding to the pet whose identity is consistent with the pet information included in the pet health analysis instruction can be screened out from the videos collected by the multiple camera devices. In this way, it is possible to avoid interference from the toilet videos of other pets on the pet health analysis result and obtain a more accurate pet health analysis result.
[0036] In some embodiments, the process of collecting the pet toilet video includes: Obtain the environmental information around the pet toilet device, and perform pet proximity detection based on the environmental information to obtain a detection result; When it is determined that the pet enters the pet toilet device based on the detection result, multiple camera devices in the pet toilet device are triggered to collect pet toilet videos.
[0037] In the embodiment of the present application, before the camera device in the pet toilet device collects the pet toilet video, the pet approach detection can be performed first. When it is detected that the pet approaches and it is determined that the pet enters the pet toilet device, multiple camera devices set in the pet toilet device are triggered to collect the pet toilet video. In this way, the resource waste caused by keeping the camera device always on can be avoided, and the collection of noisy videos can also be reduced, improving the execution efficiency of the downstream pet health analysis task.
[0038] Specifically, the environmental information around the pet toilet device can be obtained first. For example, infrared detection can be performed near the pet toilet device to determine whether an animal approaches. When it is detected that an animal approaches, it is then detected whether the animal enters the pet toilet device. Or, video collection can be performed near the pet toilet device, and it is judged whether a pet approaches through the collected video. Or, a gravity sensing device is set in the pet toilet device, and the pet approach detection is performed through the gravity sensing device.
[0039] Step S202, identify the pet toilet behavior category of the pet toilet video to obtain the toilet behavior category information.
[0040] Among them, after the pet toilet video is obtained, the pet toilet behavior category of the pet can be identified based on the pet toilet video to obtain the toilet behavior category information. Among them, the pet toilet behavior category can be one or more of a variety of preset toilet behavior categories related to pet health information. For example, the pet behavior category can include defecation behavior, urination behavior, constipation behavior (difficult defecation), dysuria, frequent urination, and pain manifestation, etc. Among them, some pet toilet behavior categories can be obtained through the analysis of the pet posture in the pet toilet video, some pet behavior categories can be obtained through the analysis of the pet posture and pet expression in the pet toilet video, and some pet behavior categories can be obtained through the analysis of the pet toilet video sequence and the collection time of each video in the sequence.
[0041] In some embodiments, to identify the pet behavior category based on the pet toilet video, a pet toilet behavior category recognition model can be used for identification. Specifically, the pet behavior category recognition model can be a multi-classification model, and the base of the model can be a recurrent neural network model for processing image sequence data. First, a large number of pet behavior videos can be used to pre-train the model without supervision, enabling the model to learn the ability to extract pet behavior features. Then, some pet toilet video samples or sequences of pet toilet video samples can be constructed, and each pet toilet video sample or sequence of pet toilet video samples corresponds to a single toilet behavior category label. Furthermore, the constructed training sample data can be used to fine-tune the multi-classification model with supervision, thereby obtaining the pet toilet behavior category recognition model. Next, the pet toilet behavior category recognition model can be deployed to the pet management application server. When it is necessary to identify the pet toilet behavior category based on the pet toilet video, the pet toilet behavior category recognition model can be called to identify the toilet behavior category information corresponding to the pet toilet video.
[0042] Among them, in some embodiments, to identify the pet toilet behavior category of the pet toilet video and obtain the toilet behavior category information, it includes: Frame-split the multi-angle video data included in the pet toilet video respectively to obtain a video frame sequence corresponding to each video data; Perform frame-level time alignment on multiple video frame sequences, and combine the multiple video frame sequences into a target video frame sequence according to the alignment result; Perform toilet behavior category recognition based on the target video frame sequence to obtain the toilet behavior category information.
[0043] In the related art, when identifying the pet toilet behavior category based on the pet toilet video, generally, the behavior category is identified for the videos collected from different angles respectively, and then the confidence levels of the multiple behavior category recognition results are evaluated, and finally the final toilet behavior category information is determined based on the confidence level evaluation result. However, this method is difficult to accurately associate the pet toilet behavior videos taken from different angles, resulting in low accuracy of pet behavior category recognition.
[0044] In the embodiments of the present application, after collecting pet toilet videos taken from multiple angles, the video data of multiple angles can be first split into frames respectively, so as to obtain a video frame sequence corresponding to each video data. Each frame in the video frame sequence corresponds to a collection time information. Then, based on the collection time information corresponding to each video frame, the video frames in multiple video frame sequences can be aligned at the frame level in terms of time. Specifically, the video frames in different video frame sequences corresponding to the same collection time are determined as a video frame group corresponding to the collection time, and then multiple video frames in the video frame group are combined to obtain a target video frame corresponding to the collection time. Among them, combining multiple video frames in the video frame group can specifically be to crop the overlapping parts in the video frames and only keep one copy, then determine the splicing direction of multiple video frames, and then splice these multiple video frames to obtain a complete target video frame. The target video frame contains pet behavior information of the pet photographed from multiple shooting angles corresponding to the collection time. Then, for each collection time, the corresponding target video frame is spliced according to the above method, and then a target video frame sequence is obtained.
[0045] Further, the toilet behavior category can be identified according to the target video frame sequence, and accurate toilet behavior category information can be obtained.
[0046] Since the pet toilet video includes multiple videos taken from multiple angles, before identifying the pet toilet behavior category based on these videos, the multiple videos can be first aligned at the image frame level in terms of time, and then the pet toilet behavior category can be identified based on the pet toilet video after time alignment, so as to obtain a more accurate pet behavior category recognition result.
[0047] Step S203: Extract the excrement image frames from the pet toilet video, and perform physiological index analysis on the excrement image frames to obtain a physiological index analysis result.
[0048] Among them, the excrement image frames extracted from the pet toilet video can include the image frames of the excrement taken during the pet's excretion process, or the image frames of the position of the excrement taken after the pet toilet device buries the excrement after the pet's excretion ends.
[0049] To extract the excrement image frames from the pet toilet video, specifically, the image frames containing the pet excrement or the excrement burial object can be first extracted from the pet toilet video, and then the part of the pet excrement or the excrement burial object can be intercepted from the image frame to obtain the excrement image frame.
[0050] In some embodiments, the pet toilet device includes a color-changing excrement burial object. Extracting the excrement image frames from the pet toilet video and performing physiological index analysis on the excrement image frames to obtain a physiological index analysis result includes: Extract an excrement image frame sequence from the pet toilet video; Perform color development analysis on the excrement image frame sequence to obtain a color development analysis result; Determine the physiological index analysis result based on the color development analysis result and the key image frames extracted from the image frame sequence.
[0051] In the embodiments of the present application, a discolorable pet excrement burial object, such as discolorable cat litter, is also provided. When a pet uses the toilet and the excrement comes into contact with the discolorable pet excrement burial object, the discolorable pet excrement burial object will change color in different ways according to the different properties of the chemical substances in the excrement. For example, the pH value in the pet excrement will cause different color changes in the discolorable pet burial object. The discolorable pet excrement burial object in the pet toilet device can be replaced regularly to ensure the color development effect.
[0052] In this way, excrement image frames can be extracted from the pet toilet video. These excrement image frames can be the excrement image frame sequence obtained by photographing the position of the excrement within a certain period of time after the pet toilet device automatically buries the excrement after the pet uses the toilet. This certain period of time can be determined according to the reaction time of the color development reaction. For example, a period of time is set to 5 minutes. After obtaining the excrement image frame sequence, the excrement image frame sequence can be further subjected to color development analysis to obtain a color development analysis result.
[0053] Specifically, the present application can provide an intelligent color analysis system based on computer vision. The system uses RGB color space analysis and combines with the HSV conversion algorithm to accurately identify the color change of the excrement burial object. The system can also configure a fill light group for the camera device to ensure that the color change of the excrement burial object can be accurately captured under different lighting conditions. After capturing the color change of the excrement burial object, the pH value, occult blood index, protein index, etc. of the pet excrement can be inferred based on this.
[0054] In some embodiments, to determine the pH value, occult blood index, protein index, etc. based on the image color change of the excrement burial object, a deep learning model can be used for determination. The deep learning model can be trained with multiple different color image samples with labels such as pH value, occult blood index, protein index, etc., so that the deep learning model learns the ability to infer indicators such as pH value, occult blood index, protein index, etc. based on the image color.
[0055] In addition, in addition to performing color development analysis based on the excrement image frame sequence in the pet toilet video, key image frames can also be extracted from the pet toilet video, and then abnormal detection of the excrement can be performed based on these key image frames. Among them, the key image frames can include the image frames of the excrement taken during the pet's toilet process, the image frames of the excrement taken before the pet toilet device starts automatic burial after the pet finishes using the toilet, and the image frames of the location of the excrement taken after the toilet device automatically buries. Then, based on the shape, color, etc. of the excrement captured in the key frames, analyze whether there are abnormal excrements such as bloody stools, soft stools, loose stools, and containing parasites, and determine the complete physiological index analysis result based on these abnormalities and pH value, occult blood index, protein index, etc.
[0056] Among them, for the specific analysis process of analyzing whether there are abnormal excrements such as bloody stools, soft stools, loose stools, and containing parasites based on the shape, color, etc. of the excrement captured in the key frames, one or more pre-trained image recognition models can be called for analysis. For example, a bloody stool recognition model can be called, which identifies whether there is a bloody stool abnormality in the excrement by detecting the color of the excrement in the key frames; and a soft stool recognition model can be called, which can identify whether there is an abnormality of soft stools or loose stools in the excrement by detecting the shape of the excrement in the key frames; in addition, a parasite recognition model can also be called, which can identify whether there are parasites in the pet's excrement by combining the recognition of different-color points in the key frames and the inter-frame difference. Among them, there is a high probability that the different-color points in the excrement in the key frames are parasites, and since some parasites are still alive when excreted, the inter-frame difference between the key frames can be used to determine whether there are surviving parasites, and the surviving parasites will cause an inter-frame difference between the key frames due to wriggling.
[0057] Step S204, based on the toilet behavior category information and the physiological index analysis result, call a vision-language model for health analysis to obtain the pet health analysis result.
[0058] Based on the pet toilet video, the pet toilet behavior category information is recognized and the pet physiological index analysis result is analyzed. Then, based on the recognized toilet behavior category information and the detected physiological index analysis result, health analysis can be further performed to quickly obtain an accurate health analysis result. Specifically, in the embodiments of the present application, in order to improve the accuracy of the pet health analysis result, a trained vision-language model can be used to generate the pet health analysis result.
[0059] Among them, the vision-language model is a multi-modal neural network model that can process both image and text data simultaneously. Its core task is to understand and generate language information related to visual content through cross-modal representation learning, or to generate corresponding visual content from language information. Compared with traditional single-modal models, the advantage of the vision-language model is that it can combine visual and language information for various reasoning and generation. For example, in the image captioning task, the vision-language model needs to understand information such as objects, scenes, and actions in the image and generate accurate descriptions through language. The working principle of the vision-language model: The vision-language model usually achieves cross-modal understanding through the following steps: 1. Image encoding: Use models such as convolutional neural networks (CNNs) or Vision Transformers (ViTs) to extract features from the image. These features are usually low-dimensional representations of the image.
[0060] 2. Text encoding: Use pre-trained language models, such as BERT or GPT, to convert the input text into high-dimensional vector representations. These text representations capture the grammar, semantics, and context information of the language.
[0061] 3. Cross-modal fusion: Fuse the representations of the image and text. Common methods include using two-stream neural networks, attention mechanisms, etc. In this way, the model can learn the mutual relationship between the image and text.
[0062] 4. Reasoning and generation: Based on the fused representation, the model performs reasoning (e.g., image question answering) or generation (e.g., image captioning). Different tasks use different output strategies.
[0063] Please refer to Figure 3 , which is a schematic structural diagram of the vision-language model provided by this application. As shown in the figure, the vision-language model provided by this application includes an image encoder 310, a text encoder 320, a feature fusion layer 330, and a content generation layer 340.
[0064] In the embodiments of this application, based on the toilet behavior category information and the physiological index analysis results, the vision-language model is called for health analysis to obtain the pet health analysis results, including: Obtain the pet health analysis prompt text, which indicates the generation of the pet health analysis results; Input the pet health analysis prompt text, the toilet behavior category information, and the physiological index analysis results into the vision-language model to obtain the pet health analysis results output by the vision-language model.
[0065] That is, in the embodiments of the present application, after identifying the toilet behavior category information and analyzing the physiological index analysis results of the pet, a prompt text for prompting the visual language model to perform health analysis based on the toilet behavior category information and the physiological index analysis results can be further obtained. Then, the pet health analysis prompt text, the toilet behavior category information, and the physiological index analysis results are input into the visual language model for analysis, and the pet health analysis results output by the visual language model are obtained.
[0066] Specifically, the pet health analysis prompt text can be "Please analyze whether there is a health abnormality in this pet?"; the toilet behavior category information can include behavior category texts, such as "abnormal defecation posture", "frequent urination", and can also include video clips related to the behavior category texts, and the clip can be a video clip extracted from the pet toilet video based on the identified toilet behavior category; the physiological index analysis results can also include index texts, such as "high pH value", "urine occult blood", etc., and can also include images corresponding to the physiological index analysis results, such as the image of the discolored excrement burial object and the urine image extracted from the pet toilet video.
[0067] Then, the pet health analysis prompt text, the behavior category text, and the index text can be input into the text encoder 320 for text feature extraction, and the video clip related to the behavior category and the image corresponding to the physiological index analysis result are input into the image encoder 310 for image feature extraction. Then, the extracted text features and image features are fused by the feature fusion layer 330 to obtain fused features, and the content generation layer is further used to generate content based on the fused features, so as to obtain the pet health analysis results.
[0068] In some embodiments, the training process of the visual language model includes: Obtaining a plurality of first pet behavior video samples and pre-training the visual language model based on the plurality of first pet behavior video samples; Obtaining a plurality of second pet behavior video samples, where the second pet behavior video samples include a plurality of pet toilet video samples, the pet health analysis prompt text samples corresponding to each pet toilet video sample, and the corresponding pet health problem text labels; Fine-tuning the visual language model according to the plurality of second pet behavior video samples to obtain the trained visual language model.
[0069] In the embodiments of the present application, a vision-language model is trained by adopting an unsupervised pre-training combined with supervised fine-tuning approach to improve the ability of the vision-language model to analyze pet health. Specifically, a large number of first pet behavior video samples can be obtained first, and then these first pet behavior video samples are used to pre-train the vision-language model. Among them, the first pet behavior video samples can not only include various behavior videos of pets, such as toileting videos, eating videos, playing videos, and resting videos, etc., but also include sample texts of pet behavior categories, pet health evaluation texts, etc. Using these samples to pre-train the vision-language model can improve the feature extraction and processing capabilities of the vision-language model.
[0070] Then, multiple second pet behavior video samples are obtained. Here, the second pet behavior video samples include pet toileting video samples of multiple pets, pet health analysis prompt text samples corresponding to each pet toileting video sample, and corresponding pet health problem text labels. Then, the vision-language model is further fine-tuned using the multiple second pet behavior video samples to obtain the trained vision-language model. Among them, when using the second pet behavior video samples to fine-tune the vision-language model, data such as pet toileting behavior categories and physiological index analysis results can be extracted from the pet toileting video samples in the second pet behavior video samples first, and then the image data and text data in these data are respectively input into the image encoder and the text encoder for feature extraction, and then the extracted features are respectively processed through the feature fusion layer and the content generation layer to obtain the pet health problem prediction result. Further, based on the pet health problem prediction result and the pet health problem text label, the loss function value is calculated, and then the model parameters of the vision-language model are iteratively updated until convergence according to the loss function value, thereby completing the fine-tuning of the vision-language model.
[0071] Please refer to Figure 4, which is a schematic diagram of the data flow in the pet health analysis process of this application embodiment. As shown in the figure, in this application embodiment, a variety of sensors and camera devices are installed in the pet toilet device to collect multimodal data. When the multimodal data during the pet's toilet use is collected, the video stream in the multimodal data can be used to identify the toilet behavior category to obtain the toilet behavior category data, and the color display image in the multimodal data can be subjected to color display analysis to obtain the pet's physiological indicators. Then, the behavior category data, physiological indicators, and auxiliary data in the multimodal data can be input into the vision-language model for health analysis to obtain the health analysis result. The auxiliary data includes the pet's weight data, age data, historical health data, and some expert knowledge in pet health, etc. Specifically, the real-time weight data and text data of the age data of the pet can be obtained from the pet's basic attribute information, and the text data corresponding to the medical record data and corresponding medical data of the pet in the recent year can be obtained. In addition, the text data corresponding to the expert knowledge related to the treatment of pet diseases and pet health management can be obtained. Then, these text data are combined with the pet health analysis prompt text, toilet behavior category text, and physiological indicator text as the text input of the vision-language model, and input into the vision-language model for pet health analysis to obtain the health analysis result output by the vision-language model.
[0072] In some embodiments, after obtaining the pet health analysis result by calling the vision-language model based on the toilet behavior category information and physiological indicator analysis result, it further includes: When the pet health analysis result indicates that the pet has a health abnormality, obtain the treatment plan corresponding to the health abnormality; Generate a health report based on the health abnormality and the treatment plan, and send the health report to the pet management application client corresponding to the pet.
[0073] In this application embodiment, when the pet management application server analyzes the pet's health analysis result based on the obtained pet toilet video, it can determine whether the pet has a health abnormality based on the pet health analysis result. If the pet does not have a health abnormality, the server can generate and save a health analysis log for subsequent pet owners to view through the pet management application. If the pet has a health abnormality, the treatment plan corresponding to the health abnormality can be further obtained, and then a corresponding health report is generated based on the health abnormality and the treatment plan, and the health report is sent to the pet management application client corresponding to the pet.
[0074] Among them, the specific process of obtaining a treatment plan corresponding to a health abnormality based on the health abnormality can be to first generate an abnormality description text of the health abnormality, extract a video stream or key image frames corresponding to the health abnormality from the pet's toilet video, and then input the abnormality description text of the health abnormality and the corresponding video stream or key image frames into a large language model for pet medicine for analysis to obtain a treatment plan output by the large language model for pet medicine.
[0075] Pet medical expenses account for a very high proportion of pet-raising expenses, and currently, pet hospitals charge exorbitantly. If pet owners lack the corresponding basic knowledge of pet medicine and treatment experience, then every time a pet has a health abnormality, they have to go to a pet hospital for treatment, resulting in high pet-raising costs. The pet health analysis method provided in the embodiments of the present application can, after quickly and accurately analyzing the pet based on pet toilet video data to obtain a health analysis result, also give accurate and effective treatment suggestions immediately when it detects that the pet's health is abnormal. On the one hand, it can ensure that the pet's health problems can be discovered in time, and it can also ensure that the pet receives timely treatment and temporary treatment, avoiding small health abnormalities of the pet from being dragged into big problems, and can also reduce the cost of pet medicine and improve the pet-raising experience.
[0076] Please refer to Figure 5 , which is a schematic diagram of an interface of the pet management application provided by the present application. As shown in the figure, when the pet management application client receives a health report sent by the pet management application server, it can display a health report display interface 500 in the application interface. In the health report display interface 500, one or more health abnormality labels 510 can be displayed. The health abnormality labels can include health abnormality information and corresponding treatment plans. Pet owners can further click on the health abnormality label 510 to switch to the health treatment interface, and the health treatment interface can display detailed data of the health abnormality, including video or image data related to the health abnormality, judgment basis data for judging the health abnormality, and detailed recommended treatment measure information. In this way, if it is necessary to go to a pet hospital for treatment, these video or image data related to the health abnormality can be displayed to provide detailed pathological materials for the doctor to make an accurate diagnosis and treatment. In the detailed recommended treatment measure information, data such as recommended pet hospitals, recommended medicines, and purchase links can also be included.
[0077] In some embodiments, generating a health report based on the health abnormality and the treatment plan and sending the health report to the pet management application client corresponding to the pet includes: Obtaining the historical health analysis results of the pet within a preset time period; Generating a health report according to the historical health analysis results, the health abnormality, and the treatment plan, and sending the health report to the pet management application client corresponding to the pet.
[0078] In an embodiment of the present application, to further improve the accuracy of the generated health report, after generating a health analysis result of a pet based on the pet's toilet video and determining a treatment plan corresponding to the health analysis result of the pet, the historical health analysis result of the pet within a certain period of time can be further obtained, and then a health report can be generated according to the historical health analysis result, health abnormality, and treatment plan.
[0079] Specifically, for example, obtaining the historical health analysis result of the pet within a week. If the historical health analysis result of the pet indicates that there is no health abnormality, it is determined that the health abnormality of the pet is a newly emerged health abnormality, and thus a treatment plan can be determined separately according to the health abnormality. If the historical health analysis result indicates that the pet has had a health abnormality before, then the health change situation of the pet (such as in recovery, deterioration, etc.) can be judged based on the historical health abnormality and the currently identified health abnormality. In this way, the treatment plan can be updated according to the health abnormality and the health change situation, and then a health report can be generated based on the updated treatment plan and health change data.
[0080] In some embodiments, after generating a health report based on the health abnormality and the treatment plan and sending the health report to the pet management application client corresponding to the pet, it further includes: Receiving health feedback information for the health abnormality returned by the pet management application client; Updating the sample data for fine-tuning the vision-language model based on the health feedback information, and fine-tuning the vision-language model based on the updated sample data.
[0081] In an embodiment of the present application, when the health report is displayed in the pet management application, a health feedback control can be further displayed in the health report display interface 500. The pet owner can touch the health feedback control to display the health feedback interface, and then the pet owner can fill in the health feedback information in the health feedback interface. Among them, the health feedback information can be the recognition of the displayed health report or the correction of the displayed health report. For example, when the pet owner undergoes a health check by a professional pet doctor and makes an accurate diagnosis and treatment judgment, and finds that the diagnosis and treatment judgment of the pet doctor is inconsistent with the health analysis result identified by the pet management application server, the diagnosis and treatment judgment result of the pet doctor can be uploaded to correct the health analysis result identified by the pet management application server.
[0082] After correcting the health analysis result based on the health feedback information, the sample data for training the vision-language model can be updated based on the health feedback information and the pet's toilet video, and then the vision-language model can be periodically updated and trained based on the updated sample data, so as to ensure that the vision-language model can be continuously optimized, and further improve the accuracy of pet health analysis.
[0083] In summary, the pet health analysis method proposed in the embodiments of the present application obtains a pet toileting video, which includes video data captured from multiple perspectives inside the pet toileting device; identifies the pet toileting behavior category of the pet toileting video to obtain toileting behavior category information; extracts excrement image frames from the pet toileting video, and analyzes the physiological indicators of the excrement image frames to obtain physiological indicator analysis results; and calls a vision-language model based on the toileting behavior category information and the physiological indicator analysis results for health analysis to obtain a pet health analysis result.
[0084] It can be seen from this that the pet health analysis method provided in the embodiments of the present application collects a toileting video based on a camera device provided in the pet toileting device when the pet is toileting, then analyzes the pet's toileting behavior and the state of the excrement according to the toileting video, and then infers the pet's health status based on the vision-language model for the toileting behavior and the state of the excrement. This method does not require manual participation and can automatically perform real-time health analysis on the pet based on the pet's toileting behavior and the state of the excrement when the pet is toileting, which can improve the efficiency of pet health analysis; in addition, this method performs pet health analysis through a trained vision-language model, which can also greatly improve the accuracy of pet health analysis.
[0085] In some embodiments, when training the vision-language model, the pet toileting video and the pet health analysis prompt text can also be directly used as the input data of the model, and the pet health label can be used as the training label of the model to fine-tune the vision-language model. In this way, when using the trained vision-language model for pet health analysis, the obtained pet toileting video can be directly input into the image encoder of the vision-language model for feature extraction, and the pet health analysis prompt text can be input into the text encoder for feature extraction, and then the features extracted respectively are fused and inferred to generate a pet health analysis result. In some embodiments, historical pet toileting videos and health analysis data can also be obtained and then input into the vision-language model together with the pet toileting video obtained this time for pet health analysis to obtain a more accurate health analysis result.
[0086] Refer to Figure 6 , in some embodiments, the embodiments of the present application also provide a pet health analysis device 600, and the pet health analysis device 600 includes: An acquisition unit 610, configured to acquire a pet toileting video, where the pet toileting video includes video data captured from multiple perspectives inside the pet toileting device; An identification unit 620, configured to identify the pet toileting behavior category of the pet toileting video to obtain toileting behavior category information; An extraction unit 630 is configured to extract excrement image frames from pet toileting videos, and perform physiological index analysis on the excrement image frames to obtain a physiological index analysis result; An analysis unit 640 is configured to call a vision-language model based on the toileting behavior category information and the physiological index analysis result for health analysis to obtain a pet health analysis result.
[0087] Optionally, in some embodiments, the analysis unit includes: A first acquisition subunit is configured to acquire a pet health analysis prompt text, where the pet health analysis prompt text instructs to generate a pet health analysis result; An input subunit is configured to input the pet health analysis prompt text, the toileting behavior category information, and the physiological index analysis result into the vision-language model to obtain the pet health analysis result output by the vision-language model.
[0088] Optionally, in some embodiments, the present application further provides a model training device, including: A pre-training subunit is configured to acquire a plurality of first pet behavior video samples, and perform pre-training on the vision-language model based on the plurality of first pet behavior video samples; A second acquisition subunit is configured to acquire a plurality of second pet behavior video samples, where the second pet behavior video samples include a plurality of pet toileting video samples, pet health analysis prompt text samples corresponding to each pet toileting video sample, and corresponding pet health problem text labels; A first fine-tuning subunit is configured to fine-tune the vision-language model according to the plurality of second pet behavior video samples to obtain a trained vision-language model.
[0089] Optionally, in some embodiments, the extraction unit includes: An extraction subunit is configured to extract a sequence of excrement image frames from a pet toileting video; An analysis subunit is configured to perform color development analysis on the sequence of excrement image frames to obtain a color development analysis result; A determination subunit is configured to determine the physiological index analysis result based on the color development analysis result and key image frames extracted from the image frame sequence.
[0090] Optionally, in some embodiments, the pet health analysis device provided by the present application further includes: A third acquisition subunit is configured to acquire a treatment plan corresponding to a health abnormality when the pet health analysis result indicates that the pet has a health abnormality; A generation subunit is configured to generate a health report based on the health abnormality and the treatment plan, and send the health report to the pet management application client corresponding to the pet.
[0091] Optionally, in some embodiments, the generating subunit includes: An obtaining module, configured to obtain the historical health analysis result of the pet within a preset time period; A sending module, configured to generate a health report according to the historical health analysis result, health abnormality and treatment plan, and send the health report to the pet management application client corresponding to the pet.
[0092] Optionally, in some embodiments, the pet health analysis device provided in this application further includes: A receiving subunit, configured to receive the health feedback information for the health abnormality returned by the pet management application client; A second fine-tuning subunit, configured to update the sample data for fine-tuning the vision-language model based on the health feedback information, and fine-tune the vision-language model based on the updated sample data.
[0093] Optionally, in some embodiments, this application further provides a video acquisition device, including: A fourth obtaining subunit, configured to obtain the environmental information around the pet toilet device, and perform pet approach detection based on the environmental information to obtain a detection result; An acquisition subunit, configured to trigger multiple camera devices in the pet toilet device to collect pet toilet videos when it is determined based on the detection result that the pet enters the pet toilet device.
[0094] Optionally, in some embodiments, the recognition unit includes: A splitting subunit, configured to respectively perform frame splitting on the multi-angle video data included in the pet toilet video to obtain a video frame sequence corresponding to each video data; An alignment subunit, configured to perform frame-level time alignment on the multiple video frame sequences, and combine the multiple video frame sequences into a target video frame sequence according to the alignment result; A recognition subunit, configured to perform toilet behavior category recognition based on the target video frame sequence to obtain toilet behavior category information.
[0095] Refer to Figure 7 , Figure 7 illustrates the hardware structure of an electronic device in another embodiment. The electronic device includes: A processor 701, which can be implemented in a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of this application; The memory 702 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 702 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 702, and the processor 701 is used to call and execute the pet health analysis method of the embodiments of this application; The input / output interface 703 is used to implement information input and output; The communication interface 704 is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.); The bus 705 transmits information between various components of the device (such as the processor 701, the memory 702, the input / output interface 703, and the communication interface 704); Among them, the processor 701, the memory 702, the input / output interface 703, and the communication interface 704 achieve communication connections with each other inside the device through the bus 705.
[0096] The embodiments of this application also provide a storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, it implements the pet health analysis method provided in this application.
[0097] The embodiments of this application also provide a computer program product. The computer program product includes a computer program. The processor of the computer device reads and executes the computer program, so that the computer device executes and implements the above-mentioned pet health analysis method.
[0098] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification of this disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of this disclosure described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "contain" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0099] In addition, in each of the embodiments of the present disclosure, each functional unit may be integrated into a processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0100] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a storage medium. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present disclosure. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0101] It should also be understood that the various embodiments provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.
[0102] The above is a specific description of the embodiments of the present disclosure, but the present disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present disclosure.
Claims
1. A pet health analysis method, characterized in that, The method includes: Obtaining a pet toileting video, where the pet toileting video includes video data captured from multiple perspectives inside a pet toileting device; Identifying the pet toileting behavior category for the pet toileting video to obtain toileting behavior category information; Extracting excrement image frames from the pet toileting video and performing physiological index analysis on the excrement image frames to obtain a physiological index analysis result; Based on the toileting behavior category information and the physiological index analysis result, invoking a vision-language model for health analysis to obtain a pet health analysis result.
2. The method according to claim 1, characterized in that, The invoking the vision-language model for health analysis based on the toileting behavior category information and the physiological index analysis result to obtain a pet health analysis result includes: Obtaining a pet health analysis prompt text, where the pet health analysis prompt text instructs to generate a pet health analysis result; Inputting the pet health analysis prompt text, the toileting behavior category information, and the physiological index analysis result into the vision-language model to obtain the pet health analysis result output by the vision-language model.
3. The method according to claim 2, wherein The training process of the vision-language model includes: Obtaining a plurality of first pet behavior video samples and pre-training the vision-language model based on the plurality of first pet behavior video samples; Obtaining a plurality of second pet behavior video samples, where the second pet behavior video samples include a plurality of pet toileting video samples, pet health analysis prompt text samples corresponding to each pet toileting video sample, and corresponding pet health problem text labels; Fine-tuning the vision-language model according to the plurality of second pet behavior video samples to obtain a trained vision-language model.
4. The method according to claim 1, wherein The pet toileting device includes a color-changing excrement burial material, and the extracting excrement image frames from the pet toileting video and performing physiological index analysis on the excrement image frames to obtain a physiological index analysis result includes: Extracting an excrement image frame sequence from the pet toileting video; Performing color development analysis on the excrement image frame sequence to obtain a color development analysis result; Determining the physiological index analysis result based on the color development analysis result and key image frames extracted from the image frame sequence.
5. The method according to claim 1, characterized in that, After the invoking the vision-language model for health analysis based on the toileting behavior category information and the physiological index analysis result to obtain a pet health analysis result, it further includes: When the pet health analysis result indicates that the pet has a health abnormality, obtaining a treatment plan corresponding to the health abnormality; Generating a health report based on the health abnormality and the treatment plan and sending the health report to the pet management application client corresponding to the pet.
6. The method according to claim 5, wherein The generating a health report based on the health abnormality and the treatment plan and sending the health report to the pet management application client corresponding to the pet includes: Obtaining the historical health analysis results of the pet within a preset time period; Generating a health report according to the historical health analysis results, the health abnormality, and the treatment plan and sending the health report to the pet management application client corresponding to the pet.
7. The method according to claim 5, wherein After generating a health report based on the health abnormality and the treatment plan and sending the health report to the pet management application client corresponding to the pet, the following steps are further included: Receiving health feedback information for the health abnormality returned by the pet management application client; Updating the sample data for fine-tuning the vision-language model based on the health feedback information, and fine-tuning the vision-language model based on the updated sample data.
8. The method according to claim 1, wherein The process of collecting the pet's toilet video includes: Obtaining the environmental information around the pet toilet device, and performing pet approach detection based on the environmental information to obtain a detection result; When it is determined based on the detection result that the pet enters the pet toilet device, triggering a plurality of camera devices in the pet toilet device to collect the pet's toilet video.
9. The method according to claim 1, wherein The identification of the pet's toilet behavior category for the pet toilet video to obtain toilet behavior category information includes: Performing frame splitting on the multi-angle video data included in the pet toilet video respectively to obtain a video frame sequence corresponding to each video data; Performing frame-level time alignment on the multiple video frame sequences, and combining the multiple video frame sequences into a target video frame sequence according to the alignment result; Performing toilet behavior category identification based on the target video frame sequence to obtain toilet behavior category information.
10. A pet health analysis device, characterized in that, The pet health analysis device includes: An acquisition unit for acquiring a pet toilet video, where the pet toilet video includes video data captured from multiple perspectives inside the pet toilet device; An identification unit for identifying the pet's toilet behavior category for the pet toilet video to obtain toilet behavior category information; An extraction unit for extracting excrement image frames from the pet toilet video and performing physiological index analysis on the excrement image frames to obtain a physiological index analysis result; An analysis unit for performing health analysis by invoking a vision-language model based on the toilet behavior category information and the physiological index analysis result to obtain a pet health analysis result.
11. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the pet health analysis method according to any one of claims 1 to 9.
12. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the pet health analysis method according to any one of claims 1 to 9.
13. A computer program product, the computer program product comprising a computer program, characterized in that, The computer program is read and executed by the processor of the computer device, so that the computer device executes the pet health analysis method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Health mapping method and system in cat litter state
CN115956518A
Structural health diagnosis vision-language basic model and multi-mode interaction system establishing method
CN117390151A
Panda health assessment method and system based on panda excrement identification
CN117423042A
Pet health monitoring method and system based on prediction model
CN119650074A
Health condition estimation system
JP2024152294A
Cited By
Pet ectoparasite type identification method based on multi-feature data fusion
CN121600506A
Pet body surface parasite species identification method based on multi-feature data fusion
CN121600506B