Pet digital twinning generation and interaction method and system based on multi-modal data and user definition

Through multimodal data processing and custom generation modules, combined with dynamic interaction and scenario synthesis, the dual-mode generation of the virtual pet system is realized, solving the problem of limited application scenarios in the existing technology, meeting the needs of different users, and providing a virtual pet experience with high-fidelity restoration and personalized generation.

CN120335614APending Publication Date: 2025-07-18威海港都貿易有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510494234.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-20
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing virtual pet system cannot simultaneously realize high-fidelity restoration based on real pet data and personalized generation based on user-defined needs, resulting in limited application scenarios and cannot meet the needs of different user groups.

Method used

The multimodal data acquisition and processing module is used to combine the customized pet generation module to generate highly restored digital twins through the multimodal data before death, and supports user-defined parameters to generate personalized virtual pets, and emotional connection is realized through dynamic interaction and situational synthesis modules to build a dual-modal generation architecture.

Benefits of technology

It realizes a seamless integration of high-fidelity restoration of real pets and user-defined virtual pets, covering the full-scene needs from emotion to cloud pet raising, and enhances the realism and personalized experience of virtual pets.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention discloses a pet digital twinning generation and interaction method and system based on multi-modal data and user definition, and a virtual pet with high-fidelity restoration and personalized creation functions is constructed through multi-modal data processing or user parameter input. The technology covers the following two core scenes. In the digital twinning mode, a highly restored virtual body (such as the appearance, behavior and emotional response of a pet 'Rifu') is generated based on pet pre-birth data (such as images, videos and voices), and the emotional accompanying requirement of an owner is met; the user-defined mode supports a user to create a special pet (such as a cat image and a small ball playing habit of'seven-seven ') through a preset template or natural language description, and the requirements of people with limited pet keeping conditions are met. The system realizes immersive pet accompanying experience of full-scene users through a dual-mode architecture, dynamic emotion response and virtual and real scene synthesis, and breaks through the application limitation of the traditional technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence (AI), digital twin technology, multimodal data processing, and human-computer interaction. Specifically, it is a method and system for supporting the generation of a digital twin of a real pet and the creation of a user-defined virtual pet by the user. This technology can generate a high-fidelity digital twin through the data of the pet during its lifetime (suitable for emotional companionship of pet owners), and can also generate a personalized virtual pet based on parameters such as breed, appearance, and personality input by the user (suitable for people who want to keep pets but are restricted by conditions), realizing immersive functions such as feeding, interaction, and scenario simulation. Background Art 1. Current Situation of the Existing Technology

[0002] Limitations of traditional virtual pets: Most of the virtual pets such as electronic chickens and desktop pets have preset images and only provide a fixed interaction mode of "feeding - cleaning - playing", which cannot meet the user's needs for "personalized pets", such as restoring the gait and expressions of real pets, or customizing the ideal pet breed and personality.

[0003] Technical gaps related to pets: Existing pet health management apps (such as TTcare) focus on disease diagnosis, and biological cloning technology depends on the survival of real pets. There is a lack of effective solutions for the scenario of "wanting to generate a virtual pet through AI without a physical pet" (such as renters, students, and allergy sufferers). 2. Defects of the Existing Technology

[0004] Insufficient coverage of the user group: Existing solutions either target "owners who have real pets and need data cloning" (such as emotional compensation after the death of a pet), or only provide virtual pets with preset images, and do not cover the wide range of people who "want to create a personalized pet through AI" (such as users who hope to "generate a golden retriever puppy that can shake hands" or "simulate raising a corgi that is afraid of thunder").

[0005] Single generation mode: It is impossible to simultaneously achieve "high-fidelity restoration based on real pet data" and "AI generation based on user-defined needs", resulting in limited application scenarios of the technology (for example: the owner who has lost a pet needs to rely on historical data, while users without pets cannot create personalized virtual pets). Summary of the Invention

[0006] Provide a virtual pet system that is compatible with the digital twin of a real pet and user-defined generation.

[0007] 1. For pet owners: Generate a highly restored digital twin through multimodal data (videos, photos, voices) during the pet's lifetime to continue the pet's companionship (such as when the pet passes away or is separated due to age or health reasons).

[0008] 2. For potential pet owners: Support the generation of personalized virtual pets through preset templates, natural language descriptions, etc., to solve the pain points of "wanting to keep a pet but being restricted by living environment, time, and health conditions". Technical Solution I. Multimodal Data Acquisition and Processing Module (Data-Driven Mode)

[0009] (1) The input data is as follows (taking the pet puppy "Laifu" as an example).

[0010] 1. Images / videos: 30 photos of "Laifu" from different angles (including features such as erect ears and curly tail), and 5 daily videos (recording behaviors such as eating, running, and wagging the tail when being stroked).

[0011] 2. Voice data: The barks of "Laifu" (such as short barks when seeing the owner and whimpers when hungry), and the sound of the collar bell.

[0012] 3. Text data: User annotations such as "Laifu is afraid of thunder", "likes to be petted on the head", and "takes a nap at 3 pm every day".

[0013] (2) The processing steps are as follows.

[0014] 1. Feature extraction: Extract the limb key points (20 joint points) of "Laifu" running through OpenPose, and use ResNet to identify the coat color (ginger), body type (medium-sized dog), and facial pattern (black nose mirror).

[0015] 2. Temporal sequence modeling: Use LSTM to analyze the video frame sequence and find the correlation between the wagging frequency of "Laifu"'s tail and the happiness value (when the happiness value > 80, the tail wagging frequency ≥ 2 times / second).

[0016] 3. Emotional labeling: Combine the Mel-spectrogram of the barks with user annotations to establish an "emotion - behavior" mapping (such as "afraid → tucking the tail and hiding under the table", "happy → running in circles"). II. Customized Pet Generation Module (Parameter-Driven Mode) 1. User input methods

[0017] Preset template screening: The user selects "pet cat" from the breed library, "short hair, yellow" from the appearance library, and "lively, affectionate" from the personality library.

[0018] Natural language description: The user inputs "Generate a virtual kitten named 'Qiqi' that likes to play with balls and will meow at the owner when hearing the command 'Qiqi'". 2. Core technologies

[0019] Build a pet feature parameter library: It includes more than 200 breed biometrics (such as the "face washing" behavior of kittens and the "licking paws" of puppies), and more than 100 behavior models corresponding to personality tags (such as "friendly" corresponding to "actively rubbing against the user's leg when the user approaches").

[0020] Use ControlNet to generate 3D models and behavior sequences that meet the parameters (such as when "Qiqi" hears the command "eat", it triggers the preset actions of "running to the food bowl + meowing"). III. Dynamic Interaction and Scenario Synthesis Module

[0021] Status management: Define "hunger level", "happiness value", and "health value". Feeding the exclusive dog food for "Laifu" by the user (recommended by the preference model trained with historical data) can reduce the hunger level and increase the happiness value.

[0022] Scenario synthesis: The user uploads a selfie taken by the sea, and the system uses Stable Diffusion to generate a composite picture of "Laifu" running on the beach, and its actions match the "running gait when excited" during its lifetime.

[0023] AR interaction: Based on the mobile phone camera, the user can "pet" the virtual "Laifu" and trigger its reactions in historical data (such as squinting and purring when being petted on the head). Innovation Points

[0024] Dual-mode generation architecture: For the first time, it realizes the seamless integration of "data cloning" and "AI customization", supports both the digital twin of real pets (such as restoring the unique behavior habits of "Laifu") and the creation of ideal virtual pets by users (such as the customized personality of the pet kitten "Qiqi"), covering the full-scenario needs from "emotional continuation" to "cloud pet-keeping".

[0025] Intelligent simulation of biometrics: For the data-driven mode, train a pet-specific behavior model through multi-modal data (such as the hiding location preference of "Laifu" when it is afraid of thunder); for the parameter-driven mode, generate behaviors that conform to biological characteristics based on the breed knowledge base (such as the instinct simulation of a golden retriever "retrieving the ball"), improving the realism of virtual pets.

[0026] Full-scenario emotional connection: The dynamic emotion calculation model adjusts the pet's state in real time according to user interactions (such as when "Laifu" has not been fed for 3 consecutive days, it enters a "depressed" state and lies in the corner; when "Qiqi" is petted by the user, the happiness value increases and triggers affectionate actions), meeting the emotional companionship needs of different users. Specific Implementation Modes

[0027] Example 1: The digital twin generation process of the real pet puppy "Laifu".

[0028] 1. User uploads data: The pet owner uploads 10 videos of the pet puppy "Laifu" before its death (including scenes such as playing, eating, and being sick), 50 photos (front view, side view, running posture), and 5 voice recordings of barking (including three emotions: happy, hungry, and scared) through the App.

[0029] 2. Data preprocessing: After video frame extraction, MMPose extracts limb key points for each frame to generate an action sequence matrix of "Laifu" walking, wagging its tail, etc.; for photos, RetinaFace detects facial features and extracts more than 300 appearance parameters such as ear shape (erect ears) and eye color (amber).

[0030] 3. Model training: Appearance generation: StyleGAN3 generates high-resolution images (1024×1024 pixels) of "Laifu" with different expressions (happy, alert, wronged) based on the photo dataset, with an accuracy rate ≥95%; Behavior model: LSTM analyzes the action sequence and user annotations, and outputs the behavior probabilities of "Laifu" in different states (e.g., the probability of "pawing at the food bowl" when hungry is 80%, and the probability of "hiding under the sofa" during a thunderstorm is 90%).

[0031] 4. Digital twin output: Integrate the 3D model with the appearance and behavior data generated by AI, and render an interactive "Laifu" in the App, supporting users to feed, stroke, and simulate scenarios (such as "taking 'Laifu' for a walk in the park").

[0032] Example 2: Customize the creation process of the virtual pet kitten "Qiqi".

[0033] 1. User input requirements: Users without a real pet enter in the App: "I want a virtual kitten named 'Qiqi', 3 months old, with yellow fur, likes to play with balls, and meows when hearing the name 'Qiqi'".

[0034] 2. AI generation processing: NLP parses keywords: breed = kitten, age = kitten, appearance = short yellow fur, behavior = play with balls, name response; calls ControlNet to generate a 3D model of "Qiqi" (erect ears, long tail, yellow fur), and generates an active behavior sequence according to the "kitten" attribute (e.g., the height of the front paws in the air is ≥20 cm when running and jumping); binds voice interaction: When the Whisper model recognizes the user calling "Qiqi", it triggers a feedback of "staring at you + meowing".

[0035] 3. Virtual pet output: "Qiqi" is presented in the App or mini-program, and users can increase its happiness value through interactions such as feeding and throwing balls, and unlock new clothes (such as the "pinball hand shape" themed dressing).

Claims

1. A method for generating a pet digital twin based on multimodal data, characterized in that, It includes: Collect the images, videos, voices and habit data of real pets, and generate high-fidelity digital twins through feature extraction and time series modeling.

2. The method according to claim 1, wherein The real pet data includes: Multi-angle photos and video frames for appearance restoration; Daily activity videos and call recordings for behavior modeling; User-defined habit descriptions for emotion annotation.

3. A method for generating a pet digital twin based on user-defined, characterized in that, It includes: Receive parameters such as pet breed, appearance, and personality input by the user, and generate customized virtual pets through a conditional generation model.

4. The method according to claim 3, characterized in that, The user input parameters include: Standardized parameters selected from preset templates; Personalized requirements described in natural language.

5. A pet digital twin interaction system, characterized in that, It includes: A dual-mode generation module, and the dual-mode generation module includes: A data-driven generation sub-module for executing the pet digital twin generation method based on multi-modal data as described in claim 1; A parameter-driven generation sub-module for executing the pet digital twin generation method based on user customization as described in claim 3; A dynamic interaction module: realizing real-time interactions such as feeding, petting, and voice commands based on an emotion computing model; A scenario synthesis module: fusing the user's real photos with virtual pets to generate an immersive interaction scenario.

6. The system according to claim 5, characterized in that The dynamic interaction module supports: Triggering corresponding behaviors based on the pet's state; Adjusting the pet's mood according to the user's operation.