Space-time backtracking personal simulation image generation method and system based on real person reality

By constructing a dual-model spatiotemporal coupling system based on real data and combining it with conditional GAN ​​for progressive optimization, the problem of insufficient spatiotemporal correlation between people and environment in existing technologies has been solved. This has enabled high-fidelity historical scene reconstruction and continuous optimization, improving the realism of generated content and user experience.

CN120852602APending Publication Date: 2025-10-28GUANGZHOU SMARTHOME TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510837671.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing generative AI technology cannot achieve high-fidelity reconstruction of specific individuals in specific historical scenarios in the field of smart homes. It severs the spatiotemporal relationship between individuals and their environment, lacks physical constraints, and lacks a dynamic optimization mechanism driven by user feedback, resulting in generated content that deviates from real-world patterns or cannot continuously evolve.

Method used

By collecting real-life biometric data and architectural spatial data through laser scanning and other means, a digital base with spatiotemporal coordinates is constructed. A physiologically driven inverse time algorithm and a scene engine are used to perform spatiotemporal coupling of dual models. Conditional GANs are combined for progressive evolution, strictly following physical rules and biometric characteristics to establish a real-digital verification closed loop.

Benefits of technology

It achieves high-precision reconstruction of personal historical scenes, and the generated content can be verified as a historical digital twin, which improves the effectiveness of memory repair, emotional inheritance and privacy protection, and significantly reduces cognitive confusion and anxiety index.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852602A_ABST
    Figure CN120852602A_ABST
Patent Text Reader

Abstract

The invention relates to a real person reality-based space-time backtracking type personal simulation image generation system and method, which initiate a physical anchoring generation normal form to solve the imaginary risk of the traditional generation type AI in historical scene reconstruction. The system receives current biological characteristic scanning of a target natural person and related building space live-action data (the laser point cloud precision is smaller than or equal to 1.5%), and a digital base with space-time coordinates is constructed; the method comprises the following steps: (1) decoupling a physiological senescence base by a face backtracking model through an inverse time diffusion algorithm, and (2) dynamically injecting environmental elements by a scene rendering model in combination with historical astronomical data and a social event rule base. And the generated content needs to meet the physical verifiability. A user supplements a reality material to trigger a progressive evolutionary mechanism, and only a difference region is finely adjusted after key point comparison, so that the SSIM similarity is improved to more than 0.93. According to the method, the generative AI is pushed from probability sampling to physical rule driving, and a medical-level tool is provided for personal memory repair and emotional inheritance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of generative artificial intelligence and digital twin technology, specifically involving a spatiotemporal retrospective personal simulated image generation system and method (or AI agent) based on real-life scenarios. It combines physical world scan data with dual-model collaborative computation to achieve high-precision reconstruction of personal historical scenes. Unlike "smart home device control" (such as Midea CN119292091A): this invention does not involve IoT device linkage, but focuses on personal historical image generation; unlike "public historical reconstruction" (such as museum digital twins): this invention is strictly limited to private home scenarios, ensuring privacy through biometric encryption and localized processing; unlike "general image-generated video" (such as Stable Diffusion): this invention requires input to be physical scan data. Its main applications include: smart home derivative scenarios, including digitization of family memories (such as the restoration of childhood living room scenes) and personal life journey retrospection (such as the generation of learning images in university libraries); medical and health assistance fields, including memory training for Alzheimer's patients (improving orientation scores) and intervention for emotional disorders (reconstruction of images of the deceased to reduce anxiety levels); privacy-sensitive personal services, including closed-loop family data management (different from reconstruction of public historical figures) and encrypted storage of biometric data (local processing of facial scan data), etc. Background Art

[0002] In recent years, the application of generative artificial intelligence in the smart home field has experienced explosive growth, especially in areas such as scene automation control and device linkage optimization, where significant breakthroughs have been achieved. For example, CN118861385A, a matching-based automatic scene generation method and a computer device (Super Smart Home), proposes an automatic scene generation method based on matching rules. This method uses neural networks to analyze user behavior preferences and enables one-click switching of environmental elements such as lighting and music. CN119292091A, a home control method and home control scene learning method and system (Wuhu Midea Kitchen & Bath Appliances), utilizes a large language model to process multimodal data and autonomously generate home control strategies, reducing the complexity of user operations. Haier Group further applies the GPT model to home control (CN120181071A / CN120181074A), generating device control commands through text parsing to improve the naturalness of interaction. These technologies mark the evolution of smart homes from "basic control" to "scene adaptation." However, their core logic still focuses on optimizing the "current environment," lacking the ability to trace and reconstruct personal historical spatiotemporal data.

[0003] The three core flaws of existing technology: (1) Generation logic of separating people from the scene. Although current patented technologies can generate scenes, they sever the spatiotemporal relationship between people and the environment. For example, CN117910081A, a smart home AI design optimization generation system (Guanglin, Guangxi), relies on the combination of general templates and material libraries uploaded by designers to generate design schemes. Its output is essentially a "stylized collage" and cannot achieve accurate restoration of specific people in specific historical scenes. When users try to reconstruct a childhood living room scene, the system can only call the preset 1990s furniture model library, but cannot combine the architectural structure of the user's real residence (such as the angle of sunlight caused by the orientation of the windows) and the biological characteristics of family members (such as the stage of deciduous teeth growth) for coupling calculation, resulting in the generated content being detached from physical reality; (2) The risk of fictitious creation due to lack of physical constraints. Mainstream generative AI technologies (such as Diffusion models) rely on probabilistic sampling, which can easily produce "illusions" that violate the laws of reality. Take the pet scene generation of most current generative AI as an example: the system can identify pets and automatically generate videos, but if the user inputs a blurry photo of the pet in its infancy, the system will freely fill in the missing toys or furniture, and cannot limit the generation boundary through structural mechanics constraints (such as the location of load-bearing walls) or biological rules (skull growth curve). Such unconstrained generation may lead to cognitive confusion in serious application scenarios (such as memory training for Alzheimer's patients); (3) Static models cannot continuously evolve. Existing solutions generally lack a dynamic optimization mechanism driven by user feedback. The "Smart Home Appliances and AIGC White Paper" points out that large industry models need to have "continuous learning and autonomous evolution capabilities," but current patents such as CN115905688A, a reference information generation method based on artificial intelligence and smart homes (Nanjing Dingshan), only replace all neurons in the neural networks of the multi-layer neural convolutional network model with capsules in the capsule network to finely extract the initial user personal scenario information and obtain accurate user personal scenario information; based on the accurate user personal scenario information, a personalized scenario preference profile of the user is established, but a verification closed loop between real physical data and generation results is not established.

[0004] The technological gaps revealed by patent citations. By comparing recently granted patents, the innovative direction of this invention can be clearly defined: Table 1: Comparison of the innovative direction of this invention with other patents Comparison Dimensions Wuhu Midea CN119292091A Haier Group CN120181071A This invention requires Data input Real-time sensor data (temperature, humidity, human movement) User text commands Historical spatiotemporal coordinates + biometric scanning Generate core Equipment control strategy GPT output character sequence A physical twin of human-scene-time coupling Verification mechanism none Keyword matching Reality anchoring (e.g., reverse verification of deciduous teeth X-rays) Evolutionary ability Model offline update Vocabulary expansion User-driven local fine-tuning .

[0005] Summary of technological gaps: The existing patent system has solved the problem of "automatic generation of current scenes" in the field of smart homes, but has failed to overcome the key challenge of "high-fidelity backtracking of personal historical time and space." In particular, there are insurmountable technological gaps in areas such as the verifiable correlation between generated content and the real world (e.g., verifying the war scene through bullet hole locations) and the continuity of biometric features across time domains (e.g., inferring acne distribution from current nasolabial folds).

[0006] The profound impact of industry pain points. This technological gap directly restricts the realization of high-value scenarios: (1) Memory repair scenario: Alzheimer's patients need to stimulate autobiographical memories through real historical scenes, but the existing system-generated "1980s wedding room" lacks physical constraints (such as the size error of the five-drawer chest > 5cm), which exacerbates the spatial and temporal orientation disorder. (2) Emotional inheritance scenario: The reconstruction of the deceased's image relies on blurry video clips. The general generation model will fabricate details that do not conform to the laws of physics (such as the height of the stove flame deviating from the gas pressure equation), which will cause emotional alienation of the family members.

[0007] Therefore, there is an urgent need for a spatiotemporal backtracking technology that strictly adheres to physical rules and supports gradual evolution, pushing generative AI from a "probabilistic sampling" model to a new paradigm of "reality anchoring." This is precisely the core mission of this invention. Summary of the Invention

[0008] This invention innovatively constructs a spatiotemporal retrospective image generation system based on real-life scenarios. Its core lies in strictly constraining generative artificial intelligence within the objective laws of the physical world. For example... Figure 1 As shown, the system collects real-life biometric features (such as facial geometry) and architectural spatial data (such as a study room with a point cloud density of 2000 points / ㎡) using laser scanning and photogrammetry, constructing a digital base with spatiotemporal coordinates. This base is not a virtual creation platform, but a digital mirror of the real world—the human image topology network contains 526 biomechanical anchor points, the error of the building BIM model is controlled within ≤1.5%, and the environmental physical rule base is accurate to 0.5L / time of exhaled nebulization in northern winters. It is this absolute adherence to physical laws (material elastic modulus, light attenuation coefficient, etc.) that fundamentally distinguishes the system from general image-based video tools: the generated content is not a wild imagination, but a verifiable historical digital twin.

[0009] The soul of the system lies in the dual-model spatiotemporal coupling mechanism (see...) Figure 2The facial reconstruction engine employs a physiologically driven inverse-time algorithm. For example, when generating an image of a university library for professional person A, it infers the density of acne distribution based on the current depth of nasolabial folds. Simultaneously, the scene engine calculates the 14:00 sunlight angle (58°) based on the window position photo. The two are dynamically linked through pupil reflection intensity (ambient illuminance × 0.7). Even more unique is the adaptive labeling of social events: when the user specifies a 2020 scene, the system automatically adds a mask and epidemic prevention broadcast (sound pressure level 65dB), while generating a 1978 wedding room scene restores the enamel basin pattern based on historical data. This triple coupling capability of person-scene-time demonstrates amazing results in Example 4—accurately locating the position of furniture in an old house from 1955 (with an error of only 0.3m) using the voiceprint characteristics (fundamental frequency 120Hz) of an elderly person saying "The wardrobe is here."

[0010] To achieve continuous optimization, the system is designed with a progressive evolutionary mechanism. Figure 3 (Revealing its principle). When the user supplements with real-world materials (such as the deciduous tooth X-ray in Example 2), the difference detection module locks the local fine-tuning area (the oral cavity area accounts for 12% of the image area) and performs targeted optimization through conditional GAN, improving the accuracy of the deciduous tooth gap from the initial 3.5mm to 3.1mm (consistent with medical records). This evolution is by no means a complete overhaul, but rather a precise calibration under physical constraints—in Example 5, the initial velocity of the fireworks' ascent was corrected using mobile phone video recording (28m / s→29.1m / s), adjusting only the motion parameters while preserving the reflective properties of the down jacket material. Each evolution forms a traceable version chain, such as in Example 3, where the height of the kitchen wall cabinet was optimized three times, reducing the error from 7cm to 0.8cm (SSIM similarity 0.96).

[0011] In real-world applications, the system demonstrates transformative value. Figure 4 The case study of image reconstruction of the deceased, using only a 2-minute blurry video and a survey map of the old house, generated a 3-minute continuous scene of making dumplings: the father rolling out the dough to a thickness of 1.3mm conformed to mechanical equations, and the color temperature of the fireworks outside the window was 5500K, verified by a local fireworks composition database, significantly reducing the anxiety index of the deceased's family. Meanwhile... Figure 5 This focused Alzheimer's intervention, using voiceprint mapping to spatial coordinates (with a positioning error of 0.3m for the "wardrobe") and combining it with a 1978 wedding photo to recreate the grain pattern of the dresser, significantly improved patients' orientation scores (e.g., a +3.2 point on the Montreal Cognitive Assessment (MoCA-B) scale). These cases demonstrate the systemic medical value—not only as an emotional support tool, but also as a clinically validated cognitive intervention.

[0012] The innovative breakthrough of this invention lies in three leaps: it is the first time that generative AI has achieved a paradigm shift from "probabilistic sampling" to "physical rule-driven"; it establishes a two-way verification closed loop between "reality and digital" (such as the GPS reverse verification of the study desk location in Example 1); and it opens up new applications in the serious medical field (Example 4 reduces the incidence of attack behavior). All of this stems from a reverence for the real world—when technology is rooted in the laws of physics and human emotions, virtual images can become a beacon illuminating memories.

[0013] This system has cross-platform compatibility in its application design: (1) Independent AI intelligent agent: It can operate independently through devices with screen display such as smartphones, tablets, PCs, smart photo frames, smartwatches, smart glasses, and home robot interactive screens. Users can directly upload physical scan data and generate images. (2) Embedded functional components: integrated into third-party applications (such as health management APP, family cloud photo album) as SDK, and call the core engine through API.

[0014] Key interactive features of this system: (1) Mobile optimization: The smart glasses support real-time scanning of building structures by LiDAR (automatic adaptation of point cloud density ≥500 points / cm²); the smartwatch can generate the scene by voice command (e.g., "Generate a living room scene from 1998"). (2) Privacy protection mechanism: Biometric data is encrypted locally (e.g., iPhone Secure Enclave); building scan data only stores topological relationships (absolute geographic coordinates are deleted). Attached Figure Description

[0015] Figure 1 Three-layer constraint architecture diagram. This diagram reveals the core constraint architecture of the system. The physical world layer acquires data through high-precision scanning (such as the pre-demolition survey of the old house in Example 3). The digital base layer embeds three types of constraints: biological constraints (the child's skull height is 15% higher than that of an adult in Example 2), physical constraints (the elastic modulus of the down jacket is 3 GPa in Example 5), and event constraints (automatically adding masks in 2020). The backtracking engine performs spatiotemporal coupling calculations: when generating the library scene in Example 1, the illuminance of the study table (85 lux) is calibrated in conjunction with the intensity of pupil reflection of the characters. User feedback (such as supplementing mobile phone video in Example 5) triggers the evolution engine, which only fine-tunes the difference areas (such as the deflection angle of the fireworks trajectory) to ensure that the generated content is always anchored to reality.

[0016] Figure 2: Flowchart of the dual-model spatiotemporal coupling. This diagram illustrates the dual-model collaborative mechanism. The facial retrospective model employs a physiologically driven algorithm: Example 2 uses the gaps between deciduous teeth to infer the jawbone state, and Example 4 reconstructs a youthful face based on the fundamental frequency of the voiceprint (120Hz). The scene rendering model performs physical calculations: In Example 3, the height of the stove flame (16cm) conforms to the gas pressure equation. The coupling controller enables human-scene interaction: When generating the library scene in Example 1, the reflectivity of the book paper (42 nits) strictly matches the 14:00 sunlight angle (58°); in Example 4, the rocking chair oscillation frequency (0.8Hz) is correlated with the speech energy spectrum. The output frame undergoes multiple constraint verifications: For example, in Example 5, the location of the firework sparks needs to be simulated to ensure that the virtual content conforms to the laws of physics.

[0017] Figure 3 : Diagram of the gradual evolution mechanism. This diagram reveals the principle of gradual evolution. User-provided real-world materials trigger intelligent evolution: In Example 2, after uploading a deciduous tooth X-ray, the system compares the old and new data (3.1mm gap between deciduous teeth vs. the initially generated 3.5mm), delineating a local fine-tuning area (modifying only the oral cavity). Fine-tuning is performed using conditional GAN: the modification range is limited to within 15% of the image area to avoid excessive changes. In Example 5, the initial velocity of the rising fireworks is calibrated using mobile phone video recording (28m / s → 29.1m / s), adjusting only the motion trajectory parameters while preserving the original material and lighting. In Example 3, the height of the hanging cabinet is corrected using a neighbor's photo (error reduced from 7cm to 0.8cm). The evolutionary process is quantifiable: after fine-tuning, the SSIM structural similarity increases from 0.82 to 0.96, and each evolution retains version traceability, forming a verifiable optimization chain.

[0018] Figure 4 : Schematic diagram of deceased person image generation (Example 3). This diagram illustrates the core technology implementation of Example 3. Input blurry video (320p resolution) is used to extract biometric features through voiceprint extraction: the father's fundamental frequency of 85Hz corresponds to a vocal cord length of 18mm, and the mother's 180Hz corresponds to 12mm. Combined with the old house demolition drawings (scale 1:50), the physical environment of the kitchen is reconstructed: the fog density of the window (1200 water droplets per square centimeter) is calculated based on the indoor and outdoor temperature difference (12℃). Dual-engine collaborative output: the facial reconstruction module generates the lip-sync trajectory of "the filling is too salty" through voiceprint-lip-sync mapping, and the scene module simulates the stove flame (liquefied gas calorific value 11000kcal / m³). Key physical anchor points include: 1) the pressure of the rolling pin keeps the dough thickness at 1.3±0.2mm (compliant with mechanical structure); 2) the color temperature of fireworks outside the window is 5500K, verified by the local fireworks composition database. The generated content is verified by reverse verification through reality: the neighbor confirms that the position error of the hanging cabinet hook is ≤1cm.

[0019] Figure 5Memory Training Application Diagram (Example 4). This diagram illustrates the medical value of Example 4. Input a 15-second video of a nursing home to extract key voiceprints: the resonance frequency of the voice "wardrobe" is 120Hz, bound to spatial coordinates (error 0.3m). Combined with a 1978 wedding photo (faded 30%), the parameters of the five-drawer chest are reconstructed: dimensions 110×50×85cm (based on receipt stubs), and the surface wood grain direction matches the annual ring growth model. The generated scene incorporates cognitive training elements: 1) paste viscosity 22Pa·s (triggers tactile memory); 2) rocking chair oscillation frequency 0.8Hz (vestibular stimulation). The output video has been clinically validated: after viewing, patients can correctly identify the location of the real wardrobe (success rate 78%), and the MoCA-B scale orientation score increases from 1.8 to 5.0. After 6 weeks of continuous use, the incidence of aggressive behavior decreased by 41%, demonstrating the unique value of this system in Alzheimer's disease intervention. Detailed Implementation

[0020] Example 1: University Library Time; (1) Required materials. People's photos: A University graduation photo (including student ID photo); Building photos: High-resolution image of the library facade, screenshot of entrance surveillance, panoramic view of the study area taken with a mobile phone; Related description: "Studying every Wednesday afternoon on the south side of the 3rd floor by the window in 2018"; (2) Key points of data processing. Facial retrospection: Based on the current workplace photo, the characteristics of the student period were reversed (removing nasolabial folds and increasing the distribution of acne). The hairstyle was reconstructed according to the 2018 campus fashion database (sideburn length 7.2±0.3cm); Scene restoration: The sunlight angle was reversed based on the window position photo (14:00 solar altitude angle 58°). The book model was matched with the cover of the school's 2018 textbook (the seventh edition of "Advanced Mathematics"). (3) Generate image description: 30-second short video. 0:00-0:10: Carrying books into the library (the access control system lights up green); 0:11-0:20: Swiping card to select a seat (the machine displays "3F-A07"); 0:21-0:30: Taking notes at the desk (the reflection of the pen matches the lighting conditions of the cloudy day). (4) Beneficial effects. Memory recall: 92% accuracy in key scene recognition; Spatiotemporal accuracy: The position of the study table was verified by alumni with an error of ≤0.5m; Emotional value: The satisfaction of returning to the alma mater 5 years after graduation was very high.

[0021] Example 2: Childhood Living Room; (1) Required materials. People's photos: current high-definition photo of the child + blurry photo of the child at age 6 (must include dental features); Architectural scene: 360° panoramic scan of the current state of the living room; Related description: "The whole family discussed school planning on the eve of the Spring Festival in 1998"; (2) Data processing focus: cross-age reconstruction, inferring the developmental status of the jawbone through the gap between deciduous teeth (average 3.1 mm), and tracing the initial pattern of sofa texture according to wear marks; dynamic restoration, the swing cycle of an old-fashioned wall clock (1.2 seconds / time), and the frequency of parents' gestures (12 times per minute during conversation). (3) Generate image description: 90-second short video, 0:30: the child points to the report card and asks a question (the missing front tooth matches the dental record); 1:05: the father pats the sofa armrest (the spring deformation amplitude is 3mm); 1:20: the mother hands over an apple (the fruit peel has a reflectivity of 0.4). (4) Beneficial effects. Memory repair: Corrects 70% of family misconceptions (such as sofa location); Biological accuracy: Matches medical records with the growth stage of deciduous teeth; Parent-child relationship: Increases family intimacy index by 35% (psychological scale assessment).

[0022] Example 3: Review of images of the deceased before their death; (1) Required materials. People's images: 3 blurry videos of the parents before their death (total duration ≤ 2 minutes); Architectural scenes: survey drawings of the old house before its demolition + photos of the neighbor's house with the same layout; Related description: "Scene of making dumplings in the kitchen on New Year's Eve 2019"; (2) Data processing focus. Facial restoration: Extracting voiceprint features (frequency range 85-180Hz) based on blurred video, matching apron oil stain pattern with database of cooking habits before death; Scene reconstruction: Simulating stove flame height (16cm blue flame on liquefied gas stove), and window fog condensation density (12℃ temperature difference between indoor and outdoor); (3) Generate image description. 3-minute short video: 0:45: Father's rolling force (dough thickness 1.3mm±0.2mm); 1:30: Mother complains "the filling is too salty" (mouth movement trajectory verification); 2:10: Fireworks outside the window illuminate his profile (color temperature 5500K conforms to local fireworks specifications); (4) Beneficial effects. Emotional comfort: User anxiety index decreased by 62% (GAD-7 scale); Historical reconstruction: Kitchen wall cabinet height error ≤1cm (compared with demolition drawings); Technological breakthrough: Generate 3 minutes of coherent action from only 2 minutes of material.

[0023] Example 4: Memory training for the elderly; (1) Required materials. People: 15-second video shot at the nursing home (including audio "This is my hometown"); Architectural photos: exterior photos of residences from various periods (1955-2020); Related description: "Mark the location of the wedding room in 1978"; (2) Key points of data processing. Soundscape fusion: Extract current voiceprint features (fundamental frequency 120Hz) to reconstruct a youthful voice and bind the voice keyword "hometown" to spatial coordinates; Spatiotemporal anchoring: Reconstruct the dimensions of the five-drawer chest (length, width and height 110×50×85cm) according to the furniture receipts of 1978, and calculate the orientation of the window by the reflection of the wedding photo frame; (3) Generate video descriptions. 5 video segments totaling 15 minutes, segment 3: pasting up the double happiness characters on the wedding night (simulated paste viscosity value 22 Pa·s); segment 5: breastfeeding scene in 1980 (rocking chair swing frequency 0.8 Hz); narration: "The wardrobe is here... your grandfather hit it" (sound source localization error 0.3 m); (4) Beneficial effects. Cognitive improvement: The MoCA-B scale orientation score increased by 3.2 points; Memory activation: 78% of the time, the correct recall of the location of objects after viewing was achieved; Emotional stability: The incidence of aggressive behavior decreased by 41%.

[0024] Example 5: Fireworks for the 2023 Spring Festival; (1) Required materials. People's photos: family group photo taken with mobile phones (including EXIF ​​geographic information); Building scene: courtyard drone scan point cloud (accuracy 5mm); Related description: "Set off 'Golden Chrysanthemum' fireworks at 20:15 on New Year's Eve"; (2) Key points of data processing. Dynamic binding: Children's head circumference growth model (annual increase of 0.8cm), physical simulation of down jacket wrinkles (elastic modulus of nylon fabric 3GPa); environmental calculation: fireworks rising trajectory (initial velocity 28m / s, northerly wind level 3), snow footprint depth (weight 65kg sinks 2.3cm). (3) Generate video description. 40-second short video: 0:12: Fireworks "Golden Chrysanthemum" bloom (diameter 4.2m, conforming to packaging instructions); 0:25: Child covers ears and backs away (snow sliding friction coefficient 0.15); 0:38: Sparks splash onto scarf (cotton ignition point 210℃ safety simulation); (4) Beneficial effects. Physical accuracy: The trajectory of the Mars fall is verified by hydrodynamics; Safety warning: Potential danger areas are marked (marked in red if less than 2m from the epicenter); Family value: Videos are generated to replace lost cell phone recordings.

[0025] The technical value comparison of the embodiments is shown in Table 2: Table 2: Comparison of Core Innovations and Data Evolution Mechanisms of the Five Implementation Examples Case Core Innovation Points Data evolution mechanism library Spatiotemporal coupling of campus scenes Added graduation certificate scanning and optimized book details Childhood Living Room Cross-age craniofacial retrograde Supplementing primary tooth X-rays improves accuracy Images of the deceased Voiceprint-driven facial reconstruction Neighbor's photos help adjust kitchen layout Older person's memory Voice-to-spatial coordinates Family members provide additional furniture dimensions. Spring Festival fireworks Dynamic labeling of dangerous areas Mobile phone video recording calibration of fireworks trajectory .

[0026] Summary of the essence of innovation. This patent constrains generative AI within real physical rules through a closed loop of "reality anchoring-generation-feedback": (1) Input must be from the real world (photo / scanned data); (2) The generation process is limited by physical laws (lighting / mechanics / materials); (3) User-supplemented materials trigger local evolution (only the difference area is modified); (4) The essential difference from traditional image-based video is that the output is a verifiable historical digital twin, rather than an artistic creation.

Claims

1. A method for generating spatiotemporal retrospective simulated images based on real-life scenarios, characterized in that, include: a) Step S1: Receive input data from the physical world scanning device, including: (1) Current biometric data of the target natural person (including at least 3 sets of multi-angle facial scans); (2) Multimodal real-scene data of associated building space (laser point cloud or photogrammetric data with an accuracy of ≤1.5%). (3) Spatiotemporal coordinate commands (including historical dates and times); b) Step S2: Construct a digital base based on the input data. The base includes: (1) Parametric human image topology network (integrating ≥200 biomechanical constraint points); (2) Building BIM model (preserving the topological relationships of doors and windows and material properties); c) Step S3: Perform spatiotemporal backtracking calculations using the dual-model collaborative engine: (1) Facial retrospective model: The physiological-driven inverse time diffusion algorithm is used to decouple the physiological aging base and the era style vector in the current appearance; (2) Scene rendering model: The lighting angle is calculated based on historical astronomical data, and environmental elements are dynamically added by calling the social event rule library; d) Step S4: Output simulated historical images, the content of which must meet the following requirements: (1) The changes in the faces of the people conform to the anthropological aging curve (error ≤ 3%). (2) Matching the position of objects in the scene with the building scanning data (error ≤ 2cm); e) Step S5: When the user adds real-world materials, the progressive evolution mechanism is triggered: (1) Locate the difference region by key point comparison algorithm (modify area < 15% of image); (2) Use conditional GAN ​​to fine-tune the local region until the structural similarity SSIM≥0.

93.

2. The method according to claim 1, characterized in that, The facial retracing model is specifically executed as follows: (1) Reconstructing historical facial features based on voiceprint features: Mapping the speech fundamental frequency range to the vocal cord length (formula: L=1 / 4×v / f, where v is the speed of sound and f is the fundamental frequency). (2) Cross-age craniofacial retrospection: The developmental status of the jawbone was calculated by measuring the interdental space (3.1±0.3mm). According to the method of claim 1, the scene rendering model comprises: (1) Social event adaptive tagging module: When a specific historical period is detected, environmental markers are automatically injected (e.g., a mask model is added in 2020). (2) Dynamic marking of dangerous areas: Mark safe distances for explosives, high-temperature objects, etc. (e.g., mark within 2m of the firework's core in red).

3. The method according to claim 1, characterized in that, The gradual evolutionary mechanism includes: (1) Difference convergence control: The fine-tuning amplitude is constrained by the set mode, Δmax≤0.05; (2) Version traceability: Record the hash value of the input material and the coordinates of the modified area for each optimization.

4. A spatiotemporal retrospective simulated image generation system, characterized in that, include: (1) Data acquisition module: configured to acquire the following data through a mobile scanning terminal: facial biometric point cloud (density ≥ 500 points / cm²) and building structure topology map (including load-bearing wall mechanical parameters); (2) Digital base construction module: configured to generate: a parametric model of human figures with a timeline (supporting ±50 years of retrospection), and a building BIM that integrates a historical building materials database (such as the compressive strength of blue bricks in the 1950s of 3.5MPa). (3) Dual-core engine module: configured to perform: appearance-scene lighting coupling calculation (pupil reflection intensity = ambient illuminance × 0.7 ± 0.05), social event rule matching (calling localized event library); (4) Evolutionary Feedback Module: Configured to: parse supplementary materials uploaded by users (such as X-ray films and old drawings) and activate the local fine-tuning channel (only modify the grid vertices in the difference area).

5. The system according to claim 5, characterized in that, The dual-core engine module includes: a physics rule verifier: which refuses to generate images that violate real-world laws (such as images generated before a person's birth date).

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1-4, and includes: (1) Physical constraint database (preset material elastic modulus, light attenuation coefficient); (2) User evolution log (records the Δ value and SSIM score for each fine-tuning).

7. Claim 8: An application deployment configuration of the system as described in claim 5, characterized in that, The system operates in any of the following ways: (1) Standalone application. It runs independently on the interactive screen of smartphones, smart glasses, PCs, smartwatches or home robots, and is configured to: acquire multi-angle facial scans (tilt angle ±30°) through the device camera, call the built-in sensors to generate building space data (such as point cloud scanned by mobile phone LiDAR), and perform dual-model calculations locally (device computing power ≥5 TFLOPS); (2) Embedded functional components. Integrate into third-party platforms in the form of SDK, and receive the following through API interface: encrypted biometric hash value (256 bits in length); building topology map (excluding GPS information); and access token for generating images (valid for ≤24 hours).

8. Claim 9: The application deployment configuration according to claim 8, characterized in that, When running on smart glasses: Historical images are superimposed onto the real scene using a waveguide display screen (spatial positioning error ≤ 1cm). Enable voiceprint binding: When a user looks at a specific object and says "here" (base frequency detection range 80-180Hz), the object is bound to the coordinates of the historical scene.

Citation Information

Patent Citations

  • Reference information generation method based on artificial intelligence and smart home

    CN115905688A

  • Home control method and home control scene learning method and system

    CN119292091A

  • Operation control method of generative pre-training GPT model and electronic device

    CN120181071A

  • Operation control method of generative pre-training GPT model and electronic device

    CN120181074A