AI intelligent child camera system
By introducing cognitive understanding, associative memory, creative generation, and learning evolution engines into the children's smart camera system, the problem of lack of deep understanding and personalized creation in existing technologies has been solved. This enables a deep understanding of the children's world and personalized image management, providing a personalized creative generation experience.
Patent Information
- Application Number
- CN202511193034.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-08-25
AI Technical Summary
Existing smart camera systems for children lack deep understanding capabilities, failing to comprehend complex scene relationships and diverse event dynamics. Their image data management is fragmented and lacks creative experience, resulting in poor collaboration and linkage between functional modules and an inability to provide personalized content generation and artistic re-creation capabilities.
Employing a cognitive understanding engine, an associative memory engine, a creative generation engine, and a learning evolution engine, the system performs preliminary and deep recognition through a dynamically switching edge-cloud collaborative cognitive system. It constructs a semantic association graph database for event-based storage and uses a three-stage generation framework and decomposable stream matching to generate animated short films and artistic images.
It achieves a deep understanding of the children's world and personalized image management, can generate meaningful memory networks and dynamic stories, provides a personalized creative experience, and forms an adaptive intelligent image interaction system.
Smart Images

Figure CN120730154B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of children's cameras, in particular to an AI intelligent children's camera system. BACKGROUND
[0002] At present, as a new technology consumer product, the core technology of children's intelligent cameras is mainly based on basic artificial intelligence applications. The products on the market generally have simple object recognition, interesting selfie stickers, basic voice question and answer, and regular album management functions. The technical architecture is essentially a linear superposition of sensors, basic algorithm models and application functions, and each functional module is independent of each other, lacking depth of cooperation and linkage.
[0003] This technical status leads to significant limitations: first, in terms of interaction, the understanding ability of artificial intelligence is superficial, and can only complete the identification task of "what is this", and cannot deeply understand complex scene relationships and diverse event dynamics. Secondly, in terms of content management, image data is stored in the form of time flow, and the large number of fragmented images taken by children that carry precious memories are seriously ignored and wasted in terms of their inherent narrative value and emotional connection. Finally, in terms of creative experience, the functions are mostly limited to applying preset filters or templates, and children are not given the ability to generate personalized content and artistic re-creation. These bottlenecks collectively limit the product from being a "photographing tool" to a "smart companion". SUMMARY
[0004] The AI intelligent children's camera system provided by the present application can build an adaptive intelligent image interaction system that can deeply understand the individual world, intelligently reconstruct image narratives, actively cooperate with content creation, and grow together with children.
[0005] The AI intelligent children's camera system provided by the present application includes: a cognitive understanding engine configured to utilize a dynamically switched end-cloud collaborative cognitive system to perform preliminary identification and deep identification on a first target photo collected by the camera, and display the preliminary identification result and the deep identification result through the camera; an associated memory engine configured to, when a new photo is input by the cognitive understanding engine, store the new photo in association with an existing event; a creation generation engine configured to obtain a child voice instruction and a corresponding second target photo, analyze the child voice instruction based on a three-stage generation framework, generate an animation short film corresponding to the second target photo based on the child voice instruction, and obtain a third target photo, generate an artistic image based on the third target photo by using a decomposable flow matching; and a learning evolution engine configured to utilize a quadrature fine-tuning technology to adjust an AI model in the AI intelligent children's camera system on the camera, so that the AI intelligent children's camera system is adapted to the child.
[0006] The cognitive understanding engine comprises: a device-side recognition module, configured to perform preliminary recognition on the first target photo collected by the camera and display the preliminary recognition result through the camera; and a cloud-side deep recognition module, configured to perform deep recognition on the first target photo when the preliminary recognition result is selected, to obtain a deep recognition result and feed back to the camera; wherein the deep recognition result comprises a target in the first target photo and derivative information of the target.
[0007] The cognitive understanding engine further comprises: a wide-area space perception and three-dimensional pose reconstruction module, configured to perform target recognition on the fourth target photo collected by the camera to obtain at least one target, and select a corresponding correction model for shape correction according to the position and area size of the target.
[0008] The association memory engine comprises: an event-based storage module of a semantic association graph database, configured to store a new photo in association with an existing event when the cognitive understanding engine inputs the new photo; and an intelligent event clustering module, configured to extract a high-latitude feature vector from each photo to be clustered, and cluster a plurality of high-latitude feature vectors.
[0009] The event-based storage module of the semantic association graph database comprises: an analysis entity and relationship unit, configured to extract all key entities and relationships in a new photo when the cognitive understanding engine inputs the new photo; an event attribution determination unit, configured to call the intelligent event clustering module to determine whether the new photo corresponds to an existing event subgraph; and a dynamic update memory network, configured to add the new photo as a new node to the corresponding event subgraph when the new photo corresponds to an existing event subgraph, or create a new event subgraph when the new photo does not correspond to an existing event subgraph.
[0010] The intelligent event clustering module comprises: a high-dimensional feature extraction unit, configured to extract a high-latitude feature vector from each photo to be clustered; and a Laplacian kernel clustering unit, configured to cluster a plurality of high-latitude feature vectors based on a mean shift algorithm of a Laplacian kernel.
[0011] The creative generation engine comprises: a multi-modal narrative video generation module based on a three-stage generation framework, configured to obtain a child voice instruction and a corresponding second target photo, and generate an animated short based on the three-stage generation framework and the second target photo corresponding to the child voice instruction; and an artistic image generation module based on a decomposable flow matching, configured to obtain a third target photo, and generate an artistic image based on the decomposable flow matching and the third target photo.
[0012] The multi-modal narrative video generation module based on a three-stage generation framework comprises: a scene and character selection unit configured to obtain a second target photo and selected protagonist information; a voice wish unit configured to receive a child voice instruction; an intelligent preview and generation unit configured to start a three-stage generation process according to the protagonist information, the second target photo and the child voice instruction, analyze the child voice instruction to generate an action script, reconstruct a three-dimensional model according to the second target photo, and further generate an animation short film corresponding to the child voice instruction according to the animation script and the three-dimensional model; and an intelligent dubbing and music unit configured to perform intelligent dubbing and music for the animation short film.
[0013] The artistic image generation module based on a decomposable flow matching comprises: a style selection unit configured to provide a style selection function when a third target photo is obtained, and receive a target style selected by a child; and a decomposable flow matching unit configured to decompose the third target photo to obtain different layers, redraw each layer according to the target style, and combine all the redrawn layers to generate an artistic image corresponding to the target style.
[0014] The AI intelligent child camera system is the cumulative sum of all the perception, memory, creation and evolution processes experienced from the beginning to the present moment.
[0015] The beneficial effects of the present application are: different from the prior art, the AI intelligent child camera system provided by the present application comprises: a cognitive understanding engine, which is used for preliminary identification and deep identification of a first target photo collected by a camera by using a dynamically switched end-cloud collaborative cognitive system, and displays the preliminary identification result and the deep identification result through the camera; an associative memory engine, which is used for associatively storing a new photo with an existing event when the cognitive understanding engine inputs the new photo; a creation generation engine, which is used for obtaining a child voice instruction and a corresponding second target photo, analyzing the child voice instruction based on a three-stage generation framework, generating an animation short film corresponding to the second target photo based on the child voice instruction, and obtaining a third target photo, generating an artistic image based on the third target photo by using a decomposable flow matching; and a learning evolution engine, which is used for adjusting an AI model in the AI intelligent child camera system on the camera by using an orthogonal fine-tuning technology, so that the AI intelligent child camera system is adapted to the child. In the above-mentioned manner, the current child intelligent camera can break through the fragmentation of functional experience, the superficialization of intelligent interaction, the inefficiency of image management, and the homogenization of creative generation, and construct a self-adaptive intelligent image interaction system that can deeply understand the individual world, intelligently reconstruct the image narrative, actively cooperate with content creation, and grow together with the child. Instead of regarding the camera as a collection of isolated functional modules, a new and organic technical ecosystem is designed. It needs to have cognitive ability beyond simple identification, can automatically organize massive and disordered photos into structured and meaningful memory networks, can provide creative ability from static recording to dynamic story, and deeply personalize this ability for each unique child user, and finally realize a self-evolving and continuously learning intelligent interaction closed loop. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor. Among them:
[0017] Figure 1 is a structural schematic diagram of an embodiment of the AI intelligent child camera system provided by the present application;
[0018] Figure 2 is a structural schematic diagram of an embodiment of the cognitive understanding engine provided by the present application;
[0019] Figure 3 is a structural schematic diagram of an embodiment of the associative memory engine provided by the present application;
[0020] Figure 4is a structural schematic diagram of an embodiment of an eventized storage module of a semantic association graph database provided by the present application.
[0021] Figure 5 is a structural schematic diagram of an embodiment of an intelligent event clustering module provided by the present application.
[0022] Figure 6 is a structural schematic diagram of an embodiment of a creative generation engine provided by the present application.
[0023] Figure 7 is a structural schematic diagram of an embodiment of a multi-modal narrative video generation module based on a three-stage generation framework provided by the present application.
[0024] Figure 8 is a structural schematic diagram of an embodiment of an artistic image generation module based on a decomposable flow matching provided by the present application. DETAILED DESCRIPTION
[0025] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. It can be understood that the specific embodiments described herein are only used to explain the present application, rather than limit the present application. In addition, it should be noted that, for the convenience of description, only parts related to the present application are shown in the drawings, rather than all structures. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0026] Reference herein to "embodiments" means that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears at various places in the specification does not necessarily all refer to the same embodiments, nor is it necessarily mutually exclusive of other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0027] Reference is made to Figure 1 , Figure 1 is a structural schematic diagram of an embodiment of an AI intelligent child camera system provided by the present application. The AI intelligent child camera system comprises a cognitive understanding engine 10, an association memory engine 20, a creative generation engine 30, and a learning evolution engine 40.
[0028] In some embodiments, the cognitive understanding engine 10 is configured to utilize a dynamically switched end-cloud collaborative cognitive system to perform preliminary identification and deep identification on a first target photo collected by the camera, and display the preliminary identification result and the deep identification result through the camera.
[0029] Reference is made to Figure 2The cognitive understanding engine 10 comprises a device-side recognition module 11, a cloud-side deep recognition module 12, and a wide-area space perception and three-dimensional pose reconstruction module 13.
[0030] The device-side recognition module 11 is configured to perform preliminary recognition on the first target photo captured by the camera and display the preliminary recognition result through the camera.
[0031] The cloud-side deep recognition module 12 is configured to perform deep recognition on the first target photo when the preliminary recognition result is selected, obtain a deep recognition result, and feed back to the camera; wherein the deep recognition result comprises a target in the first target photo and derivative information of the target.
[0032] The wide-area space perception and three-dimensional pose reconstruction module 13 is configured to perform target recognition on the fourth target photo captured by the camera, obtain at least one target, and select a corresponding correction model for shape correction according to the position and area size of the target.
[0033] In some embodiments, the cognitive understanding engine 10 is the perception cornerstone of the entire AI child camera intelligent system, and its core mission is to go beyond the traditional camera's "pixel recording" category and achieve accurate, dynamic, and spatially related deep understanding of the physical world in which the child is located. It is not just a simple "label generator", but a complex "scene analyst" that provides high-quality, structured raw information for subsequent memory, creation, and evolution. The engine is mainly composed of two core technical modules: a hybrid end-cloud collaborative cognitive architecture and a wide-area space perception and three-dimensional pose reconstruction.
[0034] 1. The hybrid end-cloud collaborative cognitive architecture is introduced as follows:
[0035] Technical principle:
[0036] Efficiency-first model: This type of model is highly simplified, has small computational load, and has extremely low energy consumption, and can provide a generally correct but limited-precision reference direction very quickly. Its core value lies in ensuring the stability and immediate response of the basic functions of the system.
[0037] Precision-first model: This type of model contains accurate mathematical descriptions of complex details, has a large computational load, and requires longer processing time, but can provide high-precision data. Its core value lies in supporting high-level tasks that require deep analysis and precise decision-making.
[0038] This application discloses a core principle for intelligent design in resource-constrained systems: dynamic, task-driven trade-offs and collaboration between "efficiency" and "precision" must be made, rather than pursuing a single extreme.
[0039] Specific implementation and effects in the camera:
[0040] The present application constructs a dynamically switched end-cloud collaborative cognitive system to solve the contradiction between limited device computing power and unlimited children's desire for knowledge.
[0041] Device end - real-time lightweight identification (efficiency first):
[0042] The camera deploys a deeply optimized lightweight identification model on the local chip. Its design goal is not to pursue a comprehensive understanding of everything in the world, but to quickly and low-power complete preliminary perception and interaction triggering.
[0043] Workflow: When the child holds up the camera and aims at a butterfly, the model can complete the calculation within milliseconds, track the butterfly in the frame on the screen in real time, and give a preliminary judgment with a high probability, such as "insect" or "butterfly".
[0044] Technical effect: This ensures zero delay in interaction for children during their exploration of the world. Every curious move they make can immediately get a response from the camera, and this immediate positive feedback is crucial to maintaining children's attention and exploration enthusiasm.
[0045] Cloud - deep knowledge graph analysis (precision first):
[0046] When the child shows further interest in the preliminary identification result on the screen (for example, "butterfly") and clicks the "What is this?" interaction button, the camera will upload the high-definition image of the frame to the cloud server through the network.
[0047] Workflow: The cloud deploys a large and accurate multi-modal knowledge graph model. The model can not only accurately identify the result from "butterfly" to "is a golden-winged butterfly of the Papilionidae family", but also rely on a large knowledge base to provide rich derivative information, such as its life cycle, distribution area, related myths, and even play a video of it flying or audio of it calling.
[0048] Technical effect: This meets the children's deep learning needs from "what is it" to "why" and "what else".
[0049] Through this hybrid architecture, the camera realizes a new two-stage cognitive process: the immediate response of the device end ignites children's curiosity, and the deep analysis of the cloud end satisfies children's desire for knowledge, perfectly combining efficiency and accuracy.
[0050] 2. Wide-area spatial perception and three-dimensional pose reconstruction are introduced as follows:
[0051] Technical thought source and principle elaboration:
[0052] Children's world is full of big scenes - family gatherings, playing with friends, running and jumping in the living room. Traditional standard lenses are difficult to capture the full picture of these scenes, while wide-angle lenses, although with a wide field of view, will bring serious geometric distortion, especially in the edge area of the image, objects will be severely stretched and distorted.
[0053] This distortion is ineffective for traditional AI algorithms trained based on the pinhole camera model. A person stretched at the edge of the image may no longer be a "person" in the eyes of the algorithm, and their posture and behavior are out of the question. Related research provides a complete solution to this problem:
[0054] Multi-model distortion correction library: Instead of relying on a single mathematical model, the research integrates a series of advanced camera models that can describe different degrees of non-linear distortion (such as equirectangular projection model, double spherical model, etc.).
[0055] Content-based dynamic model selection: The most critical innovation is that the research proposes a heuristic, content-based dynamic model selection method. This means that the algorithm will first analyze the position and size of each target (for example, each person) in the image.
[0056] Then, it intelligently determines which mathematical model in the library can be used to correct the current target to get the closest result to the real-world shape.
[0057] Implementation and effect in the camera:
[0058] Our camera is equipped with a high-quality wide-angle lens and deeply integrates the dynamic distortion correction and three-dimensional posture reconstruction technology mentioned above.
[0059] Intelligent adaptive correction:
[0060] When a child takes a family photo containing multiple people, the cognitive engine will analyze each family member in the picture one by one.
[0061] For the father standing in the center of the picture with less distortion, the system may call a relatively simple correction model.
[0062] While for the mother at the far left edge of the picture, whose body shape has been severely stretched, the system will automatically switch to a more complex double spherical model to accurately correct the image area where she is located.
[0063] Technical effect: The final result is that no matter where each family member is located in the picture, their shape will be accurately restored, laying the foundation for subsequent recognition and understanding.
[0064] From "pixels" to "spatial relationship" understanding:
[0065] On the basis of accurate correction, the next step of the engine is to reconstruct the three-dimensional spatial bone posture of each person in the picture. The effect brought by this is revolutionary:
[0066] The camera is no longer just "seeing" a bunch of pixels, but can "understand" the real scene in three-dimensional space. It "knows" that Grandpa is sitting on the sofa and leaning forward to hold the grandson in front of him. It can "distinguish" whether the two children are standing side by side or chasing each other.
[0067] Technical effect: This accurate understanding of spatial position, body orientation, and limb interaction is a decisive step for the camera to go from "seeing objects" to "understanding events". Instead of isolated "face" or "body" labels, it generates structured scene data containing rich interactive semantics. These data will be seamlessly passed to the next core engine, the associative memory engine 20, as the highest quality input, allowing the latter to build a truly meaningful, event and story-based memory network on this basis.
[0068] In some embodiments, the associative memory engine 20 is configured to store a new photo in association with an existing event when the cognitive understanding engine 10 inputs the new photo.
[0069] In some embodiments, referring to Figure 3 , the associative memory engine 20 includes an eventized storage module 21 of a semantic association graph database and an intelligent event clustering module 22.
[0070] The eventized storage module 21 of the semantic association graph database is configured to store a new photo in association with an existing event when the cognitive understanding engine 10 inputs the new photo.
[0071] In some embodiments, referring to Figure 4 , the eventized storage module 21 of the semantic association graph database includes an entity and relationship parsing unit 211, an event attribution determination unit 212, and a dynamic update memory network 213.
[0072] The entity and relationship parsing unit 211 is configured to extract all key entities and relationships in a new photo when the cognitive understanding engine 10 inputs the new photo.
[0073] The event attribution determination unit 212 is configured to call the intelligent event clustering module 22 to determine whether the new photo corresponds to an existing event subgraph.
[0074] The dynamic update memory network 213 is configured to add the new photo as a new node to the corresponding event subgraph when the new photo corresponds to an existing event subgraph, and to create a new event subgraph when the new photo does not correspond to an existing event subgraph.
[0075] The intelligent event clustering module 22 is configured to extract a high-dimensional feature vector for each photo to be clustered, and to cluster a plurality of high-dimensional feature vectors.
[0076] In some embodiments, referring to Figure 5 The intelligent event clustering module 22 includes a high-dimensional feature extraction unit 221 and a Laplacian kernel clustering unit 222.
[0077] The high-dimensional feature extraction unit 221 is configured to extract a high-dimensional feature vector for each photo to be clustered.
[0078] The Laplacian kernel clustering unit 222 is configured to cluster a plurality of high-dimensional feature vectors based on a mean shift algorithm of a Laplacian kernel.
[0079] In some embodiments, the associative memory engine 20 is configured to implement intelligent reconstruction of stories from photos.
[0080] The associative memory engine 20 is the core hub for AI kids camera to realize the transition from “passive recording” to “active storytelling”. Its mission is to solve the fundamental problem faced by all current photo album applications: how to automatically and intelligently organize hundreds of seemingly isolated, but actually related fragmented photos taken by users into “memory stories” that are rich in emotion, logic and narrative value. The core of this engine is to break the traditional storage mode based on “files”, and introduce two revolutionary technologies to realize the deep reconstruction of image memory.
[0081] 1. Event-based storage based on high-order semantic association graph database.
[0082] Source of technical idea and principle elaboration:
[0083] The core idea of this technology comes from the comparative study of high-order graph database (HO-GDB) and traditional graph database. The association in the real world is complex and diverse, and the traditional graph database has natural defects in expressing this complexity.
[0084] Limitations of traditional graph databases: Traditional graph databases (e.g. Neo4j) model the world based on "node-edge-node" dyadic relationships. For example, it can well represent "Xiao Ming knows Xiao Hong". But when we need to represent a collective event, such as "Xiao Ming, Xiao Hong, Xiao Gang jointly attended Mr. Wang's birthday party", the traditional model is not up to the task. It can only break down this event into a series of dyadic relationships (Xiao Ming-attend-party, Xiao Hong-attend-party...), and the core collective semantics of "jointly attend" and the identity of "birthday party" as a whole are lost in this breakdown.
[0085] Revolutionary breakthrough of high-order graph databases: High-order graph databases completely change this situation. They introduce advanced modeling capabilities such as "hyperedge" and "subgraph as a node".
[0086] Hyperedge: The multi-dimensional interaction of "Xiao Ming, Xiao Hong, Xiao Gang jointly attending" can be directly connected by a "hyperedge", which fully preserves the core semantics of "jointly".
[0087] Subgraph as a node: More importantly, the entire "Mr. Wang's birthday party", including all participants, photos taken, time and location of the event, and all related elements, can be encapsulated as a "subgraph". This "subgraph" itself is given an independent identity in the database and can be queried, connected, and given attributes like a normal node. It elevates the event itself to a first-class citizen in the data world.
[0088] Implementation and effects in the camera:
[0089] We have designed a new event-based image memory storage system for the camera based on the principles of high-order graph databases.
[0090] Workflow: When the cognitive understanding engine 10 transmits a newly taken photo, the associative memory engine 20 no longer simply stores it as a file in the album. It will:
[0091] 1. Analyze entities and relationships: Extract all key entities (face nodes, object nodes, location nodes, timestamp) and relationships (e.g. the "hug" three-dimensional pose interaction relationship from the cognitive engine) from the photo.
[0092] 2. Event attribution determination: through the intelligent event clustering algorithm (see next section for details), determine which existing "event subgraph" this new photo belongs to (e.g., find that it highly coincides with the "weekend beach trip" subgraph in terms of time, location, and participating characters), or whether it marks the beginning of a brand new event.
[0093] 3. Dynamically updating the memory network 213: if the photo belongs to an existing event, add it as a new node to the corresponding subgraph and update the properties of the subgraph (such as extending the time span of the event). If it is a new event, create a brand new subgraph, such as "first visit to the zoo".
[0094] Technical effects: This storage method brings a revolutionary change:
[0095] From "finding photos" to "asking stories": the interaction mode of camera albums has undergone a qualitative change. Parents and children no longer need to laboriously scroll through the timeline to find a certain photo, but can directly ask questions in natural language: "Help me find those photos and videos of me and dad building sandcastles at the beach last summer". The system can directly understand complex event elements such as "beach", "with dad", and "building sandcastles", and accurately locate the "beach trip" event subgraph, returning all relevant images.
[0096] Structured narrative foundation: More importantly, this memory network provides a perfect, structured material library for the subsequent creative generation engine 30. When generating a short film about "family", the system can directly extract all event subgraphs with "family outing" or "family gathering" as the theme from this network as creative material, greatly improving the relevance and emotional depth of content generation.
[0097] 2. Intelligent event clustering based on Laplacian kernel.
[0098] Source of technical ideas and principle elaboration:
[0099] In order to realize the above-mentioned "event attribution determination", that is, how to automatically cluster a large number of unlabeled photos into different events, we introduce the classic Mean Shift clustering algorithm. However, the key to the success of the algorithm lies in the choice of kernel function.
[0100] Related research compares the performance of different kernel functions under different data distributions, especially Gaussian kernel and Laplacian kernel.
[0101] Problem of Gaussian kernel: Gaussian kernel function tends to over-smooth when dealing with "large bandwidth" (i.e. sparse and large-span data points distribution), which easily leads to all data points converging to the same or a few cluster centers. This is disastrous for scenarios that require fine-grained differentiation of different events.
[0102] Advantage of Laplacian kernel: Laplacian kernel function exhibits a "spike" property, giving extremely high weight to nearby points and rapidly decaying weight to distant points. Studies have shown that in "large bandwidth" scenarios, this property makes it more robust in maintaining the independence of different data clusters, effectively preventing the false merging of cluster centers, and thus finding a cluster number and cluster division closer to the true situation.
[0103] Specific implementation and effects in the camera:
[0104] Children's shooting behavior is a typical "large bandwidth" data pattern by nature: they may have taken a few photos in the park in the morning and a few at home in the afternoon, with a great leap in theme and scene. Therefore, we creatively apply the Laplacian kernel function to the event clustering task of the camera.
[0105] Workflow:
[0106] 1. High-dimensional feature extraction: For each photo to be clustered, the system first extracts a high-dimensional feature vector, which contains all the information provided by the cognitive engine: recognized object labels, face embedding vectors, geographic location, shooting time, color histogram, etc.
[0107] 2. Laplacian kernel clustering: The system uses these feature vectors to run a mean shift algorithm based on the Laplacian kernel in the background. The algorithm will automatically find density peaks in these high-dimensional data points and divide them into different clusters.
[0108] Technical effects:
[0109] Accurate event boundaries: Due to the superior properties of the Laplacian kernel, the algorithm can extremely accurately cluster photos into different, meaningful events. For example, it can clearly distinguish between "morning park outing on Saturday" and "birthday party in the afternoon on Saturday" as two completely independent events, even though they are very close in time.
[0110] Fully automatic album organization: This means that the user's album is automatically organized by events from the first photo. There is no need to manually create albums or add tags. Each clustering result automatically corresponds to an "event subgraph" in the high-order graph database. This unprecedented automation and intelligent organization capability completely liberates the user from the tedious album management.
[0111] Through the deep combination of the two technologies, the associative memory engine 20 successfully transforms the camera from a simple image storage device into a "smart memory housekeeper" that can deeply understand and intelligently reconstruct a child's personal history. The event-based and networked memory library it builds not only provides great convenience itself, but also lays a solid and rich narrative foundation for subsequent personalized creation and intelligent interaction.
[0112] In some embodiments, the creative generation engine 30 is configured to obtain a child voice instruction and a corresponding second target photo, parse the child voice instruction based on a three-stage generation framework, generate an animation short film corresponding to the child voice instruction based on the second target photo, and obtain a third target photo, generate an artistic image based on the third target photo based on decomposable flow matching.
[0113] In some embodiments, referring to Figure 6 , the creative generation engine 30 includes a multi-modal narrative video generation module 31 based on a three-stage generation framework and an artistic image generation module 32 based on decomposable flow matching.
[0114] The multi-modal narrative video generation module 31 based on a three-stage generation framework is configured to obtain a child voice instruction and a corresponding second target photo, parse the child voice instruction based on a three-stage generation framework, and generate an animation short film corresponding to the child voice instruction based on the second target photo.
[0115] In some embodiments, referring to Figure 7 The multi-modal narrative video generation module 31 based on a three-stage generation framework includes a scene and character selection unit 311, a voice wish unit 312, an intelligent preview and generation unit 313, and an intelligent dubbing and music unit 314.
[0116] The scene and character selection unit 311 is configured to obtain a second target photo and selected main character information.
[0117] The voice wish unit 312 is configured to receive a child voice instruction.
[0118] The intelligent preview and generation unit 313 is configured to start a three-stage generation process according to the main character information, the second target photo, and the child voice instruction, parse the child voice instruction to generate an action script, reconstruct a three-dimensional model according to the second target photo, and then generate an animation short film corresponding to the child voice instruction according to the animation script and the three-dimensional model.
[0119] The intelligent dubbing and music unit 314 is configured to perform intelligent dubbing and music for the animation short film.
[0120] In some embodiments, the artistic image generation module 32 based on decomposable flow matching is configured to obtain a third target photo, generate an artistic image based on the third target photo based on decomposable flow matching.
[0121] In some embodiments, referring to Figure 8 , the artistic image generation module 32 based on decomposable flow matching includes a style selection unit 321 and a decomposable flow matching unit 322.
[0122] The style selection unit 321 is configured to provide a style selection function when the third target photo is obtained, and receive a target style selected by the child.
[0123] The decomposable flow matching unit 322 is configured to decompose the third target photo to obtain different layers, redraw each layer according to the target style, and combine all the redrawn layers to generate an artistic image corresponding to the target style.
[0124] In some embodiments, the creative generation engine 30 is used to jump from recording to creation.
[0125] The creative generation engine 30 is the core driving force for AI children's camera to realize the transition from a "recording tool" to a "creative partner". It shoulders the mission of transforming the child's imagination, static image memory, and understanding of the world into dynamic, interesting, and artistic works. Instead of being satisfied with "decorative" creations such as stickers or filters, the engine introduces the most advanced generative artificial intelligence technology to achieve true "creation out of nothing" and "turning stones into gold".
[0126] 1. Multi-modal narrative video generation based on a three-stage generation framework.
[0127] Technical idea source and principle elaboration:
[0128] The core technology of this function deeply integrates two complementary cutting-edge generation models: one is a framework focusing on human-scene interaction video generation (GenHSI), and the other is a model focusing on audio-video synchronization generation (Kling-Foley).
[0129] The essence of the human-scene interaction video generation framework (GenHSI):
[0130] Traditional video generation models often struggle to control the precise interaction between characters and scenes in the generated content, leading to physically unrealistic phenomena such as "feet not touching the ground" and "hands passing through walls". The GenHSI framework creatively breaks down the complex video generation task into three logically clear stages by simulating the real-world film production process, thereby solving this problem:
[0131] i. Script Writing: Utilizing large language models, the system automatically parses user-inputted vague, high-level instructions (e.g., "make my teddy bear dance on the table") into a sequence of specific, executable atomic actions (e.g., "walk to the table," "jump on the table," "start waving hands and dancing").
[0132] ii. Pre-visualization: This is the most crucial step. Before generating the video, the model first "pre-visualizes" in three-dimensional space. It reconstructs three-dimensional models of key objects (e.g., a table) from the inputted scene pictures and then calculates and generates a series of three-dimensional keyframes (3D Keyframes) that contain physical contact between characters and objects. These keyframes ensure that the subsequently generated animation is physically reasonable and logically consistent at core interaction points.
[0133] iii. Animation: With these three-dimensional keyframes as "skeletons," the model uses video diffusion models to perform smooth, visually coherent interpolation between these keyframes, ultimately generating a complete, fluid, and soundless video.
[0134] Capabilities of the Audio-Video Synchronization Generation Model (Kling-Foley):
[0135] Kling-Foley focuses on solving the "silent film" problem. It can analyze the content and dynamics of video pictures and automatically generate high-quality audio that is highly synchronized and semantically matched, including environmental sounds, sound effects, and background music.
[0136] Implementation and Effects in the Camera (the "Magic Theater" Function):
[0137] We seamlessly integrated these two frameworks and launched the "Magic Theater" function in the camera, providing children with an unprecedented storytelling platform.
[0138] Workflow:
[0139] i. Scene and Character Selection: Children first take a photo of their room, which will serve as the "stage" for the story. Then, they can choose a main character from the camera's built-in 3D character library (e.g., a cartoon astronaut) or use a three-dimensional model of their favorite toy generated by the Learning Evolution Engine 40.
[0140] ii. Voice Wishes (Script Creation): Children speak their wishes in voice: "I want the astronaut to fly down from my bookshelf and land on the sofa."
[0141] iii. Intelligent pre-visualization and generation (pre-visualization and animation): The system background immediately starts the three-stage generation process.
[0142] The GenHSI module first parses the instructions into an action script of "on the bookshelf" -> "take off" -> "fly in the air" -> "land on the sofa".
[0143] Next, it quickly reconstructs simplified 3D models of the bookshelves and sofa based on photos of the room. Then, it calculates two core 3D keyframes: one showing the astronaut standing on the bookshelf, and the other showing the astronaut landing steadily on the sofa. These two frames ensure that the beginning and end of the story are physically plausible.
[0144] Finally, the animation module generates a smooth animation of the astronaut's flight between these two frames.
[0145] iv. Intelligent voiceover and background music: The generated silent video is immediately sent to the Kling-Foley module. This module will add a "whoosh" sound to the astronaut's takeoff, a technologically advanced background music to the flight process, and a soft "plop" sound to the landing.
[0146] Technical effects:
[0147] A direct path from imagination to image: Children's imaginations are transformed directly into a vivid and complete animated short film through the most natural means (speech). This is an immersive and highly free creative experience.
[0148] Physically believable virtual interaction: Thanks to the introduction of a 3D pre-visualization stage, the generated videos completely eliminate the "floating" and "clipping" problems of traditional generated videos. Every interaction between characters and scenes is constrained by physical plausibility, giving the virtual story a realistic feel.
[0149] 2. Artistic image generation based on decomposable stream matching.
[0150] Origin and Principles of the Technology:
[0151] This feature aims to transform children's everyday photos into highly artistic works. Its core technology stems from research on the Decomposable Flow Matching (DFM) generation framework.
[0152] Traditional image generation models (such as GANs or early diffusion models) typically generate images in a "one-step" manner. When dealing with high-resolution images with complex details, this often makes it difficult to balance global structure and local texture, easily leading to detail distortion or structural collapse.
[0153] The DFM framework, inspired by the logic of human painting, adopts a generation strategy that progresses from coarse to fine, layer by layer:
[0154] Multi-scale decomposition: It first decomposes an image into multiple layers of different frequencies (e.g., through Laplacian pyramid decomposition). The lowest layer represents the most essential contours and structures of the image, the middle layer represents colors and lighting, and the highest layer represents the finest textures and details.
[0155] Layered generation and matching: When generating a new image, the model does not generate all pixels at once, but starts from the bottom structure layer under the guidance of the target style, and generates layer by layer upwards. Each layer of generation is based on the results of the previous layer, ensuring that the final image maintains global structural harmony while having rich local details.
[0156] Implementation and effects in the camera ("Little Artist" function):
[0157] The camera has a "Little Artist" function, which makes every photo a masterpiece.
[0158] Workflow:
[0159] i. A child takes a photo, for example, a photo of playing in the park.
[0160] ii. Then he selects a target style from the style library, for example, "ink landscape painting".
[0161] iii. The DFM model in the system background starts immediately. It first decomposes the park photo into structure layer, color layer and texture layer.
[0162] iv. Under the guidance of "ink landscape painting" style, the model starts to redraw layer by layer:
[0163] In the structure layer, it retains the basic contours of the figures and trees, but restructures them with freehand strokes.
[0164] In the color layer, it replaces the original bright colors with black, white and gray ink tones, and adds white space.
[0165] In the texture layer, it adds the texture of the paper to the picture.
[0166] v. Finally, all layers are recombined into a new work of art that not only looks like the original prototype, but also has the charm of ink painting.
[0167] Technical effects:
[0168] High-quality artistic re-creation: The characteristics of DFM layer-by-layer generation ensure the high quality of the generated works. It avoids the stiffness and loss of details caused by simple "style filters", and realizes true artistic re-creation.
[0169] Fast preview and efficient interaction: Since the generation process is layer by layer, the system can complete and display the bottom layer structure sketch in a very short time, allowing children to preview the general generation effect in real time. This fast feedback greatly improves the user experience on portable devices with limited hardware resources.
[0170] Through the application of these two generation technologies, the creation generation engine 30 greatly expands the functional boundaries of the camera from simple "recording" and "beautification" to "narration" and "art" in new dimensions, providing children with an unprecedented creative platform that perfectly combines reality and imagination, technology and art.
[0171] In some embodiments, the learning evolution engine 40 is used to adjust AI models in the AI smart child camera system on the camera using orthogonal fine-tuning techniques, making the AI smart child camera system suitable for children.
[0172] In some embodiments, the learning evolution engine 40 is used to achieve true "thousand faces".
[0173] The learning evolution engine 40 is the fundamental guarantee for the AI child camera to achieve deep personalization and continuous growth. Its core mission is to break the "thousand faces" dilemma of all users using the same set of general AI models, allowing the camera to learn and adapt to each child's unique personal world, including the people around them, cherished objects, and unique behavior preferences, just like a true partner. By introducing the industry's most advanced parameter-efficient fine-tuning (PEFT) technology, especially the scalable orthogonal fine-tuning method, the engine solves the core technical problem of continuous model evolution on resource-constrained edge devices.
[0174] Source of technical ideas and principle explanation:
[0175] Personalization dilemma in the era of large models:
[0176] The powerful capabilities of modern AI stem from the massive size of its models. A high-performance face recognition or object recognition model may have billions or even more parameters. It is completely unrealistic to retrain or fully fine-tune such a large model on each user's device for their personal use, in terms of computational resources, storage space, and energy consumption. This has led most consumer AI products to provide services based on general models, rather than deep individualized adaptation.
[0177] The dawn of Parameter-Efficient Fine-Tuning (PEFT):
[0178] To solve this dilemma, the academic and industrial communities have developed PEFT technology. The core idea is that when performing personalized adaptation, we freeze the main parameters of the large pre-trained model and only introduce and train a small number of additional, pluggable "adapter" parameters, usually less than 1% of the total number of parameters. This way, we can leverage the powerful general capabilities of large models while achieving adaptation to specific tasks or data at a very small cost.
[0179] The unique advantages of Orthogonal Fine-Tuning (QOFT):
[0180] Among the many PEFT methods (such as the classic LoRA), we have chosen Orthogonal Fine-Tuning (Orthogonal Fine-Tuning) and its optimized version for quantized models (Quantized Model) QOFT. This technology stems from a deep understanding of the geometric properties of neural networks and exhibits unique advantages over other methods:
[0181] Higher parameter efficiency and stability: Traditional PEFT methods (such as LoRA) achieve adaptation by introducing low-rank matrices, while orthogonal fine-tuning learns an orthogonal transformation matrix to rotate and adjust the feature space within the neural network. Mathematically, orthogonal transformation is a "rigid transformation" that preserves vector length and angle, allowing it to adjust model behavior in a more stable and less parameter-intensive manner, effectively avoiding "catastrophic forgetting" of original model knowledge during fine-tuning.
[0182] Natural affinity for quantized models: To run on edge devices, large models often need to be quantized (e.g., from 32-bit floating-point numbers to 4-bit integers) to reduce size and computational load. Many PEFT methods do not work well when fine-tuning quantized models, as their adaptation methods (such as matrix addition) can disrupt the distribution of quantized data. The "rotation" operation of orthogonal fine-tuning can well maintain the stability of this distribution, making it an ideal choice for efficient fine-tuning of quantized models.
[0183] Implementation and effects in the camera:
[0184] The learning evolution engine 40 deeply integrates the QOFT technology into the bottom layer of the camera operating system and applies it to all AI model modules that need personalization (including face recognition, specific object recognition, and specific voice in voice recognition).
[0185] Workflow: An example of "recognizing mom".
[0186] a. First encounter and labeling: After the child takes a photo of mom with the camera, the parent can assist in a simple labeling operation in the album, such as inputting the label "mom" under an automatically detected face frame.
[0187] b. Background lightweight fine-tuning task starts: This labeling behavior triggers the learning evolution engine 40 to start a QOFT fine-tuning task in the background. The system freezes all parameters of the general face recognition large model built into the camera.
[0188] c. Learning the characteristics of "mom": The system loads the photo (or multiple photos) labeled as "mom" and begins to train a very small orthogonal transformation matrix associated with the face recognition module. The learning goal of this matrix is very clear: rotate all internal representations in the neural network that match the "mom" face features into a completely new, labeled "mom" exclusive subspace.
[0189] d. Model evolution complete: This fine-tuning process may only take a few tens of seconds to a few minutes and is completed silently in the background, completely unaffected by the child's normal use. After completion, the trained orthogonal matrix representing the "mom" characteristics is saved.
[0190] e. Achieve accurate recognition: From now on, when the camera "sees" mom again, the features extracted by the general face recognition model will be transformed by this exclusive orthogonal matrix, and the system can determine with high confidence that "it's mom", not just "a female face".
[0191] Technical effects:
[0192] True "thousand faces for a thousand people": Through this mechanism, each child's camera will gradually establish a unique and private "cognitive model library". It recognizes not only general "cats" and "dogs", but also "our cat, Xiao Hua" and "our neighbor's dog, Dahuang". It recognizes not only "male voices", but also "dad's voice". This deep personalization brings unparalleled warmth and emotional connection.
[0193] Continuous learning and co-growth: Learning evolution is continuous. As the child takes more and more photos, and the annotated entities become richer, the QOFT fine-tuning task in the camera background will be periodically performed, constantly optimizing and updating these personalized "adapters". The camera thus has the ability to co-grow with the child. It not only records the child's growth, but also its own "intelligence" evolves with the child's accompaniment, and understands his / her world better and better.
[0194] Privacy and security protection: Since the core fine-tuning process is completely completed on the end-side device, all training data involving personal biological characteristics (such as human faces) do not need to be uploaded to the cloud. This maximizes the protection of the privacy and security of children and families, and solves the core concerns of users about the data security of AI products.
[0195] In some embodiments, the state of the AI smart child camera system at any given moment is the cumulative sum of all the perception, memory, creation and evolution processes it has experienced since its inception.
[0196] The application defines the final form of the AI smart child camera system as an evolving "intelligent companion" system. Its state at any "present" moment is the cumulative sum of all the perception, memory, creation and evolution processes it has experienced since its inception. This process can be described by a path integral, which integrates all the "narratives" generated by the system on the child's entire "growth path".
[0197] Among them, the definitions of each core parameter and operator are as follows:
[0198] : represents the overall state of intelligence and individuality that the AI camera has reached at the "present" moment. It is a complex system that contains all its memories, abilities and personalities.
[0199] This is a path integral. It represents the final state of the system, which is formed by continuously generating and accumulating small "narrative" units on the child's entire "growth path" (from the first use to the present). This embodies the nature of the system's dynamics, historical dependence and continuous evolution.
[0200] : represents the creative generation engine 30 at time t.
[0201] Operator It itself symbolizes the generation ability of transforming structured data (memory and imagination) into vivid works (video, picture), and its internal implementation corresponds to the multi-modal narrative video generation framework and decomposable flow matching technology.
[0202] Subscript represents the generative behavior of the creation engine is deeply influenced by the personalized model parameters . This means that the created work is not only based on the general model, but also carries the personal color and preference of the child.
[0203] represents the high-order event memory graph constructed by the associative memory engine 20 at time t. It is not static, but constantly enriched over time, and is defined as:
[0204] ;
[0205] represents a structured accumulation (direct sum). It means that the memory graph is a structured sum of all past time perception results, rather than a simple pile-up.
[0206] is an association operator, whose core is a mean shift clustering algorithm based on Laplacian kernel, responsible for intelligently organizing and integrating new perception data into existing event subgraphs.
[0207] is a cognitive operator, representing the hybrid cognitive architecture of the end-to-cloud collaborative cognitive understanding engine 10, which converts the "world" into structured, machine-understandable data.
[0208] represents function composition, indicating that the association process acts on the output of the cognitive process.
[0209] represents the creative instructions input by the child through voice or text at time t .
[0210] The union operation of this set elegantly expresses the essence of creation - it is the fusion of past memories and current inspiration (imagination).
[0211] represents the state of personalized model parameters achieved by the learning evolution engine 40 at time t. These parameters make the general AI model become the child's exclusive model. Its evolution is also an integral process:
[0212] ;
[0213] is the general model parameter when the camera is shipped.
[0214] is the evolution operator, whose core is the Quasi-Orthogonal Fine-Tuning (QOFT) technique.
[0215] A nonlinear model update operator, denoted as , efficiently acts on the large base model .
[0216] Feedback is the feedback loop of the whole system. It indicates that the evolution of the model (the update of ) depends on the accumulation of user feedback (such as likes, shares, and usage time) on all the works created in the past ( ).
[0217] Formula interpretation
[0218] The creativity of this formula system lies in the use of advanced mathematical language such as nested integrals, operator composition, path integrals, and recursive dependence, which deeply reveals the inseparable organic relationship between the various parts of the system, forming a complete logical closed loop:
[0219] 1. The input of cognition is the cornerstone of memory: memory atlas G 高阶 The formation of t is the result of the continuous action of cognitive operator Φ in the long river of time.
[0220] 2. The sedimentation of memory is the source of creation: every creation Ψ is derived from the invocation and reorganization of historical memory atlas G 高阶 . t
[0221] 3. The output of creation is the food of evolution: the driving force of every model evolution comes from the user feedback on past creative works.
[0222] 4. The result of evolution feeds back to creation: the personalized parameters Θ after evolution ( t ) will in turn guide future creations, making them more personalized.
[0223] In summary, this formula describes a complex adaptive system that is self-referential, self-organizing, and self-evolving. It perfectly explains how AI child camera evolves from a fixed tool to a "narrative companion" full of wisdom, memory, and personality through continuous interaction with children. This systematic and emergent intelligent effect is beyond the reach of any single technology or simple combination, fully embodying the deep creativity of this technical solution.
[0224] Further, the four levels of technical means are deeply combined with the "AI wisdom narrative companion system general model", and the integrated scheme is described in detail. How does this single technology stack achieve unexpected technical effects that are deeply creative and non-obvious.
[0225] Unexpected technical effects: from "smart tools" to "wisdom narrative companions".
[0226] The effects of traditional AI products are usually linear and predictable: a better recognition algorithm brings more accurate labels, and a stronger generation model produces more realistic pictures. These are "quantitative" improvements. The "AI wisdom narrative companion system general model" we propose aims to bring about a systematic and non-linear "qualitative change" through the deep coupling and recursive cycle of the four engines. The ultimate effect is not simply the sum of the functions of the four engines, but a new and emergent form of intelligence - the wisdom narrative companion.
[0227] The core qualities of this "companion" are reflected in the following aspects, each of which corresponds closely to a specific structure in the general model formula and exhibits its unexpected depth and wisdom.
[0228] 1. Emergent proactive narrative intelligence: from "passive response" to "active guidance"
[0229] Traditional effect: Traditional children's cameras are passive. They only respond with "this is a flower" when asked by children (e.g., pointing a camera at a flower). Their intelligence is passive, isolated, and one-question-one-answer.
[0230] Emergent effect of this scheme: This scheme's camera has the ability to actively initiate narratives. It can actively make creative suggestions based on a deep understanding of the child's world, just like a thoughtful partner.
[0231] Behind the model and principles:
[0232] This initiative comes from the dynamic link between the memory graph G 高阶 ( t ) and the creativity engine Ψ in the general model formula.
[0233] Pattern discovery of memory: correlation operator Γ in the formula 拉普拉斯 Not only in clustering, it also continuously analyzes high-order event memory graphs G 高阶 t It might find a high-order pattern: "In the past three weekends, the location node 'beach' and the face node 'dad' frequently co-occur in the same 'event subgraph', and the interaction labels of these subgraphs are mostly 'play' and 'happy'".
[0234] From "discovery" to "inspiration": This discovered pattern is no longer a cold data conclusion, but a high-level semantic input that triggers the creative engine Ψ. Instead of waiting for the child to ask questions, the system initiates a dialogue: "I found that you and dad often go to the beach to play recently. How about we create a magical animation about your treasure hunt on the beach together?"
[0235] Unexpected place: This suggestion is not pre-set in the program, but emerges from the system's autonomous learning and pattern mining on the child's personalized and structured memories. It demonstrates the system's qualitative change from a simple "data processor" to a partner who can "understand interests and inspire creativity". The structure of the formula Ψ( G 高阶 ( t ),…) means that creativity is based on the entire memory history, and when a certain pattern emerges, creative behavior can be triggered proactively.
[0236] 2. Deeply emotional co-creation experience: from "general template" to "personal memory sublimation"
[0237] Traditional effect: Traditional content generation, whether it is a video template or an artistic filter, is general. The generated work lacks deep connection with the child's personal life and has weak emotional connection.
[0238] Emergent effect of this solution: Each creation of this solution is an artistic sublimation of the child's personal and real memories. The generated work is full of unique personal imprint and emotional temperature.
[0239] Behind the model and principle:
[0240] This effect comes from the input term of the creative engine in the total model formula G 高阶 ( t ) ∪ {imagination( t )} and the deep involvement of personalized parameters Θ( t ).
[0241] Seamless material calling: When the child accepts the camera's "create treasure hunt animation" suggestion, the creative engine Ψ will seamlessly call all relevant and structured memory elements from the memory graph G 高阶 ( t ).
[0242] It will extract real beach photos in the "beach trip" event subgraph, and construct them as a 3D stage for the animation, through the 3D reconstruction capability of the cognitive engine Ψ.
[0243] It will invoke the learning evolution engine 40 After training, the personalized character models accurately identified as "I" and "Dad" will be the protagonists of the animation.
[0244] The fusion of memory and imagination: {imagination( t )} represents the new instruction "we found a glowing shell!" input by the child at this time. The task of the creation engine Ψ is to let the protagonists of real identity realize this imaginative plot on the stage constructed by real memory.
[0245] Unexpected: This is no longer a simple "photo video", but a perfect fusion and artistic re-creation of the child's past (memory) and present (imagination). The generated work is no longer a strange video for the child, but "a magical story happening to me". This strong sense of identity and emotional resonance is completely unmatched by general template generation, which turns the creation process itself into a deep emotional experience.
[0246] 3. Self-improving wisdom growth closed loop: from "solidified product" to "evolving life"
[0247] Traditional effect: The function and intelligence level of traditional AI products are basically fixed at the time of manufacture. Except for regular cloud model unified updates, it cannot be personalized and self-optimized for each user's unique use habits.
[0248] Emergent effect of this solution: The camera of this solution is a living wisdom body that can grow together with the child. Its "wisdom" and "personality" will evolve over time and become more and more in line with its "little master".
[0249] Behind the model and principle:
[0250] The core of this effect lies in the recursive integral equation that describes the evolution of personalized parameters in the total model formula: ;
[0251] This formula constructs a perfect wisdom growth closed loop.
[0252] Evolution of counter-nurturing: Every piece of work created by the child in collaboration with the camera (Ψ' s output), and every interaction with the camera (e.g., the child's great interest in the "treasure hunt" story, watching it five times), is recorded by the system as "feedback" data.
[0253] Feedback-driven learning: This "feedback" data (Ψ' s output) is what drives the learning evolution engine 40 to perform the next round of orthogonal fine-tuning. "nourishment" for the next round of orthogonal fine-tuning.
[0254] Evolution guiding the future: Based on these feedbacks, the engine updates the personalized model parameters . For example, it might increase the semantic weights associated with "treasure hunt," "ocean," and "father-son interaction." This updated , in turn, influences the attention preferences of the future cognition engine Ψ (perhaps more sensitive to shells on the beach), the clustering logic of the correlation engine Γ (perhaps more likely to group beach activities as a single important event), and the proactive suggestion direction of the creation engine (perhaps more likely to suggest stories with an adventure theme in the future).
[0255] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative, and the division of the modules or units is merely a logical function division. In actual implementation, another division manner can be used, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0256] The integrated units in the above other embodiments, if implemented in the form of software function units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the present application or the essential part or the whole or part of the technical solutions that make contributions to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method described in the embodiments of the present application. The foregoing storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0257] The above merely describes the embodiments of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation, or direct or indirect application in other related technical fields, which is made according to the content of the present application specification and drawings, is also included in the patent protection scope of the present application.
Claims
1. An AI intelligent child camera system, characterized in that, The AI intelligent child camera system comprises: a cognitive understanding engine configured to utilize a dynamically switched end-cloud collaborative cognitive system to perform preliminary identification and deep identification on a first target photo collected by a camera, and display the preliminary identification result and the deep identification result on the camera; an associated memory engine configured to, when a new photo is input by the cognitive understanding engine, store the new photo in association with an existing event subgraph determined by an intelligent event clustering module through an eventization storage module of a semantic association graph database; wherein the intelligent event clustering module is configured to extract a high-dimensional feature vector for each photo to be clustered, and perform clustering on a plurality of the high-dimensional feature vectors based on a mean shift algorithm of a Laplacian kernel; a creation generation engine configured to obtain a child voice instruction and a corresponding second target photo, and analyze the child voice instruction through a multi-modal narrative video generation module comprising three stages of script writing, pre-visualization, and animation generation, to generate an animation short film corresponding to the child voice instruction and the second target photo; an artistic image generation module based on decomposable flow matching, configured to obtain a third target photo, perform multi-scale decomposition on the third target photo, and redraw and merge different frequency levels of the decomposed third target photo layer by layer under target style guidance, to generate an artistic image from the third target photo; a learning evolution engine configured to utilize an extensible orthogonal fine-tuning technique to adjust an AI model in the AI intelligent child camera system on the camera, to start a lightweight fine-tuning task in the background, freeze the main parameters of a pre-trained model, and only train a small orthogonal transformation matrix associated with a specific face, so as to realize personalized adaptation of the model without affecting normal use of the child. 2.The AI intelligent child camera system of claim 1, wherein, The cognitive understanding engine comprises: a device-side identification module configured to perform preliminary identification on a first target photo collected by a camera, and display the preliminary identification result on the camera; a cloud-side deep identification module configured to, when the preliminary identification result is selected, perform deep identification on the first target photo to obtain a deep identification result, and feed back the deep identification result to the camera; wherein the deep identification result comprises a target in the first target photo and derivative information of the target. 3.The AI intelligent child camera system according to claim 1 or 2, characterized in that, The cognitive understanding engine further comprises: a wide-area space perception and three-dimensional pose reconstruction module configured to perform target identification on a fourth target photo collected by the camera to obtain at least one target, and select a corresponding correction model for shape correction according to the position and area size of the target. 4.The AI intelligent child camera system of claim 1, wherein, The multi-modal narrative video generation module based on a three-stage generation framework comprises: a scene and character selection unit configured to obtain the second target photo and selected main character information; a voice wish unit configured to receive the child voice instruction; The intelligent pre-rehearsal and generation unit is configured to start a three-stage generation process according to the main character information, the second target photo and the child voice instruction, analyze the child voice instruction to generate an action script, reconstruct a three-dimensional model according to the second target photo, and further generate an animation short film corresponding to the child voice instruction according to the action script and the three-dimensional model. The intelligent dubbing and music matching unit is configured to perform intelligent dubbing and music matching for the animation short film. 5.The AI intelligent child camera system of claim 1, wherein, The artistic image generation module based on the decomposable flow matching includes: The style selection unit is configured to provide a style selection function when the third target photo is obtained, and receive a target style selected by a child. The decomposable flow matching unit is configured to decompose the third target photo to obtain different layers, redraw each layer according to the target style, and combine all the redrawn layers to generate the artistic image corresponding to the target style.
Citation Information
Patent Citations
Text transformation method and device
CN103838866A
Selective visual display system
US20250094690A1