Child multi-modal data output method and system

By constructing a cross-wheel semantic chain and introducing a synchronization lock mechanism, the timing disorder and adaptability of recommended content in complex interactions of the agent system is solved, and the dynamic scheduling and structured output of multimodal content is realized, which improves the coherence and adaptability of children's multimodal data output.

CN120296704AActive Publication Date: 2025-07-11JIANGSU PROVINCIAL HEALTH DEV RES CENT

Patent Information

Application Number
CN202510756469.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-11
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

When facing complex continuous dialogues and multiple rounds of requests, existing agent systems lack the ability to construct cross-wheel semantic chains, the recommended content lacks deep adaptation mechanism, and the multi-modal content output lacks synchronous control, resulting in timing disorders and repeated pushes, interfering with the concentration and understanding ability of infants and young children.

Method used

By constructing a cross-wheel semantic chain, generating interactive state data, and introducing a synchronous lock mechanism to schedule the recommended content output queue, perform semantic label annotation and screening of infant stage rules, construct structured content fragments, and realize dynamic scheduling of multimodal content.

Benefits of technology

It improves the timing coherence, semantic correlation and age adaptability of recommended content, ensures that the output content matches children's attention and equipment resources, and reduces the delay and conflict of content output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296704A_ABST
    Figure CN120296704A_ABST
Patent Text Reader

Abstract

The invention discloses a child multi-modal data output method and system, and relates to the technical field of data output. The method comprises the following steps: constructing a semantic chain containing parent interaction requests and historical round input contents, and generating interaction state data reflecting real demands and context changes of parents; then, a synchronous lock mechanism is introduced to schedule a recommended content output queue, synchronous output control data is generated, then semantic label labeling is carried out on key frame actions, adaptive action fragments are screened in combination with infant stage rules, and structured content and interactive prompts are constructed; and finally, dynamic scheduling of multi-modal output contents is realized in combination with interactive prompt data and a synchronous control signal, so that the time sequence coherence, the semantic correlation and the child age adaptability of recommended contents are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data output, in particular to a method and system for outputting multi-modal data of children. Background Art

[0002] With the rapid development of artificial intelligence, big data, and multi-modal interaction technologies, children's content recommendation and intelligent parenting assistance systems have gradually become an important part of the digital services for infant and toddler families. In recent years, many typical intelligent agent applications have emerged. By integrating language models, emotion recognition, and content recommendation mechanisms, they can not only achieve question-and-answer interactions and service recommendations for parents, but also provide more contextually adaptable parenting support services by combining multi-modal data.

[0003] Although some current intelligent agent systems already have semantic recognition and multi-modal push capabilities, there are still significant deficiencies in the face of complex continuous conversations, multi-round requests, and the needs of infant and toddler behavior adaptation. First, traditional intelligent interaction systems mostly output content based on single-round semantic understanding, lacking the ability to construct cross-round semantic chains and difficult to accurately capture the potential needs expressed progressively by parents in continuous interactions. Second, the recommended content often consists of static resource collections, lacking a deep adaptation mechanism for parameters such as the child's age stage, cognitive level, and behavior characteristics. In addition, multi-modal content output often lacks a synchronization control mechanism, resulting in temporal disorder or repeated pushing among recommended contents, thus interfering with the concentration and comprehension ability of infants and toddlers. Especially in the scenario of joint output of multi-source content, the lack of effective support for core links such as semantic scheduling, key frame screening, and structured presentation severely restricts the improvement space of content personalization and response coherence. Summary of the Invention

[0004] In view of the problems existing in the above background art, the present invention is proposed.

[0005] To solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides a method for outputting multi-modal data of children, which includes collecting an interaction request input by a parent in a dialogue window, constructing a cross-round semantic chain by combining the input content of previous rounds to generate interaction status data; based on the interaction status data, applying a synchronization lock mechanism to schedule a recommended content output queue to generate synchronization output control data; according to the synchronization output control data, performing semantic label annotation on key frame actions in a recommended video, screening and adapting action segments according to infant and toddler stage rules to construct structured content segments and interaction prompt data; using the interaction prompt data and the synchronization output control data to perform dynamic scheduling on multi-modal content output.

[0006] As a preferred solution of the multi-modal data output method for children according to the present invention, wherein: the generation of the interaction state data includes: generating a semantic context correlation matrix based on the cross-turn semantic chain through a convolutional attention factor; generating interaction state data according to the semantic context correlation matrix, integrating the context information of the current sentence and the historical interaction, and extracting sub-topic identifiers highly relevant to the current expression as interaction trigger data.

[0007] As a preferred solution of the multi-modal data output method for children according to the present invention, wherein: the generation process of the synchronous output control data includes: receiving the interaction state data, and generating a content scheduling index in combination with the current device state parameters; calling a preset content feature mapping table, performing modal marking and structural splitting on the content to be recommended, disassembling each content into the smallest output unit, and constructing a recommended content output queue; based on the content scheduling index and the recommended content output queue, calling a synchronization lock mechanism to assign dynamic priority tags; generating synchronous output control data according to the priority tags and the device output ability parameters.

[0008] As a preferred solution of the multi-modal data output method for children according to the present invention, wherein: the assignment of dynamic priority tags includes: selecting key output units from the recommended content output queue based on the content scheduling index to construct a content candidate set; calling a time series window mechanism to extract the response delay estimation value of the output unit according to the content candidate set to generate a delay weighting factor; calculating a fusion score value for each output unit with reference to the current device output ability parameters; performing interval mapping on the fusion score value to label the corresponding dynamic priority tag.

[0009] As a preferred solution of the multi-modal data output method for children according to the present invention, wherein: the semantic label annotation includes: calling the selected recommended video segment based on the synchronous output control data and extracting a continuous key frame sequence, aligning the time axis of each key frame image data with the corresponding voice command; using a natural language processing model to perform bidirectional matching on the keywords and key frame image content in the semantic chain, and annotating the semantic labels of the corresponding actions at the frame granularity.

[0010] As a preferred solution of the multi-modal data output method for children according to the present invention, wherein: the screening of the adapted action segment includes: confirming the structural adaptability of the action based on the rules in the infant stage, and screening out the action frame segments that are stage-matched.

[0011] As a preferred solution of the multi-modal data output method for children according to the present invention, wherein: the bidirectional matching includes: constructing a time series embedding vector based on the sorted keywords and sub-topic identifiers in the semantic chain; performing image block division on the continuous key frames and extracting the semantic features of the local image blocks; calculating the similarity between the time series embedding vector and the semantic features of the image blocks, and marking the matching relationship.

[0012] As a preferred solution of the multi-modal data output method for children according to the present invention, wherein: the dynamic scheduling of the multi-modal content output includes: extracting a content prompt unit corresponding to the current semantic chain node in the main module according to the semantic tags included in the interaction prompt data, and loading the corresponding time window identifier and semantic tag; combining the output timing and content priority in the synchronous output control data, calling the content frame segment of the auxiliary module synchronized with the main module, and setting an output delay window for the low-priority module.

[0013] As a preferred solution of the multi-modal data output method for children according to the present invention, wherein: the dynamic scheduling of the multi-modal content output further includes: during the multi-modal content output process, performing session turn matching detection on the current output state, and when it is detected that the correlation degree between the historical semantic chain node and the new input instruction is lower than the threshold, reconstructing the semantic chain node and redirecting the output mapping relationship between the main and auxiliary modules; merging the output content of the main module and the delayed content of the auxiliary module into a modal collaborative output sequence, and associating the semantic chain node where it is located.

[0014] In a second aspect, the present invention provides a multi-modal data output system for children, which includes: a semantic chain construction module that collects the interaction requests input by parents in the dialogue window, constructs a cross-turn semantic chain in combination with the input content of the historical turns, extracts the current session state vector, and synchronously generates interaction trigger data; a synchronous scheduling module that schedules the recommended content output queue based on the session state vector and the interaction trigger data by applying a synchronous lock mechanism to generate synchronous output control data; a label screening module that performs semantic label annotation on the key frame actions in the recommended video according to the synchronous output control data, screens and adapts the action segments in combination with the rules in the infant stage to construct a structured content segment and interaction prompt data; a content output module that performs dynamic scheduling on the multi-modal content output by using the interaction prompt data and the synchronous output control data.

[0015] In a third aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and wherein: when the computer program instructions are executed by the processor, the steps of the multi-modal data output method for children as described in the first aspect of the present invention are implemented.

[0016] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and wherein: when the computer program instructions are executed by the processor, the steps of the multi-modal data output method for children as described in the first aspect of the present invention are implemented.

[0017] The beneficial effects of the present invention are as follows: By constructing a semantic chain that includes parental interaction requests and input content of previous rounds, the present invention generates interaction status data that reflects the true needs of parents and context changes; subsequently, a synchronization lock mechanism is introduced to schedule the recommended content output queue to generate synchronization output control data, and then semantic label annotation is performed on key frame actions. Combined with rules for the infant stage, suitable action segments are selected to construct structured content and interaction prompts; finally, combined with interaction prompt data and synchronization control signals, dynamic scheduling of multi-modal output content is achieved, thereby improving the temporal coherence, semantic relevance, and age adaptability of the recommended content. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0019] Figure 1 It is a flowchart of a method for outputting multi-modal data for children.

[0020] Figure 2 It is a structural diagram of a system for outputting multi-modal data for children. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention with reference to the accompanying drawings of the specification.

[0022] In the following description, many specific details are set forth to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0023] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that mutually excludes other embodiments.

[0024] As mentioned in the above background art, although some current intelligent agent systems already have semantic recognition and multi-modal push capabilities, there are still significant deficiencies in the face of complex continuous conversations, multi-round requests, and the need for adaptability to infant and toddler behaviors. First, traditional intelligent interaction systems mostly output content based on single-round semantic understanding and lack the ability to construct cross-round semantic chains, making it difficult to accurately capture the potential needs progressively expressed by parents in continuous interactions. Second, the recommended content often consists of static resource collections and lacks a deep adaptation mechanism for parameters such as the child's age stage, cognitive level, and behavioral characteristics. In addition, multi-modal content output often lacks a synchronous control mechanism, resulting in temporal disorder or repeated pushing among the recommended contents, thus interfering with the concentration and comprehension abilities of infants and toddlers. Especially in the scenario of joint output of multi-source content, the lack of effective support for core links such as semantic scheduling, key frame screening, and structured presentation severely restricts the improvement space of content personalization and response coherence. Therefore, a child multi-modal data output scheme is needed.

[0025] Figure 1 A flowchart of a child multi-modal data output method according to an embodiment of the present invention. As Figure 1 shown, in the child multi-modal data output method, it includes, S1: Collect the interaction requests input by the parent in the dialogue window, combine the input content of the historical rounds to construct a cross-round semantic chain, and generate interaction status data.

[0026] In the prior art, most dialogue systems only respond to the input text of the current round and lack the ability to model continuous semantic structures, resulting in poor understanding accuracy when the system processes ellipsis, anaphora, rhetorical questions, or euphemistic expressions. However, the present invention effectively improves the system's understanding ability of complex interaction contexts by constructing a cross-round semantic chain and fusing the current input with the historical semantic states. Through the present invention, the continuous semantic structure can be understood and the understanding ability can be improved.

[0027] S1.1: Obtain the original request data input by the parent through voice or text, and synchronously extract its pronunciation features, speech rate rhythm, and punctuation pause timing to construct a pragmatic feature vector.

[0028] In a conventional speech recognition system, only the accuracy of speech transcription is often concerned, while the pragmatic feature information hidden in the speech, such as speech rate, intonation, stress, and pause positions, is ignored. These features are often important bases for semantic supplementation and emotion discrimination. Therefore, in the present invention, while collecting the original request data, high-precision speech processing is performed, multi-dimensional signal analysis is carried out on the audio content, and pragmatic features including pronunciation features, speech rate rhythm, and punctuation pause sequence are extracted. And by constructing a vectorized representation, a pragmatic feature vector with a consistent structure and fusibility is formed, so as to provide auxiliary information at the speech level for subsequent semantic reconstruction. Especially in scenarios with euphemistic expressions or significant changes in tone, this pragmatic vector can effectively make up for the context information that cannot be perceived by traditional text models and grasp the non-explicit semantics more accurately.

[0029] S1.2: Based on the original request data and the pragmatic feature vector, perform local reconstruction on semantic gaps, incomplete instructions, or euphemistic expressions to generate a standardized semantic expression.

[0030] Specifically, first perform dependency syntactic parsing and semantic role annotation on the original text, and combine the information of the pragmatic feature vector to determine whether there are situations such as subject omission, object absence, verb ambiguity, or semantic ambiguity in the sentence. For instructions determined to be incomplete or semantically ambiguous, by constructing an ontology word graph and context word meaning transformation rules, dynamically fill in the missing semantic parts, so as to generate a standardized semantic expression with a standard syntactic structure and clear semantic reference.

[0031] S1.3: Perform multi-level association on the standardized semantic expression and the previous-round semantic expression in the historical interaction, construct a cross-round semantic chain, and generate a semantic context association matrix through a convolutional attention factor.

[0032] In the embodiment of the present invention, based on the currently generated standardized semantic expression, perform multi-level semantic association analysis on it and the previous-round or multi-round semantic expressions in the historical interaction record, and then construct a cross-round semantic chain. It should be noted that semantic understanding is mostly limited to the content of a single-round conversation and ignores the influence of historical context. Therefore, in the present invention, the constructed semantic chain not only records the semantic content of each round, but also is based on multi-dimensional connection methods such as semantic nesting relationship, pragmatic pointing relationship, and keyword co-occurrence relationship.

[0033] To further improve the association quality of multi-round semantics, a convolutional attention factor mechanism is introduced. Its core idea is to slide a convolutional kernel in the semantic chain graph to extract the local structural features of adjacent nodes, and introduce an adaptive attention factor for weighted fusion, and obtain the context association strength between nodes through weighted calculation; after summarizing the convolution attention factor modeling results between all nodes, construct a semantic context association matrix.

[0034] This matrix can not only reflect the influence weight of historical rounds on the current semantic judgment, but also guide the system to perform weighted processing on phenomena such as possible instruction reference and context connection, so as to construct a semantic structure with global context awareness ability.

[0035] S1.4: Generate interaction status data based on the semantic context correlation matrix, integrate the context information of the current statement and historical interactions, extract sub-topic identifiers highly relevant to the current expression, and use them as interaction trigger data to provide basic data support for subsequent recommended content scheduling.

[0036] Specifically, when constructing the interaction status data, first calculate the semantic aggregation weight according to the correlation strength between the semantic chain structure and the current input, and perform context fusion processing based on this weight to form a structured and temporally consistent interaction status representation. At the same time, further perform topic clustering analysis on the current expression content and historical semantic nodes to identify sub-topic identifiers with high frequency or high weight.

[0037] This sub-topic identifier is not only used to represent the semantic core direction of the current interaction, but also serves as the activation basis for triggering recommended content scheduling. Through the above method, personalized content scheduling can be carried out based on the comprehensive context information, making the recommended content more in line with the true intention expressed by the parents.

[0038] S2: Based on the interaction status data, apply a synchronization lock mechanism to schedule the recommended content output queue to generate synchronization output control data.

[0039] Since the recommended content is asynchronously generated and distributed in a multi-modal environment, and the interaction status between the user and the system usually has characteristics of dynamics, suddenness and non-linear evolution. Therefore, without a unified scheduling mechanism, problems such as chaotic output rhythm, redundant content invocation, and terminal resource overload may occur. To solve this problem, the present invention introduces a synchronization lock mechanism to incorporate the content scheduling operation into the synchronization management system, ensuring that during the concurrent execution of multiple threads, the invocation, priority calculation, etc. of the recommended content can maintain consistency and response efficiency, thereby greatly improving the response coordination ability of the content recommendation system under high interaction frequencies. The specific steps are as follows: S2.1: Receive the interaction status data and generate a content scheduling index in combination with the current device status parameters. The device status parameters include screen status, the number of concurrent tasks, and network stability indicators.

[0040] S2.2: Invoke a preset content feature mapping table to perform modal marking and structural splitting on the content to be recommended, disassemble each piece of content into the smallest output unit, and construct a recommended content output queue.

[0041] Traditional recommendation systems mostly regard content as an indivisible whole and process it as a whole content unit when outputting. However, this approach cannot flexibly adapt to the performance differences of terminal devices and the fluctuations in user interaction frequencies. The present invention realizes the reconstruction of the minimum output unit at the structural level of the content by constructing a feature mapping table.

[0042] Specifically, the content feature mapping table is a preset data structure, which is essentially a three-dimensional mapping table of modality-attribute-content fields, where modalities include but are not limited to text, image, video, audio, etc.; attributes include content complexity, loading delay, user attention, etc.; the content field refers to a specific recommended content segment or its representation vector. Through the table lookup operation, the input content to be recommended can be efficiently classified and labeled.

[0043] Specifically, first, the content to be recommended is used as the input, and based on the content feature mapping table, the modality of the content is recognized. Taking video content as an example, it is recognized as "video modality", and key parameters such as its duration, frame rate, and number of audio tracks are marked. Secondly, based on the content structure analysis rules, the structure splitting operation is performed on the content with modality markings to obtain the minimum output unit. The minimum output unit refers to the smallest content segment that can be completely loaded, rendered, and played in a single recommendation step, which is a paragraph, a picture, or a sound command.

[0044] After the above-mentioned recommended content output queue is constructed, it will serve as a candidate resource pool for content invocation in S3 and perform scheduling and invocation according to the synchronous output control data.

[0045] S2.3: Based on the content scheduling index and the recommended content output queue, a synchronization lock mechanism is called to allocate dynamic priority tags.

[0046] It should be noted that the synchronization lock mechanism is essentially a thread-level mutual exclusion control means, which is used to ensure the consistency of the output sequence and scheduling rules of the recommended content and avoid resource contention or overwrite conflicts in the parallel output scenario.

[0047] During the operation of the synchronization lock mechanism, based on the foregoing content scheduling index and the recommended content output queue, a dynamic priority tag needs to be assigned to each output unit: (a) Based on the content scheduling index, key output units are selected from the recommended content output queue to construct a content candidate set.

[0048] Preferably, the screening process can be calculated based on algorithms such as TF-IDF semantic weight, context cosine similarity, or topic model inference, and this embodiment does not make a unique limitation.

[0049] (b) According to the content candidate set, a timing window mechanism is called to extract the response delay estimation value of the output unit and generate a delay weighting factor.

[0050] Optionally, the delay weighting factor can be obtained by methods such as extreme value removal, mean smoothing, or median extraction from the response time data in the sample, and there is no unique limitation here.

[0051] (c) Referring to the output capability parameters of the current device, calculate the fusion score value for each output unit. This fusion score value represents the scheduling priority after comprehensively considering semantic relevance, delay impact, and device performance, and is expressed by a weighted linear formula.

[0052] (d) Perform interval mapping on the fusion score value, label the corresponding dynamic priority tags, such as high priority, medium priority, and low priority, and complete the label assignment.

[0053] S2.4: Generate synchronous output control data based on the priority tag and the device output capability parameters, and append this control data to the recommended content output queue for subsequent timely invocation and switching of multi-modal content.

[0054] During the generation process of the synchronous output control data, first perform a fusion modeling on the priority tag and the device output capability parameters. The device output capability parameters cover indicators such as the number of concurrent modalities that the device can currently support, graphics rendering capabilities, and audio and video decoding capabilities. The fusion process uses a logical rule tree or a neural network structure for modeling to convert discrete priority tags into a set of control signals.

[0055] The set of control signals includes output rhythm signals (such as output rate, start and end times), output mode signals, modality switching signals, etc. The generated control data structure is a structured data object, including control instruction fields, target content identification fields, scheduling time fields, etc.

[0056] Finally, append this synchronous output control data to the recommended content output queue, perform label binding on each output unit, and achieve a one-to-one correspondence between control instructions and content units.

[0057] The present invention dynamically controls the timing and priority of content output according to the current context and device status, ensuring that in a children's interaction scenario, multi-modal output can match the children's attention rhythm and device resource status, thereby improving the timeliness and accuracy of responses. Different from the traditional system where content recommendation and device output scheduling are often designed separately, lacking consideration of device response capabilities, network fluctuations, or task concurrency, resulting in problems such as inconsistent content output delays or stuttering modality switching, the present invention introduces a synchronization lock mechanism, fuses semantic priorities and device status parameters, realizes dynamic scheduling of the recommended output queue, and generates synchronous output control data that can be used for coordinated control.

[0058] S3: According to the synchronous output control data, the key frame actions in the recommended video are semantically labeled, and the adapted action segments are selected based on the infant and toddler stage rules to construct structured content segments and interactive prompt data.

[0059] S3.1: Based on the synchronous output control data, the selected recommended video clip is called and a continuous key frame sequence is extracted, and each frame of image data is time-aligned with the corresponding voice command.

[0060] According to the synchronous output control data, the target video segment is called from the multimodal content library, and a key frame extraction operation is performed.

[0061] Among them, the key frame is the image frame with the most significant image changes and the most obvious action transitions in the video. It is highly representative and is used to express the core process of the action. In order to improve the integrity of the action expression, it is preferred to use a key frame extraction algorithm based on the combination of the color histogram change rate and the optical flow intensity change threshold, that is, when it is detected that the image histogram difference between frames exceeds the threshold and the local optical flow motion intensity exceeds the set ratio, the frame is included in the key frame set.

[0062] The extracted key frame sequence needs to be time-aligned with the voice command stream. This alignment process is based on the audio and video synchronization information (PTS / DTS markers) of the video file. However, considering the natural offset between the voice content and the image content at the trigger time point, the present invention optimizes the alignment accuracy based on the sliding time window mechanism.

[0063] Specifically, a time synchronization window Δt is defined, and the context signal of the voice command is retrieved within this time period and the correlation with the inter-frame behavior unit is measured to achieve frame-level voice-image matching. For example, for the "wave" voice command, the key frame with the strongest change in action posture is found within ±Δt seconds of its appearance time as the corresponding frame to complete the establishment of the anchor point from voice to image.

[0064] S3.2: Using a natural language processing model to perform bidirectional matching between the keywords in the semantic chain and the key frame image content, and marking the semantic labels of the corresponding actions at the frame granularity, including the following steps: (1) Based on the sorted keywords and subtopic identifiers in the semantic chain, a temporal embedding vector is constructed. That is, the keywords are vectorized and encoded through a language model (such as BERT or RoBERTa), and the subtopic identifier is introduced as an attention mask to strengthen the semantic features related to the core direction of the current interaction. Combined with its order in the sentence, a temporal semantic embedding is formed through a position encoding mechanism. Each embedding vector not only represents a semantic unit, but also carries its stage information in the interaction process, thus having temporal perception capabilities.

[0065] (2) Divide the continuous key frames into image blocks and extract the semantic features of the local image blocks.

[0066] Specifically, for the key frame image content, a local image block partitioning algorithm is used, specifically: each frame image is evenly divided into a number of image blocks of a fixed size (for example, 32×32 pixels), and then the semantic feature representation of each image block is extracted respectively.

[0067] This feature extraction relies on a pre-trained visual encoding model to output a high-dimensional semantic embedding vector for each image block.

[0068] (3) Calculate the similarity between the temporal embedding vector and the semantic features of the image block, and mark the matching relationship. Filter the matching pairs according to the similarity threshold, and add the corresponding semantic labels to the image blocks in the keyframe according to the maximum matching criterion. It should be noted that when calculating the similarity between the temporal embedding vector and the semantic features of the image block, the semantic units related to the sub-topic identifier are matched first.

[0069] S3.3: Based on the infant and toddler stage rules, the structural adaptability of the action is confirmed and the action frames that match the stage are selected.

[0070] In an embodiment of the present invention, for different developmental stages of infants and young children, the annotated frames are adaptively screened according to a predefined age-action type mapping table. The age-action type mapping table is a two-dimensional mapping relationship table predefined based on medical research data, in which rows represent age intervals, columns represent allowed action types, and each cell is marked as allowed (1) or prohibited (0). The action type label is extracted from the annotated semantic label, and the frame segments belonging to the allowed set are filtered out.

[0071] S3.4: Structural encapsulation of the filtered action frames is performed to construct a content segment structure including semantic tags, frame segment indexes and rhythm control factors, and interactive prompt data is generated in combination with the triggering statements in the semantic chain.

[0072] In order to improve the interactive friendliness of the system and the accuracy of knowledge response, the action frames that have been screened are structured and encapsulated to generate a content fragment structure that can be called by the control module. The content fragment structure mainly includes: the start and end indexes of the frame segment, the associated semantic label sequence, the action rhythm control factor (such as the recommended playback speed, dwell time, etc.), etc. Among them, the rhythm control factor is used to adjust the playback rhythm when playing the action segment to make it consistent with the human operation learning rhythm or the attention span of infants and young children. For example, in the "Pumpkin Mash" frame segment, the playback speed can be set to 0.75 times and paused for 1.5 seconds at the key nodes of the action.

[0073] In addition, the present invention extracts the attribution position of each structured content segment in the semantic chain, performs semantic mapping with the original parental conversation request, and generates interactive prompt data based on natural language. The interactive prompt data includes: the original semantic chain trigger word, the semantic summary of the recommended action segment, the explanatory note for the adapted age group, and possible auxiliary suggestions. For example, it is recommended not to add granular accessories to the pumpkin puree prepared for 6-month-old babies, and the following steaming and mashing steps can be referred to. The interactive prompt can be presented in the form of text, voice, or a mixture of text and graphics.

[0074] S4: Use the interactive prompt data and the synchronous output control data to perform dynamic scheduling on the multi-modal content output.

[0075] S4.1: According to the semantic tags included in the interactive prompt data, extract the content prompt unit corresponding to the current semantic chain node in the main module, and load the corresponding time window identifier and semantic tag.

[0076] The extraction of the content prompt unit includes: by parsing the semantic tags included in the interactive prompt data, locating the content prompt unit corresponding to the current semantic chain node in the main module, and each content prompt unit includes a set of predefined output candidate information.

[0077] S4.2: Combine the output timing and content priority in the synchronous output control data, call the content frame segment of the auxiliary module synchronized with the main module, and set an output delay window for the low-priority module.

[0078] In the specific operation process, first read the main output timing in the synchronous output control data, and use this timing as the reference benchmark to call the frame segment of the auxiliary module synchronized with the main module content segment. The auxiliary module is essentially an auxiliary modal output to enhance the main content information.

[0079] In the scheduling strategy, if the priority corresponding to the auxiliary module content is lower than the set threshold, then when scheduling the output of this frame segment, set an output delay window for it. The delay window is realized by appending a predefined offset after the main output timing. The output delay window can be adjusted using a static offset or a dynamic adaptive algorithm.

[0080] Through the above scheduling mechanism, it is possible to reasonably arrange the output rhythm and timing according to the content type and priority differences, and avoid cognitive interference or playback conflicts caused by the simultaneous output of multi-modal content.

[0081] S4.3: During the multi-modal content output process, perform session turn matching detection on the current output state. When it is detected that the correlation between the historical semantic chain node and the new input instruction is lower than the threshold, reconstruct the semantic chain node and redirect the output mapping relationship of the main and auxiliary modules.

[0082] During the operation, it continuously monitors the semantic consistency between the semantic chain nodes corresponding to the current output content and the user's current instruction. This consistency is calculated through a semantic similarity model and implemented using a vector embedding and semantic distance measurement mechanism.

[0083] When it is detected that the correlation between the current input and the historical semantic chain nodes is lower than the preset threshold, it is determined that a semantic transition has occurred in the current session, and the semantic chain needs to be reconstructed.

[0084] S4.4: Merge the output content of the main module and the delayed content of the auxiliary module into a modal collaborative output sequence, and associate the semantic chain nodes where they are located.

[0085] In specific operations, an output queue structure is constructed according to the timestamp, content type, and priority label of each frame of content. This queue takes the output frame of the main module as the reference frame, and inserts the delayed frames of the auxiliary module into the queue in a Δt offset manner. In addition, the unique identifier of the current semantic chain node is associated and marked with each output frame to form an output content-semantic node-priority triple structure, so as to maintain the traceability and controllability of the content at the output layer.

[0086] It can be seen that the present invention not only realizes the synchronous output of the main and auxiliary contents, but also ensures the complete coordination of the output content in terms of structure, rhythm, and semantics, avoiding information fragmentation or playback conflicts.

[0087] Furthermore, this embodiment also provides a children's multi-modal data output system, including: A semantic chain construction module 100 that collects the interaction requests input by parents in the dialogue window, constructs a cross-round semantic chain in combination with the input content of the historical rounds, extracts the current session state vector, and synchronously generates interaction trigger data; A synchronization scheduling module 200 that, based on the session state vector and interaction trigger data, applies a synchronization lock mechanism to schedule the recommended content output queue to generate synchronization output control data; A label screening module 300 that, according to the synchronization output control data, performs semantic label annotation on the key frame actions in the recommended video, filters and adapts the action segments in combination with the rules in the infant stage, and constructs structured content segments and interaction prompt data; A content output module 400 that uses the interaction prompt data and synchronization output control data to perform dynamic scheduling on the multi-modal content output.

[0088] This embodiment also provides a computer device applicable to the case of the children's multi-modal data output method, including a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the children's multi-modal data output method proposed in the above embodiment.

[0089] The computer device may be a terminal, which includes a processor, a memory, a communication interface, a display screen, and an input device connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or may be a button, a trackball, or a touchpad provided on the outer shell of the computer device, or may also be an external keyboard, touchpad, or mouse, etc.

[0090] This embodiment also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for realizing children's multimodal data output as proposed in the above embodiment.

[0091] In summary, the present invention generates interaction state data reflecting the real needs of parents and context changes by constructing a semantic chain containing parent interaction requests and historical round input content; then introduces a synchronization lock mechanism to schedule the recommended content output queue to generate synchronization output control data, and further performs semantic label annotation on key frame actions, and combines action segments adapted by rules in the infant stage to construct structured content and interaction prompts; finally, combines the interaction prompt data with the synchronization control signal to realize the dynamic scheduling of multimodal output content, thereby improving the temporal coherence, semantic relevance, and age adaptability of recommended content for young children.

[0092] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. Method for outputting multi-modal data of children, characterized in that: Including: Collect the interaction requests input by parents in the dialogue window, construct a cross-round semantic chain in combination with the input content of historical rounds, and generate interaction status data; Based on the interaction status data, apply a synchronization lock mechanism to schedule the recommended content output queue and generate synchronization output control data; According to the synchronization output control data, perform semantic label annotation on the key frame actions in the recommended video, screen the adapted action segments in combination with the rules in the infant stage, and construct structured content segments and interaction prompt data; Use the interaction prompt data and the synchronization output control data to perform dynamic scheduling on the multi-modal content output.

2. The method for outputting multi-modal data of children according to claim 1, characterized in that: The generation of the interaction status data includes: based on the cross-round semantic chain, generating a semantic context correlation matrix through a convolutional attention factor; Generating interaction status data according to the semantic context correlation matrix, integrating the context information of the current statement and historical interactions, and extracting the sub-topic identifier highly relevant to the current expression as interaction trigger data.

3. The method for outputting multi-modal data of children according to claim 1, wherein: The generation process of the synchronization output control data includes: Receiving the interaction status data and generating a content scheduling index in combination with the current device status parameters; Invoking a preset content feature mapping table, performing modal marking and structural splitting on the content to be recommended, disassembling each piece of content into the smallest output unit, and constructing a recommended content output queue; Based on the content scheduling index and the recommended content output queue, invoking a synchronization lock mechanism to assign dynamic priority tags; Generating synchronization output control data according to the priority tags and the device output ability parameters.

4. The method for outputting multi-modal data of children according to claim 3, wherein: The assignment of dynamic priority tags includes: Based on the content scheduling index, selecting key output units from the recommended content output queue to construct a content candidate set; according to the content candidate set, invoking a time series window mechanism to extract the response delay estimation value of the output unit and generate a delay weighting factor; referring to the current device output ability parameters, calculating a fusion score value for each output unit; performing interval mapping on the fusion score value and marking the corresponding dynamic priority tag.

5. The method for outputting multi-modal data of children according to claim 1, wherein: The semantic label annotation includes: Based on the synchronization output control data, invoking the selected recommended video segment and extracting a continuous key frame sequence, aligning the time axis of each key frame image data with the corresponding voice command; Using a natural language processing model to perform bidirectional matching on the keywords in the semantic chain and the key frame image content, and annotating the semantic labels of the corresponding actions at the frame granularity.

6. The method for outputting multi-modal data of children according to claim 5, wherein: The screening of the adapted action segments includes: based on the rules in the infant stage, confirming the structural adaptability of the actions and screening out the action frame segments that match the stage.

7. The method for outputting multimodal data of children according to claim 6, wherein: The bidirectional matching includes: Constructing a time series embedding vector based on the sorted keywords and sub-topic identifiers in the semantic chain; Dividing the continuous key frames into image blocks and extracting the semantic features of the local image blocks; Calculating the similarity between the time series embedding vector and the semantic features of the image blocks and marking the matching relationship.

8. The method for outputting multi-modal data of children according to claim 1, wherein: The dynamic scheduling of the multi-modal content output includes: According to the semantic labels included in the interaction prompt data, extracting the content prompt units corresponding to the current semantic chain nodes in the main module, and loading the corresponding time window identifiers and semantic labels; Combined with the output timing and content priority in the synchronous output control data, call the content frame segment of the secondary module synchronized with the primary module, and set an output delay window for the low-priority module.

9. The method for outputting multi-modal data of children according to claim 8, characterized in that: The dynamic scheduling of the multi-modal content output further includes: During the multi-modal content output process, perform session turn matching detection on the current output state. When the correlation degree between the historical semantic chain node and the new input instruction is lower than the threshold, reconstruct the semantic chain node and redirect the output mapping relationship of the primary and secondary modules; Merge the output content of the primary module and the delayed content of the secondary module into a modal collaborative output sequence, and associate the semantic chain node where it is located.

10. A child multi-modal data output system, based on the child multi-modal data output method according to any one of claims 1 to 9, characterized in that: It also includes: A semantic chain construction module that collects the interaction requests input by the parent in the dialogue window, constructs a cross-turn semantic chain in combination with the input content of the historical turn, extracts the current session state vector, and synchronously generates interaction trigger data; A synchronous scheduling module that, based on the session state vector and interaction trigger data, applies a synchronization lock mechanism to schedule the recommended content output queue to generate synchronous output control data; A label screening module that, according to the synchronous output control data, performs semantic label annotation on the key frame actions in the recommended video, screens and adapts the action segments in combination with the rules of the infant stage, and constructs structured content segments and interactive prompt data; A content output module that uses the interactive prompt data and synchronous output control data to perform dynamic scheduling on the multi-modal content output.

Citation Information

Patent Citations

  • Image-text interaction thinking machine control system and method

    CN119376957A

  • Cache management system during large model reasoning

    CN119918632A

  • Touch and talk pen intelligent children education method based on emotion recognition

    CN119992624A

  • Children MPP auxiliary diagnosis system based on multi-modal time series data modeling

    CN120015296A

  • Scheduling tools with queue time constraints

    US7463939B1

Cited By

  • Multi-user mixed interactive data processing method and system based on large model

    CN121118904A