Method and System for Outputting Multimodal Data of Children
By building a cross-wheel semantic chain and synchronization lock mechanism, the timing disorder and adaptability of recommended content in the intelligent parenting system in complex interactions is solved, the timing coherence of multimodal output and the age adaptability of young children are realized, and the system's response and coordination ability and content personalization are improved.
Patent Information
- Application Number
- CN202510756469.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-06-09
AI Technical Summary
When facing complex continuous dialogues and multiple rounds of requests, the existing intelligent parenting system lacks the ability to construct cross-wheel semantic chains, the recommended content lacks deep adaptation to children's age, cognitive level and behavioral characteristics, and the multimodal content output lacks synchronous control, resulting in timing disorders and repeated pushes, interfering with infants' concentration and understanding abilities.
By constructing a cross-wheel semantic chain, generating interactive state data, using the synchronization lock mechanism to schedule the recommended content output queue, perform semantic label annotation, and filtering adaptive action fragments in combination with infant stage rules to achieve dynamic scheduling of multimodal content.
The timing coherence, semantic correlation and age adaptability of recommended content are improved, ensuring that multimodal output matches children's attention and equipment resources, and improving the response and coordination ability of interaction and personalization of content.
Smart Images

Figure CN120296704B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data output, in particular to a method and system for outputting multi-modal data for children. Background Art
[0002] With the rapid development of artificial intelligence, big data, and multi-modal interaction technologies, child content recommendation and intelligent parenting assistance systems have gradually become an important part of digital services for infant and toddler families. In recent years, many typical intelligent agent applications have emerged. By integrating language models, emotion recognition, and content recommendation mechanisms, they can not only achieve question-and-answer interactions and service recommendations for parents, but also provide more contextually adaptable parenting support services by combining multi-modal data.
[0003] Although some current intelligent agent systems already have semantic recognition and multi-modal push capabilities, there are still significant deficiencies in the face of complex continuous conversations, multi-round requests, and the need for adaptability to infant and toddler behaviors. First, traditional intelligent interaction systems mostly output content based on single-round semantic understanding, lacking the ability to construct cross-round semantic chains and difficult to accurately capture the potential needs gradually expressed by parents in continuous interactions. Second, the recommended content often consists of static resource collections, lacking a deep adaptation mechanism for parameters such as the child's age stage, cognitive level, and behavior characteristics. In addition, multi-modal content output often lacks a synchronization control mechanism, resulting in temporal disorder or repeated pushing between recommended contents, thus interfering with the concentration and comprehension abilities of infants and toddlers. Especially in the scenario of joint output of multi-source content, the lack of effective support for core links such as semantic scheduling, key frame screening, and structured presentation seriously restricts the improvement space of content personalization and response coherence. Summary of the Invention
[0004] In view of the problems existing in the above background art, the present invention is proposed.
[0005] To solve the above technical problems, the present invention provides the following technical solutions:
[0006] In a first aspect, the present invention provides a method for outputting multi-modal data for children, which includes collecting an interaction request input by a parent in a dialogue window, constructing a cross-round semantic chain by combining the input content of previous rounds to generate interaction status data; based on the interaction status data, applying a synchronization lock mechanism to schedule a recommended content output queue to generate synchronization output control data; according to the synchronization output control data, performing semantic label annotation on the key frame actions in a recommended video, screening adapted action segments in combination with infant and toddler stage rules, constructing structured content segments and interaction prompt data; using the interaction prompt data and the synchronization output control data to perform dynamic scheduling on multi-modal content output.
[0007] As a preferred solution of the multi-modal data output method for children according to the present invention, wherein: the generation of the interaction state data includes: based on the cross-turn semantic chain, generating a semantic context association matrix through a convolutional attention factor; generating interaction state data according to the semantic context association matrix, integrating the context information of the current sentence and the historical interaction, and extracting sub-topic identifiers highly relevant to the current expression as interaction trigger data.
[0008] As a preferred solution of the multi-modal data output method for children according to the present invention, wherein: the generation process of the synchronous output control data includes: receiving the interaction state data, and generating a content scheduling index in combination with the current device state parameters; calling a preset content feature mapping table, performing modal marking and structural splitting on the content to be recommended, disassembling each content into the smallest output unit, and constructing a recommended content output queue; based on the content scheduling index and the recommended content output queue, calling a synchronization lock mechanism to assign dynamic priority tags; generating synchronous output control data according to the priority tags and the device output ability parameters.
[0009] As a preferred solution of the multi-modal data output method for children according to the present invention, wherein: the assignment of dynamic priority tags includes: based on the content scheduling index, selecting key output units from the recommended content output queue to construct a content candidate set; according to the content candidate set, calling a time series window mechanism to extract the response delay estimation value of the output unit, and generating a delay weighting factor; referring to the current device output ability parameters, calculating a fusion score value for each output unit; performing interval mapping on the fusion score value to label the corresponding dynamic priority tag.
[0010] As a preferred solution of the multi-modal data output method for children according to the present invention, wherein: the semantic label annotation includes: based on the synchronous output control data, calling the selected recommended video segment and extracting a continuous key frame sequence, aligning the image data of each key frame with the corresponding voice command on the time axis; using a natural language processing model to perform bidirectional matching on the keywords and the content of the key frame images in the semantic chain, and annotating the semantic labels of the corresponding actions at the frame granularity.
[0011] As a preferred solution of the multi-modal data output method for children according to the present invention, wherein: the screening of the adapted action segment includes: based on the rules in the infant stage, confirming the structural adaptability of the action, and screening out the action frame segments that are stage-matched.
[0012] As a preferred solution of the multi-modal data output method for children according to the present invention, wherein: the bidirectional matching includes: constructing a time series embedding vector based on the sorted keywords and sub-topic identifiers in the semantic chain; dividing the continuous key frames into image blocks, and extracting the semantic features of the local image blocks; calculating the similarity between the time series embedding vector and the semantic features of the image blocks, and marking the matching relationship.
[0013] As a preferred solution of the multi-modal data output method for children according to the present invention, wherein: the dynamic scheduling of the multi-modal content output includes: according to the semantic tags included in the interaction prompt data, extracting the content prompt unit corresponding to the current semantic chain node in the main module, and loading the corresponding time window identifier and semantic tag; combining the output timing and content priority in the synchronous output control data, calling the content frame segment of the auxiliary module synchronized with the main module, and setting an output delay window for the low-priority module.
[0014] As a preferred solution of the multi-modal data output method for children according to the present invention, wherein: the dynamic scheduling of the multi-modal content output further includes: during the multi-modal content output process, performing session turn matching detection on the current output state, and when it is detected that the correlation degree between the historical semantic chain node and the new input instruction is lower than the threshold, reconstructing the semantic chain node and redirecting the output mapping relationship of the main and auxiliary modules; merging the output content of the main module and the delayed content of the auxiliary module into a modal collaborative output sequence, and associating the semantic chain node where it is located.
[0015] In a second aspect, the present invention provides a multi-modal data output system for children, which includes: a semantic chain construction module that collects the interaction requests input by parents in the dialogue window, constructs a cross-turn semantic chain in combination with the input content of the historical turns, extracts the current session state vector, and synchronously generates interaction trigger data; a synchronous scheduling module that schedules the recommended content output queue based on the session state vector and the interaction trigger data by applying a synchronous lock mechanism to generate synchronous output control data; a label screening module that, according to the synchronous output control data, performs semantic label annotation on the key frame actions in the recommended video, filters and adapts the action segments in combination with the rules in the infant stage, and constructs a structured content segment and interaction prompt data; a content output module that uses the interaction prompt data and the synchronous output control data to perform dynamic scheduling on the multi-modal content output.
[0016] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and wherein: when the computer program instructions are executed by the processor, the steps of the multi-modal data output method for children as described in the first aspect of the present invention are implemented.
[0017] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and wherein: when the computer program instructions are executed by the processor, the steps of the multi-modal data output method for children as described in the first aspect of the present invention are implemented.
[0018] The beneficial effects of the present invention are as follows: By constructing a semantic chain containing parental interaction requests and input content of historical rounds, the present invention generates interaction status data reflecting the true needs of parents and context changes; subsequently, a synchronization lock mechanism is introduced to schedule the recommended content output queue, generating synchronization output control data, and then semantic label annotation is performed on key frame actions, and appropriate action segments are screened in combination with infant stage rules to construct structured content and interaction prompts; finally, by combining the interaction prompt data and the synchronization control signal, dynamic scheduling of multi-modal output content is achieved, thereby improving the temporal coherence, semantic relevance, and age adaptability of the recommended content to infants. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 It is a flowchart of a method for multi-modal data output for children.
[0021] Figure 2 It is a structural diagram of a multi-modal data output system for children. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will make a detailed description of the specific embodiments of the present invention in conjunction with the drawings in the specification.
[0023] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0024] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or selectively exclusive embodiment from other embodiments.
[0025] As described in the above background art, although some current intelligent agent systems already have semantic recognition and multi-modal push capabilities, there are still significant deficiencies in the prior art when facing complex continuous conversations, multi-round requests, and the need for adaptability to infant and toddler behaviors. First, traditional intelligent interaction systems mostly output content based on single-round semantic understanding and lack the ability to construct cross-round semantic chains, making it difficult to accurately capture the potential needs gradually expressed by parents in continuous interactions. Second, the recommended content often consists of static resource collections and lacks a deep adaptation mechanism for parameters such as the child's age stage, cognitive level, and behavioral characteristics. In addition, multi-modal content output often lacks a synchronous control mechanism, resulting in temporal disorder or repeated pushing among recommended contents, thus interfering with the concentration and comprehension abilities of infants and toddlers. Especially in the scenario of joint output of multi-source content, there is a lack of effective support for core links such as semantic scheduling, key frame screening, and structured presentation, severely restricting the improvement space of content personalization and response coherence. Therefore, a child multi-modal data output solution is needed.
[0026] Figure 1 It is a flowchart of a child multi-modal data output method according to an embodiment of the present invention. As Figure 1 shown, in the child multi-modal data output method, it includes,
[0027] S1: Collect the interaction requests input by the parent in the dialogue window, construct a cross-round semantic chain in combination with the input content of previous rounds, and generate interaction status data.
[0028] In the prior art, most dialogue systems only respond to content based on the input text of the current round and lack the ability to model continuous semantic structures, resulting in poor understanding accuracy when the system processes ellipsis, anaphora, rhetorical questions, or euphemistic expressions. However, the present invention effectively improves the system's understanding ability of complex interaction contexts by constructing a cross-round semantic chain and performing fusion processing on the current input and historical semantic states. Through the present invention, the continuous semantic structure can be understood and the understanding ability can be improved.
[0029] S1.1: Obtain the original request data input by the parent through voice or text, and synchronously extract its pronunciation features, speech rate rhythm, and punctuation pause timing to construct a pragmatic feature vector.
[0030] In a conventional speech recognition system, attention is often only paid to the accuracy of speech transcription, while ignoring the pragmatic feature information implicit in the speech, such as speech rate, intonation, stress, and pause positions. These features are often important bases for semantic complementation and emotion discrimination. Therefore, in the present invention, while collecting the original request data, high-precision speech processing is performed, multi-dimensional signal analysis is carried out on the audio content, and pragmatic features including pronunciation features, speech rate rhythm, and punctuation pause sequence are extracted. And by constructing a vectorized representation, a pragmatic feature vector with a consistent structure and being fusible is formed, thereby providing auxiliary information at the speech level for subsequent semantic reconstruction. Especially in scenarios with euphemistic expressions or significant tone changes, this pragmatic vector can effectively compensate for the context information that cannot be perceived by traditional text models and grasp non-explicit semantics more accurately.
[0031] S1.2: Based on the original request data and the pragmatic feature vector, perform local reconstruction on semantic gaps, incomplete instructions, or euphemistic expressions to generate a standardized semantic expression.
[0032] Specifically, first perform dependency syntactic parsing and semantic role annotation on the original text, and combine the information of the pragmatic feature vector to determine whether there are situations such as subject omission, object absence, verb ambiguity, or semantic ambiguity in the sentence. For instructions determined to be incomplete or semantically ambiguous, by constructing an ontology word graph and context word meaning transformation rules, dynamically fill in the missing semantic parts, thereby generating a standardized semantic expression with a standard syntactic structure and clear semantic reference.
[0033] S1.3: Perform multi-level association on the standardized semantic expression and the previous round semantic expression in the historical interaction, construct a cross-round semantic chain, and generate a semantic context association matrix through a convolutional attention factor.
[0034] In an embodiment of the present invention, based on the currently generated standardized semantic expression, perform multi-level semantic association analysis on it and the previous round or multiple rounds of semantic expressions in the historical interaction record, and then construct a cross-round semantic chain. It should be noted that semantic understanding is mostly limited to the content of a single-round conversation and ignores the influence of historical context. Therefore, in the present invention, the constructed semantic chain not only records the semantic content of each round, but also is based on multi-dimensional connection methods such as semantic nesting relationship, pragmatic reference relationship, and keyword co-occurrence relationship.
[0035] To further improve the association quality of multi-round semantics, a convolutional attention factor mechanism is introduced. Its core idea is to slide a convolutional kernel in the semantic chain graph to extract local structural features of adjacent nodes, and introduce an adaptive attention factor for weighted fusion, and obtain the context association strength between nodes through weighted calculation; after summarizing the convolutional attention factor modeling results between all nodes, construct a semantic context association matrix.
[0036] This matrix can not only reflect the influence weight of historical rounds on the current semantic judgment, but also guide the system to perform weighted processing on phenomena such as possible instruction reference and context connection, so as to construct a semantic structure with global context awareness ability.
[0037] S1.4: Generate interaction status data based on the semantic context correlation matrix, integrate the context information of the current statement and historical interactions, extract sub-topic identifiers highly relevant to the current expression, and use them as interaction trigger data to provide basic data support for subsequent recommended content scheduling.
[0038] Specifically, when constructing the interaction status data, first calculate the semantic aggregation weight according to the correlation strength between the semantic chain structure and the current input, and perform context fusion processing based on this weight to form a structured and time-sequence consistent interaction status representation. At the same time, further perform topic clustering analysis on the current expression content and historical semantic nodes to identify sub-topic identifiers with high frequency or high weight.
[0039] This sub-topic identifier is not only used to represent the semantic core direction of the current interaction, but also serves as the activation basis for triggering recommended content scheduling. Through the above method, personalized content scheduling can be carried out based on the comprehensive context information, making the recommended content more in line with the true intention expressed by the parents.
[0040] S2: Based on the interaction status data, apply a synchronization lock mechanism to schedule the recommended content output queue and generate synchronization output control data.
[0041] Since the recommended content is asynchronously generated and distributed in a multi-modal environment, and the interaction status between the user and the system usually has characteristics of dynamics, suddenness and non-linear evolution. Therefore, without a unified scheduling mechanism, problems such as chaotic output rhythm, redundant content invocation, and terminal resource overload may occur. To solve this problem, the present invention introduces a synchronization lock mechanism to incorporate the content scheduling operation into the synchronization management system, ensuring that during the concurrent execution of multiple threads, the invocation, priority calculation, etc. of the recommended content can maintain consistency and response efficiency, thereby greatly improving the response coordination ability of the content recommendation system under high interaction frequencies. The specific steps are as follows:
[0042] S2.1: Receive the interaction status data and generate a content scheduling index in combination with the current device status parameters. The device status parameters include screen status, the number of concurrent tasks, and network stability indicators.
[0043] S2.2: Call a preset content feature mapping table to perform modal marking and structural splitting on the content to be recommended, decompose each content into the smallest output unit, and construct a recommended content output queue.
[0044] Traditional recommendation systems mostly regard content as an indivisible whole and process it as a whole content unit when outputting. However, this method cannot flexibly adapt to the performance differences of terminal devices and the fluctuations in user interaction frequencies. The present invention realizes the reconstruction of the minimum output unit of content at the structural level by constructing a feature mapping table.
[0045] Specifically, the content feature mapping table is a preset data structure, which is essentially a three-dimensional mapping table of modality-attribute-content fields. The modalities include, but are not limited to, text, image, video, audio, etc.; the attributes include content complexity, loading delay, user attention, etc.; and the content fields refer to specific recommended content segments or their representation vectors. Through table lookup operations, the input content to be recommended can be efficiently classified and labeled.
[0046] Specifically, first, the content to be recommended is used as the input. Based on the content feature mapping table, the modality of the content is recognized. Taking video content as an example, it is recognized as "video modality", and key parameters such as its duration, frame rate, and number of audio tracks are marked. Secondly, based on the content structure analysis rules, the structure splitting operation is performed on the content with modality markings to obtain the minimum output unit. The minimum output unit refers to the smallest content segment that can be completely loaded, rendered, and played in a single recommendation step, which is a paragraph, a picture, or a sound instruction.
[0047] After the above-mentioned recommended content output queue is constructed, it will serve as a candidate resource pool for content invocation in S3 and perform scheduling and invocation according to the synchronous output control data.
[0048] S2.3: Based on the content scheduling index and the recommended content output queue, a synchronization lock mechanism is called to assign dynamic priority tags.
[0049] It should be noted that the synchronization lock mechanism is essentially a thread-level mutual exclusion control means, which is used to ensure the consistency of the output sequence and scheduling rules of recommended content and avoid resource contention or overwrite conflicts in parallel output scenarios.
[0050] During the operation of the synchronization lock mechanism, dynamic priority tags need to be assigned to each output unit based on the aforementioned content scheduling index and the recommended content output queue:
[0051] (a) Based on the content scheduling index, key output units are selected from the recommended content output queue to construct a content candidate set.
[0052] Preferably, the screening process can be calculated based on algorithms such as TF-IDF semantic weights, context cosine similarity, or topic model inference, and this embodiment is not limited to a single one.
[0053] (b)Based on the described content candidate set, call the time series window mechanism to extract the response delay estimation value of the output unit and generate a delay weighting factor.
[0054] Optionally, the delay weighting factor can be obtained by methods such as extreme value removal, mean smoothing, or median extraction from the response time data in the sample, and there is no unique limitation here.
[0055] (c)Refer to the current device output capability parameters and calculate the fusion score value for each output unit. This fusion score value represents the scheduling priority after comprehensively considering semantic relevance, delay impact, and device performance, and is expressed by a weighted linear formula.
[0056] (d)Perform interval mapping on the fusion score value, label the corresponding dynamic priority tags, such as high priority, medium priority, low priority, and complete the label assignment.
[0057] S2.4: Generate synchronous output control data based on the priority tag and the device output capability parameters, and attach this control data to the recommended content output queue for subsequent timely calling and switching of multi-modal content.
[0058] During the generation process of the synchronous output control data, first perform fusion modeling on the priority tag and the device output capability parameters. The device output capability parameters cover indicators such as the number of concurrent modalities currently supported by the device, graphics rendering ability, audio and video decoding ability, etc. The fusion process uses a logical rule tree or a neural network structure for modeling to convert the discrete priority tag into a control signal set.
[0059] The control signal set includes output rhythm signals (such as output rate, start and end time), output mode signals, modality switching signals, etc. The generated control data structure is a structured data object, including a control instruction field, a target content identification field, a scheduling time field, etc.
[0060] Finally, attach this synchronous output control data to the recommended content output queue, perform label binding on each output unit, and achieve one-to-one correspondence between control instructions and content units.
[0061] The present invention dynamically controls the timing and priority of content output based on the current context and device status, ensuring that in children's interactive scenarios, multimodal output can match the child's attention rhythm and device resource status, thereby improving the timeliness and accuracy of responses. Unlike traditional systems, where content recommendation and device output scheduling are often separated and lack consideration of device responsiveness, network fluctuations, or task concurrency, resulting in inconsistent content output delays or mode switching issues, the present invention introduces a synchronization lock mechanism, integrates semantic priority and device status parameters, and achieves dynamic scheduling of recommendation output queues, generating synchronous output control data that can be used for coordinated control.
[0062] S3: Based on the synchronous output control data, semantically label the key frame actions in the recommended video, filter the adapted action segments based on the infant and toddler stage rules, and construct structured content segments and interactive prompt data.
[0063] S3.1: Based on the synchronous output control data, call the selected recommended video clip and extract a continuous key frame sequence, and align the time axis of each frame of image data with the corresponding voice command.
[0064] According to the synchronous output control data, the target video segment is called from the multimodal content library and the key frame extraction operation is performed.
[0065] Keyframes are the most significant image changes and action transitions in a video. They are highly representative and represent the core process of an action. To enhance the integrity of action representation, a keyframe extraction algorithm based on a combination of the color histogram change rate and an optical flow intensity change threshold is preferred. Specifically, if the inter-frame image histogram difference exceeds a threshold and the local optical flow motion intensity exceeds a set ratio, the frame is included in the keyframe set.
[0066] The extracted keyframe sequence must be time-aligned with the voice command stream. This alignment is based on the video file's inherent audio and video synchronization information (PTS / DTS markers). However, given the natural offset between the triggering time of voice and image content, this invention optimizes alignment accuracy using a sliding time window mechanism.
[0067] Specifically, a time synchronization window Δt is defined within which contextual signals of the voice command are retrieved and correlations are measured with inter-frame behavioral units to achieve frame-level speech-to-image matching. For example, for the "wave" voice command, the keyframe with the strongest gesture change within ±Δt seconds of its occurrence is found as the corresponding frame, completing the establishment of the speech-to-image anchor.
[0068] S3.2: Use a natural language processing model to perform bidirectional matching between the keywords in the semantic chain and the content of key-frame images, and label the semantic tags of corresponding actions at the frame granularity, including the following steps:
[0069] (1) Based on the sorted keywords and sub-topic identifiers in the semantic chain, construct a temporal embedding vector. That is, vectorize and encode the keywords through a language model (such as BERT or RoBERTa), introduce the sub-topic identifier as an attention mask to strengthen the semantic features related to the current interaction core direction, and combine its order in the sentence to form a temporal semantic embedding through a position encoding mechanism. Each embedding vector not only represents a semantic unit but also carries the stage information in the interaction process, thus having the ability of temporal perception.
[0070] (2) Divide the continuous key frames into image patches and extract the semantic features of local image patches.
[0071] Specifically, for the content of key-frame images, use a local image patch division algorithm, specifically: evenly divide each frame of the image into a number of fixed-size image patches (such as 32×32 pixels), and then extract the semantic feature representations of each image patch respectively.
[0072] This feature extraction relies on a pre-trained visual coding model to output high-dimensional semantic embedding vectors for each image patch.
[0073] (3) Calculate the similarity between the temporal embedding vector and the semantic features of the image patches, and mark the matching relationship. Filter out the matching pairs according to the similarity threshold, and add the corresponding semantic tags to the image patches in the key frames according to the maximum matching criterion. It should be noted that when calculating the similarity between the temporal embedding vector and the semantic features of the image patches, the semantic units related to the sub-topic identifier are preferentially matched.
[0074] S3.3: Based on the rules for the infant stage, confirm the structural fitness of the actions and screen out the action frame segments that match the stage.
[0075] In the embodiment of the present invention, for different development stages of infants, perform adaptability screening on the labeled frames according to a pre-defined month-old - action type mapping table. Among them, the month-old - action type mapping table is a two-dimensional mapping relationship table pre-defined based on medical research data, where the rows represent month-old intervals, the columns represent allowed action types, and each cell is marked as allowed (1) or prohibited (0). Extract the action type tags from the labeled semantic tags and filter out the frame segments belonging to the allowed set.
[0076] S3.4: Structurally encapsulate the screened action frame segments, construct a content segment structure including semantic tags, frame segment indexes, and rhythm control factors, and generate interactive prompt data in combination with the trigger statements in the semantic chain.
[0077] To improve the interaction friendliness of the system and the accuracy of knowledge response, the screened action frame segments are structurally encapsulated to generate a content segment structure that can be called by the control module. The content segment structure mainly includes: the start and end indexes of the frame segment, the associated semantic tag sequence, action rhythm control factors (such as recommended playback speed, dwell time, etc.). Among them, the rhythm control factor is used to adjust the playback rhythm when playing this action segment to make it conform to the human operation learning rhythm or the attention duration of infants and young children. For example, in the "press pumpkin puree" frame segment, a playback speed of 0.75 times can be set and a 1.5-second pause can be made at the key nodes of the action.
[0078] In addition, the present invention extracts the attribution position of each structured content segment in the semantic chain, performs semantic mapping with the original parent's dialogue request, and generates interactive prompt data based on natural language. The interactive prompt data includes: the original semantic chain trigger word, the semantic summary of the recommended action segment, the explanatory note for the adapted age group, and possible auxiliary suggestions. For example, it is recommended not to add granular accessories to the pumpkin puree prepared for 6-month-old babies, and the following steaming and pressing steps can be referred to. This interactive prompt can be presented in the form of text, voice, or a mixture of text and graphics.
[0079] S4: Use the interactive prompt data and the synchronous output control data to perform dynamic scheduling on the multi-modal content output.
[0080] S4.1: According to the semantic tags included in the interactive prompt data, extract the content prompt unit corresponding to the current semantic chain node in the main module, and load the corresponding time window identifier and semantic tag.
[0081] The extraction of the content prompt unit includes: by parsing the semantic tags included in the interactive prompt data, locating the content prompt unit corresponding to the current semantic chain node in the main module, and each content prompt unit contains a set of predefined output candidate information.
[0082] S4.2: Combine the output timing and content priority in the synchronous output control data, call the content frame segment of the auxiliary module synchronized with the main module, and set an output delay window for the low-priority module.
[0083] In the specific operation process, first read the main output timing in the synchronous output control data, and use this timing as a reference benchmark to call the frame segment of the auxiliary module synchronized with the content frame segment of the main module. The auxiliary module is essentially an auxiliary modal output to enhance the main content information.
[0084] In terms of the scheduling strategy, if the priority corresponding to the content of the auxiliary module is lower than the set threshold, when scheduling the output of this frame segment, an output delay window is set for it. The delay window is achieved by appending a predefined offset after the main output timing. The output delay window can be adjusted using a static offset or a dynamic adaptive algorithm.
[0085] Through the above scheduling mechanism, it is possible to reasonably arrange the output rhythm and timing according to the content type and priority differences, avoiding cognitive interference or playback conflicts caused by the simultaneous output of multi-modal content.
[0086] S4.3: During the output process of multi-modal content, perform session turn matching detection on the current output state. When the correlation degree between the historical semantic chain node and the new input instruction is lower than the threshold, reconstruct the semantic chain node and redirect the output mapping relationship of the main and auxiliary modules.
[0087] During the operation, it continuously monitors the semantic consistency between the semantic chain node corresponding to the current output content and the user's current instruction. This consistency is calculated through a semantic similarity model and implemented using a vector embedding and semantic distance measurement mechanism.
[0088] When it is detected that the correlation degree between the current input and the historical semantic chain node is lower than the preset threshold, it is determined that a semantic transition has occurred in the current session, and the semantic chain needs to be reconstructed.
[0089] S4.4: Merge the output content of the main module and the delayed content of the auxiliary module into a modal collaborative output sequence and associate the semantic chain nodes where they are located.
[0090] In specific operations, according to the timestamp, content type, and priority label of each frame of content, an output queue structure is constructed. This queue uses the output frame of the main module as the reference frame, and inserts the delayed frames of the auxiliary module in a Δt offset manner into the queue. In addition, the unique identifier of the current semantic chain node is associated and marked with each output frame, forming an output content-semantic node-priority triple structure, so as to maintain the traceability and controllability of the content at the output layer.
[0091] It can be seen that the present invention not only realizes the synchronous output of the main and auxiliary contents, but also ensures the complete coordination of the output content in terms of structure, rhythm, and semantics, avoiding information fragmentation or playback conflicts.
[0092] Furthermore, this embodiment also provides a children's multi-modal data output system, including:
[0093] A semantic chain construction module 100, which collects the interaction requests input by parents in the dialogue window, constructs a cross-turn semantic chain in combination with the historical turn input content, extracts the current session state vector, and synchronously generates interaction trigger data;
[0094] The synchronization scheduling module 200 schedules the recommended content output queue based on the session state vector and interaction trigger data, and applies a synchronization lock mechanism to generate synchronization output control data.
[0095] The label screening module 300 performs semantic label annotation on the key frame actions in the recommended video according to the synchronization output control data, combines the rules in the infant stage to screen and adapt the action segments, and constructs structured content segments and interaction prompt data.
[0096] The content output module 400 performs dynamic scheduling on the multi-modal content output by using the interaction prompt data and the synchronization output control data.
[0097] This embodiment also provides a computer device applicable to the case of the child multi-modal data output method, including a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions to implement the child multi-modal data output method proposed in the above embodiment.
[0098] This computer device may be a terminal. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, a touchpad, or a mouse, etc.
[0099] This embodiment also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the child multi-modal data output method proposed in the above embodiment.
[0100] In summary, the present invention generates interaction status data reflecting the real needs of parents and context changes by constructing a semantic chain including parent interaction requests and historical round input content; subsequently, a synchronization lock mechanism is introduced to schedule the recommended content output queue to generate synchronization output control data, and then semantic label annotation is performed on key frame actions, and appropriate action segments are screened and adapted in combination with the rules in the infant stage to construct structured content and interaction prompts; finally, dynamic scheduling of multi-modal output content is achieved by combining the interaction prompt data and the synchronization control signal, thereby improving the temporal coherence, semantic relevance, and age adaptability of the recommended content.
[0101] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for outputting multimodal data of children, characterized by: include: Collect the interaction requests entered by parents in the dialogue window, build a cross-turn semantic chain based on the input content of historical rounds, and generate interaction status data; Based on the interaction state data, a synchronization lock mechanism is applied to schedule the recommended content output queue to generate synchronization output control data; Based on the synchronous output control data, semantic tags are annotated for key frame actions in the recommended videos. Adaptive action segments are selected based on infant and toddler stage rules to construct structured content segments and interactive prompt data. Dynamically scheduling multimodal content output using the interactive prompt data and the synchronous output control data; The generation of the interaction state data includes: generating a semantic context association matrix based on the cross-round semantic chain through the convolution attention factor; generating the interaction state data based on the semantic context association matrix, integrating the context information of the current sentence and the historical interactions, and extracting the sub-topic identifier that is highly relevant to the current expression as the interaction trigger data; The process of generating the synchronous output control data includes: receiving the interaction state data and generating a content scheduling index based on current device state parameters; calling a preset content feature mapping table to perform modal tagging and structural decomposition on the recommended content, breaking each content into minimum output units, and constructing a recommended content output queue; and calling a synchronization lock mechanism to assign dynamic priority tags based on the content scheduling index and the recommended content output queue. generating synchronous output control data according to the priority tag and the device output capability parameter; The semantic tagging includes: based on the synchronous output control data, calling the selected recommended video clip and extracting a continuous key frame sequence, aligning the time axis of each key frame image data with the corresponding voice command; using a natural language processing model to perform bidirectional matching between the keywords in the semantic chain and the key frame image content, and tagging the semantic tags of the corresponding actions at the frame granularity; The dynamic scheduling of multimodal content output includes: extracting the content prompt unit corresponding to the current semantic chain node in the main module according to the semantic tag contained in the interactive prompt data, and loading the corresponding time window identifier and semantic tag; combining the output timing and content priority in the synchronous output control data, calling the auxiliary module content frame segment synchronized with the main module, and setting the output delay window for the low-priority module.
2. The method for outputting multimodal data of children according to claim 1, wherein: The dynamic priority label allocation includes: Based on the content scheduling index, key output units are selected from the recommended content output queue to construct a content candidate set. Based on the content candidate set, the timing window mechanism is called to extract the response delay estimate of the output unit and generate a delay weighting factor. With reference to the current device output capability parameters, a fusion score value is calculated for each output unit. The fusion score value is interval-mapped and labeled with the corresponding dynamic priority label.
3. The method for outputting multimodal data of children according to claim 2, wherein: The screening of the adapted action segments includes: confirming the structural adaptability of the action based on infant and toddler stage rules, and screening out stage-matched action frames.
4. The method for outputting multimodal data of children according to claim 1, wherein: The two-way matching includes: Construct temporal embedding vectors based on the sorted keywords and subtopic identifiers in the semantic chain; Divide the continuous key frames into image blocks and extract the semantic features of the local image blocks; Calculate the similarity between the temporal embedding vector and the semantic features of the image patch, and mark the matching relationship.
5. The method for outputting multimodal data of children according to claim 1, wherein: The dynamically scheduling the multimodal content output further includes: During the multimodal content output process, the current output state is tested for conversation turn matching. When the correlation between the historical semantic chain node and the new input instruction is detected to be lower than the threshold, the semantic chain node is reconstructed and the output mapping relationship between the main and auxiliary modules is redirected. The output content of the main module and the delayed content of the auxiliary module are merged into a modal collaborative output sequence and associated with the semantic chain nodes.
6. A children's multimodal data output system, based on the children's multimodal data output method according to any one of claims 1 to 5, characterized in that: Also includes: The semantic chain building module collects the interaction requests entered by parents in the dialogue window, builds a cross-turn semantic chain based on the input content of historical rounds, extracts the current session state vector, and simultaneously generates interaction trigger data; A synchronization scheduling module, based on the session state vector and the interaction trigger data, applies a synchronization lock mechanism to schedule the recommended content output queue and generates synchronization output control data; The label screening module, based on the synchronous output control data, semantically labels the key frame actions in the recommended videos, selects the appropriate action segments based on the infant stage rules, and constructs structured content segments and interactive prompt data; The content output module uses the interactive prompt data and the synchronous output control data to perform dynamic scheduling on the multimodal content output.