Context-based multi-modal data generation method and device, equipment and medium

By using a pre-set semantic recognition model to perform semantic segmentation on the context text and automatically extracting target semantic tags, the system achieves accurate matching between text content and multimedia materials, generating multimodal data that integrates text, audio, and video. This solves the problems of difficulty in understanding and cumbersome data entry for elderly users in intelligent voice interaction systems, and improves interaction efficiency.

CN121168431APending Publication Date: 2025-12-19PING AN HEALTH CLOUD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511115765.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing intelligent voice interaction systems lack emotional interaction in the fields of fintech and healthcare and elderly care, leading to difficulties in understanding for elderly users, cumbersome data entry, and low interaction efficiency.

Method used

By using a pre-set semantic recognition model to perform semantic segmentation on the context text, the target semantic tags are automatically extracted, achieving accurate matching between text content and multimedia materials, and generating multimodal data integrating text, images, audio, and video. Users only need to engage in natural dialogue without complicated operations.

Benefits of technology

It improves the interaction efficiency between the intelligent voice interaction system and users, enhances the understanding of elderly users and the convenience of data entry, and improves the interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121168431A_ABST
    Figure CN121168431A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of speech semantics, and discloses a context-based multi-modal data generation method, device and equipment and a medium, and the method comprises the steps: carrying out the semantic segmentation of context text information through a semantic recognition model, and determining a target semantic tag; multimedia materials are obtained, association is carried out according to the multimedia materials and the context text information, and initial multi-modal data are generated; and generating target multi-modal data according to the initial multi-modal data and the multi-modal data template. According to the mode, the target semantic label of the context text is extracted through the preset semantic recognition model, the text subjected to voice transcription is dynamically associated with the multimedia material according to the semantic label, and the multi-modal data integrating the image-text, the audio and the video is generated, so that a user only needs natural dialogue and does not need typewriting or complex operation, and the user experience is improved. In the business fields of financial science and technology, medical health, old-age care and the like, the interaction efficiency between the intelligent voice interaction system and the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech and semantic technology, and in particular to a method, apparatus, device and medium for generating multimodal data based on context. Background Technology

[0002] In the field of fintech business, existing banks, insurance companies and securities firms only provide standardized forms for "asset statements" and "pension annual statements", which lack storytelling and timeline presentation. Elderly customers find it difficult to intuitively understand the life events and meaning behind asset changes. High-net-worth elderly customers need to repeatedly fill out lengthy questionnaires or go to the branch for face-to-face signing when making family trusts or wealth transfers, and there is a lack of voice guidance and emotional interview methods.

[0003] In the field of medical and health care and elderly care, traditional electronic medical records are mainly based on structured terminology and lack the patient's first-person voice, family perspective and life scene reproduction, which is not conducive to palliative care, spiritual comfort and family joint decision-making. Existing health management terminals have small fonts and deep menus, making it difficult for elderly users to upload image data and enter medical history, resulting in data loss or errors.

[0004] Therefore, improving the interaction efficiency between intelligent voice interaction systems and users in business areas such as fintech, healthcare, and elderly care has become an urgent technical problem to be solved. Summary of the Invention

[0005] This application provides a context-based multimodal data generation method, apparatus, device, and medium to improve the interaction efficiency between intelligent voice interaction systems and users.

[0006] In a first aspect, this application provides a context-based multimodal data generation method, the method comprising:

[0007] Based on the target topic, collect the contextual speech information of the target user, and convert the contextual speech information into contextual text information;

[0008] The context text information is semantically segmented using a preset semantic recognition model to determine at least one target semantic label for the context text information.

[0009] The multimedia materials of the target user are acquired, and the multimedia materials and the contextual text information are associated according to the target semantic tags to generate initial multimodal data;

[0010] A multimodal data template is determined based on the target topic, and target multimodal data is generated based on the initial multimodal data and the multimodal data template.

[0011] Secondly, this application also provides a context-based multimodal data generation apparatus, the apparatus comprising:

[0012] The context text information conversion module is used to collect the context speech information of the target user according to the target topic, and convert the context speech information into context text information;

[0013] The target semantic label determination module is used to perform semantic segmentation on the context text information through a preset semantic recognition model, and determine at least one target semantic label of the context text information.

[0014] An initial multimodal data generation module is used to acquire multimedia materials of the target user and associate the multimedia materials with the contextual text information according to each target semantic tag to generate initial multimodal data;

[0015] The target multimodal data generation module is used to determine a multimodal data template based on the target topic, and to generate target multimodal data based on the initial multimodal data and the multimodal data template.

[0016] Thirdly, this application also provides a computer device, the computer device including a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and, when executing the computer program, implement the context-based multimodal data generation method as described above.

[0017] Fourthly, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement the context-based multimodal data generation method described above.

[0018] This application discloses a context-based multimodal data generation method, apparatus, device, and medium. The context-based multimodal data generation method includes: collecting contextual speech information of a target user based on a target topic and converting the contextual speech information into contextual text information; semantically segmenting the contextual text information using a preset semantic recognition model to determine at least one target semantic tag for the contextual text information; acquiring multimedia materials of the target user and associating the multimedia materials with the contextual text information according to each target semantic tag to generate initial multimodal data; determining a multimodal data template based on the target topic and generating target multimodal data based on the initial multimodal data and the multimodal data template. Through the above method, this application performs semantic segmentation of contextual text using a preset semantic recognition model, automatically extracts target semantic tags, achieves accurate matching between text content and multimedia materials, dynamically associates speech-transcribed text with multimedia materials according to semantic tags, and generates integrated multimodal data of text, audio, and video. Users only need natural dialogue, without typing or complex operations, improving the interaction efficiency between intelligent voice interaction systems and users in business fields such as fintech, healthcare, and elderly care. Attached Figure Description

[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a schematic flowchart of a context-based multimodal data generation method provided in an embodiment of this application;

[0021] Figure 2 A schematic block diagram of a context-based multimodal data generation apparatus provided for embodiments of this application;

[0022] Figure 3 A schematic block diagram of the structure of a computer device provided for an embodiment of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0025] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0026] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0027] This application provides a context-based multimodal data generation method, apparatus, device, and medium. The context-based multimodal data generation method can be applied to intelligent voice interaction systems. It performs semantic segmentation of contextual text using a preset semantic recognition model, automatically extracts target semantic tags, and achieves accurate matching of text content with multimedia materials. The speech-transcribed text and multimedia materials are dynamically associated according to semantic tags, generating integrated multimodal data encompassing text, audio, and video. Users only need to engage in natural dialogue, without typing or complex operations, thus improving the interaction efficiency between intelligent voice interaction systems and users in business fields such as fintech, healthcare, and elderly care.

[0028] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0029] Please see Figure 1 , Figure 1 This is a schematic flowchart illustrating a context-based multimodal data generation method provided in an embodiment of this application. This context-based multimodal data generation method can be applied to intelligent voice interaction systems, improving the interaction efficiency between the intelligent voice interaction system and the user in business areas such as fintech, healthcare, and elderly care.

[0030] like Figure 1 As shown, the context-based multimodal data generation method specifically includes steps S10 to S40.

[0031] Step S10: Based on the target topic, collect the contextual speech information of the target user and convert the contextual speech information into contextual text information;

[0032] Specifically, this embodiment will be described using specific business scenarios in the field of fintech. The numerical values ​​are for illustrative purposes only and do not represent any limitations.

[0033] On senior-friendly smart terminals or smart speakers in banks or financial institutions, users can select target topics by voice, such as "My First Investment," "30-Year Pension Account Story," or "Family Trust Establishment Review."

[0034] The system activates a directional microphone array to collect the user's contextual voice information (such as "I bought my first fund at the bank counter in 1998, and the handling fee was 1.5% at that time"), and converts it into contextual text information in real time through a preset speech conversion model; it also performs high-precision financial entity recognition for keywords such as amount, date, and institution name.

[0035] In practical applications within the healthcare and elderly care sector, users select target themes such as "My 30 Days of Post-Operative Rehabilitation," "20 Years of Hypertension Management," or "Early Care Records for Alzheimer's Disease."

[0036] The system activates a medical-grade noise-canceling microphone array to collect user contextual voice information (such as "On the third day after surgery, I walked 50 meters with a walker, my heart rate was 110, and my blood pressure was 135 / 85..."), and converts it into contextual text information in real time through a speech conversion model enhanced with medical corpus; it also performs high-confidence recognition of medical entities (symptoms, signs, drug names, dosages).

[0037] Step S20: Semantically segment the context text information using a preset semantic recognition model to determine at least one target semantic label for the context text information;

[0038] Specifically, a pre-set semantic recognition model, fine-tuned from financial corpus, is invoked to perform semantic segmentation on the context text, automatically extracting and outputting several target semantic tags, such as timestamps (e.g., 1998-05-12), trading instruments (e.g., "open-ended equity funds"), transaction amounts (e.g., "10,000 yuan"), or related events (e.g., "preparation for college expenses").

[0039] Step S30: Obtain the multimedia materials of the target user, and associate the multimedia materials with the contextual text information according to each target semantic tag to generate initial multimodal data;

[0040] Specifically, the system prompts target users to upload multimedia materials (such as old photos, audio recordings, and videos) related to the current topic through voice guidance. Image recognition, speech recognition, and text recognition technologies are used to analyze the content of the materials, extract key information such as people, places, times, and events, and map them into corresponding target semantic tags.

[0041] The parsed tags are semantically matched and associated with the contextual text information generated during the interview (including paragraphs, keywords, and timestamps transcribed from user speech), and the materials are automatically inserted into the corresponding positions in the text, thereby generating initial multimodal data containing images, text, audio, and video that accurately correspond to the semantic tags.

[0042] Specifically, the system uploads historical photos (such as family photos or medical records) via a terminal device's camera / album and connects to wearable devices to obtain health monitoring videos (such as rehabilitation training videos). It also scans paper documents to generate digitized materials (such as prescriptions and examination reports). The system analyzes the shooting time in the photo metadata and associates it with text events within the same time period. If there is no time metadata, it infers the era characteristics through image recognition.

[0043] In a specific embodiment, photos, videos, audio files, and PDFs are automatically extracted from the user's mobile phone album, chat history, and authorized electronic files from hospitals, banks, and elderly care platforms. Based on semantic tags such as "account opening year," "pain level 7," and "first pension payment," intelligent matching is performed using timestamps, text recognition technology, and facial and scene recognition. For example, a scanned copy of the 1998 account opening application is pasted into the "account opening year" paragraph, and a photo of the wound today is embedded below the "pain level 7" description, while retaining the patient's original voice "It still hurts today" as an audio node. All materials and text are automatically aggregated according to a three-dimensional index of tags, time, and emotion to form an initial multimodal data package that can be directly rendered.

[0044] Step S40: Determine a multimodal data template based on the target topic, and generate target multimodal data based on the initial multimodal data and the multimodal data template.

[0045] Specifically, based on the user's selected life stage theme (such as "childhood growth"), the system automatically retrieves the corresponding multimodal data template. This template pre-sets fields for text, audio, and video, along with their semantic tags. The system uses the user's initial voice-to-text responses, uploaded photos, and video clips as initial multimodal data. This data is then semantically matched and aligned with the fields in the template. Based on this, text paragraphs, images, audio clips, and video clips are automatically embedded into the corresponding positions in the template. After layout, background music, and subtitle processing, structured target multimodal memoir data is generated with a single click.

[0046] In specific business scenarios, based on the target theme (such as "30 days of postoperative recovery" or "pension investment journey"), the corresponding multimodal data template is called from the template library. The template predefines timeline segments, key indicator display positions, emotional original sound insertion points, and compliant desensitization areas.

[0047] The generated initial multimodal data (aligned text + images / audio / curves) is mapped to a unified feature space according to template rules, and the feature vectors of different modalities are aligned to a shared high-dimensional space through linear projection. Then, a fusion network is used to weight and stitch the aligned features to form a fusion feature map, which is then divided into several sub-regions according to the slice windows in the template. Each sub-region corresponds to a visualization component (such as a line chart, wound comparison image, and original sound playback button). Finally, the system automatically renders and typesets the data according to preset layout parameters (number of partitions, information level, update frequency, and display precision), and outputs a target multimodal data package that can be directly previewed, edited, and shared in compliance with regulations on smart terminals.

[0048] This application discloses a context-based multimodal data generation method. The method includes: collecting contextual speech information of a target user based on a target topic and converting the contextual speech information into contextual text information; semantically segmenting the contextual text information using a preset semantic recognition model to determine at least one target semantic tag for the contextual text information; acquiring multimedia materials from the target user and associating the multimedia materials with the contextual text information according to each target semantic tag to generate initial multimodal data; determining a multimodal data template based on the target topic and generating target multimodal data based on the initial multimodal data and the multimodal data template. Through this method, this application performs semantic segmentation of contextual text using a preset semantic recognition model, automatically extracts target semantic tags, achieves accurate matching between text content and multimedia materials, dynamically associates speech-transcribed text with multimedia materials according to semantic tags, and generates integrated multimodal data encompassing text, audio, and video. Users only need to engage in natural dialogue without typing or complex operations, improving the interaction efficiency between intelligent voice interaction systems and users in business fields such as fintech, healthcare, and elderly care.

[0049] based on Figure 1 In the illustrated embodiment, step S10 includes:

[0050] Based on the target topic, retrieve the target dialogue template from the preset dialogue template library;

[0051] Based on the target dialogue objective, guide the target user to interact and generate the contextual voice information;

[0052] The contextual speech information is converted into the contextual text information based on a preset time series.

[0053] Specifically, in fintech scenarios, the system first retrieves the corresponding "target dialogue template" from a pre-set dialogue template library based on the user's selected "target topic" (such as "personal financial planning" or "investment risk assessment").

[0054] The intelligent voice interaction system uses the structured questions in this template as a script to guide users to describe their asset status, risk preferences and investment experience through real-time voice calls. During the process, follow-up questions can be asked in real time and contextual voice information can be recorded.

[0055] The intelligent voice interaction system sends continuous contextual voice information into a financial-grade voice recognition engine according to a preset time sequence (such as slices every 3 seconds). Combined with financial knowledge graphs and real-time market data, it converts the information into time-stamped contextual text information, providing structured input for the subsequent generation of personalized financial reports or risk assessment documents.

[0056] based on Figure 1 In the illustrated embodiment, step S20 includes:

[0057] The independent event description text and / or emotional expression text of the contextual text information are identified by a preset semantic recognition model;

[0058] Convert the independent event description text and / or the sentiment expression text into subtext feature vectors;

[0059] The subtext feature vector is matched with each preset tag feature vector in the preset tag feature vector library;

[0060] The preset labels corresponding to each preset label feature vector that matches the subtext feature vector are determined as the target semantic labels.

[0061] Specifically, in the fintech business field, intelligent voice interaction utilizes a pre-set semantic recognition model trained specifically for financial scenarios to extract events and perform sentiment analysis on "contextual text information," identifying independent event description text and sentiment expression text.

[0062] The two types of text are encoded into sub-text feature vectors using domain word vectors and pre-trained models respectively. Then, these sub-text feature vectors are matched with the financial tag feature vectors covered in the preset tag feature vector library using cosine similarity. When the similarity exceeds a set threshold (such as 0.85), the corresponding preset tag is determined as the target semantic tag of the paragraph, and a personalized financial memoir or risk assessment report is subsequently generated.

[0063] Specifically, in the healthcare business field, the preset semantic recognition model first scans the context text, automatically segments the independent event description text and sentiment expression text, and marks them with "event" or "sentiment" type tags respectively.

[0064] Independent event description texts are generated into 128-dimensional vectors by a medical entity encoder, highlighting key features such as operation, location, and time.

[0065] The emotional expression text generates a 64-dimensional vector through an emotion-symptom joint encoder, which quantifies the intensity of emotions such as anxiety and pain tolerance.

[0066] The cosine similarity between the two types of sub-text feature vectors and the preset label feature vector library is calculated. When the similarity is greater than the set threshold (e.g., 0.85), the corresponding label is locked as the target semantic label.

[0067] based on Figure 1 In the illustrated embodiment, step S30 includes:

[0068] Extract the metadata features of the multimedia material;

[0069] Calculate the similarity matrix between each of the metadata features and each of the target semantic tags;

[0070] Based on the similarity scores in the similarity matrix, the correspondence between the metadata features and the target semantic tags is determined.

[0071] Based on the corresponding relationships, the multimedia materials and the contextual text information are associated and fused to generate the initial multimodal data.

[0072] Specifically, metadata features are extracted from each uploaded image, video clip, or audio file, including timestamps, location, identified text, image visual vectors, and audio spectral features. These features are then uniformly encoded into metadata feature vectors. These metadata feature vectors are then compared pairwise with the preset label feature vectors corresponding to the identified target semantic tags to calculate cosine similarity and construct a "metadata feature-semantic tag" similarity matrix.

[0073] The similarity matrix is ​​thresholded (e.g., items with similarity > 0.8 are retained) to determine the target semantic tag that best matches each metadata feature, establish the correspondence between "materials and tags", and based on these correspondences, multimedia materials are automatically inserted into the context text at positions that match their semantic tags, forming initial multimodal data that aligns text, audio, and video content.

[0074] based on Figure 1 In the illustrated embodiment, step S40 includes:

[0075] The target topic is matched with the preset topics corresponding to each preset data template in the preset data template library;

[0076] The preset data template corresponding to the preset theme that matches the target theme is determined as the multimodal data template.

[0077] Specifically, the system uses the user's selected "target topic" (such as "childhood development" or "career experience") as the query key to send a matching request to the preset data template library.

[0078] Each preset data template in the preset data template library is bound to a "preset topic" field. The matching degree between the "target topic" and each "preset topic" is compared by string similarity or keyword matching algorithms (such as semantic embedding), and the template with the highest similarity is selected as the "multimodal data template".

[0079] In a specific embodiment, step S40 further includes:

[0080] The structured placeholder regions of the initial multimodal data are parsed, wherein the structured placeholder regions include text container regions, image slot regions, and original sound anchor point regions;

[0081] The target multimodal data is generated based on the target semantic tags and each of the structured placeholder regions.

[0082] Specifically, the initial multimodal data (which already includes contextual text, images, audio, and video materials) is loaded, and three types of structured placeholder areas are identified and marked: text container area, image slot area, and original sound anchor point area.

[0083] The text container area is used to hold narrative text that has been categorized by semantic tags;

[0084] The image slot area is used to embed wound photos, examination reports, or rehabilitation demonstration images that match the labels;

[0085] The original audio anchor point area is used to locate the play button for the patient's original audio clip.

[0086] In the fields of medical care, health and elderly care, based on target semantic tags such as "pain relief", the corresponding text paragraphs, images and audio are automatically filled into the above areas: the text container area is inserted with the transcribed sentence "Today's pain has dropped to 3 points", the image slot area is inserted with the wound comparison image of the day, and the original audio anchor point area is inserted with the voice node of the patient's self-statement "I feel much better". The three are bidirectionally bound through tag IDs and finally rendered into complete target multimodal data that can be directly shared with family members.

[0087] In a specific embodiment, the target multimodal data is generated based on the target semantic tags and each of the structured placeholder regions, including:

[0088] The initial multimodal data is decomposed into text data, image data, and voice data based on the target semantic tags;

[0089] The text data is filled into the text container area, the image data is filled into the image slot area, and the audio data is filled into the original audio anchor point area to generate data to be rendered;

[0090] Based on preset rendering rules and the data to be rendered, the target multimodal data is generated.

[0091] Specifically, the initial multimodal data is split into ternary segments based on the target semantic labels:

[0092] Text data: All text paragraphs under the corresponding tag;

[0093] Image data: All images, video keyframes, and their metadata under the corresponding tags;

[0094] Voice data: Original audio segments and metadata (timestamp, duration) under the corresponding tags.

[0095] Iterate through the text data and write it into the parsed text container area in paragraph order;

[0096] Image data is inserted sequentially into image slots with the same name based on timestamps, and the size is automatically adjusted and text is generated.

[0097] Embed the voice data into the original audio anchor area of ​​the same name tag.

[0098] In the fields of medical care, health, and elderly care, based on target semantic tags such as "pain relief," the initial multimodal data is divided into three categories:

[0099] Desensitized text data, such as "Today's pain score: 3 points";

[0100] Image data corresponding to specific time points, such as wound photos and thumbnails of test reports;

[0101] Patient's original voice data, such as a 30-second recording of "I feel much better".

[0102] Next, according to the structured placeholder rules, the three types of materials are written into the "text container area" to generate paragraphs; image data is adaptively filled into the "image slot area" at a ratio of 1:1 or 16:9 and alt text is automatically added; and voice data is embedded into the "original sound anchor point area" to generate a waveform icon with a playback bar.

[0103] Please see Figure 2 , Figure 2 This application provides a schematic block diagram of a context-based multimodal data generation apparatus, which is used to execute the aforementioned context-based multimodal data generation method. The context-based multimodal data generation apparatus can be configured on a server.

[0104] like Figure 2 As shown, the context-based multimodal data generation apparatus 400 includes:

[0105] The context text information conversion module 410 is used to collect the context speech information of the target user according to the target topic, and convert the context speech information into context text information;

[0106] The target semantic label determination module 420 is used to perform semantic segmentation on the context text information through a preset semantic recognition model, and determine at least one target semantic label of the context text information.

[0107] The initial multimodal data generation module 430 is used to acquire the multimedia materials of the target user and associate the multimedia materials with the contextual text information according to each target semantic tag to generate initial multimodal data;

[0108] The target multimodal data generation module 440 is used to determine a multimodal data template based on the target topic, and to generate target multimodal data based on the initial multimodal data and the multimodal data template.

[0109] Furthermore, the context text information conversion module 410 includes:

[0110] The target dialogue template invocation submodule is used to invoke the target dialogue template from the preset dialogue template library according to the target topic;

[0111] The context speech information generation submodule is used to guide the target user to interact and generate the context speech information based on the target dialogue objective.

[0112] The context text information determination submodule is used to convert the context speech information into the context text information based on a preset time series.

[0113] Furthermore, the target semantic tag determination module 420 includes:

[0114] The contextual text information recognition submodule is used to recognize the independent event description text and / or sentiment expression text of the contextual text information through a preset semantic recognition model;

[0115] The subtext feature vector conversion submodule is used to convert the independent event description text and / or the sentiment expression text into subtext feature vectors;

[0116] The matching submodule is used to match the subtext feature vector with each preset tag feature vector in the preset tag feature vector library;

[0117] The target semantic label determination submodule is used to determine the preset labels corresponding to each preset label feature vector that matches the subtext feature vector as the target semantic label.

[0118] Furthermore, the initial multimodal data generation module 430 includes:

[0119] The metadata feature extraction submodule is used to extract the metadata features of the multimedia material;

[0120] A similarity matrix calculation submodule is used to calculate the similarity matrix between each of the metadata features and each of the target semantic tags;

[0121] The correspondence determination submodule is used to determine the correspondence between the metadata features and the target semantic tags based on the similarity scores in the similarity matrix.

[0122] The initial multimodal data generation submodule is used to associate and fuse the multimedia materials and the contextual text information according to the corresponding relationships to generate the initial multimodal data.

[0123] Furthermore, the target multimodal data generation module 440 includes:

[0124] The preset topic matching submodule is used to match the target topic with the preset topics corresponding to each preset data template in the preset data template library;

[0125] The multimodal data template determination submodule is used to determine the preset data template corresponding to the preset theme that matches the target theme as the multimodal data template.

[0126] Furthermore, the target multimodal data generation module 440 includes:

[0127] The structured placeholder region parsing submodule is used to parse the structured placeholder region of the initial multimodal data, wherein the structured placeholder region includes a text container region, an image slot region, and an original sound anchor point region;

[0128] The target multimodal data generation submodule is used to generate the target multimodal data based on the target semantic tags and each of the structured placeholder regions.

[0129] Furthermore, the target multimodal data generation submodule includes:

[0130] An initial multimodal data decomposition unit is used to decompose the initial multimodal data into text data, image data, and voice data according to the target semantic label;

[0131] The data to be rendered generation unit is used to generate data to be rendered by filling the text data into the text container area, filling the image data into the image slot area, and filling the voice data into the original sound anchor point area.

[0132] The target multimodal data generation unit is used to generate the target multimodal data based on preset rendering rules and the data to be rendered.

[0133] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the above-described apparatus and modules can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0134] The aforementioned apparatus can be implemented as a computer program, which can be used in, for example... Figure 3 It runs on the computer device shown.

[0135] Please see Figure 3 , Figure 3 This is a schematic block diagram illustrating the structure of a computer device according to an embodiment of this application. The computer device may be a server.

[0136] See Figure 3 The computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.

[0137] Non-volatile storage media can store operating systems and computer programs. These computer programs include program instructions that, when executed, cause the processor to perform any context-based multimodal data generation method.

[0138] The processor provides computing and control capabilities, supporting the operation of the entire computer device.

[0139] Internal memory provides an environment for the execution of computer programs in non-volatile storage media. When the computer program is executed by the processor, it enables the processor to perform any context-based multimodal data generation method.

[0140] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0141] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.

[0142] In one embodiment, the processor is configured to run a computer program stored in memory to perform the following steps:

[0143] Based on the target topic, collect the contextual speech information of the target user, and convert the contextual speech information into contextual text information;

[0144] The context text information is semantically segmented using a preset semantic recognition model to determine at least one target semantic label for the context text information.

[0145] The multimedia materials of the target user are acquired, and the multimedia materials and the contextual text information are associated according to the target semantic tags to generate initial multimodal data;

[0146] A multimodal data template is determined based on the target topic, and target multimodal data is generated based on the initial multimodal data and the multimodal data template.

[0147] In one embodiment, the contextual speech information of the target user is collected according to the target topic, and the contextual speech information is converted into contextual text information to achieve the following:

[0148] Based on the target topic, retrieve the target dialogue template from the preset dialogue template library;

[0149] Based on the target dialogue objective, guide the target user to interact and generate the contextual voice information;

[0150] The contextual speech information is converted into the contextual text information based on a preset time series.

[0151] In one embodiment, the contextual text information is semantically segmented using a preset semantic recognition model to determine at least one target semantic label of the contextual text information, for the purpose of:

[0152] The independent event description text and / or emotional expression text of the contextual text information are identified by a preset semantic recognition model;

[0153] Convert the independent event description text and / or the sentiment expression text into subtext feature vectors;

[0154] The subtext feature vector is matched with each preset tag feature vector in the preset tag feature vector library;

[0155] The preset labels corresponding to each preset label feature vector that matches the subtext feature vector are determined as the target semantic labels.

[0156] In one embodiment, the multimedia material and the contextual text information are associated according to the target semantic tags to generate initial multimodal data for the purpose of:

[0157] Extract the metadata features of the multimedia material;

[0158] Calculate the similarity matrix between each of the metadata features and each of the target semantic tags;

[0159] Based on the similarity scores in the similarity matrix, the correspondence between the metadata features and the target semantic tags is determined.

[0160] Based on the corresponding relationships, the multimedia materials and the contextual text information are associated and fused to generate the initial multimodal data.

[0161] In one embodiment, a multimodal data template is determined based on the target topic to achieve:

[0162] The target topic is matched with the preset topics corresponding to each preset data template in the preset data template library;

[0163] The preset data template corresponding to the preset theme that matches the target theme is determined as the multimodal data template.

[0164] In one embodiment, target multimodal data is generated based on the initial multimodal data and the multimodal data template, for the purpose of:

[0165] The structured placeholder regions of the initial multimodal data are parsed, wherein the structured placeholder regions include text container regions, image slot regions, and original sound anchor point regions;

[0166] The target multimodal data is generated based on the target semantic tags and each of the structured placeholder regions.

[0167] In one embodiment, the target multimodal data is generated based on the target semantic tags and each of the structured placeholder regions, for the purpose of:

[0168] The initial multimodal data is decomposed into text data, image data, and voice data based on the target semantic tags;

[0169] The text data is filled into the text container area, the image data is filled into the image slot area, and the audio data is filled into the original audio anchor point area to generate data to be rendered;

[0170] Based on preset rendering rules and the data to be rendered, the target multimodal data is generated.

[0171] The embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions, and the processor executing the program instructions to implement any of the context-based multimodal data generation methods provided in the embodiments of this application.

[0172] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.

[0173] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A context-based multi-modal data generation method, characterized by, include: Based on the target topic, collect the contextual speech information of the target user, and convert the contextual speech information into contextual text information; The context text information is semantically segmented using a preset semantic recognition model to determine at least one target semantic label for the context text information. The multimedia materials of the target user are acquired, and the multimedia materials and the contextual text information are associated according to the target semantic tags to generate initial multimodal data; A multimodal data template is determined based on the target topic, and target multimodal data is generated based on the initial multimodal data and the multimodal data template. 2.The context-based multi-modal data generation method of claim 1, wherein, The step of collecting contextual speech information of the target user based on the target topic and converting the contextual speech information into contextual text information includes: Based on the target topic, retrieve the target dialogue template from the preset dialogue template library; Based on the target dialogue objective, guide the target user to interact and generate the contextual voice information; The contextual speech information is converted into the contextual text information based on a preset time series. 3.The context-based multi-modal data generation method of claim 1, wherein, The step of semantically segmenting the contextual text information using a preset semantic recognition model to determine at least one target semantic label for the contextual text information includes: The independent event description text and / or emotional expression text of the contextual text information are identified by a preset semantic recognition model; Convert the independent event description text and / or the sentiment expression text into subtext feature vectors; The subtext feature vector is matched with each preset tag feature vector in the preset tag feature vector library; The preset labels corresponding to each preset label feature vector that matches the subtext feature vector are determined as the target semantic labels. 4.The context-based multi-modal data generation method of claim 1, wherein, The step of associating the multimedia material and the contextual text information according to each of the target semantic tags to generate initial multimodal data includes: Extract the metadata features of the multimedia material; Calculate the similarity matrix between each of the metadata features and each of the target semantic tags; Based on the similarity scores in the similarity matrix, the correspondence between the metadata features and the target semantic tags is determined. Based on the corresponding relationships, the multimedia materials and the contextual text information are associated and fused to generate the initial multimodal data. 5.The context-based multi-modal data generation method of claim 1, wherein, The step of determining the multimodal data template based on the target topic includes: The target topic is matched with the preset topics corresponding to each preset data template in the preset data template library; The preset data template corresponding to the preset theme that matches the target theme is determined as the multimodal data template. 6.The context-based multi-modal data generation method of claim 5, wherein, The step of generating target multimodal data based on the initial multimodal data and the multimodal data template includes: The structured placeholder regions of the initial multimodal data are parsed, wherein the structured placeholder regions include text container regions, image slot regions, and original sound anchor point regions; The target multimodal data is generated based on the target semantic tags and each of the structured placeholder regions.

7. The context-based multi-modal data generation method of claim 6, wherein, The step of generating the target multimodal data based on the target semantic tags and each of the structured placeholder regions includes: The initial multimodal data is decomposed into text data, image data, and voice data based on the target semantic tags; The text data is filled into the text container area, the image data is filled into the image slot area, and the audio data is filled into the original audio anchor point area to generate data to be rendered; Based on preset rendering rules and the data to be rendered, the target multimodal data is generated.

8. A context-based multimodal data generation device, characterized in that, include: The context text information conversion module is used to collect the context speech information of the target user according to the target topic, and convert the context speech information into context text information; The target semantic label determination module is used to perform semantic segmentation on the context text information through a preset semantic recognition model, and determine at least one target semantic label of the context text information. An initial multimodal data generation module is used to acquire multimedia materials of the target user and associate the multimedia materials with the contextual text information according to each target semantic tag to generate initial multimodal data; The target multimodal data generation module is used to determine a multimodal data template based on the target topic, and to generate target multimodal data based on the initial multimodal data and the multimodal data template.

9. A computer device, characterized in that, The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and, in executing the computer program, implement the context-based multimodal data generation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to implement the context-based multimodal data generation method as described in any one of claims 1 to 7.