Multimedia real-time interaction method and system based on generative large model mixing
By using a generative large model hybrid approach to synchronize and align multimodal data in time, a high-dimensional semantic feature cloud is constructed, which solves the problem of deep semantic alignment of multimodal data in multimedia interaction in the financial field. This enables virtual digital humans to accurately understand and logically express complex financial content, thereby enhancing the depth and professionalism of the interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI RONGSHU INFORMATION TECH CO LTD
- Filing Date
- 2026-02-26
- Publication Date
- 2026-04-24
AI Technical Summary
Existing technologies in real-time multimedia interaction in the financial field make it difficult to achieve deep semantic alignment and fusion of multimodal data. This results in virtual digital humans having a one-sided understanding of complex financial contexts, and their generated responses may deviate from the core points or lack contextual coherence, making it difficult to conduct in-depth financial logic deduction and knowledge transfer.
By using a generative large model hybridization approach, the multimodal raw data stream is synchronized and time-aligned, and distributed in parallel to a dedicated generative model processing unit in the cloud. A multimodal joint feature tensor is generated using a cross-modal dynamic alignment and fusion network to construct a high-dimensional semantic feature cloud. Through semantic manifold trajectories and dynamic semantic aggregation regions, a virtual digital human driving instruction set is generated to ensure the synchronicity and logical coherence of the interaction.
It achieves deep semantic alignment and fusion of complex multimodal information in the financial field. The virtual digital human can accurately understand professional contexts, avoid information fragmentation and misunderstanding, ensure the depth and professionalism of interaction, improve the logical relevance and naturalness of the virtual digital human's audio and video streams, and meet the immediacy requirements of real-time financial consultation and interactive training.
Smart Images

Figure CN121725116B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimedia real-time interactive technology, and in particular to a multimedia real-time interactive method and system based on generative large model hybridization. Background Technology
[0002] With the rapid development of fintech, the demand for interactive, personalized, and highly immersive multimedia real-time interaction is becoming increasingly urgent in scenarios such as knowledge dissemination, customer service, investment advisory services, and employee training in the financial sector. Traditional online financial courses, virtual customer service, or financial news broadcasts mostly use pre-recorded videos, text and image pushes, or Q&A interactions based on simple rules and scripts. Their content is static, their responses are rigid, and they lack in-depth real-time interaction and emotional exchange, making it difficult to meet users' high requirements for professionalism, real-time performance, and immersive experiences.
[0003] In recent years, breakthroughs in large-scale generative AI models (such as large language models, text-to-image models, and audio / video generation models) have provided a new technological foundation for building intelligent interactive virtual digital humans. Existing attempts often employ single or loosely combined generative models to process user input (such as text, speech, and images), thereby driving the virtual digital human to respond. However, in the highly specialized, logically-driven, and information-dense field of finance, existing technological solutions face the following challenges:
[0004] Financial explanations or consultations often involve multimodal information such as charts (candlestick charts, financial statement charts), data streams, audio explanations of professional terms, and lecturer gestures. Existing methods either process the data of each modality independently or simply splice them together, lacking in-depth cross-modal semantic alignment and fusion. This results in virtual digital humans having a one-sided understanding of complex financial contexts, and the generated responses may deviate from the core points or lack contextual coherence.
[0005] Financial topics are typically characterized by rigorous logic and closely related concepts. Existing real-time interactive systems struggle to dynamically construct and maintain a deep, structured semantic cognitive system during continuous dialogue or explanation. Virtual digital avatars' responses are often limited to the current turn of conversation, lacking organic retrospection and extension of previous discussion points, and unable to predict and construct subsequent semantic trajectories. This results in superficial interactions, hindering in-depth financial logic deduction and knowledge transfer. Summary of the Invention
[0006] The technical problem to be solved by this invention is to provide a multimedia real-time interaction method and system based on generative large model hybridization, which is a novel interaction method that can generate highly accurate and synchronized virtual digital human behavior instructions, and can improve the intelligence of multimedia real-time interaction in the financial field.
[0007] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:
[0008] Firstly, a real-time multimedia interactive method based on generative large model hybridization, the method comprising:
[0009] The multimodal raw data stream from the teaching end is synchronized and time-aligned, and decoupled according to modality type, and distributed in parallel to the corresponding dedicated generative model processing unit in the cloud;
[0010] Each dedicated generative model processing unit performs real-time parsing and processing of the input data, and outputs multiple sets of feature vectors;
[0011] Multiple sets of feature vectors are input into a cross-modal dynamic alignment and fusion network to generate a multimodal joint feature tensor;
[0012] The multimodal joint feature tensor is mapped to a high-dimensional semantic feature cloud, where each point represents a fused feature unit. Multiple semantic manifold trajectories are constructed in the high-dimensional semantic feature cloud based on the semantic correlation and spatiotemporal proximity between feature units. A dynamic semantic aggregation region is defined based on the distribution and intersection of the semantic manifold trajectories.
[0013] Multiple semantic sampling anchor points are selected at the internal core, boundary, and external associated locations of the dynamic semantic aggregation region; through the temporal progression of teaching interaction, the semantic sampling anchor points are connected to form a higher-order semantic closed manifold; the topological stability index and semantic density gradient of the higher-order semantic closed manifold at multi-dimensional scale are calculated.
[0014] Based on the multimodal joint feature tensor and the corresponding topological stability index and semantic density gradient, a set of virtual digital human driving instructions calibrated with spatial topological consistency is synthesized through a multimodal instruction generation network.
[0015] The virtual digital human driving instruction set is parsed and rendered at the edge node to generate a synchronized virtual digital human audio and video stream, which is then pushed to the interactive terminal with low latency.
[0016] Secondly, a real-time interactive multimedia system based on generative large-scale model hybridization includes:
[0017] The data preprocessing and decoupling module is used to synchronize and time-align the multimodal raw data stream from the teaching end, and decouple it according to modality type, and distribute it in parallel to the corresponding dedicated generative model processing unit in the cloud.
[0018] The feature vector output module is used to perform real-time parsing and processing of the input data through various dedicated generative model processing units, and output multiple sets of feature vectors.
[0019] The alignment, calibration, and aggregation module is used to input multiple sets of feature vectors into a cross-modal dynamic alignment and fusion network to generate a multimodal joint feature tensor.
[0020] The mapping, construction, and delimitation module is used to map the multimodal joint feature tensor into a high-dimensional semantic feature cloud, where each point in the high-dimensional semantic feature cloud represents a fused feature unit; multiple semantic manifold trajectories are constructed in the high-dimensional semantic feature cloud through the semantic correlation and spatiotemporal proximity between feature units; and a dynamic semantic aggregation region is defined based on the distribution and intersection of the semantic manifold trajectories.
[0021] The anchor point selection and index gradient calculation module is used to select multiple semantic sampling anchor points at the internal core, boundary, and external associated positions of the dynamic semantic aggregation region; through the temporal progression of teaching interaction, the semantic sampling anchor points are connected to form a high-order semantic closed manifold; the topological stability index and semantic density gradient of the high-order semantic closed manifold at multi-dimensional scale are calculated.
[0022] The instruction set generation module is used to synthesize a virtual digital human driving instruction set that has been spatially topologically consistent, based on the multimodal joint feature tensor and the corresponding topological stability index and semantic density gradient, through a multimodal instruction generation network.
[0023] The parsing, rendering, and push module is used to parse and render the virtual digital human driving instruction set at the edge node, generate a synchronized virtual digital human audio and video stream, and push it to the interactive end with low latency.
[0024] The above-described solution of the present invention has at least the following beneficial effects:
[0025] By constructing a cross-modal dynamic alignment and fusion network and a high-dimensional semantic feature cloud, deep semantic alignment and fusion of complex multimodal information (such as voice, charts, data, and gestures) in the financial field is achieved, enabling virtual digital humans to accurately and comprehensively understand professional contexts and avoid information fragmentation and misunderstanding.
[0026] By dynamically modeling real-time interactive content using semantic manifold trajectories and dynamic semantic aggregation regions, the system can automatically identify and track core topics, related concepts, and their logical relationships in financial explanations, forming a structured cognitive graph that supports deep and coherent financial logical deduction and knowledge transfer. By calculating the topological stability index and semantic density gradient of the higher-order semantic closed manifold, the system can quantitatively evaluate and ensure the stability and rationality of the semantic core during interaction. This effectively prevents semantic drift, logical contradictions, or loss of focus that easily occur in discussions of complex financial topics, enhancing the depth and professionalism of the interaction.
[0027] Based on the aforementioned deep understanding and stability analysis, the driving instruction set synthesized through the multimodal instruction generation network has undergone spatial topology consistency calibration. This makes the generated virtual digital human's audio and video streams more logically relevant and natural in expressing complex financial content, with high synchronization across multiple channels, significantly improving the expressiveness, credibility, and immersiveness of lectures or services. Employing a collaborative architecture of parallel distribution of raw data streams, dedicated cloud-based model processing, and edge node rendering and push, this architecture ensures high performance in complex analysis and computation while achieving low end-to-end latency from information input to virtual human feedback, meeting the stringent immediacy requirements of real-time financial consultation and interactive training scenarios. Attached Figure Description
[0028] Figure 1 This is a flowchart illustrating a real-time multimedia interactive method based on generative large model hybridization provided in an embodiment of the present invention.
[0029] Figure 2 This is a schematic diagram of a multimedia real-time interactive system based on generative large model hybridization provided by an embodiment of the present invention. Detailed Implementation
[0030] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0031] like Figure 1 As shown, embodiments of the present invention propose a real-time multimedia interaction method based on generative large model hybridization, the method comprising the following steps:
[0032] Step 1: Synchronize and time-align the multimodal raw data stream from the teaching end, decouple it according to modality type, and distribute it in parallel to the corresponding dedicated generative model processing unit in the cloud.
[0033] Step 2: Each dedicated generative model processing unit performs real-time parsing and processing of the input data, and outputs multiple sets of feature vectors;
[0034] Step 3: Input multiple sets of feature vectors into a cross-modal dynamic alignment and fusion network to generate a multimodal joint feature tensor;
[0035] Step 4: Map the multimodal joint feature tensor into a high-dimensional semantic feature cloud, where each point represents a fused feature unit; construct multiple semantic manifold trajectories in the high-dimensional semantic feature cloud based on the semantic correlation and spatiotemporal proximity between feature units; and define a dynamic semantic aggregation region based on the distribution and intersection of the semantic manifold trajectories.
[0036] Step 5: Select multiple semantic sampling anchor points at the internal core, boundary, and external associated locations of the dynamic semantic aggregation region; connect the semantic sampling anchor points to form a higher-order semantic closed manifold through the temporal progression of teaching interaction; calculate the topological stability index and semantic density gradient of the higher-order semantic closed manifold at multi-dimensional scale.
[0037] Step 6: Based on the multimodal joint feature tensor and the corresponding topological stability index and semantic density gradient, a set of virtual digital human driving instructions calibrated by spatial topological consistency is synthesized through a multimodal instruction generation network.
[0038] Step 7: Parse and render the virtual digital human driving instruction set at the edge node, generate a synchronized virtual digital human audio and video stream, and push it to the interactive end with low latency.
[0039] In this embodiment of the invention, the latency inconsistency problem of multi-source heterogeneous data (such as audio / video, screen streams, and sensor data) is solved, ensuring the spatiotemporal consistency foundation for subsequent processing. Modal decoupling and parallel distribution fully utilize the parallel processing capabilities of cloud computing, maximizing processing efficiency and initially reducing end-to-end latency, providing a feasible entry point for real-time interaction. The use of specialized large-scale models optimized for specific modalities (such as speech, vision, and text) for parsing fully leverages the expertise of top models in various fields, achieving high-precision, deep-level single-modal feature extraction from the raw data. This specialized division of labor provides high-quality, semantically rich raw materials (feature vectors) for subsequent fusion, a key prerequisite for ensuring the accuracy of overall understanding. Through a dynamic alignment network, not only is the alignment problem of different modal information on the time axis solved, but more importantly, deep semantic association and complementarity are achieved at the feature level. The generated multimodal joint feature tensor breaks down the information barriers between modalities, forming a unified, complete, and context-related semantic representation of the teaching content. Mapping the joint features into a high-dimensional semantic feature cloud and depicting the semantic evolution path with a manifold trajectory can intuitively and quantitatively show the logical context, core concept clusters, and dynamic evolution process of the teaching content. Defining the dynamic semantic aggregation region can automatically identify and focus on the core topic of the current explanation, enabling the system to have a human-like attention mechanism, thus ensuring the logical coherence and emphasis of the teaching process.
[0040] By constructing a high-order semantic closed manifold, the system can self-examine and evaluate the integrity and self-consistency of the semantic structure during a continuous interaction, calculate the topological stability index and semantic density gradient, and provide scientific indicators for quantitatively evaluating whether the interaction logic is compact, deviates from the core, or has logical jumps or contradictions. The multimodal instruction generation network is not only based on the fused semantic content, but more importantly, it incorporates topological stability and semantic density as constraints to perform spatial topological consistency calibration. This ensures that the generated virtual digital human-driven instructions (language, expressions, gestures) not only match the literal content, but also highly match their underlying deep logical structure, key emphasis, and emotional tone, thus producing highly realistic, professional, and expressive behavioral planning. Instruction parsing and audio / video rendering are performed at edge nodes close to the user, greatly reducing network transmission latency and cloud-based centralized rendering backhaul latency. This allows high-fidelity, synchronous virtual digital human audio / video streams to be presented to the user with extremely low latency, ensuring real-time interaction and immersion, making the user feel as if they are communicating with a real object that responds instantly.
[0041] In a preferred embodiment of the present invention, step 1 above may include:
[0042] Step 1.1: Receive the multimodal raw data stream sent by the teaching end. The multimodal raw data stream includes at least audio data packet sequences, video frame sequences, image data packets, and text data packets. Specifically, throughout the entire process of real-time financial multimedia interaction (including financial explanations, investment consultations, employee training, customer service, etc.), continuously receive the multimodal raw data stream transmitted in a continuous streaming format from the teaching end (financial lecturer end, investment advisor end, training lecturer end, etc.). This data stream fully covers the four core types of data generated during the financial interaction process: audio data packet sequences, i.e., a set of continuous data segments corresponding to the financial lecturer's explanation voice, investment consultation dialogue voice, and professional terminology explanation voice; video frame sequences, i.e., images of the financial lecturer's gestures and K-line charts. The system consists of a continuous collection of frames from the demonstration process, financial statement chart operations, and interactive scenarios; image data packets, which are image data corresponding to static financial materials such as financial PPT courseware, candlestick charts, financial statement charts, and risk control rule diagrams; and text data packets, which are text data such as financial knowledge point annotations entered by the instructor, customer inquiry messages, key financial statement data annotations, and training focus notes. All types of received data are continuously and temporarily stored in real time. Simultaneously, the integrity of each data unit is checked one by one. For any missing or damaged data units identified during the check (such as missing segments of financial statement data, broken frames of candlestick chart images, or interrupted audio segments), timely completion operations are performed to ensure that the original data entering subsequent processing stages remains continuous, complete, and undamaged.
[0043] Step 1.2: Identify the modality type of each data unit in the multimodal raw data stream, and attach a corresponding modality label and high-precision source timestamp to each data unit to obtain a set of labeled time-series data units. Specifically, this includes: for the multimodal raw data stream that has completed real-time caching, first, split the data units according to a preset independent data unit division standard. The division standard is based on the packet structure of data transmission, data length threshold, and natural separation characteristics of financial data (such as K-line chart time period separation, financial report chapter separation, and consultation dialogue round separation). The continuous data stream is decomposed into independent data units that can be processed individually. Then, according to the order of data reception, each of the split independent data units is analyzed and its modality is determined. For audio... For data units of different types, identification is achieved by recognizing their built-in sampling rate parameters, number of audio channels, and unique data header identifiers, with a focus on distinguishing financial explanations from environmental noise. For video data units, differentiation is achieved by detecting their frame structure features, frame rate parameters, and video encoding format identifiers, with a focus on identifying valid video frames containing lecturer postures and candlestick chart demonstrations. For image data units, classification is completed based on their static pixel matrix distribution features, image file format identifiers, and size parameters, with a focus on distinguishing different static financial materials such as candlestick charts, financial statement charts, and courseware images. For text data units, classification is determined by recognizing their character encoding format, text markers, and sentence delimiters, with a focus on filtering core text information such as financial terminology, financial statement data, and consultation requests.
[0044] After the modality attribution of each data unit is determined, a unique and exclusive modality identifier is attached to it. The identifier adopts the format of modality type prefix + unique sequence number (such as audio A-, video V-, image I-, text T-), ensuring that the identifiers of different modalities and different data units are not duplicated. At the same time, the identifier is directly bound to the basic information of the data unit (such as data source, financial theme). In this process, each data unit is labeled with high-precision source time information that accurately reflects its actual generation time. This time information is linked in real time to the generation behavior of the data unit (such as lecturer's explanation actions, K-line data updates, customer consultation input), avoiding multimodal misalignment caused by time deviation (such as the explanation voice and K-line demonstration being out of sync). Finally, all data units with completed modality identifiers and source time information labels are arranged in an orderly manner according to their actual generation sequence, forming a set of labeled time-series data units with clear structure, clear time sequence, and complete modality and time information.
[0045] Step 1.3: Based on the high-precision source timestamp and the preset reference clock, perform global clock synchronization on the tagged time-series data unit set. This is used to attach a unified synchronized time axis coordinate to all data units, generating a time axis-aligned multimodal data unit set. Specifically, this includes: pre-building a unified reference time system suitable for the entire process of real-time interactive financial multimedia. This system uses the cloud-based global clock as the core reference and achieves time coordination with the data acquisition equipment at the teaching end and subsequent processing nodes through a network time synchronization protocol, ensuring the uniformity and stability of the reference time. At the same time, it clarifies the unified format of the reference time, providing a unified standard for the time calibration of all data units and adapting to the high requirements of real-time performance and synchronization in financial interaction; based on the high-precision source time information carried by each data unit in the tagged time-series data unit set, the source time of each data unit is... Each data unit is compared with a unified benchmark time, and any time discrepancies (such as audio delays and image transmission lags) are corrected to complete global time calibration of all data units. After calibration, each data unit is assigned a unique and unified benchmark time axis coordinate, which precisely corresponds to the calibrated time of the data unit and is bound to the modal identifier and original data information of the data unit. Subsequently, all data units with additional benchmark time axis coordinates are strictly rearranged according to the order of the coordinates, and the overall time sequence after the arrangement is checked. The consistency of the time sequence of core financial data (such as K-line demonstrations and explanations, financial statement analysis and text annotations) is checked to ensure that there are no misalignments or confusions in the time sequence. Finally, a set of multimodal data units that are precisely aligned in the time dimension is formed, which solves the problem of inconsistent time delays in multi-source heterogeneous financial data.
[0046] Step 1.4: Based on the modal tags, decouple the time-axis aligned multimodal data unit set into independent audio data subsets, video data subsets, image data subsets, and text data subsets. Specifically, this includes: accurately reading the unique modal identifier carried by each data unit in the time-axis aligned multimodal data unit set; matching and classifying all data units in the set one by one according to the modal type prefix in the identifier; first, filtering out all data units with the same prefix, and then simultaneously performing cross-modal mixing verification; during verification, combining the inherent characteristics of the data unit itself, such as the sampling rate characteristics of audio, the frame structure characteristics of video, the financial material characteristics of images, and the financial terminology characteristics of text, to verify whether its modal identifier is consistent with the actual data characteristics. The process involves systematically identifying and eliminating erroneous data units whose identifier prefixes do not match the data features, as well as ambiguous data units whose modal attribution is unclear (such as blurry non-financial images or meaningless noise audio). After verification, qualified data units whose features and identifiers match perfectly are then integrated into independent datasets according to their baseline timeline coordinates. Through the above classification, verification, and integration operations, audio, video, image, and text datasets are ultimately separated into independent datasets with no data overlap and no redundant information. Each independent dataset retains the original baseline timeline coordinates of all data units within it, ensuring that the data units within each dataset maintain a continuous and regular time series relationship.
[0047] Step 1.5: Based on a unified synchronous timeline coordinate, dynamically allocate processing priorities and computing resource quotas to each data subset. Then, according to the input interface protocols of each dedicated generative model processing unit, distribute the corresponding data subsets in parallel to the corresponding speech processing model unit, video processing model unit, image processing model unit, and text processing model unit in the cloud. Specifically, this includes: combining the unified baseline timeline coordinates carried by each independent dataset, and based on the logical sequence of data generation and the importance of content during financial interactions (e.g., prioritizing core financial report data, key K-line nodes, and core customer consultation needs), dynamically determining the appropriate processing order for each independent dataset. Simultaneously, based on the hierarchical differences in the processing order, assign processing priorities and computing resource quotas to each data subset. The system matches the corresponding resource allocation rules, prioritizing the processing of datasets related to core financial knowledge points and customer consultation needs throughout the overall processing chain. The resource allocation rules are adjusted in real time as the financial interaction process progresses (e.g., from knowledge point explanation to case analysis, from basic consultation to in-depth deduction), always keeping pace with the processing needs of the current stage. Then, according to the input adaptation requirements of each dedicated generative model processing unit in the cloud, the audio data subset, video data subset, image dataset, and text data subset are formatted and protocol adapted respectively. The data structure, transmission logic, and information encoding method of each dataset are calibrated one by one, so that the overall form of each dataset fully matches the input standards of the corresponding processing unit.
[0048] The cloud-based generative model processing units are divided into four dedicated categories. Each category adopts a customized layered architecture adapted to the characteristics of the corresponding modality and has independent core processing functions, primarily catering to the highly specialized and logically demanding needs of the financial sector.
[0049] The speech processing model unit adopts a layered architecture of temporal feature perception and deep financial semantic analysis. The bottom layer is responsible for continuous feature regularization of a subset of audio data, integrating scattered audio information into a coherent temporal feature sequence. The middle layer performs speech content recognition and semantic decomposition in financial scenarios based on audio features, accurately capturing financial professional terms, explanation logic, tone and rhythm features in the lecture speech, and focusing on identifying core semantics such as risk control rules, investment logic, and financial statement interpretation. The upper layer performs financial-specific noise reduction and enhancement on the recognized semantic information, filtering out irrelevant environmental interference information, highlighting core financial semantics, and adapting to various financial speech scenarios such as financial lecturers' knowledge point lectures, customer consultation and Q&A, and professional terminology voice explanations. It can completely retain and extract key information related to financial interaction in the speech.
[0050] The video processing model unit adopts an architecture of parallel analysis of spatiotemporal features and modeling related to financial content: the bottom layer performs frame-level temporal and spatial feature decomposition on a subset of video data, separating the temporal evolution features and spatial image features of the video; the middle layer performs feature modeling on content such as K-line demonstration trajectory, financial statement charting process, financial lecturer's gestures, and interactive scene switching in the video, transforming the visual images into spatiotemporal feature information bound to the financial interaction process; the top layer performs association annotation of financial knowledge points on the modeled features, so that the video features correspond to the financial explanation and consultation content, adapting to video processing scenarios such as real-time K-line demonstration, financial statement charting, and financial lecturer teaching, and can accurately capture the core spatiotemporal information supporting financial expression in the video.
[0051] The image processing model unit adopts an architecture of static visual feature layered extraction and financial visual semantic parsing: the bottom layer decomposes the visual elements of the image data subset step by step, separating the basic visual features and core visual elements of the image; the middle layer performs visual semantic recognition on content such as financial courseware images, candlestick charts, financial statement charts, and risk control rule diagrams, extracting key teaching visual information such as financial data, graphic symbols, and professional annotations from the images; the top layer structures and organizes the recognized visual information, integrating scattered visual elements into complete financial visual semantic units (such as candlestick trend units and financial statement data units), adapting to scenarios such as financial courseware display, static financial material analysis, and chart knowledge point presentation, and can efficiently identify and organize all financial-related information carried in the image.
[0052] The text processing model unit adopts a financial semantic-oriented sequence encoding and knowledge logic parsing architecture: the bottom layer performs character normalization and sequence structuring on a subset of text data, organizing scattered text information into an ordered text sequence; the middle layer performs semantic decomposition and logical association analysis on content such as financial knowledge point text, lecture annotations, customer consultation messages, and financial statement summaries, clarifying the hierarchical relationship of financial knowledge, investment logic, and key consultation needs in the text; the top layer refines and classifies the text semantics to form a systematic financial knowledge semantic unit (such as risk control knowledge point unit and investment strategy unit), which is suitable for scenarios such as financial lecturer text knowledge point input, customer text consultation interaction, and financial statement data annotation, and can deeply analyze the financial knowledge system and logical context contained in the text.
[0053] All four types of dedicated generative model processing units adopt a modular and decoupled design, with each unit operating independently and without interference. Furthermore, their architecture has been optimized for the specific characteristics of financial scenarios, abandoning the generalized processing logic of universal models and focusing on the professional analysis of financial data. The advantages of using these dedicated processing units are: they can achieve refined and customized processing for financial data of different modalities, avoiding the loss of key financial information (such as financial statement data and risk control rules) and semantic deviations caused by general processing models; the modular design allows for more targeted resource allocation, effectively improving overall processing efficiency and meeting the low-latency requirements of real-time financial interaction; and they enhance the extraction and retention of core financial information, providing accurate and complete basic data for subsequent feature analysis and cross-modal fusion processes.
[0054] After completing the format adaptation of all datasets, the matching and verification with the corresponding dedicated processing units, and the final confirmation of resource allocation, each adapted independent dataset is synchronously and in parallel distributed to the corresponding speech processing model unit, video processing model unit, image processing model unit, and text processing model unit in the cloud according to a unified time reference. This ensures that each processing unit receives the dataset of the corresponding modality synchronously under the same time reference, making full use of the parallel processing capabilities of cloud computing to maximize processing efficiency and initially reduce end-to-end latency.
[0055] In a preferred embodiment of the present invention, step 2 above may include:
[0056] Step 2.1 involves receiving and caching the distributed data subsets, including audio, video, image, and text data subsets, as input data to be parsed by the corresponding processing units. Specifically, this includes receiving the corresponding data subsets distributed in parallel. The speech processing model unit receives the audio data subset, the video processing model unit receives the video data subset, the image processing model unit receives the image data subset, and the text processing model unit receives the text data subset. Each dedicated generative model processing unit stores the received data subset in its own temporary cache. During the caching process, the original baseline time axis coordinates and modal identifiers of each data unit are synchronously associated to ensure that the temporal correlation of financial data is not lost, especially ensuring the time synchronization of K-line demonstrations, financial report analysis, and explanatory audio and text annotations, laying the foundation for subsequent cross-modal semantic alignment. Simultaneously, the cached data subset is checked for integrity and temporal continuity, with a focus on verifying whether there are any issues such as missing core financial data (e.g., missing key data segments from financial reports or key frames of candlestick charts), disordered timing (e.g., the audio explanation is out of sync with the candlestick chart demonstration), or cross-modal mixing (e.g., irrelevant images or background noise). After the verification is passed, the cached data subset is determined as the input data to be parsed for each unit, ensuring the accuracy and reliability of subsequent parsing.
[0057] Step 2.2 involves implementing parallel feature encoding processes adapted to the characteristics of each modality for the received audio data subset, video data subset, image data subset, and text data subset, respectively, to obtain the original feature encoding results. Specifically, this includes: initiating parallel feature encoding processes adapted to the inherent characteristics of each modality of financial data, with the four encoding processes proceeding synchronously and without interference. During the encoding process, the reference time axis coordinates of each data unit are always associated to ensure that the temporal correlation between the features and the original financial data is not interrupted, and to ensure that the encoding efficiency meets the low latency requirements of real-time financial interaction (such as investment consultation and real-time explanation).
[0058] To address the temporal continuity of the audio data subset, its baseline time axis coordinates are first aligned, and then feature encoding is performed segment by segment according to continuous temporal sequence. The focus is on extracting raw temporal features including the pronunciation of financial terminology, tone fluctuations (such as changes in tone when emphasizing risk control points and investment risks), and the rhythm intervals of explanations, forming the raw audio feature encoding results and accurately capturing the financial semantic relationships in the speech. For the spatiotemporal dual attributes of the video data subset, the spatial dimension of intra-frame visual information (such as candlestick patterns, details of financial statement charts, and lecturer gestures) and the temporal dimension of inter-frame evolution information (such as changes in candlestick trends, steps in financial statement chart explanations, and transitions in lecturer actions) are first separated. Then, parallel encoding is performed on both types of information simultaneously, extracting spatial features such as intra-frame pixel distribution and contour structure, as well as temporal features such as inter-frame action changes and scene transitions. The system synthesizes the original spatiotemporal feature encoding results of the video, focusing on the core visual and temporal information supporting financial explanations and consultations. For the static visual attributes of the image data subset, encoding is carried out in a hierarchical manner from the whole to the part. First, basic visual features such as the overall layout and color distribution of the image are extracted, and then detailed features such as K-line indicators, financial report data, and financial symbols are extracted in depth to generate the original visual feature encoding results of the image, highlighting the core information of the static financial materials. For the sequence association attributes of the text data subset, sequence encoding is carried out according to the order of the sentences, extracting features such as the combination relationship of financial terms, sentence logic (such as investment logic and risk control rule expression), and paragraph association (such as the relationship between financial report data and interpretation conclusions) to form the original sequence feature encoding results of the text, accurately capturing the financial knowledge context and consultation needs in the text.
[0059] Step 2.3 involves performing deep semantic analysis on the original feature encoding results based on the teaching scenario context to extract and enhance key semantic information strongly related to the teaching objectives, generating scenario-enhanced feature states corresponding to each modality. Specifically, this includes: conducting deep semantic mining on various original feature encoding results based on financial scenario contexts (such as investment consulting scenarios, financial statement interpretation scenarios, and risk control training scenarios), focusing on strengthening core financial semantics, filtering out invalid redundancy, and addressing the problem of existing technologies having a one-sided understanding of complex financial contexts. For the original audio feature encoding results, the audio content is correlated with the preceding and following time sequences to analyze the financial knowledge points expressed (such as risk control rules and investment strategies), the key content emphasized by tone (such as key financial statement data and investment risk warnings), and the logical relationship between the two parties in the consultation (such as the correlation between customer questions and lecturer responses), filtering out invalid features corresponding to environmental noise; for the original video spatiotemporal feature encoding results, combined with the financial interaction process context, the market trend corresponding to the K-line trend, the data analysis logic corresponding to the financial statement chart steps, the interpretation guidance conveyed by the lecturer's body language, and the interaction transition information corresponding to scene switching (such as from knowledge point explanation to case analysis), weakening background elements and irrelevant actions. Redundant features: Based on the original visual feature encoding results of the images, and referring to the financial knowledge system (such as the K-line analysis system and the financial statement interpretation framework), the core financial knowledge points of the courseware, the logical relationship between K-line charts and financial statement charts, and the supplementary explanations corresponding to the image annotations (such as K-line indicator annotations and financial statement data notes) are analyzed to remove decorative elements and invalid information corresponding to irrelevant backgrounds; Based on the original sequence feature encoding results of the text, and combined with the financial knowledge framework, the hierarchical relationship of financial knowledge (such as the association between basic terminology and complex strategies), the expression of key conclusions (such as investment advice and financial statement conclusions), and the evolution of interactive logic (such as the progression of customer consultation needs and the logical advancement of the lecturer's interpretation) are analyzed to filter out colloquial redundant expressions and features corresponding to irrelevant symbols.
[0060] During the analysis process, key semantic information that is strongly related to financial interaction goals, knowledge transfer, and consultation response is simultaneously enhanced, ultimately forming scenario semantic enhancement feature states that accurately correspond to each modality of data, so that each modality feature focuses on the core financial needs.
[0061] Step 2.4 involves structurally encapsulating and reducing the dimensionality of the scene semantic enhancement feature states to generate standardized feature vector sets, including speech semantic feature vectors, video spatiotemporal feature vectors, image visual feature vectors, and text deep semantic feature vectors. Specifically, this includes: performing structured encapsulation on various scene semantic enhancement feature states; organizing feature data hierarchically according to feature information, reference time axis coordinates, and modal identifiers; associating and binding each set of enhanced features with the precise time coordinates of the corresponding data unit (such as the time point corresponding to K-line data, or the time sequence node of financial report interpretation) and the unique modal identifier one by one; and simultaneously establishing a bidirectional traceability index between features and original financial data to ensure traceable temporal correlation and distinguishable modal attributes, ultimately forming a well-structured feature set with complete associated information, providing a basis for subsequent semantic manifold trajectory construction. Support is provided; then targeted dimensionality reduction is carried out: First, based on the semantic priority of the financial scenario, the importance weight of features is set, and core feature dimensions that are strongly related to financial knowledge transmission, consultation response, and risk warning are selected, such as the expression features of financial professional terms in speech, the features of K-line trends and financial statement charts in video, the features of K-line indicators and financial statement data in images, and the features of investment logic and consultation needs in text, etc. Redundant and repetitive feature dimensions that are irrelevant to the financial interaction goal are eliminated, such as the features of environmental noise in audio and the features of decorative background in images, etc.; then, based on the concentrated distribution characteristics of financial-related features, the compression scale and processing method are adjusted, so as to retain the original distribution characteristics of core financial semantics while controlling the dimensions, and not destroy the inherent connection of key information (such as the connection between financial statement data and interpretation conclusions, and the connection between K-line indicators and market trends).
[0062] Finally, the dimensionality-reduced structured feature set is standardized into a unified data type, a fixed numerical range, and a consistent dimensionality, generating four types of standardized feature vectors: speech semantic feature vectors, carrying the knowledge points, tone logic, and interactive semantics of financial explanations and consultations, associated with corresponding timeline coordinates; video spatiotemporal feature vectors, integrating intra-frame visual information and inter-frame temporal evolution features, mapping financial meanings such as K-line trends, financial statement charts, and lecturer gestures; image visual feature vectors, condensing the core content of financial courseware, K-line charts, and financial statement chart logic in images; and text deep semantic feature vectors, extracting the financial knowledge hierarchy, key conclusions (such as investment advice and financial statement summaries), and interactive logic framework (such as consultation needs and response context) of the text. These four types of vectors together constitute a standardized feature vector group. The data format of the vector group is simultaneously verified to ensure compliance, dimensional uniformity, and compatibility with the input interface of the subsequent cross-modal dynamic alignment fusion network, ensuring that the verified data can directly support the smooth progress of the subsequent fusion process.
[0063] In a preferred embodiment of the present invention, step 3 above may include:
[0064] Step 3.1 involves synchronously reading standardized feature vector sets from the caches of each dedicated generative model processing unit. These include speech semantic feature vectors, video spatiotemporal feature vectors, image visual feature vectors, and text deep semantic feature vectors. Specifically, this includes initiating a synchronous feature vector reading process triggered by a set time synchronization signal. The corresponding standardized feature vectors—speech semantic feature vectors, video spatiotemporal feature vectors, image visual feature vectors, and text deep semantic feature vectors—are precisely extracted from the dedicated cache areas of the four types of dedicated generative model processing units. During the reading process, a unified time rhythm is strictly followed to ensure that the reading of the four types of feature vectors is completed synchronously. Simultaneously, the synchronous time axis coordinates and modal identifiers carried by each type of feature vector are associated one by one, establishing a traceability link between the feature vectors and the original financial data (such as candlestick data, financial statement materials, and consultation audio), providing support for subsequent semantic tracing and anomaly investigation. After reading, a multi-dimensional integrity and validity check is performed, focusing on verifying whether there are missing vectors, whether the dimensions meet the preset standards, whether the time axis coordinates are continuous (ensuring the consistency of the timing of K-line demonstrations, financial report interpretations, and audio / text), and whether the modal identifiers are accurate. Abnormal vectors found during the check (such as vectors missing core financial semantics or vectors with disordered timing) are marked and removed to ensure that the final feature vector set is complete and valid, providing high-quality raw materials for subsequent cross-modal alignment and fusion.
[0065] Step 3.2: Based on the synchronous time axis coordinates carried by each feature vector, perform fine-grained time axis alignment and resampling on multiple sets of feature vectors to generate a time-synchronized multimodal feature vector sequence. Specifically, this includes: using the synchronous time axis coordinates carried by each feature vector as a unified benchmark, and combining the real-time requirements of financial interactions (such as real-time investment consultation, financial report interpretation, and real-time K-line analysis), setting an appropriate fine-grained time division window, and decomposing the overall time axis into continuous and equal small time units. The duration of each time unit precisely matches the minimum change cycle of financial data (such as minute-level K-line updates and the time interval of financial report data interpretation). Subsequently, the four types of feature vectors—voice, video, image, and text—are precisely mapped to their corresponding time units according to their respective time axis coordinates, completing the initial time dimension alignment and ensuring that the multimodal features within the same time unit correspond to the same financial interaction node (such as a K-line inflection point or a moment of financial report data interpretation).
[0066] To address the issue of varying sampling densities across different modal feature vectors over time, targeted resampling is implemented: For feature vectors with sampling densities lower than the time unit requirements (such as text annotation features), reasonable interpolation is performed based on the financial semantic correlation between adjacent features (such as the interpretation logic of adjacent candlestick charts and the correlation of financial report data) to ensure the continuity of feature information; for feature vectors with excessively high sampling densities (such as video candlestick frame features), key feature points related to core financial semantics (such as candlestick inflection points and features corresponding to sudden changes in trading volume) are selected and retained, while redundant and repetitive information is removed. After resampling, the consistency of sampling densities and the completeness of feature information for each modality within each time unit are checked, ultimately integrating them to form a multimodal feature vector sequence that is temporally continuous, has a unified time base, and a balanced sampling density.
[0067] Step 3.3 involves inputting the time-synchronized multimodal feature vector sequence into the multi-head cross-attention alignment layer of the cross-modal dynamic alignment and fusion network. In this layer, dynamic correlation weights between different modal feature vectors are calculated, and feature recalibration and spatial projection are performed accordingly to generate a preliminary aligned cross-modal feature map. Specifically, the cross-modal dynamic alignment and fusion network originates from the Transformer architecture and is improved to address the temporal correlation (e.g., K-line time-series evolution, financial statement interpretation logic) and semantic specificity (e.g., financial terminology, investment logic, risk control rules) of multimodal data in financial scenarios. It retains the core multi-head attention mechanism of Transformer and adds a financial semantic guidance module and a dynamic temporal adaptation layer. This solves the problems of traditional Transformer in cross-modal fusion, such as neglecting financial scenario-specific semantic correlations and insufficient adaptation to dynamic temporal changes (e.g., real-time K-line fluctuations, evolving consultation needs). It also addresses the pain points of existing technologies, such as independent processing of multimodal data and insufficient semantic alignment. The specific construction and training process is as follows:
[0068] The network adopts a hierarchical architecture consisting of an input adaptation layer, a multi-head cross-attention alignment layer, a feature projection layer, and an output layer. The input adaptation layer receives speech semantic feature vectors, video spatiotemporal feature vectors, image visual feature vectors, and text deep semantic feature vectors. It unifies the features of different modalities into the same dimensional space through dimensionality normalization and embeds financial time-series identifiers (such as K-line timestamps and financial report interpretation nodes) to ensure the temporal correlation between features and the financial interaction process. The core multi-head cross-attention alignment layer is the core improvement. In addition to the regular attention heads, it adds three financial semantic-specific attention heads, which focus on the correlation of financial knowledge point descriptions, the correlation between K-line / financial report visual features and speech interpretation, and the correlation between financial courseware / charts and text annotations, respectively. This achieves accurate capture of the core semantic dimensions of finance. After parallel computation by each attention head, the correlated features are output through feature concatenation and fusion. The feature projection layer uses a 1×1 convolutional kernel to complete feature dimensionality reduction and spatial mapping, eliminating the differences in feature distribution between modalities. The output layer outputs the preliminary aligned cross-modal feature mapping, providing a foundation for subsequent synthesis layer processing. During the construction process, the parameters of each layer are initialized using the Xavier initialization strategy to ensure training stability.
[0069] The training data is selected from datasets covering multiple financial scenarios (such as stock candlestick chart interpretation, financial statement analysis, risk control training, and fund investment consulting). It includes multimodal data such as lecturer audio, candlestick chart demonstration videos, financial statement charts, and text annotations of financial knowledge points. Simultaneously, intermodal financial semantic association labels are annotated (e.g., association labels between candlestick patterns and audio interpretation, and association labels between financial statement data and textual conclusions). First, the training data is preprocessed, generating standardized feature vector sets according to steps 1 and 2, and dividing the dataset into training, validation, and test sets. Then, the training set feature vectors are input into the constructed network. The core loss functions are cross-modal feature alignment accuracy and financial semantic association recognition accuracy. A weighted sum of cross-entropy loss and mean squared error loss is used, and the Adam optimizer is employed for iterative training with an initial learning rate of 1×10⁻⁶. -4 The model is adjusted exponentially every 10 epochs. An early stopping mechanism is introduced during training: training stops when the semantic association recognition accuracy on the validation set does not improve for 5 consecutive epochs to avoid overfitting. Finally, the generalization ability is verified by testing the test set (selecting typical financial scenario data such as real-time interpretation of stock K-lines, analysis of listed company financial reports, and explanation of risk control rules). The network parameters are fine-tuned for the feature differences of different financial scenarios to ensure that the model can achieve accurate cross-modal alignment in various financial scenarios.
[0070] The time-synchronized multimodal feature vector sequence is input into the multi-head cross-attention alignment layer of the cross-modal dynamic alignment and fusion network constructed and trained above. Each attention head (including 3 financial semantic-specific attention heads) simultaneously conducts multi-dimensional correlation analysis on the feature vectors of different modalities: the conventional attention head focuses on temporal correlation and feature distribution correlation, while the financial semantic-specific attention head focuses on the core needs of financial scenarios, accurately analyzing the semantic correlation between voice financial knowledge point descriptions and video K-line / financial report presentations, the logical correlation between image financial charts and text annotations, and other financial-specific semantic dimensions. Based on the analysis results, the dynamic correlation weights between various modal features are calculated. The weight values directly reflect the degree of correlation between different modal features when conveying the same financial semantics (such as explaining the K-line inflection point of a stock or a core indicator of a financial report). For example, when explaining the K-line inflection point of a stock, the feature correlation weight between the voice semantic features (inflection point interpretation) and the K-line presentation in the video will be significantly improved.
[0071] Based on dynamic correlation weights, feature recalibration is performed on the feature vectors of each modality, adjusting the weight ratio of feature components to strengthen features strongly correlated with core financial semantics, such as financial terminology in speech, K-line trends / financial statement charts in videos, K-line indicators / financial statement data in images, and investment logic in text, while weakening irrelevant and redundant features, such as environmental noise in speech, background clutter in videos, and decorative elements in images. Simultaneously, through spatial projection operations of the feature projection layer, feature vectors of different modalities are mapped from their respective original feature spaces to a unified cross-modal feature space, completely eliminating feature distribution differences between modalities and enabling each modality's features to have a unified basis for comparability and fusion. Finally, a preliminary aligned cross-modal feature mapping is generated, in which each modality's features achieve accurate matching at both the temporal and financial semantic levels, solving the problem of existing technologies' one-sided understanding of complex financial contexts.
[0072] This network, leveraging the advantages of an improved Transformer architecture and customized construction and training for financial scenarios, offers the following advantages compared to traditional cross-modal fusion models (such as simple feature splicing and fixed-weight fusion models): First, it dynamically adapts to the temporal changes of financial scenarios, accurately capturing the dynamic correlation of modal features during financial interactions (such as the real-time semantic correlation when a lecturer interprets a candlestick chart while demonstrating), thus meeting the low-latency requirements of real-time financial interactions. Second, it strengthens the identification of core financial semantic associations by using a dedicated attention head design to avoid omissions or misjudgments of specific financial scenario semantics by general cross-modal models, thereby improving the semantic accuracy of multimodal fusion. Third, it has strong generalization capabilities; after training and fine-tuning with data from various financial scenarios, it can adapt to the fusion needs of various financial scenarios such as stocks, funds, financial reports, and risk control, without the need for repeated modeling for a single scenario. Fourth, it provides high-quality fusion features for subsequent deep understanding of financial semantics and intelligent interactive responses, ensuring that subsequent steps can quickly and accurately extract core financial information and support the realization of interactive functions such as virtual digital human financial consultation and knowledge point explanation.
[0073] Step 3.4 involves inputting the initially aligned cross-modal feature maps into the feature synthesis layer of the cross-modal dynamic alignment fusion network. Through cascading and nonlinear transformation operations, the multimodal feature maps are aggregated and compressed into a unified, high-dimensional multimodal joint feature tensor. Specifically, this includes inputting the initially aligned cross-modal feature maps into the feature synthesis layer of the cross-modal dynamic alignment fusion network, first performing an ordered cascading operation, and following the time-unit order of the synchronous time axis coordinates, combined with the cross-modal financial semantic correlation tightness (such as the correlation between K-line features and speech interpretation, and the correlation between financial report data and text annotations), concatenating and integrating the feature maps of different modalities along the feature dimension to form a temporally continuous, cross-modal information interconnected full-modal integrated feature set, ensuring that the correlation of the core financial semantics is not lost.
[0074] Subsequently, a nonlinear transformation adapted to the financial scenario is performed on the set. The response intensity of the feature components is adjusted in a targeted manner through dedicated activation processing to enhance the consistency of features expressing the same financial semantics in different modalities (such as the consistency between K-line trends and trend judgments in speech, and between financial report data and text conclusions). This resolves semantic expression conflicts between modalities, thereby optimizing the semantic coupling effect between modalities and further improving the feature's ability to accurately represent core financial semantics such as financial knowledge point descriptions, K-line / financial report presentations, and consultation responses. Next, targeted aggregation and lightweight dimensional compression are carried out. First, based on the financial semantic priorities preset in conjunction with the needs of financial knowledge transmission and consultation response (such as core financial data, key K-line indicators, and investment risk warnings), the core semantics of each modality and cross-modal correlation features in the integrated features are weighted, condensed, and aggregated to focus on retaining high-value financial information. Then, the feature dimensions are precisely screened, and redundant dimensions that repeatedly represent the same financial information or have low semantic value are eliminated through lightweight compression. The compression process follows the principle of preserving the core and controlling the scale, and while reasonably controlling the overall size of the feature tensor, its high-dimensional representation ability is fully preserved, ensuring that the inherent correlation of the core financial semantics is not destroyed (such as the correlation between financial data and interpretation conclusions, and between K-line indicators and market trends).
[0075] After compression, the semantic integrity and dimensional compliance of the features are checked simultaneously to confirm that no core financial information is lost and that the dimensions meet the preset standards. Finally, a unified, high-dimensional multimodal joint feature tensor is generated. This tensor fully integrates the core financial semantics and cross-modal correlation information of four modalities: voice, video, image, and text. It breaks down the information barriers between modalities and forms a unified, complete, and context-related semantic representation of financial interaction content. Its data format, dimensional scale, and feature distribution are highly compatible with the input requirements of subsequent stages such as deep understanding of financial semantics and generation of virtual digital human-driven instructions.
[0076] In a preferred embodiment of the present invention, step 4 above may include:
[0077] Step 4.1: Input the multimodal joint feature tensor into a high-dimensional space projection layer. Through nonlinear dimensionality reduction mapping, each fused feature unit in the multimodal joint feature tensor is represented as a coordinate point in a high-dimensional semantic space, thereby generating a high-dimensional semantic feature cloud. Specifically, this includes: inputting the multimodal joint feature tensor, which integrates the core financial semantics of four modalities (speech, video, image, and text), into a high-dimensional space projection layer designed specifically for financial scenarios. This projection layer is based on a multilayer perceptron architecture in deep learning and has been deeply optimized for the high-dimensional representation characteristics and cross-modal correlation characteristics of financial semantics. Specifically, the projection layer adopts a three-level serial architecture consisting of a feature preprocessing sublayer, a nonlinear mapping sublayer, and a semantic calibration sublayer. The feature preprocessing sublayer is responsible for normalizing the input multimodal joint feature tensor, standardizing the feature values to a uniform range, and embedding financial time series identifiers and modal association labels to provide semantic and time series anchors (such as K-line timestamps, financial data, etc.) for subsequent mapping. The parameter initialization of this sub-layer (the anchor point for interpreting nodes) adopts the He initialization strategy adapted to the distribution of financial features to ensure the stability of the processing. The core nonlinear mapping sub-layer consists of three fully connected network layers. Each layer introduces an activation function adapted to financial semantics. It achieves accurate mapping of high-dimensional features through piecewise nonlinear transformation. At the same time, each layer sets a Dropout layer to suppress overfitting. The Dropout probability is dynamically adjusted according to the size of the financial dataset. For example, for a financial sample dataset of 100,000 levels (including K-line, financial report, consultation dialogue and other multimodal data), the probability is set to 0.2. This sub-layer learns the distribution pattern of financial semantics through previous training and can accurately capture the hierarchical relationship of financial knowledge points and the evolution of interaction links (such as from K-line recognition to trend judgment, from financial report data to conclusion derivation) and other core semantic features. The semantic calibration sub-layer calls the preset financial semantic rule library to fine-tune and calibrate the mapped features, correct the semantic deviation caused by nonlinear transformation, and ensure the semantic integrity of the features.
[0078] After the projection layer completes the architecture adaptation, a targeted nonlinear dimensionality reduction mapping process is initiated: First, the preprocessing sublayer normalizes and embeds anchor points into the multimodal joint feature tensor, eliminating numerical scale differences between different modal features and strengthening temporal and semantic association labels; then, the nonlinear mapping sublayer performs hierarchical nonlinear transformations on the feature tensor based on the financial semantic distribution patterns learned during training. While gradually compressing redundant and low-value dimensions, it fully preserves the financial knowledge point associations, cross-modal semantic coupling information, and temporal evolution logic between each fused feature unit through feature recombination (such as the association between K-line indicators and market trends, and the coupling between financial report data and voice interpretation), avoiding the loss of core financial information; then, the semantic calibration sublayer combines the financial semantic rule base (such as K-line MA) with the financial semantic rule base (such as K-line MA). The linkage rules between CD indicators and moving averages, the correlation logic between net profit and revenue in financial reports, and the hierarchical relationship of risk control rules are used to calibrate the mapped features unit by unit to correct local semantic deviations. Finally, each independent fused feature unit that has undergone three levels of processing is accurately mapped to a unique coordinate point in the high-dimensional semantic space based on its semantic attributes and temporal correlation labels. Feature units with high semantic similarity and temporal proximity (such as K-line upward signals under different modalities and multimodal representations of the same financial report indicator) will have their coordinate points clustered in the high-dimensional space, while those with low similarity will have their coordinate points dispersed. All coordinate points are naturally clustered according to financial semantic logic, ultimately generating a high-dimensional semantic feature cloud that can comprehensively and accurately reflect the core semantic distribution characteristics of finance and has clear temporal correlations.
[0079] Step 4.2: In the high-dimensional semantic feature cloud, calculate the semantic correlation degree and spatiotemporal proximity metric between any two feature unit coordinate points. Based on the metric values, construct edges representing the connection relationships between feature units to form a high-dimensional feature relationship graph. Specifically, this includes: in the generated high-dimensional semantic feature cloud, initiating a dual-dimensional correlation metric calculation process adapted to financial scenarios to ensure that the metric results accurately match the financial semantic correlation requirements, thus overcoming the one-sidedness of multimodal semantic correlation analysis in existing technologies; for any two feature unit coordinate points, first calculate the semantic correlation degree: extract key information such as the core financial knowledge points (e.g., K-line indicators, financial report data, risk control rules), financial interaction types (e.g., K-line interpretation, financial report analysis, investment consultation, risk warning), and financial semantic expression logic (e.g., investment logic, risk control derivation logic) corresponding to the two feature units, and obtain the matching degree of financial knowledge points (denoted as a) and the fit of financial interaction (denoted as b) through quantification. The three indicators are financial semantic logic consistency (denoted as c). The semantic relevance is calculated using a simple weighted summation method. The weighted formula is: Semantic Relevance = 0.4 × a + 0.3 × b + 0.3 × c, where the weight coefficient is set based on the financial scenario adaptation (this basic weight is common to scenarios such as interpretation of core financial knowledge points and routine consultation, and can be fine-tuned as needed). The higher the semantic fit, the closer the value is to 1. For example, two feature units that both represent the MACD golden cross of K-line have a semantic relevance close to 1. Then, the spatiotemporal proximity metric is calculated: based on the original synchronous time axis coordinates of the two feature units, the time difference is calculated. Combined with the minimum time granularity requirements of real-time financial interaction (such as minute-level updates of K-line and time intervals of financial report interpretation), the time difference is normalized and transformed into a spatiotemporal proximity parameter in the range of 0 to 1. The closer the time, the closer the parameter value is to 1. For example, the spatiotemporal proximity parameter of K-line demonstration and voice interpretation feature units at the same time point is close to 1.
[0080] When combining the measurement results from both dimensions, a dynamic weighting strategy is adopted: the weight ratio is adjusted according to the current financial interaction stage. For example, when interpreting core financial indicators (such as net profit in financial reports and key inflection points in K-lines), the weight of semantic relevance is increased, and when demonstrating K-line trends and financial report charts in real time, the weight of spatiotemporal proximity is increased to obtain the final relevance strength. A threshold for relevance strength based on financial data statistics is set, such as by analyzing the minimum strength of effective semantic relevance in historical financial interaction data. Only feature unit pairs with relevance strength exceeding the threshold are connected by edges, and the edges have a clear temporal direction, that is, from the feature unit in earlier time to the unit in later time (such as K-line pattern features first, followed by the corresponding voice interpretation features). The weight of the edge is determined by the final relevance strength after min-max normalization. All feature units and the connecting edges with temporal direction and weight together form a high-dimensional feature relationship graph with a complete structure and clear semantic and temporal relevance, clearly presenting the evolutionary relevance of financial semantics and the spatiotemporal correspondence.
[0081] Step 4.3: In the high-dimensional feature graph, multiple core connection paths are searched and extracted based on edge weights. Each path consists of a series of feature units and connecting edges in sequence. Finally, each core connection path is defined and smoothed into a semantic manifold trajectory. Specifically, this includes: in the constructed high-dimensional feature graph, a three-order path extraction strategy of core anchor point localization, hierarchical progressive search, and path selection optimization is adopted to ensure that the extracted paths accurately reflect the evolution of core financial semantics and solve the problem that existing technologies cannot dynamically construct a structured financial semantic system. First, core anchor point localization is performed: based on the distribution density of the high-dimensional semantic feature cloud, the feature units with the highest semantic density are selected as core anchor points. These units usually correspond to the multimodal fusion representation of core financial knowledge points (such as the multimodal fusion features of key K-line inflection points, core financial statement indicators, and core risk control rules). The process begins with a path search starting from the core anchor point. A hierarchical, progressive search is then conducted: edges are divided into three levels based on their weight (high weight, medium weight, and low weight). During the search, high-weight edges are prioritized for extension (high-weight edges correspond to feature units closely related to financial semantics, such as the connection between candlestick patterns and voice interpretation, or financial report data and textual conclusions). When the extension of a high-weight edge is blocked (no subsequent edge meets the criteria), medium-weight edges are then attempted, with low-weight edges serving only as a supplement. A loop detection mechanism is implemented during the extension process to determine in real time whether there are duplicate nodes in the current path, avoiding the formation of invalid loops. A path length threshold is also set (corresponding to the reasonable duration of a single segment of financial interaction content, such as the duration of a candlestick interpretation or a financial report indicator analysis). Extension is terminated if the threshold is exceeded. Finally, multiple core connection paths covering the main evolutionary context of financial semantics are extracted from the graph.
[0082] Each extracted core connection path is further filtered and optimized to determine the semantic correlation between feature units on the path and the path's rationality, i.e.:
[0083] For each feature unit on the core connection path, since each feature unit is arranged in an orderly manner on the path according to the evolution of financial semantics, and each feature unit corresponds to a unique coordinate in the high-dimensional semantic space, a temporal sliding selection method is adopted. Four adjacent feature units are selected in sequence according to the path extension direction as a group. During the selection process, the sliding continuity is maintained. That is, after completing the selection of a group, one feature unit is moved forward to form a new group, ensuring that all continuous financial semantic segments on the path (such as feature units corresponding to continuous interactions such as K-line interpretation, financial report data derivation, and risk control rule explanation) are covered, avoiding omission of key semantic segments or selection gaps. After selection, the coordinates of the four feature units in the high-dimensional semantic space are used as the basis for constructing the triangular pyramid structure. The construction rules are as follows: the coordinates of the feature unit with the earliest time sequence among the four feature units are used as the vertex of the triangular pyramid. This vertex corresponds to a core financial semantic node (such as K-line inflection point marker, core node of financial report data, or starting node of risk control rules). The coordinates of the remaining three temporally adjacent feature units are used as the three vertices of the base triangle. These three vertices correspond to the subsequent multimodal supplementary features of the core semantic node (such as the coordinates of feature units corresponding to voice interpretation, image demonstration, and text annotation). Then, by connecting the vertex to the three vertices of the base triangle, the three lateral edges of the triangular pyramid are formed. At the same time, by connecting the three vertices of the base triangle, a complete base is formed. Finally, a triangular pyramid structure corresponding one-to-one with the set of feature units is constructed in the high-dimensional semantic space. Each triangular pyramid structure accurately corresponds to a continuous financial semantic segment on the core connection path. The size of the triangular pyramid volume is then used to determine the degree of financial semantic association and path coherence among the four feature units. The smaller the volume, the higher the aggregation of the four feature units in the high-dimensional semantic space, the closer the corresponding financial semantic association, and the better the coherence of the path segments. They belong to the same core financial semantic segment (such as the multimodal representation of the same K-line indicator, the logical derivation of the same financial report data, or the interpretation segment of the same risk control rule). The feature units and their corresponding path segments are retained. The larger the volume, the more dispersed the distribution of the four feature units in the high-dimensional semantic space, the looser the corresponding financial semantic association, and the poorer the coherence of the path segments. They may contain redundant or irrelevant semantic information or path deviations. In this case, the path segments corresponding to the feature units are removed, or intermediate transitional feature units are added, the path direction is adjusted, and the semantic coherence and rationality of the path are optimized to ensure that each segment of the core connection path after screening focuses on the core financial semantics, is free from redundant interference, and the path direction conforms to the evolution logic of financial semantics.
[0084] After completing the screening and optimization based on the triangular pyramid volume algorithm, targeted smoothing optimization is performed on each core connection path: a moving average algorithm adapted to the financial time series granularity is adopted, and the sliding window size is set, such as the window duration matching the number of feature units corresponding to a 5-second financial interaction segment. The coordinate values of each feature unit in the path and the weight values of the edges are calculated by moving average to eliminate path deviations caused by local coordinate fluctuations and weight abrupt changes, and to correct the local irregularities of the path. After smoothing, each path is defined as a continuous and smooth semantic manifold trajectory. The direction of the trajectory directly corresponds to the evolution of financial semantics, such as the complete process from K-line pattern recognition and indicator analysis to market trend judgment, and the evolution from financial report data extraction and logical deduction to conclusions about the company's operating conditions. The curvature change of the trajectory reflects the intensity of semantic transitions. For example, a large curvature corresponds to the switching of financial interaction links (such as switching from K-line interpretation to risk warning), and a small curvature corresponds to the continuous explanation of the same financial knowledge point (such as continuously interpreting the calculation and analysis of a certain financial report indicator).
[0085] Step 4.4 analyzes the spatial distribution density of all semantic manifold trajectories in the high-dimensional semantic space and the topological structure of the intersection points between trajectories. Based on preset density and topological rules, a dynamic semantic aggregation region with continuous or discrete boundaries is delineated and output. Specifically, this includes: performing a progressive analysis of all semantic manifold trajectories in the high-dimensional semantic space through global scanning, local refinement, and rule matching to ensure that the dynamic semantic aggregation region accurately covers the core financial semantics, enabling the system to have a human-like attention mechanism and focus on the core content of financial interactions. First, a global density scan is performed: an adaptive sliding window is used, with the window size dynamically adjusted according to the trajectory distribution density of the current region. The window shrinks in dense trajectory areas and expands in sparse areas. The spatial distribution density of trajectories is statistically analyzed region by region to generate a density heatmap. The heatmap accurately identifies high-density areas with densely clustered trajectories (corresponding to core semantic ranges such as explanations of core financial knowledge points and important interactive links, such as K-line trend analysis and interpretation of core financial data) and low-density areas with sparse trajectories (corresponding to non-core semantic areas such as environmental interference and irrelevant interactions). At the same time, the boundary range, clustering intensity, and time series span of the high-density areas are recorded.
[0086] Subsequently, a detailed analysis of the topological structure was conducted: focusing on the intersection points of different semantic manifold trajectories within a high-density area, and analyzing the topological attributes of the intersection points, including the spatial coordinates of the intersection points, the number of connected trajectories, and the branching / aggregation patterns of the trajectories at the intersection points. Among them, intersection points with a large number of connected trajectories and obvious aggregation characteristics usually correspond to core financial knowledge points jointly represented by multiple modalities, such as the intersection of voice explanations, video demonstrations, and text annotations of a certain K-line turning point, or key nodes of financial interaction, such as the node of switching from indicator analysis to investment advice. These key intersection points were marked as key points.
[0087] Finally, dynamic semantic aggregation regions are defined based on preset rules: preset density thresholds are used, based on global statistics of historical data from multiple financial scenarios, combined with the semantic distribution characteristics of different financial scenarios such as stocks, funds, financial reports, and risk control, as well as the semantic aggregation characteristics of different interactive links such as indicator interpretation, case analysis, and consultation and Q&A, to differentiate the minimum density value of core semantic regions, specifically for accurately distinguishing core semantic regions from non-core semantic regions; preset topology matching rules are used: priority is given to retaining intersections that connect three or more semantic manifold trajectories and are verified by the financial semantic rule base as corresponding core financial knowledge points, while requiring such intersections to be mainly in the form of aggregation, eliminating weakly related branching intersections that are merely simple switching in financial interaction links, specifically for accurately screening effective intersection point types that can represent core financial semantics; firstly, high density thresholds are used to define the minimum density value of core semantic regions. The initial boundary of the dense region is used to delineate the area's scope. The boundary is then adjusted based on the locations of key marked intersections to ensure the region fully encompasses the core semantic cluster and key intersections. The final output dynamic semantic aggregation region exhibits a boundary shape that flexibly changes with the distribution of financial semantics: in scenarios where core financial knowledge points are continuously explained (e.g., continuously interpreting the K-line trend of a stock), the trajectory shows a continuous clustering state, and the region boundary is a continuous smooth curve; when multiple independent financial interaction links (e.g., K-line interpretation + financial statement analysis + investment consulting) are distributed at intervals, the trajectory shows a segmented clustering state, and the region boundary consists of multiple discrete closed curves. This region fully carries the core semantic information of financial interaction, providing a precise semantic focus range for subsequent in-depth understanding of financial semantics and virtual digital human interactive responses, ensuring that the virtual digital human's response focuses on the core and does not deviate from the key points of financial interaction.
[0088] In a preferred embodiment of the present invention, step 5 above may include:
[0089] Step 5.1: Based on the geometric center, density contour, and correlation strength with external feature units of the dynamic semantic aggregation region in the high-dimensional semantic space, determine multiple candidate point sets at the internal core, boundary, and external correlation positions of the dynamic semantic aggregation region. Specifically, this includes: firstly, conducting refined quantitative analysis of the core attributes of the dynamic semantic aggregation region in the high-dimensional semantic space: determining the geometric center by calculating the weighted average of the coordinates of all feature units within the region, with the weight being the semantic correlation strength of each unit, avoiding interference from low semantic value units at the edges (such as feature units corresponding to irrelevant environmental interference). This center is the concentrated representation area of core financial semantics (such as interpretation of core indicators, risk control rules, and key K-line trends); based on the previously generated density heatmap, and analyzing the density contour... A continuous gradient curve from high to low density is plotted to determine the precise boundaries of the high-density core area (i.e., the peak area of financial semantic aggregation, corresponding to core financial knowledge points and key interaction nodes), the medium-density transition area (i.e., the financial semantic connection area, corresponding to knowledge point transitions and process switching), and the low-density edge area (i.e., the financial semantic decay area, corresponding to non-core supplementary content). At the same time, the correlation strength between each feature unit on the boundary of the region and the external feature unit is calculated: the correlation strength is still calculated using the dynamic weighted formula of semantic correlation degree × 0.6 + spatiotemporal proximity × 0.4. Boundary units with correlation strength exceeding the preset threshold are selected. These units are usually the connection nodes between the aggregation area and the external core semantics (such as financial knowledge points that connect before and after, semantics across interaction processes, and related market conditions).
[0090] Based on the above analysis, candidate point sets are precisely determined in three categories: The internal core candidate point set is selected from a high-density core area within 5 units surrounding the geometric center, prioritizing feature units with high semantic correlation and verified as core financial knowledge points (such as key K-line inflection points, core financial statement indicators, and core risk control rules) through matching with the financial semantic rule base, ensuring the point set focuses on the core semantics of financial interaction; the boundary candidate point set is selected from the gradient boundary line of the density contour, selecting feature units using an equal-interval sampling method, focusing on retaining units that can represent the transition of financial semantics (such as the switching between K-line pattern recognition and trend judgment, from financial statement data extraction to conclusion derivation, and from indicator analysis to risk warning); the external related candidate point set is selected from boundary units with high correlation strength and their corresponding external related units, while eliminating points with low semantic value and irrelevant to the core logic of financial interaction (such as irrelevant market noise and units corresponding to non-core supplementary information), ensuring the point set can effectively connect the core semantic links inside and outside the aggregation area, supporting the logical coherence of financial interaction.
[0091] Step 5.2: Select the coordinate points with the highest semantic purity and discriminative power from each candidate point set, mark them as semantic sampling anchor points, and obtain the coordinate set of all semantic sampling anchor points in the high-dimensional semantic space. Specifically, this includes: conducting a dual quantitative evaluation of semantic purity and discriminative power for each coordinate point in the three categories of candidate point sets: internal core, boundary, and external relation. The semantic purity evaluation is achieved through two-step verification: the first step is to match the financial semantic rule base to determine whether the point accurately corresponds to a single core financial knowledge point or a clear financial interaction link, such as MACD golden cross interpretation, financial statement net profit analysis, and risk control rule derivation, and record the semantic orientation. The first step is to score the uniqueness of the point; the second step is to combine the financial knowledge framework (such as K-line analysis system, financial statement interpretation framework, risk control logic system) to verify whether the semantics of the point conforms to the logical progression of financial interaction, and eliminate fuzzy points with logical conflicts and multiple knowledge points mixed together (such as units that are associated with multiple unrelated financial indicators at the same time); the discrimination evaluation adopts the point-to-point comparison + cluster analysis method: first, calculate the semantic difference value between the point and the candidate points of the same type, the difference value is calculated in reverse based on the semantic relevance, and then use cluster analysis to determine whether the point can form an independent semantic cluster center to avoid semantic duplication with other points (such as units that repeatedly represent the same K-line pattern).
[0092] After the assessment, a three-step screening process is adopted: initial screening, fine screening, and review. Initial screening directly eliminates coordinate points whose semantic purity and discrimination scores do not reach the set thresholds (determined based on historical financial interaction data statistics). Fine screening, combined with financial interaction needs (such as investment consultation, indicator explanation, and risk control training), prioritizes the remaining points semantically, retaining coordinate points corresponding to key financial knowledge points and key interaction nodes (such as market data introduction, core indicator interpretation, risk warning, investment advice, and Q&A conclusion). Review involves manual annotation and sample calibration to correct biases during the assessment and ensure that the screening results match actual financial interaction needs. Finally, the coordinate points with the highest semantic purity and discrimination are marked as semantic sampling anchors, and the precise spatial coordinates and corresponding core semantic labels of each anchor are recorded simultaneously (such as knowledge point: MACD indicator application, stage: K-line trend judgment; knowledge point: financial statement net profit, stage: enterprise operating status analysis), integrating them to form a semantically clear and reasonably distributed set of semantic sampling anchor coordinates.
[0093] Step 5.3: Based on the temporal evolution of the multimodal joint feature tensor during the teaching interaction, determine the activation time sequence of each anchor point in the semantic sampling anchor point coordinate set, and connect the anchor points sequentially according to this temporal sequence to form a spatiotemporally closed trajectory, i.e., a higher-order semantic closed manifold. Specifically, this includes: first tracing the temporal evolution trajectory of the multimodal joint feature tensor in the complete financial interaction process. This trajectory is formed by the dynamic changes of the feature tensor as the financial interaction progresses, and corresponds one-to-one with the original synchronous time axis coordinates of each feature unit in the high-dimensional semantic space (e.g., real-time K-line updates, financial report interpretation progress, and consultation dialogue progression). Based on this correspondence, the original timestamp corresponding to each anchor point in the semantic sampling anchor point coordinate set is matched in reverse to confirm the activation time of each anchor point in the financial interaction process. Based on the activation time, the activation order of all semantic sampling anchor points is determined. This order strictly follows the natural process logic of financial interaction. For example, starting from the anchor point in the financial topic introduction stage (such as market overview), it will sequentially transition to the anchor point for explaining core knowledge points (such as K-line indicator interpretation, financial report data analysis), the anchor point for case analysis (such as individual stock trend deduction, corporate financial report case interpretation), the anchor point for risk warning, and the anchor point for summary and suggestions.
[0094] All semantic sampling anchors are connected in an orderly manner according to their activation order. During the connection process, a financial semantic logic verification mechanism is added: for every two adjacent anchors connected, it is verified whether the semantics of the two correspond to the progressive relationship of financial knowledge. For example, after the indicator definition anchor, it is necessary to connect to the indicator application anchor, rather than directly connecting to the case application anchor; after the financial report data extraction anchor, it is necessary to connect to the data interpretation anchor, rather than directly connecting to the investment advice anchor. If there is a logical deviation, the connection order is adjusted or intermediate transition anchors are added (such as adding indicator calculation anchors or data logic derivation anchors). After the connection is completed, the earliest activated anchor and the latest activated anchor are connected to form a spatiotemporal closed trajectory that combines temporal continuity and spatial correlation. This trajectory fully covers the core semantic process from the beginning to the end of financial interaction, which is a high-order semantic closed manifold, solving the pain point that existing technologies cannot dynamically construct a structured financial semantic system.
[0095] Step 5.4 involves performing topological analysis on the high-order semantic closed manifold at multiple preset scales to calculate structural durability parameters across different scales, which serve as topological stability indicators. Specifically, this includes: pre-setting three sets of differentiated analysis scales adapted to financial scenarios, and determining the temporal granularity and spatial scope of each scale: Fine-grained scales correspond to detailed explanations of a single core financial knowledge point (such as interpretation of a candlestick indicator or analysis of a financial report data), with a time granularity of 1 to 3 minutes (meeting the low-latency requirements of real-time financial interaction), and a spatial scope focusing on the semantic region corresponding to a single financial knowledge point; medium-grained scales correspond to interactive modules composed of multiple related financial knowledge points (such as candlestick combination indicator analysis or multi-data linkage interpretation of financial reports), with a time granularity of 10 to 15 minutes, and a spatial scope covering all core semantic regions within the module; coarse-grained scales correspond to complete financial interaction segments or explanations. (Such as comprehensive analysis of individual stocks, training on certain risk control rules), with a time granularity of 1 lesson or less (within 30 minutes), and a spatial scope covering the entire dynamic semantic aggregation area; at each scale, targeted topological structure analysis is conducted on high-order semantic closed manifolds: fine-grained scale focuses on analyzing the local node connection methods of the manifold (such as the aggregation connection of different modal features within a single financial indicator, the semantic association of indicator details interpretation), and the smoothness of local trajectories; medium-grained scale focuses on the trajectory bifurcation and aggregation of the manifold, such as the aggregation of multiple sub-indicators to the core indicator, the bifurcation of core knowledge points to different cases, and the bifurcation of indicator interpretation to risk warnings; coarse-grained scale analyzes the topological morphology of the overall core semantic region, such as the overall outline of the closed manifold, the distribution density of core financial semantic nodes, and records the core topological features at different scales, such as the number of nodes, the distribution of connection edge weights, and the topological complexity of the closed region.
[0096] By comparing the overlap and variation of core topological features across different scales, a structural durability parameter is calculated: the higher the overlap and the smaller the variation, the higher the parameter value, indicating greater stability of the topological structure during scale changes. This parameter serves as a topological stability indicator, and its core function is to assess the stability of the financial semantics carried by higher-order semantic closed manifolds. For example, a high parameter value indicates that the semantic representation of core financial knowledge points (such as the application of K-line MACD indicators and the interpretation of net profit in financial reports) is consistent across different time granularities, which can support subsequent stable financial interaction responses (such as coherent interpretation of virtual digital humans and deep logical deduction). A low parameter value indicates that there are fluctuations or deviations in semantic logic, which need to be calibrated in a timely manner to ensure the rigor of financial interactions.
[0097] Step 5.5: While performing topological structure analysis, calculate the rate of change of semantic feature density in the local space traversed by the higher-order semantic closed manifold, and fit the semantic density gradient. Specifically, this includes: while conducting topological structure analysis, deploying an adaptive sliding window along the trajectory direction of the higher-order semantic closed manifold. The initial size of the window is set based on the minimum semantic change cycle of the financial scenario, such as the number of feature units corresponding to a 1-minute financial interaction segment. Subsequently, it is dynamically adjusted according to the semantic density of the current trajectory segment. When the semantic density is detected to be higher than a set threshold (such as the interpretation of core financial indicators or risk warning nodes), the window size is reduced to half of the initial value to ensure accurate capture of local semantic details, such as the logical progression of indicator interpretation and risk warning nodes. Detailed explanation of the risks: When the semantic density is below the threshold (such as non-core supplementary information or transitional statements), the window size is increased to twice the initial value to avoid missing semantic connections across segments (such as the connection between different indicators or the connection between cases and knowledge points); the trajectory is scanned segment by segment through a sliding window to count the semantic feature density in each window. The density is the number of core semantic feature units (such as units corresponding to financial knowledge points or key interactive nodes) in the window divided by the window space range. Then, the semantic feature density values of each of the two adjacent sliding windows are obtained in turn. The density value of the previous window is used as the benchmark to calculate the relative change of the density value of the next window relative to the previous window, so as to obtain the semantic feature density change rate between the two adjacent windows.
[0098] A larger rate of change indicates a more significant difference in the financial semantics carried by the two trajectories, typically corresponding to a shift in financial semantics, such as a switch in financial knowledge points (from candlestick analysis to financial statement interpretation) or a change in interactive elements (from indicator interpretation to case analysis, or from case analysis to risk warnings). A smaller rate of change indicates continuous and stable semantics, corresponding to continuous explanations of the same financial knowledge point (such as continuously interpreting the application scenarios of a particular candlestick indicator or the analytical logic of a particular financial statement data). Based on the density change rate data of all windows, a piecewise smoothing fitting algorithm is used to generate a continuous semantic density gradient curve: linear fitting is used in semantically stable segments to ensure curve smoothness, while piecewise fitting is used in semantically transitional segments to retain density abrupt change points, avoiding over-smoothing that masks key semantic changes (such as semantic abrupt changes in risk warning nodes). The final generated semantic density gradient curve can accurately reflect the distribution pattern and intensity of change of financial semantics on a high-order semantic closed manifold.
[0099] In a preferred embodiment of the present invention, step 6 above may include:
[0100] Step 6.1 involves concatenating the multimodal joint feature tensor with the topological stability index and semantic density gradient to form a topologically semantically enhanced fusion feature vector. Specifically, this includes: first, performing dimensional adaptation and standardization preprocessing on the multimodal joint feature tensor, topological stability index, and semantic density gradient: transforming the single value of the topological stability index into a vector with the same dimension as the multimodal joint feature tensor through dimensional expansion; extracting key feature points from the semantic density gradient (i.e., curve features) through feature sampling and transforming it into a vector of the same dimension; and simultaneously normalizing all three to eliminate numerical scale differences. The concatenation is performed in the order of core features first, followed by enhanced features. Specifically, the multimodal joint feature tensor is concatenated first to carry core financial semantics and cross-modal fusion information; then, the topological stability index vector is concatenated to supplement semantic stability features, and the semantic density gradient vector is concatenated to supplement semantic change intensity features. This ultimately forms a topologically semantically enhanced fusion feature vector that combines financial semantic representation capabilities with topological structure perception capabilities. This vector provides more comprehensive and accurate feature support for subsequent virtual digital human-driven instruction generation, ensuring that instruction generation aligns with the core needs of financial interaction.
[0101] Step 6.2: Input the topologically enhanced fusion feature vector into the instruction decoder of the multimodal instruction generation network; the instruction decoder generates an initial sequence of virtual digital human multimodal driving parameters in parallel. The multimodal driving parameters include at least speech synthesis parameters, facial motion parameters, limb skeleton parameters, and whiteboard rendering parameters. Specifically, the multimodal instruction generation network is based on the classic Transformer encoder-decoder architecture and is customized and improved to meet the needs of virtual digital human multimodal driving in the financial scenario of this invention. The cross-modal dynamic alignment and fusion network mentioned above also originates from the Transformer architecture, focusing on the temporal alignment and semantic fusion of multimodal financial features. This multimodal instruction generation network focuses on the mapping and generation of financial topological semantic features to virtual digital human driving instructions. The two form a complementary technical link. The network is constructed in three core layers: a semantic feature encoding layer, a financial semantic adaptation layer, and a multi-branch parallel instruction decoding layer. The semantic feature encoding layer uses linear transformation and residual connection units to perform deep encoding on the input topological semantic enhancement fusion feature vector, fully preserving the correlation information between multimodal joint features, topological stability indicators, and semantic density gradients. The financial semantic adaptation layer has built-in prior mapping sub-modules for financial scenarios such as financial report interpretation, K-line analysis, risk control explanation, and investment consulting, and pre-stores the benchmark rules for driving parameters corresponding to different financial semantics (such as the benchmark parameters for risk warning semantics corresponding to serious tone and gesture warning). The multi-branch parallel instruction decoding layer sets up four independent decoding branches that share encoding features, corresponding to the generation tasks of speech synthesis parameters, facial motion parameters, limb skeleton parameters, and interface interaction parameters, respectively. Each branch is configured with dedicated parameter mapping neurons and temporal alignment units. The parameters of each layer of the network are initialized using a He normal distribution strategy to adapt to the distribution characteristics of financial features and ensure training stability.
[0102] The network was trained using a proprietary financial multimodal dataset, encompassing various financial interaction samples such as stocks, funds, financial reports, and risk control. Each sample was simultaneously labeled with standard driving parameter tags for the virtual digital human that matched the financial semantics (e.g., gesture parameters for K-line interpretation and voice parameters for financial report explanation). During training, the topological semantic enhancement fusion feature vector generated according to the previous steps of this invention was used as the network input, and the manually labeled standard driving parameter sequence was used as the supervision label. The AdamW optimizer was used for iterative training, employing a composite loss function combining mean squared error loss and cross-entropy loss to constrain the generation accuracy of continuous and discrete driving parameters, respectively. During training, validation samples from typical financial scenarios (e.g., real-time interpretation of individual stock K-lines and analysis of listed company financial reports) were used at fixed intervals to verify the effect. An early stopping mechanism was introduced to avoid model overfitting, and the network parameters were gradually fine-tuned to learn the accurate mapping relationship between financial topological semantics and virtual digital human driving parameters.
[0103] The multimodal instruction generation network described above offers several advantages that align with this invention: its improved Transformer-based architecture ensures the stability of encoding and generating long-term financial semantic features, meeting the needs of continuous explanation and logical deduction in financial interactions; its multi-branch parallel structure enables the synchronous and efficient generation of four types of driving parameters, satisfying the low-latency requirements of real-time financial interactions; and the prior knowledge embedding in the financial semantic adaptation layer prevents the driving parameters from becoming disconnected from financial semantics, generating driving parameter benchmarks that fit financial logic in scenarios such as K-line trend derivation, financial report data interpretation, and risk control rule explanation. Compared to general instruction generation networks, it is more adaptable to the professional needs of financial scenarios, and the generated parameter sequences are more closely aligned with actual financial interaction scenarios (such as consultation and Q&A, and indicator explanation).
[0104] The multimodal instruction generation network, after completion of construction and training convergence, is put into inference operation. The fused feature vector with topological semantic enhancement is input into the network's semantic feature encoding layer. After deep encoding and rule matching with the financial semantic adaptation layer, the features are transmitted to the multi-branch parallel instruction decoding layer. Four independent decoding branches simultaneously start inference generation: the speech synthesis parameter branch generates initial speech synthesis parameters including intonation (e.g., a serious tone for risk warnings, a gentle tone for routine explanations), speech rate, pauses, and emphasis (e.g., emphasis on financial terms and core indicators); the facial action parameter branch generates parameters including eye gaze (e.g.,...). The initial facial motion parameters include the gaze direction when pointing to the K-line / financial report interface, mouth movements, and facial muscle states (such as serious or calm). The limb skeleton parameter branch generates initial limb skeleton parameters including torso posture, arm swing, and hand trajectory (such as gestures pointing to K-line inflection points or financial report data). The interface interaction parameter branch generates initial interface interaction parameters including financial interface switching, K-line annotation, financial report data highlighting, and risk warning marking. After each branch is generated, they are integrated in an orderly manner according to a unified financial interaction time frame order to finally obtain the initial sequence of virtual digital human multimodal driving parameters.
[0105] Step 6.3: Based on the topological stability index and semantic density gradient, a spatial topological consistency calibration function is constructed. The calibration function is used to fine-tune the amplitude and temporal sequence of the multimodal driving parameters in the initial sequence frame by frame, generating a calibrated driving parameter sequence. Specifically, this includes: constructing a spatial topological consistency calibration function based on the topological stability index and semantic density gradient. The core is to accurately match the changes in driving parameters with the stability and intensity of changes in financial semantics. Its mathematical expression is: Pcal = Pinit × [α × T + β × (1-G)]; where the meaning of each character is briefly explained: Pcal represents the calibrated driving parameters, and Pinit represents the initial driving parameters. T represents the topological stability index (range 0 to 1), G represents the semantic density gradient (range 0 to 1), α and β are weighting coefficients, α+β=1, determined based on financial scenario statistics (e.g., α is 0.6 and β is 0.4 in the core indicator explanation scenario, and α is 0.4 and β is 0.6 in the interactive switching scenario); the core logic of this expression is: the higher the value of the topological stability index T, the more stable the semantics (e.g., continuous explanation of core financial indicators), the closer the adjustment coefficient is to α, and the smaller the change in the driving parameters; the higher the value of the semantic density gradient G, the more significant the semantic transition (e.g., switching from K-line analysis to risk warning), the closer the adjustment coefficient is to β, and the smoothness of the parameter time sequence connection can be appropriately optimized.
[0106] To more clearly illustrate the calculation logic, a numerical calculation example is given in a real financial scenario: Interpreting core financial indicators (semantic stability, weak transitions). Assuming the initial speech rate parameter Pinit = 140 words / minute, the current topological stability index T = 0.9 (high stability), and the semantic density gradient G = 0.2 (low transition), substituting into the formula, we calculate: Pcal = 140 × [0.6 × 0.9 + 0.4 × (1 - 0.2)] = 140 × (0.54 + 0.32) = 140 × 0.86 =120.4 words / minute; after calibration, the speaking speed slows down to match the rigorous and gentle rhythm of the core indicator explanation; for example, in the semantic transition scenario (switching from K-line analysis to risk warning), the initial gesture amplitude Pinit=8 (units), T=0.3 (low stability), G=0.8 (high transition), and the calculated Pcal=8×[0.4×0.3+0.6×(1-0.8)]=8×(0.12+0.12)=1.92. The gesture amplitude is slowed down and the connection is smoother, avoiding abrupt transitions.
[0107] The calibration function performs frame-by-frame fine-tuning of the multimodal driving parameters in the initial sequence: extracting the T and G values corresponding to the current parameters frame by frame, substituting them into the expression to calculate the calibrated parameter values, and simultaneously verifying the matching degree between the calibrated parameters of each frame and the current financial semantics (such as interpretation of core indicators, switching of interactive links, and risk warnings), correcting parameter amplitude deviations and temporal misalignments: for example, when explaining core financial indicators, a high T value reduces the variation in speech rate and body movements to maintain the stability of interface interaction (such as fixing the K-line interface); when semantic transitions occur (such as switching from indicator interpretation to risk warnings), a high G value strengthens the temporal coherence of facial movements, body gestures, and interface interactions in adjacent frames, avoiding stuttering or disconnection in driving actions; finally, a calibrated driving parameter sequence with consistent spatial topology and highly adapted to financial semantics is generated to ensure that the virtual digital human's driving actions conform to the financial interaction logic.
[0108] Step 6.4 involves packaging and timing the calibrated drive parameter sequence according to the set instruction set encapsulation protocol, ultimately outputting a virtual digital human drive instruction set calibrated for spatial topology consistency. Specifically, this includes: the virtual digital human adapted in this step is a dedicated image for financial interaction scenarios (such as a financial consultant or indicator explainer), possessing facial expression simulation, natural body language presentation, and financial interface interaction capabilities tailored to financial scenarios; and operating according to the set financial scenario-specific virtual digital human drive instruction set encapsulation protocol. This protocol is compatible with the standard interface of mainstream virtual digital human real-time drive systems and specifies core requirements such as parameter type identification specifications, millisecond-level timing stamp format, financial scenario-based transmission priority, and structured encapsulation templates.
[0109] First, the calibrated driving parameter sequences are categorized and organized into four types: speech synthesis, facial motion, limb skeleton, and interface interaction. A unique contextual type identifier is added to each type of parameter (e.g., TTS-FIN for speech parameters, UI-FIN for interface interaction parameters). Invalid and redundant parameter frames (such as stuttering or meaningless motion parameters) are removed from the sequence to ensure the validity of the parameter data. Then, for each frame of data for each type of parameter, a high-precision millisecond-level time stamp is added, which strictly corresponds to the original timeline of the financial interaction. This time stamp is matched one-to-one with the time nodes of the previous financial interaction process (e.g., K-line updates, financial report interpretation, and Q&A nodes). Then, frame synchronization verification is initiated based on the time stamp of the speech synthesis parameters. Micro-calibration is performed on the time sequence of each frame of facial, limb, and interface interaction parameters to correct micro-time sequence deviations within ±10 milliseconds. This ensures that multiple types of driving parameter frames at the same time node are accurately aligned without misalignment or delay, meeting the low-latency requirements of real-time financial interaction.
[0110] Next, based on the actual needs of financial interaction, a financial scenario-based transmission priority assignment is performed, setting a three-level transmission priority: the first priority is for interface interaction parameters and speech synthesis parameters, ensuring real-time synchronization between core financial interface operations (such as K-line annotation and financial report highlighting) and explanatory speech, avoiding disconnection of core interactions; the second priority is for facial motion parameters, ensuring accurate adaptation of expressions to financial semantics (such as seriousness and calmness); the third priority is for limb skeletal parameters, adapting to the parsing needs of non-core posture movements; the priority information will be embedded in the header identifier of each parameter frame, making it convenient for the drive system to prioritize the parsing of core parameters and improve the drive response speed.
[0111] Following the structured encapsulation template defined in the protocol, the four types of driving parameter sequences, after classification, timing calibration, and priority assignment, are integrated into a unified instruction data packet. This data packet contains three parts: header information, main parameter data, and a tail checksum. The header information records the protocol version, financial scenario identifiers (such as candlestick chart interpretation and financial statement analysis), and the total number of parameter frames. A CRC32 cyclic redundancy checksum is added to the tail to verify the integrity of the data packet. Finally, the instruction data packet undergoes a final verification of its integrity and accuracy, eliminating corrupted or incorrectly formatted packets. After successful verification, the final output is a virtual digital human driving instruction set that has undergone spatial topology consistency calibration, precise timing synchronization, and a complete and standardized structure. This instruction set can directly interface with the standard interface of the virtual digital human's real-time driving end without additional format conversion. It can drive the virtual digital human to complete multimodal interactive actions highly adapted to financial semantics (such as candlestick chart interpretation, financial statement analysis, risk warnings, and Q&A), overcoming the pain point of existing virtual digital human financial interactions being superficial and failing to meet professional needs.
[0112] In a preferred embodiment of the present invention, step 7 above may include:
[0113] Step 7.1: The edge node receives the virtual digital human driving instruction set and performs unpacking and parsing operations on the virtual digital human driving instruction set to obtain structured multimodal rendering instruction data. Specifically, the edge node in this step is deployed on the edge side of the financial interaction network, close to the interaction end and the financial consultation terminal (explanation end). Its core function is to reduce data transmission latency and core network load, meeting the low latency requirements of real-time financial interaction (such as investment consultation and indicator explanation), and has the ability to receive low-latency data, parse instructions, and schedule local rendering. The edge node starts a dedicated instruction receiving service, which is compatible with the virtual digital human driving instruction set encapsulation protocol output in step 6.4 and has the ability to receive low-latency data and perform preliminary verification. After receiving the instruction set, the unpacking operation is performed first: according to the protocol... The data packets are split into header, body, and tail structures. The header contains information such as the protocol version and financial scenario identifiers (e.g., K-line interpretation, financial report analysis). The tail CRC32 cyclic redundancy check code is verified to confirm that the data packets are not damaged or tampered with. If the verification fails, a retransmission mechanism is triggered. Then, the body parameter data is parsed: based on the exclusive type identifiers of the four types of driving parameters (e.g., voice parameter identifier TTS-FIN, interface interaction parameter identifier UI-FIN), the structured data of speech synthesis, facial movements, limb skeleton, and interface interaction are separated. The parameter format is converted into a standard format that the edge node rendering engine can directly recognize. At the same time, the millisecond-level time stamp of each frame of data is extracted and indexed. Finally, structured multimodal rendering instruction data containing complete driving information, clear timing, and standardized format is obtained.
[0114] Step 7.2 distributes the structured multimodal rendering instruction data to the parallel graphics and audio rendering engines on the edge nodes, driving these engines to generate audio and video data frames with a unified timestamp sequence. Specifically, the edge nodes call a preset parallel graphics and audio rendering engine. This engine adopts a dual-branch parallel architecture with a shared timing synchronization module, optimized for real-time interactive scenarios in the virtual digital human financial sector, and possesses low-latency and high-synchronization rendering capabilities. The two branches are an audio rendering branch and a graphics rendering branch. The shared timing synchronization module is responsible for uniformly managing the processing progress and timestamp generation of both branches, ensuring... The audio and video frames are synchronized. The audio rendering branch has a built-in speech synthesizer and audio frame cutter adapted for financial scenarios. The speech synthesizer optimizes the clarity and naturalness of the speech in financial explanations, and can accurately reproduce the effects required for financial explanations, such as speech speed, pauses, and emphasis on professional terms. The graphics rendering branch includes four core units: a virtual digital human 3D model driving unit, a facial expression texture generator, a limb skeleton posture solver, and a financial interface renderer. Each unit works together through a data bus to efficiently complete multi-dimensional graphics rendering. The financial interface renderer is specifically responsible for rendering interactive effects of financial interfaces such as K-line annotation, financial report data highlighting, and risk warning marking.
[0115] The parsed structured multimodal rendering instruction data is distributed by type: audio-related instructions (including parameters such as tone and speed) are sent to the audio rendering branch, and graphics-related instructions (including facial movements, limb skeletons, and interface interaction parameters) are sent to the graphics rendering branch. Upon receiving the instructions, the audio rendering branch uses a speech synthesizer to generate raw audio signals conforming to financial semantics, which are then cut into fixed-length audio data frames by an audio frame cutter according to a unified timestamp sequence. The graphics rendering branch works synchronously, with the 3D model driving unit coordinating and scheduling each subunit. Specifically, it uses a facial expression texture generator to adjust facial details (such as a serious expression during risk warnings), a limb skeleton pose solver to calculate joint poses (such as gestures pointing to K-line inflection points), and a financial interface renderer to generate financial interface interaction layers. Finally, the rendered data is merged and output as video data frames containing virtual digital human explanations and financial interface interaction effects. Throughout the process, a shared timing synchronization module assigns a unique timestamp to each frame of audio and video data, ensuring precise timing synchronization between the two branches and avoiding audio-visual discrepancies.
[0116] Step 7.3: Based on the unified timestamp sequence, perform time alignment and streaming encapsulation on audio and video data frames with the unified timestamp sequence to synthesize a continuous and synchronized virtual digital human audio-visual stream. Specifically, this includes: initiating a time alignment verification process, using the unified timestamp sequence as the core benchmark, and matching the timestamp information of audio and video data frames frame by frame; fine-tuning frame data with slight time deviations: correcting deviations through audio interpolation or delayed video frame output to ensure accurate alignment of audio and video frames corresponding to the same timestamp, avoiding audio-visual desynchronization issues (such as misalignment of voice and gestures when the virtual digital human explains K-lines); after alignment, using a streaming encapsulation format adapted to real-time financial interaction (such as HLS or RTMP format), encapsulating continuous audio and video data frames into streaming data blocks according to timestamp order, retaining inter-frame temporal correlation information during encapsulation, and adding stream status identifiers (such as playback start time and frame interval), finally synthesizing a continuous, smooth, and precisely synchronized virtual digital human audio-visual stream. This audio-visual stream fully presents the collaborative effect of the virtual digital human's financial explanation actions, voice, and financial interface interaction.
[0117] Step 7.4 involves performing low-latency encoding and network adaptive optimization on the continuously synchronized virtual digital human audio and video stream, and then pushing the optimized streaming media data to the interactive end to complete real-time interaction. Specifically, this includes: using low-latency encoding standards, such as H.265 video encoding and AAC audio encoding, to encode the virtual digital human audio and video stream. During the encoding process, the bitrate control strategy is optimized: while ensuring the clarity of financial explanation images (such as the virtual digital human's posture), audio clarity, and details of the financial interface (such as K-line trends and financial report data), the redundant bitrate is reduced and the encoding latency is shortened to meet the low-latency requirements of real-time financial interaction. Subsequently, network adaptive optimization is performed: edge nodes monitor the network bandwidth, latency, and packet loss rate between the edge nodes and the interactive end (such as user terminals, financial consultation screens, and tablets) in real time, and dynamically adjust the transmission bitrate and packet size of the streaming media based on the monitoring data.
[0118] When network bandwidth is insufficient, the resolution of non-core images (such as the background of the virtual digital human) is appropriately reduced to ensure the transmission quality of core financial content (such as voice explanations and K-line / financial report interfaces). When the network is stable, the bitrate is increased to ensure image quality and interface details. After optimization, streaming media data is pushed to the interactive end through the low-latency push service of edge nodes. At the same time, a real-time feedback channel is established to receive playback status feedback from the interactive end (such as stuttering or synchronization anomalies). The push parameters are dynamically adjusted according to the feedback (such as adjusting the bitrate and packet size) to ensure that the interactive end can smoothly play the audio and video streams of the virtual digital human, fully present the collaborative effect of the virtual digital human's financial explanation and interface interaction, and finally complete the real-time interaction of the virtual digital human in the financial scenario, solving the pain points of high latency and poor experience of virtual digital human financial interaction in existing technologies.
[0119] like Figure 2 As shown, embodiments of the present invention also provide a multimedia real-time interactive system based on generative large model hybridization, including:
[0120] The data preprocessing and decoupling module is used to synchronize and time-align the multimodal raw data stream from the teaching end, and decouple it according to modality type, and distribute it in parallel to the corresponding dedicated generative model processing unit in the cloud.
[0121] The feature vector output module is used to perform real-time parsing and processing of the input data through various dedicated generative model processing units, and output multiple sets of feature vectors.
[0122] The alignment, calibration, and aggregation module is used to input multiple sets of feature vectors into a cross-modal dynamic alignment and fusion network to generate a multimodal joint feature tensor.
[0123] The mapping, construction, and delimitation module is used to map the multimodal joint feature tensor into a high-dimensional semantic feature cloud, where each point in the high-dimensional semantic feature cloud represents a fused feature unit; multiple semantic manifold trajectories are constructed in the high-dimensional semantic feature cloud through the semantic correlation and spatiotemporal proximity between feature units; and a dynamic semantic aggregation region is defined based on the distribution and intersection of the semantic manifold trajectories.
[0124] The anchor point selection and index gradient calculation module is used to select multiple semantic sampling anchor points at the internal core, boundary, and external associated positions of the dynamic semantic aggregation region; through the temporal progression of teaching interaction, the semantic sampling anchor points are connected to form a high-order semantic closed manifold; the topological stability index and semantic density gradient of the high-order semantic closed manifold at multi-dimensional scale are calculated.
[0125] The instruction set generation module is used to synthesize a virtual digital human driving instruction set that has been spatially topologically consistent, based on the multimodal joint feature tensor and the corresponding topological stability index and semantic density gradient, through a multimodal instruction generation network.
[0126] The parsing, rendering, and push module is used to parse and render the virtual digital human driving instruction set at the edge node, generate a synchronized virtual digital human audio and video stream, and push it to the interactive end with low latency.
[0127] In this embodiment, the various functional modules of the multimedia real-time interactive system based on generative large model hybridization, including the data preprocessing and decoupling module, feature vector output module, alignment, calibration and aggregation module, mapping, construction and delimitation module, anchor point selection and index gradient calculation module, instruction set generation module, and parsing rendering and push module, are implemented in software form and deployed in the computing environment of the cloud or edge nodes. They complete their corresponding functions by calling the corresponding hardware resources through program instructions.
[0128] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A real-time interactive multimedia method based on generative large model hybridization, characterized in that, The method includes: The multimodal raw data stream from the teaching end is synchronized and time-aligned, and decoupled according to modality type, and distributed in parallel to the corresponding dedicated generative model processing unit in the cloud; Each dedicated generative model processing unit performs real-time parsing and processing of the input data, and outputs multiple sets of feature vectors; Multiple sets of feature vectors are input into a cross-modal dynamic alignment and fusion network to generate a multimodal joint feature tensor; The multimodal joint feature tensor is mapped to a high-dimensional semantic feature cloud, where each point represents a fused feature unit. Multiple semantic manifold trajectories are constructed in the high-dimensional semantic feature cloud based on the semantic correlation and spatiotemporal proximity between feature units. A dynamic semantic aggregation region is defined based on the distribution and intersection of the semantic manifold trajectories. Multiple semantic sampling anchor points are selected at the internal core, boundary, and external associated locations of the dynamic semantic aggregation region; through the temporal progression of teaching interaction, the semantic sampling anchor points are connected to form a higher-order semantic closed manifold; the topological stability index and semantic density gradient of the higher-order semantic closed manifold at multi-dimensional scale are calculated. Based on the multimodal joint feature tensor and the corresponding topological stability index and semantic density gradient, a set of virtual digital human driving instructions calibrated with spatial topological consistency is synthesized through a multimodal instruction generation network. The virtual digital human driving instruction set is parsed and rendered at the edge node to generate a synchronized virtual digital human audio and video stream, which is then pushed to the interactive terminal with low latency.
2. The multimedia real-time interactive method based on generative large model hybridization according to claim 1, characterized in that, The multimodal raw data stream from the teaching end is synchronized and time-aligned, decoupled by modality type, and distributed in parallel to the corresponding dedicated generative model processing unit in the cloud, including: The system receives a multimodal raw data stream sent by the teaching terminal, wherein the multimodal raw data stream includes at least an audio data packet sequence, a video frame sequence, an image data packet, and a text data packet; Identify the modality type of each data unit in the multimodal raw data stream, and attach a corresponding modality label and a high-precision source timestamp to each data unit to obtain a set of labeled time-series data units; Based on a high-precision source timestamp and a preset reference clock, global clock synchronization is performed on the tagged time-series data unit set to attach a unified synchronization time axis coordinate to all data units and generate a time axis-aligned multimodal data unit set. Based on the modal labels, the time-axis aligned multimodal data unit set is decoupled into independent audio data subsets, video data subsets, image data subsets, and text data subsets; Based on a unified synchronous time axis coordinate, processing priorities and computing resource quotas are dynamically allocated to each data subset. According to the input interface protocol of each dedicated generative model processing unit, the corresponding data subsets are distributed in parallel to the corresponding speech processing model unit, video processing model unit, image processing model unit, and text processing model unit in the cloud.
3. The multimedia real-time interactive method based on generative large model hybridization according to claim 2, characterized in that, Each dedicated generative model processing unit performs real-time parsing and processing of the input data, outputting multiple sets of feature vectors, including: Receive and cache the distributed data subsets, including audio data subsets, video data subsets, image data subsets and text data subsets, as input data to be parsed by the corresponding processing unit; Parallel feature encoding processes adapted to the characteristics of each modality are performed on the received subsets of audio data, video data, image data, and text data respectively to obtain the original feature encoding results; The original feature encoding results are subjected to semantic deep analysis based on the teaching scenario context in order to extract and enhance key semantic information that is strongly related to the teaching objectives, and generate scenario semantic enhancement feature states corresponding to each modality; The scene semantic enhancement feature states are structured, encapsulated, and dimensionality reduced to generate standardized feature vector sets, including speech semantic feature vectors, video spatiotemporal feature vectors, image visual feature vectors, and text deep semantic feature vectors.
4. The multimedia real-time interactive method based on generative large model hybridization according to claim 3, characterized in that, Multiple sets of feature vectors are input into a cross-modal dynamic alignment and fusion network to generate a multimodal joint feature tensor, including: Synchronously read standardized feature vector sets from the caches of each dedicated generative model processing unit, including speech semantic feature vectors, video spatiotemporal feature vectors, image visual feature vectors, and text deep semantic feature vectors. Based on the synchronous time axis coordinates carried by each feature vector, fine-grained time axis alignment and resampling are performed on multiple sets of feature vectors to generate a time-synchronized multimodal feature vector sequence. The time-synchronized multimodal feature vector sequence is input into the multi-head cross-attention alignment layer in the cross-modal dynamic alignment and fusion network. In the multi-head cross-attention alignment layer, the dynamic correlation weights between different modal feature vectors are calculated, and feature recalibration and spatial projection are performed accordingly to generate a preliminary aligned cross-modal feature map. The initially aligned cross-modal feature maps are input into the feature synthesis layer of the cross-modal dynamic alignment and fusion network. Through cascading and nonlinear transformation operations, the multimodal feature maps are aggregated and compressed into a unified, high-dimensional multimodal joint feature tensor.
5. The multimedia real-time interactive method based on generative large model hybridization according to claim 4, characterized in that, The multimodal joint feature tensor is mapped to a high-dimensional semantic feature cloud, where each point represents a fused feature unit. Multiple semantic manifold trajectories are constructed in the high-dimensional semantic feature cloud based on the semantic correlation and spatiotemporal proximity between feature units. A dynamic semantic aggregation region is defined according to the distribution and intersection relationships of these semantic manifold trajectories, including: The multimodal joint feature tensor is input into a high-dimensional spatial projection layer. Through nonlinear dimensionality reduction mapping, each fused feature unit in the multimodal joint feature tensor is represented as a coordinate point in a high-dimensional semantic space, thereby generating a high-dimensional semantic feature cloud. In the high-dimensional semantic feature cloud, the semantic correlation degree and spatiotemporal proximity metric between any two feature unit coordinate points are calculated. Based on the metric, edges representing the connection relationship between feature units are constructed to form a high-dimensional feature relationship graph. In the high-dimensional feature relation graph, multiple core connection paths are searched and extracted based on edge weights. Each path consists of a series of feature units and the order of connecting edges. Finally, each core connection path is defined and smoothed into a semantic manifold trajectory. Analyze the spatial distribution density of all semantic manifold trajectories in the high-dimensional semantic space and the topological structure of the intersection points between trajectories. Based on preset density and topological rules, delineate and output a dynamic semantic aggregation region with continuous or discrete boundaries.
6. The multimedia real-time interactive method based on generative large model hybridization according to claim 5, characterized in that, Multiple semantic sampling anchor points are selected at the internal core, boundary, and external related locations of the dynamic semantic aggregation region; through the temporal progression of teaching interaction, the semantic sampling anchor points are connected to form a high-order semantic closed manifold. Calculate the topological stability index and semantic density gradient of a high-order semantically closed manifold at multiple scales, including: Based on the geometric center, density contour, and association strength with external feature units of the dynamic semantic aggregation region in the high-dimensional semantic space, multiple candidate point sets are determined at the internal core, boundary, and external association positions of the dynamic semantic aggregation region, respectively. The coordinate points with the highest semantic purity and discriminative power are selected from each candidate point set, and they are marked as semantic sampling anchor points. The coordinate set of all semantic sampling anchor points in the high-dimensional semantic space is obtained. Based on the temporal evolution of the multimodal joint feature tensor during the teaching interaction process, the temporal order of activation of each anchor point in the semantic sampling anchor point coordinate set is determined, and the anchor points are connected sequentially according to this temporal order to form a spatiotemporally closed trajectory, i.e., a higher-order semantic closed manifold. Topological analysis of higher-order semantic closed manifolds is performed at multiple preset scales to calculate structural durability parameters at different scales, which are used as topological stability indicators. While performing topological analysis, the rate of change of semantic feature density in the local space traversed by the higher-order semantic closed manifold is calculated, and the semantic density gradient is obtained by fitting.
7. The multimedia real-time interactive method based on generative large model hybridization according to claim 6, characterized in that, Based on the multimodal joint feature tensor and the corresponding topological stability index and semantic density gradient, a set of virtual digital human driving instructions calibrated with spatial topological consistency is synthesized through a multimodal instruction generation network, including: The multimodal joint feature tensor is concatenated with the topological stability index and semantic density gradient to form a topologically semantically enhanced fusion feature vector. The topologically enhanced fusion feature vector is input into the instruction decoder of the multimodal instruction generation network; the instruction decoder generates an initial sequence of virtual digital human multimodal driving parameters in parallel, the multimodal driving parameters including at least speech synthesis parameters, facial motion parameters, limb skeleton parameters and whiteboard rendering parameters; Based on topological stability index and semantic density gradient, a spatial topological consistency calibration function is constructed; the multimodal driving parameters in the initial sequence are fine-tuned frame by frame in terms of amplitude and temporal sequence through the calibration function to generate a calibrated driving parameter sequence. The calibrated drive parameter sequence is packaged and time-synchronized according to the set instruction set encapsulation protocol, and finally output as a virtual digital human drive instruction set that has been calibrated for spatial topology consistency.
8. The multimedia real-time interactive method based on generative large model hybridization according to claim 7, characterized in that, The virtual digital human driving instruction set is parsed and rendered at the edge node to generate a synchronized virtual digital human audio and video stream, which is then pushed to the interactive end with low latency, including: The virtual digital human driving instruction set is received at the edge node, and the virtual digital human driving instruction set is unpacked and parsed to obtain structured multimodal rendering instruction data. The structured multimodal rendering instruction data is distributed to the parallel graphics and audio rendering engines at the edge nodes, driving the parallel graphics and audio rendering engines to generate audio data frames and video data frames with a unified timestamp sequence. Based on the unified timestamp sequence, audio data frames and video data frames with the unified timestamp sequence are time-aligned and stream-encapsulated to synthesize a continuous and synchronized virtual digital human audio and video stream. The system performs low-latency encoding and network adaptive optimization on the continuously synchronized audio and video streams of virtual digital humans, and pushes the optimized streaming media data to the interactive end to complete real-time interaction.
9. A real-time interactive multimedia system based on generative large model hybridization, wherein the system implements the method as described in any one of claims 1 to 8, characterized in that, include: The data preprocessing and decoupling module is used to synchronize and time-align the multimodal raw data stream from the teaching end, and decouple it according to modality type, and distribute it in parallel to the corresponding dedicated generative model processing unit in the cloud. The feature vector output module is used to perform real-time parsing and processing of the input data through various dedicated generative model processing units, and output multiple sets of feature vectors. The alignment, calibration, and aggregation module is used to input multiple sets of feature vectors into a cross-modal dynamic alignment and fusion network to generate a multimodal joint feature tensor. The mapping, construction, and delimitation module is used to map the multimodal joint feature tensor into a high-dimensional semantic feature cloud, where each point in the high-dimensional semantic feature cloud represents a fused feature unit; multiple semantic manifold trajectories are constructed in the high-dimensional semantic feature cloud through the semantic correlation and spatiotemporal proximity between feature units; and a dynamic semantic aggregation region is defined based on the distribution and intersection of the semantic manifold trajectories. The anchor point selection and indicator gradient calculation module is used to select multiple semantic sampling anchor points at the internal core, boundary, and external associated locations of the dynamic semantic aggregation region. Through the sequential progression of teaching interactions, a higher-order semantic closed manifold is formed by connecting various semantic sampling anchor points; the topological stability index and semantic density gradient of the higher-order semantic closed manifold at multi-dimensional scales are calculated. The instruction set generation module is used to synthesize a virtual digital human driving instruction set that has been spatially topologically consistent, based on the multimodal joint feature tensor and the corresponding topological stability index and semantic density gradient, through a multimodal instruction generation network. The parsing, rendering, and push module is used to parse and render the virtual digital human driving instruction set at the edge node, generate a synchronized virtual digital human audio and video stream, and push it to the interactive end with low latency.
Citation Information
Patent Citations
A digital human interaction method and system based on multimodal large model
CN119761511A
AI digital human interactive response method based on large language model
CN121144484A