Multimedia interaction control method and system based on man-machine feedback closed loop

By processing multimodal feedback data in a spatiotemporal coupling manner and implementing human-machine feedback closed loop, the spatiotemporal coupling problem of multimodal interaction control in financial scenarios is solved, enabling accurate capture of user intent and dynamic adjustment of control strategies, thereby improving the accuracy and security of financial interactions.

CN121635689BActive Publication Date: 2026-04-24SHANGHAI RONGSHU INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI RONGSHU INFORMATION TECH CO LTD
Filing Date
2026-02-05
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing multimodal interaction control solutions in financial scenarios lack a spatiotemporal coupling system for multimodal feedback data, making it impossible to dynamically quantify user interaction states. This results in deviations between control parameters and the user's true intentions, failing to meet the requirements for high precision, high security, and self-adaptability.

Method used

By acquiring multimodal feedback data in real time, performing preprocessing and spatiotemporal coupling processing, multimodal fusion features are generated, feature reference points are dynamically determined, feature geometric envelope surfaces are constructed, interactive state compensation coefficients are calculated, multimedia interactive control parameters are generated, and the control strategy is updated through human-machine feedback closed loop to achieve multimedia interactive control.

Benefits of technology

Accurately capture users' financial interaction intentions and real-time status, avoid biased asset allocation suggestions and accidental triggering of trading orders, improve the accuracy of interaction instructions and the timeliness of information response, and ensure user operation security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121635689B_ABST
    Figure CN121635689B_ABST
Patent Text Reader

Abstract

The application provides a multimedia interaction control method and system based on human-computer feedback closed loop, and relates to the technical field of financial technology.The method comprises the following steps: acquiring multi-modal original feedback data of a user in an interaction environment in real time; preprocessing the multi-modal original feedback data to obtain preprocessed multi-modal feedback data; performing space-time coupling processing on the preprocessed multi-modal feedback data to generate multi-modal fusion features; in a feature space corresponding to the multi-modal fusion features, a plurality of feature reference points are dynamically determined based on feedback data of different modes; the feature reference points at least include a user pupil center coordinate point in visual feedback data, a fundamental frequency energy peak point in voice feedback data, and a pressure distribution barycenter point in tactile feedback data.The application can realize space-time collaborative processing, dynamic feature quantization and closed loop adaptive compensation of multi-modal feedback in a financial scenario, and improve the instruction accuracy and information response timeliness of financial interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of financial technology, and in particular to a multimedia interactive control method and system based on human-machine feedback closed loop. Background Technology

[0002] With the development of financial technology, multimodal interaction integrating vision, voice, and touch has become the mainstream form in scenarios such as smart investment advisory terminals, financial self-service terminals, and financial transaction screens. The industry has put forward higher requirements for the accuracy of interactive commands, response timeliness, and security of fund operations.

[0003] Current multimodal interaction control solutions in financial scenarios mostly employ a logic of independent data collection and shallow splicing and fusion of each modality, relying on fixed manually calibrated features as the basis for judgment. The technical shortcomings are as follows: most existing control architectures lack a unified spatiotemporal coupling system for multimodal feedback data, and also lack a geometric representation and quantification mechanism for dynamic feature reference points, making it impossible to standardize numerical representations of the spatiotemporal dynamic changes in user interaction states. Not only do control parameters rely solely on shallow fusion results, failing to dynamically generate compensation coefficients based on interaction states, but they also deviate from the user's true intentions and real-time state. For example, in intelligent investment advisory scenarios, when users hesitate due to investment risks, resulting in subtle coupling changes such as pupil deviation, voice frequency fluctuations, and changes in touch pressure, the system struggles to capture these changes, easily leading to configuration suggestion deviations and inaccurate prompts. Furthermore, the lack of secondary feedback closed-loop control logic means that the control strategy remains static, and deviations accumulate continuously. For instance, in financial trading dashboard scenarios, market fluctuations causing user anxiety and resulting in interaction state fluctuations, the system cannot adapt, easily leading to problems such as erroneous command triggers and delayed fund transfers.

[0004] In summary, it is impossible to achieve spatiotemporal collaborative processing, dynamic quantification, and closed-loop compensation for multimodal feedback in financial scenarios, making it difficult to meet the usage requirements of high precision, high security, and self-adaptation. Summary of the Invention

[0005] The technical problem to be solved by this invention is to provide a multimedia interactive control method and system based on human-machine feedback closed loop, which can realize spatiotemporal collaborative processing of multimodal feedback, dynamic geometric feature quantification and closed-loop adaptive compensation in financial scenarios, thereby improving the accuracy of financial interaction commands, the timeliness of information response and the security of fund operation.

[0006] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0007] Firstly, a multimedia interactive control method based on a human-machine feedback closed loop, the method comprising:

[0008] Real-time acquisition of raw multimodal feedback data from users in the interactive environment; preprocessing of the raw multimodal feedback data to obtain preprocessed multimodal feedback data;

[0009] The preprocessed multimodal feedback data is subjected to spatiotemporal coupling processing to generate multimodal fusion features. In the feature space corresponding to the multimodal fusion features, multiple feature reference points are dynamically determined based on feedback data of different modalities. The feature reference points include at least the coordinate point of the user's pupil center in the visual feedback data, the fundamental frequency energy peak point in the voice feedback data, and the pressure distribution centroid point in the tactile feedback data.

[0010] Based on multiple feature reference points, a dynamic feature geometric envelope is constructed in the feature space; the volume change rate and centroid offset vector of the feature geometric envelope are calculated at continuous interaction times; and interaction state compensation coefficients are generated based on the volume change rate and centroid offset vector.

[0011] Based on multimodal fusion features and interaction state compensation coefficients, determine multimedia interaction control parameters; and execute multimedia interaction actions according to the multimedia interaction control parameters.

[0012] After the multimedia interactive action is executed, the user's secondary interactive feedback data is obtained; based on the secondary interactive feedback data, the control deviation of the multimedia interactive control parameters is calculated.

[0013] Based on the control deviation, the interactive control strategy is updated. By adopting the updated interactive control strategy, the determination of feature reference points and the generation of interactive state compensation coefficients in the next interactive cycle are adjusted to complete the human-machine feedback closed loop.

[0014] Furthermore, real-time acquisition of raw multimodal feedback data from users in the interactive environment; preprocessing of the raw multimodal feedback data to obtain preprocessed multimodal feedback data, including:

[0015] Construct a user-centric adaptive perception field; the adaptive perception field is a dynamic, multi-dimensional data perception space, and the spatial configuration and perception granularity of the adaptive perception field change in response to the user's real-time interaction status.

[0016] Within the constructed adaptive perception field, the user's behavioral focus area and the state interference factors of the interactive environment are monitored and analyzed in real time to generate synchronous analysis results of the environment and user state. Among them, the behavioral focus area includes the user's gaze area, the voice source localization area, and the core area of ​​limb contact; the state interference factors include ambient light intensity, background noise spectrum, and dynamic parameters of irrelevant moving objects.

[0017] Based on the synchronous analysis results of the environment and user status, an adaptive data acquisition strategy is generated. The adaptive data acquisition strategy is used to dynamically configure the spatial directivity, sampling frequency and data confidence weight of each modal sensor in the adaptive sensing field.

[0018] Based on an adaptive data acquisition strategy, the adaptive sensing field is driven to collaboratively acquire the user's multimodal interaction behavior to obtain multimodal raw feedback data that matches the user's state and environment.

[0019] The original multimodal feedback data is standardized by performing time stamp alignment, feature scale normalization, and outlier filtering to obtain preprocessed multimodal feedback data.

[0020] Furthermore, the preprocessed multimodal feedback data undergoes spatiotemporal coupling processing to generate multimodal fusion features. In the feature space corresponding to the multimodal fusion features, multiple feature reference points are dynamically determined based on feedback data from different modalities. These feature reference points include at least the user's pupil center coordinates in the visual feedback data, the fundamental frequency energy peak point in the voice feedback data, and the pressure distribution centroid point in the tactile feedback data, including:

[0021] The preprocessed multimodal feedback data is calibrated and aligned in the temporal domain on a unified time reference, and spatial registration is performed based on a predefined spatial mapping relationship to generate spatiotemporal coupling features with spatiotemporal consistency, including spatiotemporal coupling features of visual modality, spatiotemporal coupling features of speech modality and spatiotemporal coupling features of tactile modality.

[0022] Multi-level feature fusion is performed on various spatiotemporal coupled features to generate a unified multimodal fusion feature that represents the user's interaction intent;

[0023] Map multimodal fusion features to a high-dimensional feature space;

[0024] Based on the spatiotemporal coupling features of each modality in the high-dimensional feature space, at least three feature reference points are dynamically determined. Specifically, based on the spatiotemporal coupling features of the visual modality, the projection coordinates of the user's pupil center in the high-dimensional feature space are determined as the spatial visual focus; based on the spatiotemporal coupling features of the speech modality, the mapping coordinates of the speech fundamental frequency energy peak in the high-dimensional feature space are determined as the acoustic feature hub; and based on the spatiotemporal coupling features of the tactile modality, the mapping coordinates of the pressure distribution centroid in the high-dimensional feature space are determined as the mechanical action center.

[0025] Furthermore, based on multiple feature reference points, a dynamic feature geometric envelope surface is constructed in the feature space; the volume change rate and centroid offset vector of the feature geometric envelope surface are calculated at continuous interaction times; based on the volume change rate and centroid offset vector, interaction state compensation coefficients are generated, including:

[0026] In the high-dimensional feature space, a dynamically changing feature geometric envelope is constructed with the spatial visual focus, acoustic feature hub and mechanical action center as vertices.

[0027] Obtain the spatial morphological sequence of the dynamically changing feature geometric envelope surface at multiple consecutive interaction moments;

[0028] Based on the spatial morphology sequence, the rate of change of the volume of the dynamically changing feature geometric envelope surface with time is calculated as the volume change rate; at the same time, the displacement vector of the centroid of the dynamically changing feature geometric envelope surface during continuous interaction time is calculated as the centroid offset vector.

[0029] The volume change rate and centroid offset vector are normalized and weighted to generate interaction state compensation coefficients that characterize the stability of user interaction state and the clarity of intent.

[0030] Furthermore, based on multimodal fusion features and interaction state compensation coefficients, multimedia interaction control parameters are determined; according to the multimedia interaction control parameters, multimedia interaction actions are executed, including:

[0031] Obtain the unified multimodal fusion features representing user interaction intent, and the interaction state compensation coefficients representing the stability of user interaction state and the clarity of intent;

[0032] Based on the current interaction scenario, the multimodal fusion features and the interaction state compensation coefficients are input into a pre-trained parameter mapping model to generate multimedia interaction control parameters.

[0033] Based on the multimedia interaction control parameters, the multimedia interaction control engine is driven to execute the corresponding multimedia interaction actions.

[0034] Furthermore, after the multimedia interactive action is executed, secondary interaction feedback data from the user is acquired; based on the secondary interaction feedback data, the control deviation of the multimedia interaction control parameters is calculated, including:

[0035] During the execution of multimedia interactive actions, the user-generated secondary multimodal raw feedback data is collected through the adaptive sensing field;

[0036] The user-generated secondary multimodal raw feedback data is preprocessed and spatiotemporally coupled to generate secondary multimodal fusion features;

[0037] The secondary multimodal fusion features and the multimedia interaction control parameters are input together into a predefined deviation calculation function;

[0038] The control deviation of the multimedia interaction control parameters relative to the current user's actual interaction intention is output through a predefined deviation calculation function.

[0039] Furthermore, based on the control deviation, the interactive control strategy is updated. By adopting the updated interactive control strategy, the determination of feature reference points and the generation of interactive state compensation coefficients in the next interaction cycle are adjusted to complete the human-machine feedback closed loop, including:

[0040] Obtain the control deviation output by the predefined deviation calculation function;

[0041] The acquired control deviation is input into the adaptive strategy to update the model;

[0042] The model is updated by an adaptive strategy. Based on the control deviation, the mapping rule of the parameter mapping model and the generation rule of the interaction state compensation coefficient are jointly optimized to obtain the updated parameter mapping model and the updated generation rule of the interaction state compensation coefficient.

[0043] The updated parameter mapping model is dynamically injected into the multimedia interaction control parameter generation process of the next interaction cycle. At the same time, the updated interaction state compensation coefficient generation rules are dynamically injected into the feature reference point determination and compensation coefficient calculation process of the next interaction cycle, thus completing the human-computer feedback closed loop.

[0044] Secondly, a multimedia interactive control system based on a human-machine feedback closed loop includes:

[0045] The data acquisition and preprocessing module is used to acquire multimodal raw feedback data from users in the interactive environment in real time; and to preprocess the multimodal raw feedback data to obtain preprocessed multimodal feedback data.

[0046] The data coupling and reference point determination module is used to perform spatiotemporal coupling processing on the preprocessed multimodal feedback data to generate multimodal fusion features. In the feature space corresponding to the multimodal fusion features, multiple feature reference points are dynamically determined based on feedback data of different modalities. The feature reference points include at least the coordinate point of the user's pupil center in the visual feedback data, the fundamental frequency energy peak point in the voice feedback data, and the pressure distribution centroid point in the tactile feedback data.

[0047] The dynamic construction and calculation module is used to construct a dynamic feature geometric envelope surface in the feature space based on multiple feature reference points; calculate the volume change rate and centroid offset vector of the feature geometric envelope surface at continuous interaction times; and generate interaction state compensation coefficients based on the volume change rate and centroid offset vector.

[0048] The parameter determination and action execution module is used to determine the multimedia interaction control parameters based on the multimodal fusion characteristics and interaction state compensation coefficients; and to execute the multimedia interaction actions based on the multimedia interaction control parameters.

[0049] The control deviation generation module is used to obtain secondary interaction feedback data from the user after the multimedia interactive action is executed; and to calculate the control deviation of the multimedia interactive control parameters based on the secondary interaction feedback data.

[0050] The human-machine feedback closed-loop module is used to update the interactive control strategy based on the control deviation. By adopting the updated interactive control strategy, the feature reference point determination and interactive state compensation coefficient generation in the next interaction cycle are adjusted to complete the human-machine feedback closed loop.

[0051] The above-described solution of the present invention has at least the following beneficial effects:

[0052] Because it employs multimodal feedback data spatiotemporal coupling processing, dynamic feature reference point geometric representation and quantitative calculation, and a human-machine closed-loop control strategy based on secondary feedback, it overcomes the technical problems of existing financial scenario interaction control architectures lacking a unified spatiotemporal coupling system, being unable to quantify dynamic changes in user interaction, and having deviations between control parameters and true intentions due to the lack of closed-loop control, which in turn affects user operation safety. This allows for precise capture of user financial interaction intentions and real-time status, avoiding issues such as biased asset allocation suggestions and accidental triggering of trading instructions. It improves the accuracy of interaction instructions and the timeliness of information response, ensures user operation safety, enhances the reliability and credibility of financial interactions, and meets the high-precision, high-security, and self-adaptive usage requirements of the financial sector. Attached Figure Description

[0053] Figure 1 This is a flowchart illustrating a multimedia interactive control method based on human-machine feedback closed loop provided in an embodiment of the present invention.

[0054] Figure 2 This is a schematic diagram of a multimedia interactive control system based on human-machine feedback closed loop provided in an embodiment of the present invention. Detailed Implementation

[0055] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0056] like Figure 1 As shown, embodiments of the present invention propose a multimedia interactive control method based on a human-machine feedback closed loop, the method comprising the following steps:

[0057] For ease of description and understanding, in the specific embodiments of this specification, several core feature reference points defined in the claims are referred to by aliases with more geometric and physical meaning. Their correspondence and substantive meaning are as follows: The user's pupil center coordinate point in visual feedback data, i.e., the dynamic projection point representing the core position of the user's visual attention in a high-dimensional feature space, is also referred to as the spatial visual focus in this specification. The fundamental frequency energy peak point in voice feedback data, i.e., the dynamic mapping point representing the hub of the user's voice command energy and intent in a high-dimensional feature space, is also referred to as the acoustic feature hub in this specification. The pressure distribution centroid point in tactile feedback data, i.e., the dynamic mapping point representing the core of the user's touch operation mechanical action in a high-dimensional feature space, is also referred to as the mechanical action center in this specification. The above aliases are only for ease of understanding their geometric roles and functions in the feature space; their technical essence is completely the same as the definitions in the claims.

[0058] Step 1: Acquire the user's raw multimodal feedback data in the interactive environment in real time; preprocess the raw multimodal feedback data to obtain preprocessed multimodal feedback data.

[0059] Step 2: Perform spatiotemporal coupling processing on the preprocessed multimodal feedback data to generate multimodal fusion features; in the feature space corresponding to the multimodal fusion features, dynamically determine multiple feature reference points based on feedback data of different modalities; the feature reference points include at least the coordinate point of the user's pupil center in the visual feedback data, the fundamental frequency energy peak point in the voice feedback data, and the pressure distribution centroid point in the tactile feedback data.

[0060] Step 3: Construct a dynamic feature geometric envelope in the feature space based on multiple feature reference points; calculate the volume change rate and centroid offset vector of the feature geometric envelope at continuous interaction times; generate interaction state compensation coefficients based on the volume change rate and centroid offset vector.

[0061] Step 4: Determine the multimedia interaction control parameters based on the multimodal fusion features and interaction state compensation coefficients; execute the multimedia interaction actions according to the multimedia interaction control parameters.

[0062] Step 5: After the multimedia interactive action is executed, obtain the user's secondary interactive feedback data; based on the secondary interactive feedback data, calculate the control deviation of the multimedia interactive control parameters.

[0063] Step 6: Update the interactive control strategy according to the control deviation amount. By adopting the updated interactive control strategy, adjust the determination of feature reference points and the generation of interactive state compensation coefficients in the next interaction cycle to complete the human-machine feedback closed loop.

[0064] In this embodiment of the invention, the spatiotemporal dynamic changes of user multimodal interactions in financial scenarios can be accurately captured. The multimodal fusion features generated through spatiotemporal coupling processing are more closely aligned with the user's actual interaction intentions. The dynamically determined feature reference points and the generated interaction state compensation coefficients can effectively compensate for the deviation between the control parameters of existing solutions and the user's real-time state, avoiding problems such as inaccurate asset allocation suggestions and misaligned transaction threshold prompts. At the same time, the human-machine feedback closed-loop mechanism can update the control strategy in real time based on secondary interaction feedback, preventing the accumulation of deviations and reducing risks such as accidental triggering of transaction instructions and delays in fund transfers. This effectively and reasonably improves the accuracy of financial interaction instructions and the timeliness of information response, ensures user operation security, enhances the reliability and credibility of financial interaction, and fully adapts to the high-precision, high-security, and self-adaptive interaction control requirements of the financial field.

[0065] In a preferred embodiment of the present invention, step 1 above may include:

[0066] Step 1.1: Construct a user-centric adaptive sensing field. This adaptive sensing field is a dynamic, multi-dimensional data sensing space, and its spatial configuration and sensing granularity change in response to the user's real-time interaction state. Specifically, it includes: using the user in various financial interaction scenarios such as intelligent investment advisory terminals, financial self-service terminals, financial trading screens, and online wealth management platforms as the core reference point in three-dimensional space; firstly, collecting basic physical parameters such as the user's limb spatial coordinates, distance from the operating interface, upper limb operating posture, and head orientation angle; then, based on the accuracy, anti-interference, and user operation security adaptation requirements of multi-modal interaction in financial scenarios, systematically integrating and functionally dividing the sensing hardware; and finally, constructing an interconnected, multi-dimensional, dynamic, adaptive sensing field. The specific process is as follows: firstly, selectively choosing sensing hardware adapted to financial scenarios. The system is modularly integrated, featuring a high-precision eye-tracking module paired with a high-definition interactive interface camera to form the core visual perception hardware. This hardware accurately captures visual features strongly correlated with the user's eye gaze trajectory and facial micro-expressions, such as those related to financial interaction intentions. A microphone array module adapted for financial scenarios integrates noise reduction and far-field sound pickup capabilities, adapting to the voice command acquisition needs of users at different operating distances while filtering out interference signals such as crowd noise and device operating noise. A tactile pressure sensing module is configured according to different scenarios, integrating high-precision pressure-sensing touch hardware in financial self-service terminals and pressure-sensing operation hardware in smart investment advisory terminals to accurately collect tactile features such as user touch pressure intensity, swipe trajectory, and press duration. An environmental perception module integrates light, noise, and motion detection functions to comprehensively collect environmental interference parameters for financial scenarios.

[0067] During the integration process, the hardware interfaces of each module are standardized and then integrated into the control terminal, breaking down the barriers of independent deployment and data fragmentation of each modality in existing technologies. Subsequently, the integrated hardware units are divided into four related subdomains according to their functions: the visual perception subdomain, with eye-tracking modules and interface cameras as its core, is responsible for collecting user visual behavior and interactive interface related features; the voice perception subdomain, with microphone arrays as its core, focuses on the accurate acquisition of user voice commands and emotion-related voice features; the tactile perception subdomain, with pressure-sensing hardware as its core, captures the mechanical features of user touch operations; and the environmental perception subdomain, with various environmental detection components as its core, collects environmental interference factors that affect interaction accuracy in real time. Each subdomain does not work independently but achieves collaboration through a set association triggering mechanism: when the visual perception subdomain detects that the user is staring at the transaction confirmation button, it simultaneously activates the tactile perception subdomain to improve sampling accuracy and triggers the voice perception subdomain to enhance sound pickup sensitivity, ensuring the synchronization of multimodal data acquisition.

[0068] Finally, based on the user's basic physical parameters, an initial sensing range is defined. A multi-dimensional, dynamically adaptive sensing field is constructed, centered on the user and covering the comfortable human interaction space. This sensing field completely abandons the static mode of pre-calibrating fixed spatial ranges, fixed sensing densities, and fixed hardware layouts in existing technologies. It reads the user's current business process in real time, distinguishing the accuracy requirements of different operational scenarios such as solution consultation, clause confirmation, instruction issuance, and information browsing. Simultaneously, it captures changes in user interaction behavior caused by decision-making concerns, process changes, and differences in operational proficiency. Using user behavior changes and business process requirements as the core adjustment basis, it dynamically shrinks or expands the overall spatial coverage boundary of the sensing field, dynamically adjusts the spatial overlap ratio and range of each modal sensing subdomain, and adjusts the field based on user operations. The precision level and the clarity of the interaction intent dynamically adjust the parsing level and feature extraction fineness of data perception. When users perform fine operations with high security requirements, such as inputting instructions and confirming terms, the focus of the perception field is tightened and the perception granularity is improved. When users perform routine operations such as browsing basic information and querying service information, the coverage of the perception field is appropriately widened and matched with an appropriate perception granularity. This allows the spatial configuration and perception granularity of the adaptive perception field to respond in real time to the user's physical operation status, business interaction links, and changes in behavior and emotions. This builds a dynamic and unified physical perception foundation for the subsequent synchronous and collaborative collection of multimodal data, unified spatiotemporal alignment, and spatiotemporal coupling processing. It makes up for the underlying defects of existing technologies, such as the lack of a unified perception space and the inability of fixed collection layouts to adapt to the dynamic interaction status of users, from the source of collection.

[0069] Step 1.2: Within the constructed adaptive perception field, monitor and analyze in real time the user's behavioral focus area and the state interference factors of the interactive environment, generating synchronous analysis results of the environment and user state. The behavioral focus area includes the user's gaze area, voice source localization area, and core area of ​​physical contact. State interference factors include ambient light intensity, background noise spectrum, and dynamic parameters of unrelated moving objects. Specifically, within the constructed adaptive perception field, simultaneously perform fine-grained capture of the user's behavioral focus area and precise collection and analysis of state interference factors: continuously capture the user's pupil movement trajectory and accurately map the trajectory to specific functional areas of the financial interaction interface, such as the solution selection area, risk clause reading area, transaction parameter input area, and operation confirmation button area, not only locking onto the core area where the gaze lingers for extended periods. It also records the frequency of eye movement, duration of gaze, and magnitude of gaze shift between different areas to accurately determine the focus of eye gaze. Simultaneously, it captures subtle changes in visual behavior caused by user hesitation due to investment risks or tension due to market fluctuations. It collects voice signals during user financial interactions, precisely locating the core area of ​​the sound source, and simultaneously records the rhythm and pauses of the voice signals. It captures the characteristics of voice command clarity fluctuations caused by emotional fluctuations, clarifying the effective range of voice interaction. Furthermore, it extracts continuous pressure distribution data, sliding trajectory coordinates, and pressing duration during user touch operations, filtering out core pressing points and interface functional areas corresponding to continuous sliding paths to complete the positioning of the core area of ​​physical contact. It also captures details of touch pressure fluctuations caused by changes in the rhythm of operation, achieving full-dimensional and dynamic capture of the user's behavioral focus area.

[0070] Simultaneously and accurately collecting three types of state interference factors affecting the accuracy of financial interaction, specifically including: First, illumination interference factors, collecting the illumination intensity values, illumination-darkness change rates, and light incident angles of different areas within the interaction scene, and quantitatively analyzing the impact of insufficient, excessive, or sudden changes in illumination intensity on the recognition accuracy of visual features (such as pupil positioning and interface element recognition); Second, noise interference factors, collecting the full-band spectrum distribution, transient peak energy, and noise duration of background noise within the scene, filtering out redundant sound signals that overlap with the frequency bands of user financial interaction voice commands such as confirming configuration and submitting transactions, and quantifying the interference level of noise on the clarity recognition of voice commands; Third, irrelevant motion interference factors, collecting the outlines, displacement speeds, movement ranges, and trajectories of moving objects that are not the user subject within the scene, determining whether they will obscure core areas of the financial interaction interface, such as transaction confirmation buttons and risk warning windows, or interfere with the perception signals of the user's focus area, while simultaneously recording the time point and duration of the interference.

[0071] Finally, the spatial coordinates and dynamic change parameters of the user's three types of behavioral focus areas, switching frequency, dwell time, offset amplitude, pressure changes, etc., are compared and integrated with the quantitative data of the three types of state interference factors under a unified time-series benchmark. This process removes pupil trajectory distortion data caused by sudden changes in illumination, redundant voice signal data caused by noise interference, and touch positioning deviation data caused by irrelevant motion occlusion. At the same time, it corrects the time-series misalignment problem of different modal data. The final result is an environmental and user state synchronous analysis that accurately reflects the user's real and effective financial interaction behavior and clearly marks the degree and period of environmental interference. This provides a core basis for subsequent targeted generation of adaptive data collection strategies, avoidance of environmental interference, and improvement of multimodal data collection quality.

[0072] Step 1.3: Based on the synchronous analysis results of the environment and user status, an adaptive data acquisition strategy is generated. This adaptive data acquisition strategy dynamically configures the spatial directivity, sampling frequency, and data confidence weights of each modal sensor within the adaptive sensing field. Specifically, it includes: using the synchronous analysis results of the environment and user status as the core basis, and combining the spatial coordinates of the user's behavioral focus, the business importance level of the interactive operation, and the influence degree of different modal interference factors, dynamically configuring the operating parameters of each modal sensor (visual, speech, and tactile) within the adaptive sensing field modally, adjusting the acquisition angle and detection orientation of each sensor to ensure its spatial directivity is accurately aligned with the user's behavioral focus area and to weaken non-focused areas. The sampling weight for invalid areas is determined by setting differentiated sampling frequencies based on the accuracy requirements of each operation. The sampling frequency is increased for key stages such as transaction instructions and risk confirmation, while the normal sampling frequency is adapted for ordinary information browsing. At the same time, the confidence weight is assigned to each modality of collected data based on the intensity of environmental interference. The weight is reduced for modalities that are more susceptible to light and noise interference, while the weight is increased for modalities that are less affected by interference and closely match the user's actual behavior. The configuration rules of spatial orientation, sampling frequency, and data confidence weight are integrated into a complete execution logic, forming an adaptive data collection strategy that can be dynamically iterated and adapted to the current user status and environmental conditions. This completely abandons the static mode of fixed collection parameters and manually calibrated weights in existing technologies.

[0073] Step 1.4: Based on the adaptive data acquisition strategy, drive the adaptive sensing field to collaboratively acquire the user's multimodal interaction behavior, obtaining multimodal raw feedback data that matches the user's state and environment. Specifically, this includes: using the adaptive data acquisition strategy as the execution benchmark, initiating a synchronization scheduling command using a unified system clock to ensure that the acquisition work of visual, voice, and tactile sensing dimensions starts and progresses collaboratively under the same timing benchmark, avoiding timing misalignment caused by independent acquisition of each modality; each sensing dimension strictly follows the spatial directional requirements set by the strategy, accurately focusing on the core area of ​​the user's gaze, the core area of ​​voice output, and the core area of ​​limb touch control for acquisition, actively avoiding non-interactive areas. Invalid data collection is eliminated to ensure that the collection focus fully matches the user's actual interaction behavior. Simultaneously, continuous time-series capture is dynamically executed according to the sampling frequency levels defined by the strategy: for detailed operations such as transaction confirmation and parameter entry, feature-intensive collection is completed according to high-frequency sampling requirements; for routine operations such as information browsing and interface switching, an adaptive sampling frequency is used to balance collection accuracy and processing efficiency; in the data transmission stage, signals are strictly filtered according to preset data confidence weights: high-confidence data that is less affected by environmental interference and can stably reflect user behavior is prioritized for transmission and temporary storage; low-confidence data that is significantly affected by light and noise interference undergoes secondary validity verification to remove distorted portions.

[0074] During the process, the system simultaneously and accurately captures subtle coupling characteristics caused by user hesitation and emotional fluctuations due to investment risks, including pupil gaze deviation trajectory, voice base frequency energy fluctuation amplitude, and dynamic changes in touch pressure distribution. Finally, through the above-mentioned collaborative collection and filtering process, redundant data that deviates from the focus of behavior, data with strong interference distortion, and invalid signals unrelated to financial interaction are automatically filtered out. This yields multimodal raw feedback data that is fully synchronized in time, precisely aligned in space, pure and valid, and can completely reflect the user's real-time financial operation status and current interaction environment. From the collection and execution level, this solves the core defects of existing technologies, such as independent collection of each modality, misaligned timing, and excessive invalid data.

[0075] Step 1.5 involves standardizing the multimodal raw feedback data, including timestamp alignment, feature scale normalization, and outlier filtering, to obtain preprocessed multimodal feedback data. Specifically, this includes: first, adding frame-by-frame matching timestamps to the visual pupil coordinate data, speech fundamental frequency energy data, and tactile pressure distribution data using a unified high-precision clock, and completing temporal calibration alignment to eliminate spatiotemporal misalignment caused by hardware delays of different modal sensors; then, mapping the original feature data with significant differences in dimensions, numerical ranges, and variation amplitudes under different modalities to a unified standard numerical range, completing feature scale normalization processing, so that each modal data has a unified benchmark for subsequent spatiotemporal coupling and feature fusion; then, identifying abnormal values ​​and discrete invalid data points caused by sudden changes in illumination, noise spikes, accidental touch operations, and device fluctuations through data validity judgment scene identification, and completing outlier filtering and local smoothing completion processing in combination with the normal behavior range of financial interaction, finally obtaining preprocessed multimodal feedback data that is completely consistent in time, has a unified scale standard, is free of abnormal interference, and can be directly used for subsequent spatiotemporal coupling processing.

[0076] In a preferred embodiment of the present invention, step 2 above may include:

[0077] Step 2.1 involves performing temporal calibration and alignment on a unified time reference for the preprocessed multimodal feedback data, and spatial registration based on predefined spatial mapping relationships to generate spatiotemporally consistent spatiotemporal coupling features, including visual modal spatiotemporal coupling features, speech modal spatiotemporal coupling features, and tactile modal spatiotemporal coupling features. Specifically, for the preprocessed visual pupil trajectory, gaze direction, speech fundamental frequency energy, command syllables, and tactile touch pressure and sliding trajectory multimodal feedback data, a global high-precision unified clock is used as the sole time reference to initiate the full-modal data temporal calibration process: traversing each time frame of each modal data, extracting the high-precision timestamp markers added in the preprocessing stage, and first distinguishing the time delay categories of different modal data. The method involves confirming the hardware acquisition latency of visual data, the transmission encoding latency of voice data, and the touch response latency of tactile data. Then, it compares the timestamp deviations of different modal data at the same interaction moment (such as a user initiating a voice command for solution consultation at a smart service interaction terminal or clicking an operation confirmation button at an offline smart service counter). For single-frame data with latency deviations, a frame-by-frame offset correction is performed using a latency type matching correction method. After correction, the error is ensured to be within a set threshold through timestamp consistency verification of the same interaction event. Ultimately, this achieves data alignment of visual pupil changes, voice command output, and tactile pressing actions at completely consistent time points, completing full-domain temporal calibration and solving the problem of temporal misalignment of multimodal data in existing technologies from a temporal perspective.

[0078] Based on the predefined three-dimensional mapping rules for the physical acquisition space, user operation space, and professional interactive interface functional space, the implementation process of these predefined rules is as follows: First, accurately define the core range and key parameters of the three spaces respectively; delineate the effective detection boundaries and detection accuracy thresholds of each sensor within the physical acquisition space; define the comfortable range of limb movements within the user operation space that is suitable for professional service scenarios, such as 0.5 to 1.2 meters in front of the intelligent service interactive terminal and 0.3 to 0.8 meters in the offline intelligent service counter operation area; and mark the precise coordinates of each core functional area in the professional interactive interface functional space, such as the solution selection area and rule prompts. The coordinates of the upper left and lower right corners of the input area and parameter input area, as well as the center coordinates and radius of the operation confirmation button, are determined. Then, typical professional service interaction scenarios such as intelligent service solution consultation, offline intelligent counter business processing, and professional interactive screen command issuance are selected for scenario testing. The sensor detection points in the physical acquisition space, the limb movement points in the user operation space, and the corresponding functional area coordinates in the professional interactive interface function space are recorded under different user operation behaviors, establishing a one-to-one correspondence between the three spatial points. Finally, the measured data is filtered and noise-reduced, solidifying the effective correspondence into standardized three-dimensional mapping rules.

[0079] Based on this rule, the spatial data of each modality are uniformly anchored to a professional interactive coordinate system: In the visual modality, the physical coordinates of the pupil center are first converted into the absolute coordinates of the interactive interface through a three-dimensional mapping rule, ensuring that the physical position of the pupil gaze can be accurately matched with specific functional areas such as the solution selection area, rule prompt area, and parameter input area; In the voice modality, the physical position of the sound source is converted into the relative operation distance coordinates between the user and the professional interactive terminal, and it is simultaneously determined whether the distance is within the effective interaction range, providing a basis for subsequent screening of effective voice signals; In the tactile modality, the physical coordinates of the touch pressure distribution are directly matched to the preset coordinates of the touch operation area of ​​the interactive interface, such as the operation confirmation button and the parameter adjustment slider, ensuring that each set of pressure data can be accurately associated with specific touch operation behavior; Through the above detailed and precise spatiotemporal calibration and registration operations, the double misalignment deviation of different modal data in the time and space dimensions is completely eliminated.

[0080] Ultimately, it generates spatiotemporal coupling features of visual modalities with triple associations of the same time node, the same interactive spatial location, and the same professional operation behavior, such as the pupil gaze feature in the solution selection area at a certain moment; spatiotemporal coupling features of voice modalities, such as the voice fundamental frequency feature of the intelligent service interaction terminal confirming the solution at 0.8 meters at a certain moment; and spatiotemporal coupling features of tactile modalities, such as the pressure feature of the operation confirmation button at a certain moment. It builds a unified spatiotemporal coupling system of multimodal data dedicated to professional service scenarios from the bottom layer, directly making up for the core deficiency of existing technologies in the lack of spatiotemporal collaborative processing, and providing accurate and collaborative basic data support for subsequent deep feature fusion and dynamic feature reference point positioning.

[0081] Step 2.2 involves multi-level feature fusion of various spatiotemporal coupling features to generate a unified multimodal fusion feature representing the user's interaction intent. Specifically, this includes: for the three types of spatiotemporal coupling features already generated for financial scenarios, following a multi-level deep fusion logic of basic feature selection, business association integration, and intent feature extraction, abandoning the shallow fusion method of simple splicing in existing technologies; the first step is the basic feature association selection at the bottom layer: performing point-by-point validity verification on the spatiotemporal coupling features of each modality, retaining basic features strongly related to financial interaction, such as the trajectory features of the gaze risk warning area in the visual modality, and asset allocation transaction confirmation in the voice modality. The first step involves identifying the fundamental frequency characteristics of the instruction and the pressure characteristics of precisely pressing buttons in the tactile modality. Redundant features without interactive significance in a single modality are eliminated, such as occasional eye movements in the visual modality, invalid syllables left by the environment in the voice modality, and slight accidental touch jitter in the tactile modality. The second step is the integration of mid-level business behaviors: linking the multimodal basic features in the same time and space with financial business processes. For example, combining the visual features of staring at the asset allocation option area, the voice features of asking about risk levels, and the tactile features of lightly touching the option slider, and integrating them into the risk confirmation behavior association features in asset allocation consultation, capturing the behavioral logic of users' continuous financial operations. The third step is to extract the high-level interaction intent: combining the core needs of financial business logic, such as asset allocation consultation, transaction instruction issuance, and risk clause confirmation, to conduct intent-oriented analysis on the mid-level related features. For example, from the combination of visual features of staring at the transaction confirmation button, voice features of voice command confirmation of the transaction, and tactile features of high-pressure long press of the button, the core features of the interaction intent to complete the transaction are extracted. Finally, through the three-layer progressive fusion processing, a multimodal fusion feature that can completely, accurately and uniformly represent the user's real financial interaction intent is generated.

[0082] Step 2.3 maps the multimodal fusion features to a high-dimensional feature space. This specifically includes: initiating a standardized high-dimensional feature space mapping process for the multimodal fusion features of financial scenarios that have undergone multi-level deep fusion; first, loading pre-defined financial interaction-specific feature mapping rules, which are pre-classified according to different financial business links such as asset allocation, transaction confirmation, and risk consultation, and accurately associating the core feature dimension requirements of each link. For example, the asset allocation link focuses on associating feature dimensions such as pupil gaze trajectory and the base frequency of the speech for inquiring about the plan, while the transaction confirmation link focuses on associating feature dimensions such as pressing pressure and confirmation command syllables. At the same time, the unique dimension index of each fusion feature in the high-dimensional space is determined to ensure that different modal features can be accurately positioned in the high-dimensional space; then, fusion features with modal differences in the low-dimensional space, such as visual features being coordinate data, speech features being energy data, and tactile features being pressure data, are projected dimension by dimension into the dedicated high-dimensional feature space according to the mapping rules.

[0083] During the projection process, the numerical range and representation form of different modal features are uniformly normalized to eliminate the dimensional differences between the pixel-level pupil coordinates and the force level of the pressure applied during the press. This ensures that all fused features and their interrelationships, such as the negative correlation between pupil deviation and reduced pressure during risk hesitation, can be presented in a standardized numerical form in a high-dimensional space. Simultaneously, a feature dynamic tracking channel is constructed in the high-dimensional space to record the dynamic changes of fused features in real time as the user's financial interaction behavior changes, such as the dynamic trajectory from hesitant asset allocation to confirmed transaction. Ultimately, this forms a unified high-dimensional space carrier that can accommodate the spatiotemporal distribution, dynamic offset trends, and interrelationships of multimodal features. This provides a standardized and computable spatial foundation for the subsequent positioning, geometric representation, and quantitative calculation of dynamic feature reference points, thus addressing the deficiency of being unable to provide standardized numerical representations of the spatiotemporal dynamic changes in the user's financial interaction state.

[0084] Step 2.4: In the high-dimensional feature space, dynamically determine at least three feature reference points based on the spatiotemporal coupling features of each modality. Specifically, based on the spatiotemporal coupling features of the visual modality, determine the projection coordinates of the user's pupil center in the high-dimensional feature space as the spatial visual focus; based on the spatiotemporal coupling features of the speech modality, determine the mapping coordinates of the speech fundamental frequency energy peak in the high-dimensional feature space as the acoustic feature hub; and based on the spatiotemporal coupling features of the tactile modality, determine the mapping coordinates of the pressure distribution centroid in the high-dimensional feature space as the mechanical action center. This includes: in the financial scenario-specific high-dimensional feature space after multimodal fusion feature projection, initiating a dynamic feature reference point localization process. This process completely abandons the static mode of fixed manual calibration used in existing technologies, using the spatiotemporal coupling features of each modality updated frame by frame as the core data source. The continuous dynamic coordinate calibration is completed through layered calculation. The first step is to locate the spatial visual focus: First, the spatiotemporal coupling features of the visual modality are retrieved, and the spatiotemporal data sequence of the pupil center in the current financial interaction of the user, such as asset allocation consultation and transaction instruction issuance, is extracted. Invalid points caused by accidental eye movement are eliminated, and valid data points of the core functional areas such as the asset allocation option area and the transaction confirmation button are retained. Then, according to the three-dimensional mapping rules preset in step 2.1 and the high-dimensional space mapping rules in step 2.3, the physical coordinates of the valid pupil center are converted into the projection coordinates of the high-dimensional feature space point by point. During the conversion process, the matching degree between the coordinates and the functional areas of the financial interface is checked simultaneously. Finally, the average value of the valid coordinates of three consecutive frames is taken as the stable coordinate at the current moment and calibrated as the spatial visual focus.

[0085] The second step is to locate the acoustic feature hub: retrieve the spatiotemporal coupling features of the speech modality, first filter out invalid frequency band data corresponding to scene environmental noise, and select the valid speech segments corresponding to the user's financial voice commands; then split the segment by time frame, calculate the fundamental frequency energy value of each frame, and select the peak region with continuous and stable energy by comparing the energy changes of adjacent frames. Take the coordinates corresponding to the energy value of the center frame of the region as candidate coordinates, and then complete the coordinate conversion by combining the high-dimensional space mapping rules. The converted stable coordinates are marked as the acoustic feature hub. The third step is to locate the mechanical action center: retrieve the spatiotemporal coupling features of the tactile modality, integrate the touch pressure distribution data of multiple consecutive frames, first remove invalid pressure points caused by accidental touch and hovering touch, and retain the pressure data corresponding to valid operations such as pressing the transaction confirmation button and sliding the asset configuration parameter bar; then count the distribution range of the valid pressure points, use the pressure value of each point as the weight, calculate the weighted average value of the coordinates of all valid points, obtain the initial coordinates of the centroid of the pressure distribution, and then complete the conversion by high-dimensional space coordinate transformation. The converted coordinates are marked as the mechanical action center.

[0086] Throughout the process, the coordinates of the three feature reference points are updated frame-by-frame according to the subtle coupling changes in the user's financial interaction state: when the user hesitates due to investment risks and causes pupil deviation, the effective pupil point screening and coordinate mean calculation are re-executed; when market fluctuations cause voice base frequency fluctuations, the voice segment energy peak calculation and coordinate conversion are re-completed; when the user's operational safety concerns cause changes in touch pressure, the pressure points are re-statistically counted and the weighted average center of gravity is calculated; all reference point coordinates are strongly bound to the user's real-time financial operation behavior, ultimately completing the determination of at least three core dynamic feature reference points, constructing the geometric representation basis of the user's interaction state in the financial scenario, realizing the accurate quantitative representation of subtle changes in user interaction, and from the computational level, making up for the core defects of existing technologies that lack dynamic feature quantification calculation mechanisms and cannot adapt to dynamic financial interactions.

[0087] In a preferred embodiment of the present invention, step 3 above may include:

[0088] Step 3.1: In the high-dimensional feature space, using the spatial visual focus, acoustic feature hub, and mechanical action center as vertices, construct a dynamically changing feature geometric envelope. Specifically, this includes: in the high-dimensional feature space, first accurately extracting the determined spatial visual focus, with coordinates as... Acoustic feature hub, with coordinates as The center of force is marked as The real-time coordinates of three core dynamic feature reference points initiate a dual-validation process for coordinate validity: First, calculate the average deviation of each coordinate from the valid coordinates of the past 10 frames. If... x , y , zIf the deviation in all three directions does not exceed 5 units, and the accuracy requirements of the financial scenario are met, then it is initially determined to be valid. The second step is to combine the current financial interaction process, such as asset allocation consultation and transaction instruction issuance, to determine whether the interaction area corresponding to the coordinate is the core functional area of ​​the process, such as the asset allocation option area or the transaction confirmation button area. If it matches, it is finally confirmed to be valid, and abnormal coordinate points caused by instantaneous sensor interference or accidental user misoperation are eliminated.

[0089] After the verification is successful, P 1 、P 2 、P 3. Using three valid coordinates as the core vertices, perform a precise point-by-point connection operation to form a basic triangular outline; then, select multimodal feature points (such as pupil offset points at adjacent time points, secondary peak points of speech energy, and edge points of touch pressure distribution) within 3 units of the core vertices around the outline as associated supplementary points, and mark the coordinates of these supplementary points as follows: , Then, it is added to the basic outline; smoothing is then achieved by calculating the average coordinates of adjacent boundary points, for example, two adjacent points on the boundary. and The coordinates of the smoothed point are: Eliminate the shape distortion caused by sharp edges; ultimately construct a system that can fully encompass all effective multimodal features in the current financial interaction scenario, and adapt accordingly. P 1. P 2. P The feature geometric envelope surface is dynamically updated in real time with 3 coordinates, realizing the geometric encapsulation of the user's financial interaction state and solving the deficiency of the lack of geometric representation in the existing technology.

[0090] Step 3.2: Obtain the spatial morphological sequence of the dynamically changing feature geometric envelope surface at multiple consecutive interaction moments. Specifically, this includes: setting an adaptive time sampling interval based on the operational rhythm differences of different interaction stages in the financial scenario: for slower-paced stages such as asset allocation consultation and risk consultation, the sampling interval is set to 500 milliseconds to balance data integrity and processing efficiency; for faster-paced stages such as transaction instruction issuance and fund transfer confirmation, the sampling interval is shortened to 200 milliseconds to ensure that sudden changes in interaction states caused by market fluctuations are not missed; according to the set sampling interval, synchronously collect the full-dimensional morphological parameters of the feature geometric envelope surface at each interaction moment, specifically including: three core vertices. P 1. P 2. P 3. Real-time coordinates; coordinates of all boundary points including supplementary points on the envelope contour; curvature values ​​of each segment of the envelope edge, obtained by calculating the angle between line segments using the coordinates of three adjacent points; distribution density (number of feature points per unit volume) and association strength of effective feature points inside the envelope; the closer the feature point is to the core vertex, the higher the association strength.

[0091] Simultaneously, a precise timestamp is added to the morphological parameters at each moment, ensuring that this timestamp is consistent with the time reference after time-series calibration in step 2.1, with a deviation of no more than 10 milliseconds. The morphological parameters collected at different moments are then sorted chronologically, and duplicate parameter records with the same timestamp are removed. For missing parameters at individual moments, linear interpolation of parameters from preceding and following moments is used, such as at time... t 1 parameter is M 1. Time t 3 parameters are M 3. Time t 2. Missing parameters The data is supplemented and completed to form a continuous, complete, and time-sequential spatial morphological sequence, which fully preserves the dynamic change trajectory of the feature geometric envelope as users' financial interaction behavior progresses, providing comprehensive morphological data support for subsequent quantitative analysis.

[0092] Step 3.3: Based on the spatial morphology sequence, calculate the rate of change of the volume of the dynamically changing feature geometric envelope surface over time, as the volume change rate; simultaneously, calculate the displacement vector of the centroid of the dynamically changing feature geometric envelope surface between consecutive interaction moments, as the centroid offset vector. Specifically, this includes: based on the acquired spatial morphology sequence, completing the calculation of core quantitative indicators in two steps to accurately characterize changes in user interaction state: the first step is to calculate the volume change rate, first targeting two adjacent interaction moments in the sequence. t 1 and t The corresponding feature geometric envelope surface is calculated using a high-dimensional space capacity integral decomposition method. Each envelope surface is decomposed into multiple directly calculable three-dimensional subspaces, such as triangular pyramids or cuboids with the core vertex as the vertices. For each three-dimensional subspace, the volume is calculated using the corresponding geometric formula, such as the volume of a cuboid = length × width × height, and the volume of a triangular pyramid = base area × height / 3. The volumes of all subspaces are then summed to obtain the final volume. t Envelope volume at time 1 V 1 and t 2-time envelope volume V 2; Calculate the volume difference Δ V = V 2- V 1. Then use the formula Obtain the rate of change of volume at adjacent time points R ;like R A positive value indicates an expansion of the asset envelope, corresponding to a strengthening of the user's financial interaction intent, such as moving from hesitation about asset allocation to a clear choice of allocation plan; if... R A negative value indicates a shrinking envelope, corresponding to a weakened interactive intent; if R A value close to 0 indicates that the interaction state is stable.

[0093] The second step is to calculate the centroid offset vector. First, for the feature geometric envelope surface at each time step, assign values ​​based on the weights of the financial interaction representations of each feature point within it, such as the weights of the tactile pressure points in the transaction confirmation process. W t =0.6, visual pupil fixation point weight W v =0.3, speech energy point weight W s =0.1; The centroid coordinates are calculated using a weighted average, and the formula for the centroid coordinates is as follows:

[0094] ;

[0095] in W i For the first i The weights of each feature point X i 、Y i 、Z i For the first i The coordinates of each feature point are calculated; t Centroid at time 2 C 2 and t Centroid at moment 1 C Coordinate difference of 1 , , According to Δ X Δ Y Δ Z The sign of the centroid determines the direction of the centroid offset, such as Δ. X A positive value corresponding to the direction of the transaction confirmation zone indicates an intention to offset towards the confirmed transaction. This is determined by calculating the offset distance. The degree of offset is determined; finally, the offset direction and distance are integrated to obtain the centroid offset vector between consecutive interaction moments, so as to achieve accurate quantification of changes in the user's financial interaction state.

[0096] Step 3.4 involves normalizing and weighting the volume change rate and centroid offset vector to generate interaction state compensation coefficients that characterize the stability of user interaction states and the clarity of intent. Specifically, this includes: first, normalizing and weighting the calculated volume change rate... R and centroid offset vector (using offset distance) L (Characteristics) Perform normalization processing separately to eliminate dimensional differences between different indicators: for volume change rate R First, we statistically analyzed a large number of interaction samples in financial scenarios. R Maximum value R max and minimum value R min Through formula WillR Mapped to the interval between 0 and 1, where R =0 corresponds to the weakest interactive intent. R =1 corresponds to the state with the strongest interactive intent; this is based on the centroid offset distance. L In the same statistical sample L maximum value L max Through formula Will L Normalized to the interval between 0 and 1, L =0 corresponds to a stable state with no offset. L =1 corresponds to the unstable state with the largest offset.

[0097] Subsequently, considering the core requirements of security and accuracy in financial transactions, differentiated weights were determined through calibration using thousands of financial interaction samples: the centroid offset vector directly reflects the directional change of the user's interaction intent, and directional deviation can easily lead to risks such as erroneous triggering of transaction instructions and incorrect fund transfers, therefore, weights were assigned. W c =0.7; The volume change rate reflects the intensity change of the interaction intent, and its impact on security is relatively indirect. (Assign weights accordingly.) W r =0.3; The comprehensive value is calculated using the weighted summation formula: compensation coefficient ; for the calculated K Perform a validity check; if K If >1, then take K =1, if K If <0, then take K =0, ensuring K It falls within a reasonable range of 0 to 1; after verification... K This is the interaction state compensation coefficient: K A value between 0 and 0.3 indicates that the user's interaction state is stable and their intent is clear. K A value between 0.3 and 0.7 indicates a slight fluctuation in the interaction state, and the intent needs further confirmation. K A value between 0.7 and 1 indicates unstable interaction and ambiguous intent; this coefficient can be directly used for dynamic adjustment of subsequent financial interaction control parameters, such as... K The larger the scale, the slower the response speed and the more confirmation steps are needed to achieve closed-loop adaptive compensation for user interaction status, thus making up for the core defects of existing technologies that lack dynamic compensation mechanisms and cannot guarantee the security of fund operations.

[0098] In a preferred embodiment of the present invention, step 4 above may include:

[0099] Step 4.1: Obtain the unified multimodal fusion feature representing the user's interaction intent, and the interaction state compensation coefficient representing the stability and clarity of the user's interaction state. Specifically, this includes: first, accurately retrieving the generated unified multimodal fusion feature representing the user's interaction intent, and simultaneously extracting the calculated interaction state compensation coefficient representing the stability and clarity of the user's interaction state. Validate these two types of core data to confirm that the multimodal fusion feature does not contain redundant or distorted information and that the interaction state compensation coefficient is within a reasonable range of 0 to 1. Then, associate and bind the two types of data that have passed the validation according to a unified time series benchmark to form a complete data set containing the user's interaction intent and real-time state, providing core data support for the subsequent generation of adapted multimedia interaction control parameters.

[0100] Step 4.2: Combining the current interaction scenario, input the multimodal fusion features and the interaction state compensation coefficients into the pre-trained parameter mapping model to generate multimedia interaction control parameters. Specifically, this includes: first determining the type of the current financial interaction scenario, confirming whether it is an asset allocation consultation scenario of a smart investment advisory terminal, a business processing scenario of an offline smart counter, or an instruction issuance scenario of a financial transaction screen, etc., to provide a scenario benchmark for model adaptation. The pre-trained parameter mapping model used in this study is an improvement based on the fully connected neural network (FCN) architecture. Addressing the needs of multimodal data collaborative processing and dynamic state adaptation in financial scenarios, it adds a scenario adaptation layer and a feature attention mechanism to the traditional FCN. Its core advantage lies in its ability to accurately capture the collaborative relationship between multimodal fusion features and interaction state compensation coefficients. Simultaneously, it can adaptively adjust the mapping logic according to the security level and interaction rhythm of different financial scenarios, effectively solving the problem of disconnect between control parameter generation and scenario / user state in existing technologies.

[0101] The specific construction and training process of this model revolves around the interaction needs of financial scenarios: The first step is model construction, which involves building a basic network framework. The input layer is set to an adaptation dimension, which can simultaneously receive multimodal fusion feature vectors and interaction state compensation coefficients, with a single-dimensional value. The intermediate layer adds three hidden layers and one scenario adaptation layer. The hidden layers achieve deep nonlinear mapping of features through activation functions. The scenario adaptation layer has built-in weight matrices for different financial scenarios. For example, the asset allocation consultation scenario emphasizes the accuracy weight of voice response parameters, while the transaction instruction issuance scenario emphasizes the real-time weight of touch feedback parameters. The output layer generates vectors for three types of control parameters: visual display, voice response, and touch feedback, ensuring that the output parameters match the core control needs of financial interaction one by one.

[0102] The second step is training data preparation, collecting a large number of real-world interaction samples from financial scenarios, covering typical scenarios such as robo-advisory asset allocation consultation, offline account opening, and instruction issuance on financial trading dashboards. Each sample contains a complete correspondence between multimodal fusion features, interaction state compensation coefficients, and optimal multimedia interaction control parameters. The optimal control parameters are labeled by financial interaction experts based on scenario requirements and user experience feedback. For example, in the asset allocation consultation scenario, when the user interaction state compensation coefficient... K When the value is in the range of 0.7 to 1 (unstable state, ambiguous intent), the optimal parameters for annotation are to slow down the speech rate by 30%, repeat the key guidance once, and increase the brightness of the prompt box in the core area of ​​the interface by 20%, so as to ensure that the training data fits the security needs of the financial scenario and the user's interaction habits.

[0103] The third step is model training and optimization. First, the labeled training data is divided into training, validation, and test sets in a 7:2:1 ratio. Mini-batch gradient descent is used to iteratively train the model. During training, the core loss objective is the sum of the absolute deviations between the model's output control parameters and the expert-annotated optimal parameters. Simultaneously, financial safety constraints are introduced, such as ensuring that the deviation of control parameters in transaction-related scenarios does not exceed a preset safety threshold. The connection weights of each hidden layer and the scene weight matrix of the scene-adaptive weight layer are continuously adjusted through backpropagation. Early stopping and L2 regularization are also introduced during training. Training is stopped when the loss value on the validation set does not decrease for 10 consecutive iterations to avoid overfitting. L2 regularization is used to suppress parameter oscillations caused by excessive weights. After training, the generalization ability is verified using a test set containing new scenario samples that were not included in the training, such as product consultation scenarios on online wealth management platforms. For deviations that occur during verification, such as excessively high voice response delays for users under high stress and unclear touch feedback for users with low proficiency, the parameters of the scene adaptive weight layer are fine-tuned to optimize the model. Finally, a pre-trained parameter mapping model with strong generalization ability, adaptability to multiple financial scenarios, and compliance with safety constraints is obtained.

[0104] Finally, the associated and bound multimodal fusion features and interaction state compensation coefficients are input into the pre-trained model. The model first loads the weight matrix of the corresponding scene based on the scene recognition results, and then strengthens the feature dimensions that are strongly related to user operation security and operation accuracy through a financial security-oriented feature attention mechanism. For example, in the transaction confirmation scenario, the tactile pressure distribution feature is strengthened, and in the risk warning scenario, the visual gaze feature is strengthened. The input data is then collaboratively analyzed and deeply mapped. During the mapping process, the parameter generation log is output simultaneously to record the basis for the generation of each control parameter, such as the 30% slowdown in speech rate due to the compensation coefficient. K=0.85+asset allocation scenario weight, ultimately generating a set of multimedia interactive control parameters including three categories: visual display, voice response, and touch feedback. Among them, the visual display parameters specify the specific location coordinate range of the interface prompts, the brightness gradient level, and the contrast adjustment threshold; the voice response parameters specify the speech rate level, volume level, frequency of repetition of key information, and timing of broadcasting; and the touch feedback parameters specify the vibration intensity level, response delay threshold, and sensitivity adjustment coefficient. This ensures that the generated parameters accurately match the user's actual interaction intent and adapt to the current scenario's security level and operating rhythm, thus solving the problem of inherent deviations between control parameters and user state and scenario requirements in the background technology from the parameter generation level.

[0105] Step 4.3: Based on the multimedia interaction control parameters, drive the multimedia interaction control engine to execute the corresponding multimedia interaction actions. Specifically, this includes: The multimedia interaction control engine is an integrated core execution unit designed specifically for multimodal interaction in financial scenarios. It has a built-in parameter parsing module, a multimodal instruction scheduling module, an execution drive module, and a status monitoring and feedback module. Its core advantage lies in its ability to achieve synchronous and coordinated execution of multimodal control instructions, real-time calibration of action accuracy, and seamless integration with subsequent feedback correction mechanisms. It can solve the core problem of disconnect between interactive actions and user status and scenario requirements, and the cumulative deviation affecting user operation safety from the execution level. In specific implementation, the multimedia interaction control parameter set is first transmitted to the engine, and the parameter parsing module first structures the parameter set. The parsing and security verification process involves breaking down control commands into subsets based on three main modalities: visual, voice, and touch. The execution priority of each command is confirmed; for example, in a financial transaction confirmation scenario, risk warning voice broadcasts and confirmation button highlighting have higher priority than adjustments to ordinary interface elements. Security verification checks whether each parameter is within preset financial security thresholds, such as ensuring the voice broadcast volume is not lower than the security warning threshold and the touch response delay does not exceed the operational safety limit, avoiding execution deviations caused by abnormal parameters. After successful verification, the multimodal command scheduling module completes the timing coordination and orchestration of commands, ensuring that interactive actions of different modalities are triggered synchronously without timing misalignment. For example, when a user presses the transaction confirmation button, the touch vibration feedback and confirmation voice broadcast are activated simultaneously, avoiding delays that could lead to user misjudgment.

[0106] Subsequently, the execution driver module drives the corresponding hardware modules to perform interactive actions: For visual control commands, the display module is driven to execute actions precisely. For example, in the intelligent investment advisory asset allocation scenario, the configuration option area prompt box that the user is looking at is adjusted to a preset brightness gradient according to the parameter command, and the distinction between the risk warning text and the background is enhanced according to the specified contrast threshold. At the same time, it is ensured that the position of the prompt box accurately matches the coordinate range of the user's pupil gaze to avoid visual guidance misalignment. For voice control commands, the audio module is driven to execute broadcast actions according to preset speech speed levels and volume levels. For example, in the financial transaction large screen scenario, when the user's state is unstable (compensation coefficient K>0.7), the speech speed is slowed down according to the parameter command, and the key safety prompt of "please check the operation before transaction confirmation" is repeated once. The timing of the broadcast is precisely connected to the user's operation pause node to avoid interfering with the user's thinking. For touch control commands, the touch feedback module is driven to execute feedback actions according to vibration intensity levels and response delay thresholds. For example, in the offline counter business handling scenario, the touch vibration feedback intensity is enhanced for users with low proficiency according to the parameter command, while the response delay is shortened, so that the user can clearly perceive the effectiveness of the operation and reduce the probability of misoperation.

[0107] During action execution, the status monitoring and feedback module collects execution data from each hardware module in real time, including the actual brightness / position of visual elements, the actual speech rate / volume of voice broadcasts, and the actual force / delay of touch feedback. The collected data is compared with the set control parameters in real time. If a deviation is found, such as the brightness of the visual prompt box not reaching the preset threshold or the voice broadcast delay exceeding 100 milliseconds, dynamic calibration is immediately performed through the execution drive module. At the same time, complete execution status data (including action parameters, deviation status, and execution sequence) is stored, and a data interface is reserved to connect with subsequent secondary interaction feedback mechanisms. This ensures that if interaction deviations are found later, the source of the problem can be traced based on the data and the control parameters can be adjusted. Ultimately, this achieves accurate, coordinated, and secure execution of multimodal interactive actions in financial scenarios.

[0108] In a preferred embodiment of the present invention, step 5 above may include:

[0109] Step 5.1: During the execution of multimedia interactive actions, secondary multimodal raw feedback data generated by the user is collected through the adaptive perception field. Specifically, this includes: throughout the entire execution of multimedia interactive actions, real-time data collection is carried out based on the constructed adaptive perception field. This perception field covers the core interactive areas of financial scenarios such as intelligent investment advisory interactive terminals, offline intelligent counters of financial platforms, and large screens for financial transactions, and simultaneously captures secondary multimodal raw feedback data generated by the user in real time. The raw data collected includes, in the visual dimension, changes in the user's pupil gaze position, facial micro-expression fluctuations, and gaze duration; in the voice dimension, user response statements, tone fluctuations, and operational questioning voices; and in the touch dimension, user pressure adjustment, changes in operation rhythm, and repeated touch trajectories. This ensures that the collected data can fully reflect the user's true feedback state to the current multimedia interactive action, providing accurate raw data support for subsequent deviation calculation.

[0110] Step 5.2 involves preprocessing and spatiotemporally coupling the user-generated secondary multimodal raw feedback data to generate secondary multimodal fusion features. Specifically, this includes: first, performing preprocessing on the collected secondary multimodal raw feedback data; at the visual data level, removing invalid frames caused by ambient light interference and retaining clear pupil and facial feature data; at the speech data level, filtering background noise, selecting effective user feedback speech segments for interactive actions, and removing irrelevant environmental noise; at the touch data level, removing interference data caused by accidental touches and hovering touches, and retaining user-initiated touch operation data; after preprocessing, the following steps are followed... The spatiotemporal coupling processing logic in section 2.1 undergoes standardized processing. In terms of timing, a globally unified time benchmark is used as a reference to calibrate the secondary feedback data of different modalities to the same time node, eliminating timing deviations caused by acquisition and transmission delays. Spatially, based on preset three-dimensional mapping rules, the feedback data of each modality is uniformly mapped to the financial interaction-specific coordinate system to ensure accurate matching between the data and the functional areas of the financial interface. Finally, through feature association and integration, the spatiotemporally calibrated multimodal feedback data is fused into a unified secondary multimodal fusion feature. This feature can accurately represent the user's current actual interaction intent, providing a reliable feature basis for deviation judgment.

[0111] Step 5.3 involves inputting the secondary multimodal fusion features and the multimedia interaction control parameters into a predefined deviation calculation function. Specifically, this includes: first, performing data association alignment and validity verification; binding the secondary multimodal fusion features and the output multimedia interaction control parameters frame-by-frame according to a globally unified timestamp to ensure that each set of secondary fusion features accurately corresponds to the control parameters that trigger the feedback; simultaneously verifying timestamp matching errors during the binding process; if the error exceeds 100 milliseconds, it is determined as an association failure, and the feature and parameter data of the corresponding interaction stage are re-traceable and supplemented; simultaneously, performing format adaptation processing on the two types of bound data, establishing a correspondence between each dimension (visual, voice, touch) of the secondary multimodal fusion features and each module (visual display, voice response, touch feedback) of the multimedia interaction control parameters, for example, linking visual gaze features with visual... The system demonstrates the correlation between parameters, voice tone features, and voice response parameters to ensure accurate data dimension matching. Subsequently, the two types of data, after correlation and adaptation, are input into a predefined deviation calculation function. This function is designed to address the closed-loop control requirements of the financial scenarios described in this invention. Its core design addresses the core deficiency in the background technology of lacking closed-loop quantitative control logic, incorporating financial operation safety standards (such as parameter deviation thresholds for transaction-related interactions) and thresholds for the interaction habits of a vast number of financial users, such as the acceptable voice speed and touch feedback intensity range for users in different scenarios. The function has a built-in data parsing module that can automatically identify the actual user needs represented by the secondary multimodal fusion features, as well as the preset interaction targets corresponding to the control parameters. Its input logic strictly adapts to the multimodal interaction architecture of this invention, and through one-to-one correspondence analysis of the two types of data dimensions, it lays the foundation for subsequent precise quantitative control of deviations.

[0112] Step 5.4: Using a predefined deviation calculation function, output the control deviation of the multimedia interaction control parameters relative to the current user's actual interaction intent. Specifically, this includes: first, analyzing the user's actual interaction intent represented by the secondary multimodal fusion features using the predefined deviation calculation function; extracting the core actual features of the three dimensions of vision, voice, and touch; then, extracting the corresponding preset features of each dimension in the multimedia interaction control parameters; and finally, calculating the control deviation using a simple weighted summation logic. Specifically, the mathematical expression of the predefined deviation calculation function can be described as:

[0113]

[0114] in, E The total control deviation represents the final output and is used to comprehensively characterize the overall degree of deviation between the multimedia interactive control parameters and the user's actual interactive intention. The value ranges from 0 to 1, and the larger the value, the more significant the deviation. W 1. W 2. W3 represents the deviation weights for visual, voice, and touch dimensions, calibrated based on security requirements and interaction priorities in financial scenarios. For example, the touch dimension weight in a transaction confirmation scenario. W 3=0.4, Voice Dimension Weight W 2=0.4, visual dimension weight W 1 = 0.2, the weight of the voice dimension in asset allocation consulting scenarios. W 2=0.4, visual dimension weight W 1=0.3, Touch dimension weight W 3 = 0.3, and satisfies W 1+ W 2+ W 3 = 1;

[0115] F v_real These are the actual feature values ​​of the visual dimension in the secondary multimodal fusion features, such as the standardized values ​​of pupil gaze duration and gaze focus clarity. F v_pre These are preset feature values ​​for the visual dimension in multimedia interactive control parameters, such as the expected user gaze duration and focus clarity threshold. F a_real These are the actual feature values ​​of the speech dimension in the secondary multimodal fusion features, such as the standardized values ​​of the user's speech rate and intonation fluctuations in the speech. F a_pre These are preset feature values ​​for the voice dimension in multimedia interactive control parameters, such as the user-adapted response speech rate corresponding to the broadcast speech rate and the tone stability threshold. F t_real These are the actual feature values ​​of the touch dimension in the secondary multimodal fusion features, such as the standardized values ​​of user pressure and operation interval duration. F t_pre The preset feature values ​​for the touch dimension in the multimedia interaction control parameters are as follows: such as the user-adapted pressing pressure corresponding to the touch feedback force and the operation rhythm threshold; the symbol || represents taking the absolute value, which is used to ensure that the deviation of each dimension is non-negative and accurately characterizes the degree of deviation; finally, the total deviation is obtained by weighted summation and integration of the deviations of each dimension.

[0116] In the specific calculation process, the difference between the actual feature value and the preset feature value of each dimension is first calculated, and the absolute value is taken to obtain the single-dimensional deviation (such as visual dimension deviation). F v_real - F v_pre | Voice dimension deviation | F a_real -F a_pre |), then multiply by the corresponding dimension weights respectively, and finally sum the weighted biases of the three dimensions to obtain the total control bias.E ; To illustrate with an example from a financial scenario: In an asset allocation consulting scenario, if the actual feature value of the voice dimension in the secondary fusion features... F a_real =0.6, corresponding to the user's hurried response due to speaking too fast and not being heard clearly; the preset feature value of the speech dimension in the control parameters. F a_pre =0.3, corresponding to the system's desired clear user response state, voice dimension weight. W If 2 = 0.4, then the weighted bias for the voice dimension is 0.4 × |0.6 - 0.3| = 0.12; if the weighted bias for the visual dimension is 0.09 and the weighted bias for the touch dimension is 0.06, then the total control bias is... E =0.09+0.12+0.06=0.27, indicating that there is a slight deviation between the current control parameters and the user's actual intention.

[0117] The final output total control deviation E It will clearly distinguish the direction of deviation (judged by the positive or negative difference between the actual and preset feature values ​​in each dimension, such as...) F a_real > F a_pre This indicates that the voice response was too hasty, corresponding to an excessively fast speaking speed, and the degree of deviation (through...). E The numerical value judgment provides a precise quantitative basis for subsequent adjustment of control parameters and optimization of interaction strategies. It solves the core defects of existing technologies, such as the lack of quantitative control logic for deviation and the continuous accumulation of deviation, from the closed-loop feedback level, and ensures the accuracy and security of financial scenario interaction.

[0118] In a preferred embodiment of the present invention, step 6 above may include:

[0119] Step 6.1: Obtain the control deviation output by the predefined deviation calculation function. This includes: first, accurately retrieving the control deviation output by the predefined deviation calculation function, and then associating the corresponding financial interaction scenario type and specific interaction stage with the deviation, such as the parameter broadcasting stage in the asset allocation consultation scenario of a smart investment advisory terminal, or the touch operation stage in the transaction confirmation scenario of a financial trading screen; then, verifying the validity of the obtained control deviation, checking whether the deviation is within a reasonable range of 0 to 1, and whether there are any abnormal extreme values ​​caused by data transmission interference or feature matching errors. If any abnormality is found, the deviation calculation process in step 5.4 is re-triggered. If the verification passes, the deviation and the corresponding scenario and stage information are organized into a standardized data package to provide accurate and reliable quantitative basis for subsequent strategy optimization.

[0120] Step 6.2 involves inputting the acquired control deviation into the adaptive strategy update model. Specifically, this includes inputting the control deviation data package, after validity verification, into the adaptive strategy update model. Simultaneously, the model reads auxiliary information such as the security level of the current financial interaction scenario and the user's operational proficiency. This information serves as the core basis for adjusting optimization weights, ensuring that subsequent optimization strategies are deeply adapted to scenario requirements and user characteristics. The adaptive strategy update model used in this study originates from the Q-Learning reinforcement learning architecture. Addressing the core requirements of this invention—high security constraints, multi-scenario adaptation, and closed-loop control in financial scenarios—it adds a financial scenario constraint layer and a multi-dimensional strategy collaborative optimization module to the traditional Q-Learning architecture. Its core advantage lies in its ability to actively learn the dynamic mapping relationship between control deviation and optimization strategy, enabling autonomous iteration of deviation feedback, strategy adjustment, and effect verification. Furthermore, it integrates financial security thresholds into the strategy update process, preventing optimized rules from causing user operational security risks and completely resolving the core defects of static and unchanging interactive control strategies in the background technology, which cannot be dynamically optimized based on deviations.

[0121] The specific construction and training process of this model fully conforms to the financial interaction closed-loop control logic of this invention, and revolves around eliminating deviations, adapting to scenarios, and ensuring security throughout: First, the model architecture is constructed, starting with a basic reinforcement learning framework based on Q-Learning. The input layer is designed as a multi-dimensional fusion structure, which can simultaneously receive core data such as the amount of control deviation, the current financial scenario type (such as intelligent investment advisory asset allocation consultation, financial transaction screen transaction confirmation), scenario security level, and user operation proficiency, ensuring that the input information completely supports strategy optimization decisions; The middle layer adds two core modules: a financial scenario constraint module and a multi-strategy collaborative optimization module. The scenario constraint module has built-in security thresholds for different financial scenarios. For example, the strategy adjustment range in transaction scenarios must not exceed the preset security range to avoid touching the security red line during the optimization process. The multi-strategy collaborative optimization module is specifically adapted to the joint optimization needs of parameter mapping model mapping rules and interaction state compensation coefficient generation rules to prevent adaptation conflicts after the optimization of the two types of rules; The output layer is designed as a dual-output structure, corresponding to the optimization parameter adjustment instructions of the two sets of rules respectively, ensuring that the optimization direction accurately points to the source of deviation.

[0122] Secondly, training data preparation involved collecting a large number of closed-loop interaction samples, focusing on typical interaction scenarios in the financial industry and common deviation types (voice speed deviation, touch pressure deviation, and visual cue deviation), covering scenarios with different security levels and user groups with different levels of operational proficiency. For each sample, a complete correspondence was constructed between the control deviation amount, current scenario attributes, user characteristics, and the optimal strategy optimization scheme. The optimal strategy optimization scheme was labeled by financial interaction experts using a dual standard of deviation correction effect and security constraint requirements. For example, in the financial transaction confirmation scenario, when the control deviation amount... EWhen the deviation is greater than 0.6 and the deviation mainly originates from the touch dimension, the optimal solution recommended by experts is to increase the weight of the touch dimension in the compensation coefficient generation rule by 0.1, correct the mapping logic of touch feedback parameters in the parameter mapping model, and clean the collected sample data to remove invalid samples caused by sensor interference and data transmission errors, and supplement missing samples in extreme scenarios (such as interaction deviations under high tension during market fluctuations) to improve the completeness and representativeness of the data.

[0123] Next comes model training and optimization. The labeled training data is first divided into training, validation, and test sets in a 7:2:1 ratio. Time-series batch training is used to iteratively train the model. During training, the reduction in control deviation corresponding to the optimized rule is used as the core reward objective, and compliance with financial scenario security constraints is used as the penalty condition. Through a reinforcement learning trial-and-error iterative mechanism, the model autonomously learns the optimal strategy optimization logic under different input conditions. Simultaneously, an early stopping mechanism and L1 regularization are introduced. Training stops when the deviation reduction on the validation set shows no improvement for eight consecutive iterations, preventing... Model overfitting is addressed by using L1 regularization to suppress policy instability caused by excessive fluctuations in optimization parameters. After training, the generalization ability is validated using a test set (including new scenario samples such as risk assessment of online wealth management platforms and offline smart counter business processing of financial platforms that were not included in the training). To address the inaccurate optimization issues that arise during validation, such as insufficient policy adjustment range in scenarios with low proficiency and excessively conservative optimization in high-security scenarios, the parameters of the scenario constraint module are fine-tuned for optimization. Ultimately, an adaptive policy update model with strong generalization ability, adaptability to multiple financial scenarios, and compliance with security constraints is obtained.

[0124] After receiving input data, the model first locates the optimal optimization dimension through scene attributes and user features. Then, based on the mapping logic learned by reinforcement learning, it initiates a joint optimization process for the parameter mapping model mapping rules and the interaction state compensation coefficient generation rules to ensure that the optimization direction is accurate and the process is safe and controllable.

[0125] Step 6.3: Update the model using an adaptive strategy. Based on the control deviation, jointly optimize the mapping rules of the parameter mapping model and the generation rules of the interaction state compensation coefficients to obtain the updated parameter mapping model and the updated interaction state compensation coefficient generation rules. Specifically, this includes: After receiving data, the adaptive strategy update model initiates a joint optimization process for the mapping rules of the parameter mapping model and the generation rules of the interaction state compensation coefficients to ensure that the two types of rules are coordinated and adapted to eliminate deviations; Optimize the mapping rules of the parameter mapping model: First, locate the main source dimension of the deviation. For example, if the voice dimension accounts for the highest proportion of the deviation, it indicates that there is an adaptation problem in the mapping logic of the voice response parameters in the current mapping rule; Then, adjust the weight allocation of each modal feature in the mapping rule in combination with the current scenario type. For example, in the asset allocation consultation scenario, if the total deviation is too high due to the deviation in the speech rate of the voice broadcast, the weight allocation of each modal feature in the mapping rule will be adjusted. The larger the deviation, the greater the weight of the voice modality features in the mapping rules. Simultaneously, the generation logic of the control parameters for this dimension is corrected, making the subsequently generated voice response parameters more closely match the user's actual acceptance ability. For the optimization of the interaction state compensation coefficient generation rules: the weight allocation and calculation threshold in the coefficient generation process are adjusted based on the control deviation. For example, if the touch dimension deviation is significant, the weight of the tactile modality spatiotemporal coupling features in the compensation coefficient calculation is increased, while the interval division threshold of the compensation coefficient is corrected, enabling the compensation coefficient to more accurately represent the stability and clarity of the user's interaction state. During the optimization process, the synergy between the two types of rules is ensured simultaneously. For example, after adjusting the weight of the voice dimension in the mapping rules, the weight of voice-related features in the compensation coefficient generation rules is correspondingly corrected to avoid adaptation conflicts between rules. Finally, an updated parameter mapping model and an updated interaction state compensation coefficient generation rule are obtained.

[0126] Step 6.4: Dynamically inject the updated parameter mapping model into the multimedia interaction control parameter generation process of the next interaction cycle. Simultaneously, dynamically inject the updated interaction state compensation coefficient generation rule into the feature reference point determination and compensation coefficient calculation process of the next interaction cycle, completing the human-machine feedback closed loop. Specifically, this includes: first, validating the updated parameter mapping model and interaction state compensation coefficient generation rule by simulating interaction data from the current financial interaction scenario to test whether the control parameters and compensation coefficients generated by the updated rule can effectively reduce the control deviation. If the verification passes, the dynamic injection operation is performed; then, the updated parameter mapping model dynamically replaces the original model and is injected into the next interaction cycle. The multimedia interactive control parameter generation process ensures that the control parameters generated in the next cycle have been integrated into the deviation correction logic. At the same time, the updated interactive state compensation coefficient generation rules are dynamically injected into the feature reference point determination and compensation coefficient calculation process of the next interactive cycle, so that the subsequent compensation coefficient generation is more accurate and adapts to the user's interactive state. After the injection is completed, the interactive data and deviation changes of the next interactive cycle are monitored in real time, the effect of the rule update is recorded, and data interfaces are reserved for possible multiple rounds of optimization in the future. Finally, the human-machine feedback closed loop is completed, which fundamentally solves the core defects of static and unchanging interactive control strategies and continuous accumulation of deviations in the background technology, and realizes adaptive closed-loop control of multimodal interaction in financial scenarios.

[0127] like Figure 2 As shown, embodiments of the present invention also provide a multimedia interactive control system based on human-machine feedback closed loop, including:

[0128] The data acquisition and preprocessing module is used to acquire multimodal raw feedback data from users in the interactive environment in real time; and to preprocess the multimodal raw feedback data to obtain preprocessed multimodal feedback data.

[0129] The data coupling and reference point determination module is used to perform spatiotemporal coupling processing on the preprocessed multimodal feedback data to generate multimodal fusion features. In the feature space corresponding to the multimodal fusion features, multiple feature reference points are dynamically determined based on feedback data of different modalities. The feature reference points include at least the coordinate point of the user's pupil center in the visual feedback data, the fundamental frequency energy peak point in the voice feedback data, and the pressure distribution centroid point in the tactile feedback data.

[0130] The dynamic construction and calculation module is used to construct a dynamic feature geometric envelope surface in the feature space based on multiple feature reference points; calculate the volume change rate and centroid offset vector of the feature geometric envelope surface at continuous interaction times; and generate interaction state compensation coefficients based on the volume change rate and centroid offset vector.

[0131] The parameter determination and action execution module is used to determine the multimedia interaction control parameters based on the multimodal fusion characteristics and interaction state compensation coefficients; and to execute the multimedia interaction actions based on the multimedia interaction control parameters.

[0132] The control deviation generation module is used to obtain secondary interaction feedback data from the user after the multimedia interactive action is executed; and to calculate the control deviation of the multimedia interactive control parameters based on the secondary interaction feedback data.

[0133] The human-machine feedback closed-loop module is used to update the interactive control strategy based on the control deviation. By adopting the updated interactive control strategy, the feature reference point determination and interactive state compensation coefficient generation in the next interaction cycle are adjusted to complete the human-machine feedback closed loop.

[0134] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A multimedia interactive control method based on human-machine feedback closed loop, characterized in that, The method includes: Step 1: Acquire the user's raw multimodal feedback data in the interactive environment in real time; preprocess the raw multimodal feedback data to obtain preprocessed multimodal feedback data. Step 2: Perform spatiotemporal coupling processing on the preprocessed multimodal feedback data to generate multimodal fusion features; in the feature space corresponding to the multimodal fusion features, dynamically determine multiple feature reference points based on feedback data of different modalities; the feature reference points include at least the coordinate point of the user's pupil center in the visual feedback data, the fundamental frequency energy peak point in the voice feedback data, and the pressure distribution centroid point in the tactile feedback data. Step 3: Construct a dynamic feature geometric envelope in the feature space based on multiple feature reference points; calculate the volume change rate and centroid offset vector of the feature geometric envelope at continuous interaction moments; generate interaction state compensation coefficients based on the volume change rate and centroid offset vector, including: constructing a dynamically changing feature geometric envelope in the high-dimensional feature space with spatial visual focus, acoustic feature hub, and mechanical action center as vertices; obtaining the spatial morphology sequence of the dynamically changing feature geometric envelope at multiple consecutive interaction moments; calculating the volume change rate of the dynamically changing feature geometric envelope over time based on the spatial morphology sequence, as the volume change rate; simultaneously calculating the displacement vector of the centroid of the dynamically changing feature geometric envelope at continuous interaction moments, as the centroid offset vector; normalizing and weighted fusion processing the volume change rate and centroid offset vector to generate interaction state compensation coefficients characterizing the stability and intent clarity of the user interaction state. Step 4: Determine the multimedia interaction control parameters based on the multimodal fusion features and interaction state compensation coefficients; execute the multimedia interaction actions according to the multimedia interaction control parameters. Step 5: After the multimedia interactive action is executed, obtain the user's secondary interactive feedback data; based on the secondary interactive feedback data, calculate the control deviation of the multimedia interactive control parameters. Step 6: Update the interactive control strategy according to the control deviation amount. By adopting the updated interactive control strategy, adjust the determination of feature reference points and the generation of interactive state compensation coefficients in the next interaction cycle to complete the human-machine feedback closed loop.

2. The multimedia interactive control method based on human-machine feedback closed loop according to claim 1, characterized in that, Real-time acquisition of raw multimodal feedback data from users in the interactive environment; preprocessing of the raw multimodal feedback data to obtain preprocessed multimodal feedback data, including: Construct a user-centric adaptive perception field; the adaptive perception field is a dynamic, multi-dimensional data perception space, and the spatial configuration and perception granularity of the adaptive perception field change in response to the user's real-time interaction status. Within the constructed adaptive perception field, the user's behavioral focus area and the state interference factors of the interactive environment are monitored and analyzed in real time to generate synchronous analysis results of the environment and user state. Among them, the behavioral focus area includes the user's gaze area, the voice source localization area, and the core area of ​​limb contact; the state interference factors include ambient light intensity, background noise spectrum, and dynamic parameters of irrelevant moving objects. Based on the synchronous analysis results of the environment and user status, an adaptive data acquisition strategy is generated. The adaptive data acquisition strategy is used to dynamically configure the spatial directivity, sampling frequency and data confidence weight of each modal sensor in the adaptive sensing field. Based on an adaptive data acquisition strategy, the adaptive sensing field is driven to collaboratively acquire the user's multimodal interaction behavior to obtain multimodal raw feedback data that matches the user's state and environment. The original multimodal feedback data is standardized by performing time stamp alignment, feature scale normalization, and outlier filtering to obtain preprocessed multimodal feedback data.

3. The multimedia interactive control method based on human-machine feedback closed loop according to claim 2, characterized in that, The preprocessed multimodal feedback data is subjected to spatiotemporal coupling processing to generate multimodal fusion features. In the feature space corresponding to the multimodal fusion features, multiple feature reference points are dynamically determined based on feedback data from different modalities. Reference points include at least the coordinates of the user's pupil center in visual feedback data, the peak point of the fundamental frequency energy in voice feedback data, and the centroid of pressure distribution in tactile feedback data, including: The preprocessed multimodal feedback data is calibrated and aligned in the temporal domain on a unified time reference, and spatial registration is performed based on a predefined spatial mapping relationship to generate spatiotemporal coupling features with spatiotemporal consistency, including spatiotemporal coupling features of visual modality, spatiotemporal coupling features of speech modality and spatiotemporal coupling features of tactile modality. Multi-level feature fusion is performed on various spatiotemporal coupled features to generate a unified multimodal fusion feature that represents the user's interaction intent; Map multimodal fusion features to a high-dimensional feature space; Based on the spatiotemporal coupling features of each modality in the high-dimensional feature space, at least three feature reference points are dynamically determined. Specifically, based on the spatiotemporal coupling features of the visual modality, the projection coordinates of the user's pupil center in the high-dimensional feature space are determined as the spatial visual focus; based on the spatiotemporal coupling features of the speech modality, the mapping coordinates of the speech fundamental frequency energy peak in the high-dimensional feature space are determined as the acoustic feature hub; and based on the spatiotemporal coupling features of the tactile modality, the mapping coordinates of the pressure distribution centroid in the high-dimensional feature space are determined as the mechanical action center.

4. The multimedia interactive control method based on human-machine feedback closed loop according to claim 3, characterized in that, Based on multimodal fusion features and interaction state compensation coefficients, the control parameters for multimedia interaction are determined. Based on the multimedia interactive control parameters, execute multimedia interactive actions, including: Obtain the unified multimodal fusion features representing user interaction intent, and the interaction state compensation coefficients representing the stability of user interaction state and the clarity of intent; Based on the current interaction scenario, the multimodal fusion features and the interaction state compensation coefficients are input into a pre-trained parameter mapping model to generate multimedia interaction control parameters. Based on the multimedia interaction control parameters, the multimedia interaction control engine is driven to execute the corresponding multimedia interaction actions.

5. The multimedia interactive control method based on human-machine feedback closed loop according to claim 4, characterized in that, After the multimedia interactive action is executed, secondary interaction feedback data from the user is acquired; based on the secondary interaction feedback data, the control deviation of the multimedia interaction control parameters is calculated, including: During the execution of multimedia interactive actions, the user-generated secondary multimodal raw feedback data is collected through the adaptive sensing field; The user-generated secondary multimodal raw feedback data is preprocessed and spatiotemporally coupled to generate secondary multimodal fusion features; The secondary multimodal fusion features and the multimedia interaction control parameters are input together into a predefined deviation calculation function; The control deviation of the multimedia interaction control parameters relative to the current user's actual interaction intention is output through a predefined deviation calculation function.

6. The multimedia interactive control method based on human-machine feedback closed loop according to claim 5, characterized in that, Based on the control deviation, the interactive control strategy is updated. By adopting the updated interactive control strategy, the determination of feature reference points and the generation of interactive state compensation coefficients in the next interaction cycle are adjusted to complete the human-machine feedback closed loop, including: Obtain the control deviation output by the predefined deviation calculation function; The acquired control deviation is input into the adaptive strategy to update the model; The model is updated by an adaptive strategy. Based on the control deviation, the mapping rule of the parameter mapping model and the generation rule of the interaction state compensation coefficient are jointly optimized to obtain the updated parameter mapping model and the updated generation rule of the interaction state compensation coefficient. The updated parameter mapping model is dynamically injected into the multimedia interaction control parameter generation process of the next interaction cycle. At the same time, the updated interaction state compensation coefficient generation rules are dynamically injected into the feature reference point determination and compensation coefficient calculation process of the next interaction cycle, thus completing the human-computer feedback closed loop.

7. A multimedia interactive control system based on human-machine feedback closed loop, wherein the system implements the method as described in any one of claims 1 to 6, characterized in that, include: The data acquisition and preprocessing module is used to acquire multimodal raw feedback data from users in the interactive environment in real time; and to preprocess the multimodal raw feedback data to obtain preprocessed multimodal feedback data. The data coupling and reference point determination module is used to perform spatiotemporal coupling processing on the preprocessed multimodal feedback data to generate multimodal fusion features. In the feature space corresponding to the multimodal fusion features, multiple feature reference points are dynamically determined based on feedback data of different modalities. The feature reference points include at least the coordinate point of the user's pupil center in the visual feedback data, the fundamental frequency energy peak point in the voice feedback data, and the pressure distribution centroid point in the tactile feedback data. The dynamic construction and calculation module is used to construct a dynamic feature geometric envelope surface in the feature space based on multiple feature reference points; calculate the volume change rate and centroid offset vector of the feature geometric envelope surface at continuous interaction times; and generate interaction state compensation coefficients based on the volume change rate and centroid offset vector. The parameter determination and action execution module is used to determine the multimedia interaction control parameters based on the multimodal fusion characteristics and interaction state compensation coefficients; and to execute the multimedia interaction actions based on the multimedia interaction control parameters. The control deviation generation module is used to obtain secondary interaction feedback data from the user after the multimedia interactive action is executed; and to calculate the control deviation of the multimedia interactive control parameters based on the secondary interaction feedback data. The human-machine feedback closed-loop module is used to update the interactive control strategy based on the control deviation. By adopting the updated interactive control strategy, the feature reference point determination and interactive state compensation coefficient generation in the next interaction cycle are adjusted to complete the human-machine feedback closed loop.

Citation Information

Patent Citations

  • Transmodal input fusion for a wearable system

    CN112424727A