Multi-modal data analysis method and system based on intelligent interaction
Through multimodal data fusion and dynamic weight allocation technology, the problem of unconscious touch misidentification and multi-person interaction conflict in the guide equipment is solved, and interaction accuracy and system resource utilization efficiency are improved.
Patent Information
- Application Number
- CN202510565602.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In multimodal data analysis, existing guided equipment has problems such as unconscious touch misidentification and response conflicts in multi-person parallel interaction scenarios, resulting in a degradation of interactive experience and system performance.
Through multimodal data fusion, touch screen instructions are used to link video images, combined with deep neural networks and Gaze360 models, a three-dimensional pose model is established, user weight index is dynamically calculated, error touch duration is set, high-value operation instructions are selected, and system resource allocation is optimized.
It improves the interaction accuracy and system resource scheduling efficiency in multi-user scenarios, reduces unconscious touch interference, and ensures the smoothness of core users' operation and system robustness.
Smart Images

Figure CN120469578A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data analysis technology, and in particular to a multimodal data analysis method and system based on intelligent interaction. Background Art
[0002] With the advancement of artificial intelligence and sensing technologies, intelligent interactive systems are transforming traditional human-computer interaction methods, which rely primarily on single-mode input. Single-mode input struggles to fully capture user behavior and environmental information, limiting the interactive experience and system performance. To address this issue, multimodal data analysis technology has recently become a key research direction in intelligent interaction.
[0003] At present, navigation equipment generally uses interactive systems for demonstrations. Through interactive means such as touch screens, users can browse content and obtain detailed information more conveniently. However, there are still some drawbacks in the use of interactive systems: on the one hand, due to the lack of a multi-dimensional verification mechanism for operation intentions, the system cannot effectively distinguish between unconscious touches and valid commands, resulting in a large number of low-complexity gestures or random operations being mistakenly identified as functional operations. On the other hand, in multi-person parallel interaction scenarios, the system adopts a uniform response strategy and does not build a dynamic weight model based on user behavior characteristics. The request priorities of different users cannot be processed differently, causing response conflicts and resource preemption. These defects jointly restrict the service accuracy and user experience of the navigation system. Therefore, at this stage, a more intelligent and efficient multimodal data analysis technology solution is needed to solve the above problems. Summary of the Invention
[0004] The purpose of the present invention is to provide a multimodal data analysis method and system based on intelligent interaction to solve the problems raised in the above background technology.
[0005] In order to solve the above technical problems, the present invention provides a multimodal data analysis method based on intelligent interaction, comprising:
[0006] S100: When the navigation device is activated, it receives a touch screen instruction and captures surrounding video images.
[0007] Touch screen commands refer to the operation commands generated by users by touching the screen of the navigation device, including operation gestures, command locations and keyword libraries.
[0008] Activation of a navigation device refers to the process by which the device begins operating or starts its tour guidance function. When the user uses some method to put the device into the tour guidance or navigation mode, it begins providing audio commentary, graphic displays, or other tour guidance services.
[0009] Gestures refer to the various actions that users perform on the touch screen with their fingers, which are used to interact with the navigation device.
[0010] Common gestures include:
[0011] Touch: Quickly press the screen with your finger, similar to a mouse click, to select or activate an item.
[0012] Double-click: Click twice quickly in succession, usually used to zoom in or select something.
[0013] Long press: Press and hold your finger on the same location for a period of time to call out the menu or enter editing mode.
[0014] Slide: Slide your finger across the screen, often used to scroll pages or drag objects.
[0015] Pinch: Place two fingers on the screen and move them inward or outward to zoom in or out.
[0016] Rotate: Use two fingers to rotate the screen to rotate an image or object.
[0017] The command position refers to the specific coordinates of the touch screen touched by the user.
[0018] The keyword library is a collection of different key fields, which is used to describe the functions that can be achieved after the system executes the corresponding operation instructions.
[0019] The number of key fields in each keyword library varies. The more key fields a keyword library has, the higher the level of functionality implemented by the corresponding operation instructions within the system framework. Each key field corresponds to a description of a basic unit function, and together they constitute the functional components of the corresponding keyword library.
[0020] By analyzing the number of key fields in the keyword library, we can indicate the complexity of the operating instructions and the system hierarchy. The keyword library serves as a mapping tool for the system functional module hierarchy, facilitating functional classification and management.
[0021] By integrating touchscreen commands with real-time video capture, we achieve the simultaneous fusion of multimodal data, aligning user actions with the spatiotemporal alignment of scene visual data. Based on the hierarchical design of the keyword library, we differentiate the functional weights of different commands, laying the foundation for subsequent command priority analysis.
[0022] S200: Identify the target object based on the video image, analyze the touch screen command and match the interactive object. Specifically including:
[0023] S201, using the target recognition algorithm to analyze the people in the video image, setting the people whose distance is less than the threshold Q and whose face is facing the navigation device as the target object, and recording the setting time TD x .
[0024] Use target recognition algorithms to screen valid users, eliminate interference from invalid bystanders, and improve recognition accuracy.
[0025] S202, intercepting the TD of each target object in the video image x The image frames at that time are used to identify the human joints of the target objects in each image frame, and the arm length of each target object is analyzed and calculated.
[0026] A physical parameter model is established through human joint analysis to provide a biometric basis for touch range judgment.
[0027] S203: When receiving the touch screen instruction I0, record the operation time TC z , capture the TC in the video image z Image frame IMG z . Analyze and calculate the matching index of each target object based on the image frame, so as to match the interactive object for the touch screen instruction I0. Specifically including:
[0028] S2031, in the image frame IMG z The joint points of each target object are marked in the image, and a deep neural network is used in combination with reference objects in the image frame to analyze the operating distance between each joint point and the screen of the navigation device.
[0029] A reference object is an object in an image with known size, stable position, and clear attributes. The existence of a reference object enables the program to convert the "pixel size" or "relative size" in the image into "real-world scale," achieving the mapping from a two-dimensional image to three-dimensional physical distance. The specific implementation process includes:
[0030] Image acquisition: The device camera acquires image frames containing joint points and reference objects.
[0031] Joint Detection: Use a human pose estimation network to identify the pixel coordinates of each joint in the image.
[0032] Depth estimation: Use deep neural networks to predict the relative depth map corresponding to image pixels.
[0033] Reference object detection and calibration: Detect the reference object and calculate its pixel size in the image. Once the actual size of the reference object is known, calculate the conversion ratio between pixels and physical units.
[0034] Absolute distance calculation: Multiply the relative depth of the pixel where the joint point is located by the conversion ratio to obtain the actual physical distance from the joint point to the navigation device screen.
[0035] S2032: Establish three-dimensional models of the target object and the navigation device respectively, and construct the relative position of the target object and the navigation device according to the operation distance.
[0036] Combining reference object calibration with 3D model construction, image pixels are mapped to physical distances to quantify user operation accessibility.
[0037] S2033. Mark a target object whose operating distance of at least one shoulder joint is no greater than the corresponding arm length, and use the Gaze360 model to analyze the three-dimensional gaze vector of each marked target object.
[0038] The Gaze360 model is used to calculate the three-dimensional gaze vector and combine it with the intersection position deviation distance to solve the "contactless mismatching" problem in multi-person scenarios.
[0039] S2034: Analyze the intersection angle angle between the three-dimensional gaze vector and the navigation device screen, as well as the intersection position on the navigation device screen, and calculate the deviation distance JL between the command position of the touch screen command I0 and the intersection position.
[0040] S2035. Calculate the matching index P of each marked target object according to the formula, and select the marked target object with the largest matching index as the interactive object of the touch screen instruction I0. The formula is as follows:
[0041]
[0042] Where α is a constant greater than 1, JL max JD is the diagonal length of the navigation device screen. ave is the average operating distance of all shoulder joints, and JD is the arm length.
[0043] S300: Establish an instruction set for each interactive object and calculate a weight index based on the instruction set. Specifically including:
[0044] S301 , matching interactive objects for each received touch screen instruction in turn, establishing an instruction set for each interactive object, and sequentially placing each touch screen instruction into the instruction set of the corresponding interactive object.
[0045] Establish instruction sets by user to track user operation sequences and support user behavior pattern mining.
[0046] S302: Analyze the reception time, operation gesture, and keyword library of each touch screen instruction in the instruction set, and calculate the weight index of each interactive object. The specific calculation steps are as follows:
[0047] S3021. Obtain all touch screen commands in the command set and arrange them in order of receipt time. Set different difficulty coefficients for different operation gestures and count the number of key fields in the keyword library for each touch screen command.
[0048] The difficulty coefficient is a numerical indicator that reflects the difficulty of executing a gesture and the probability of accidental touches. A larger value indicates a more complex gesture and a less likely to cause accidental touches, and the system will be more rigorous in identifying the gesture.
[0049] S3022. Obtain the difficulty coefficient corresponding to the operation gesture of each touch screen instruction and the number of key fields in the keyword library in order of arrangement, and substitute them into the formula to calculate the weight index QZ of the interactive object corresponding to the instruction set:
[0050]
[0051] Where m is the number of all touch screen instructions in the instruction set, CD n-1 CD is the difficulty coefficient corresponding to the n-1th touch screen instruction. ave is the average difficulty coefficient of all operation gestures, GS n-1 is the number of key fields in the n-1 touch screen instruction keyword library, GS n The number of key fields in the keyword library of the nth touch screen instruction.
[0052] The gesture difficulty coefficient and the number of keyword library fields reflect the instruction level and distinguish high-value operations.
[0053] CD ave It is the average difficulty value of all operation gestures, regardless of whether the operation gesture is in the corresponding instruction set. It does not use the average difficulty value of all touch screen instructions in the instruction set.
[0054] As the number of touch screen commands increases, the number of key fields in the corresponding keyword library gradually decreases, which means that the user's intention is becoming clearer and clearer, and a higher weight coefficient is given to avoid interference.
[0055] S400: Set the false touch duration according to the weight index, and filter out the execution instructions based on the weight index. Specifically including:
[0056] S401, dynamically calculate the weight index of each interactive object, set the limit time f, and calculate the false touch time of each interactive object according to the weight index k :
[0057]
[0058] Where, QZ sum is the sum of the weight indices of all interacting objects.
[0059] The accidental touch duration formula enables highly active users to obtain a longer operation protection window and reduce accidental touch coverage conflicts.
[0060] S402: When the interactive object generates an operation instruction, the operation time TC1 is recorded, and the false touch duration time is time starting from the operation time TC1. k The system enters the anti-accidental touch mode and automatically refuses to execute the operation instructions of other interactive objects.
[0061] Block competing instructions with lower weights to ensure smooth operations for core users.
[0062] S403: After the system cancels the accidental touch prevention mode, if it receives operation instructions from two or more interactive objects at the same time, it automatically executes the operation instruction of the interactive object with the largest weight index.
[0063] After the anti-accidental touch mode is released, instantaneous multi-command conflicts are executed with the highest weight to avoid system congestion.
[0064] The present invention also provides a multimodal data analysis system based on intelligent interaction, which includes a data acquisition module, an object recognition module, an operation analysis module and an intelligent interaction module.
[0065] When the navigation device is activated, the data acquisition module receives touch screen instructions and collects surrounding video images.
[0066] The data acquisition module uses touchscreen commands on the navigation device to capture user gestures, command location coordinates, and a keyword library. It also activates the device's camera to capture surrounding video images. Touchscreen interactions trigger device startup, capturing time-stamped operational data and visual information.
[0067] Achieve precise and synchronous collection of multimodal data, and provide structured input for subsequent analysis through hierarchical mapping of system functions through keyword libraries, ensuring data integrity and cross-modal alignment capabilities.
[0068] The object recognition module identifies the target object based on the video image, analyzes the touch screen instructions and matches the interactive object.
[0069] The object recognition module uses a target recognition algorithm and OpenPose joint detection technology to locate the target object, and combines it with a deep neural network to predict the physical distance of the joints. The Gaze360 model analyzes the user's 3D gaze vector, calculates the deviation between the command position and the gaze intersection position, and associates touchscreen commands with interactive objects using a matching index formula.
[0070] Accurately identify valid users, solve the problem of command attribution in multi-person scenarios, use depth information and human posture modeling to improve the mapping accuracy between physical space and screen operations, and reduce the mismatch rate.
[0071] The operation analysis module is used to establish an instruction set for each interactive object and calculate the weight index based on the instruction set.
[0072] The operational analysis module constructs a time-sequential set of commands for each interactive object and dynamically calculates a weighted index based on the difficulty coefficient of the gestures and the number of key fields in the keyword library. The calculation formula combines the ratio of the current command complexity to the average operation difficulty to quantify the user's operation priority.
[0073] By dynamically adjusting the processing logic through difficulty coefficient and function weight, high-value instructions can be identified, system resource allocation can be optimized, and personalized service weight allocation can be achieved.
[0074] The intelligent interaction module sets the false touch duration according to the weight index and selects the execution instructions based on the weight index.
[0075] The intelligent interaction module dynamically generates the duration of false touches based on the weight index, triggering the false touch prevention mode to block competing commands. In the event of a priority conflict, the command with the highest weight index is selected for execution, and a time window mechanism is used to reduce interference.
[0076] It effectively suppresses interference caused by multiple users' accidental touches, improves response accuracy through adaptive delay control and competition arbitration mechanism, ensures smooth operation for core users, and enhances system robustness.
[0077] Compared with the prior art, the present invention has the following beneficial effects:
[0078] Multimodal fusion verification: Through cross-validation of multi-dimensional data of operation gestures and command positions in touch screen commands and video images (such as the intersection of fingertip coordinates and three-dimensional gaze focus), unconscious touches (such as children's random scratches) are effectively filtered out, and the credibility of operation intentions is evaluated by combining eye tracking and biometrics (arm length), eliminating low-value command interference.
[0079] Dynamic weighted hierarchical response: Dynamically calculate the user weight index based on gesture complexity (double-click / pinch > light touch) and functional level (number of key fields mapped to system framework level), establish a priority differentiation management strategy, so that high-frequency / complex operations automatically obtain higher response rights, and avoid operation conflicts caused by uniform processing in multi-person scenarios.
[0080] 3D user binding: This system reconstructs the user's 3D posture model through a deep neural network, calculates the reach of limb manipulation using reference object calibration technology, and analyzes gaze angles using the Gaze360 model. This allows for precise binding of touch commands to the actual user (e.g., responding only to commands where the arm is reachable and the gaze focus matches), thus resolving the identity misassociation problem encountered in traditional image matching.
[0081] Adaptive anti-accidental touch collaboration: Dynamically generates differentiated accidental touch shielding durations based on user weight index, achieving long-term operation protection for high-weight users and efficient filtering of low-weight accidental touches. Combined with the weight arbitration mechanism under timing conflicts, it ensures that system resources are tilted towards core operations and improves service continuity when multiple users are concurrent.
[0082] Functional hierarchy mapping optimization: The number of key fields in the keyword library is used as a quantitative indicator of system functional complexity, and is strongly correlated with the priority of operation instructions. This ensures that users with clear intentions (such as instructions corresponding to functional levels gradually decreasing from general functions to specific functions) receive a higher weight index, thereby optimizing the system resource allocation logic. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0084] Figure 1 This is a flow chart of a multimodal data analysis method based on intelligent interaction according to the present invention;
[0085] Figure 2 It is a structural diagram of a multimodal data analysis system based on intelligent interaction of the present invention. DETAILED DESCRIPTION
[0086] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0087] See also Figure 1 The present invention provides a multimodal data analysis method based on intelligent interaction, comprising:
[0088] S100: When the navigation device is activated, it receives a touch screen instruction and captures surrounding video images.
[0089] Touch screen commands refer to the operation commands generated by users by touching the screen of the navigation device, including operation gestures, command locations and keyword libraries.
[0090] Activation of a tour guide device refers to the process by which the device begins operating or starts its tour guidance function. When a user activates the device through some means (such as pressing a button, scanning a QR code, proximity sensing, or entering a number), the device enters the tour guidance or navigation state and begins providing audio commentary, graphic displays, or other tour guidance services.
[0091] Gestures refer to the various actions that users perform on the touch screen with their fingers, which are used to interact with the navigation device.
[0092] Common gestures include:
[0093] Touch: Quickly press the screen with your finger, similar to a mouse click, to select or activate an item.
[0094] Double-click: Click twice quickly in succession, usually used to zoom in or select something.
[0095] Long press: Press and hold your finger on the same location for a period of time to call out the menu or enter editing mode.
[0096] Slide: Slide your finger across the screen, often used to scroll pages or drag objects.
[0097] Pinch: Place two fingers on the screen and move them inward or outward to zoom in or out.
[0098] Rotate: Use two fingers to rotate the screen to rotate an image or object.
[0099] The command position refers to the specific coordinates of the touch screen touched by the user.
[0100] The keyword library is a collection of different key fields, which is used to describe the functions that can be achieved after the system executes the corresponding operation instructions.
[0101] The number of key fields in each keyword library varies. The more key fields a keyword library has, the higher the level of functionality implemented by the corresponding operation instructions within the system framework. Each key field corresponds to a description of a basic unit function, and together they constitute the functional components of the corresponding keyword library.
[0102] By analyzing the number of key fields in the keyword library, we can indicate the complexity of the operating instructions and the system hierarchy. The keyword library serves as a mapping tool for the system functional module hierarchy, facilitating functional classification and management.
[0103] By linking touchscreen commands (gestures, position) with real-time video capture, we achieve synchronous fusion of multimodal data, aligning user actions with the spatiotemporal alignment of scene visual data. Based on the hierarchical design of the keyword library (where the number of key fields maps to the system's functional hierarchy), we differentiate the functional weights of different operational commands, laying the foundation for subsequent command priority analysis.
[0104] S200: Identify the target object based on the video image, analyze the touch screen command and match the interactive object. Specifically including:
[0105] S201, using the target recognition algorithm to analyze the people in the video image, setting the people whose distance is less than the threshold Q and whose face is facing the navigation device as the target object, and recording the setting time TD x .
[0106] Use target recognition algorithms to screen valid users (distance threshold + facial orientation), eliminate interference from invalid bystanders, and improve recognition accuracy.
[0107] S202, intercepting the TD of each target object in the video image x The image frames at that time are used to identify the human joints of the target objects in each image frame, and the arm length of each target object is analyzed and calculated.
[0108] A physical parameter model is established through human joint analysis (arm length calculation) to provide a biometric basis for touch range judgment.
[0109] S203: When receiving the touch screen instruction I0, record the operation time TC z , capture the TC in the video image z Image frame IMG z . Analyze and calculate the matching index of each target object based on the image frame, so as to match the interactive object for the touch screen instruction I0. Specifically including:
[0110] S2031, in the image frame IMG z The joint points of each target object are marked in the image, and a deep neural network is used in combination with reference objects in the image frame to analyze the operating distance between each joint point and the screen of the navigation device.
[0111] A reference object is an object in an image with known size, stable position, and clear attributes. The existence of a reference object enables the program to convert the "pixel size" or "relative size" in the image into "real-world scale," achieving the mapping from a two-dimensional image to three-dimensional physical distance. The specific implementation process includes:
[0112] Image acquisition: The device camera acquires image frames containing joint points (such as key points of the human body) and reference objects.
[0113] Joint detection: Use a human pose estimation network (such as OpenPose) to identify the pixel coordinates of each joint in the image.
[0114] Depth estimation: Use deep neural networks to predict the relative depth map corresponding to image pixels.
[0115] Reference object detection and calibration: Detect the reference object and calculate its pixel size in the image. Once the actual size of the reference object is known, calculate the conversion ratio between pixels and physical units.
[0116] Absolute distance calculation: Multiply the relative depth of the pixel where the joint point is located by the conversion ratio to obtain the actual physical distance from the joint point to the navigation device screen.
[0117] S2032: Establish three-dimensional models of the target object and the navigation device respectively, and construct the relative position of the target object and the navigation device according to the operation distance.
[0118] Combining reference object calibration with 3D model construction, image pixels are mapped to physical distances to quantify user operation accessibility.
[0119] S2033. Mark a target object whose operating distance of at least one shoulder joint is no greater than the corresponding arm length, and use the Gaze360 model to analyze the three-dimensional gaze vector of each marked target object.
[0120] The Gaze360 model is used to calculate the three-dimensional gaze vector and combine it with the intersection position deviation distance to solve the "contactless mismatching" problem in multi-person scenarios.
[0121] S2034: Analyze the intersection angle angle between the three-dimensional gaze vector and the navigation device screen, as well as the intersection position on the navigation device screen, and calculate the deviation distance JL between the command position of the touch screen command I0 and the intersection position.
[0122] S2035. Calculate the matching index P of each marked target object according to the formula, and select the marked target object with the largest matching index as the interactive object of the touch screen instruction I0. The formula is as follows:
[0123]
[0124] Where α is a constant greater than 1, JL max JD is the diagonal length of the navigation device screen. ave is the average operating distance of all shoulder joints, and JD is the arm length.
[0125] S300: Establish an instruction set for each interactive object and calculate a weight index based on the instruction set. Specifically including:
[0126] S301 , matching interactive objects for each received touch screen instruction in turn, establishing an instruction set for each interactive object, and sequentially placing each touch screen instruction into the instruction set of the corresponding interactive object.
[0127] Establish instruction sets by user to track user operation sequences and support user behavior pattern mining.
[0128] S302: Analyze the reception time, operation gesture, and keyword library of each touch screen instruction in the instruction set, and calculate the weight index of each interactive object. The specific calculation steps are as follows:
[0129] S3021. Obtain all touch screen commands in the command set and arrange them in order of receipt time. Set different difficulty coefficients for different operation gestures and count the number of key fields in the keyword library for each touch screen command.
[0130] The difficulty coefficient is a numerical indicator that reflects the difficulty of executing a gesture and the probability of accidental touches. A larger value indicates a more complex gesture and a less likely to cause accidental touches, and the system will be more rigorous in identifying the gesture.
[0131] S3022. Obtain the difficulty coefficient corresponding to the operation gesture of each touch screen instruction and the number of key fields in the keyword library in order of arrangement, and substitute them into the formula to calculate the weight index QZ of the interactive object corresponding to the instruction set:
[0132]
[0133] Where m is the number of all touch screen instructions in the instruction set, CD n-1 CD is the difficulty coefficient corresponding to the n-1th touch screen instruction. ave is the average difficulty coefficient of all operation gestures, GS n-1 is the number of key fields in the n-1 touch screen instruction keyword library, GS n The number of key fields in the keyword library of the nth touch screen instruction.
[0134] The instruction level is reflected by the gesture difficulty coefficient (double-click > light touch) and the number of keyword library fields to distinguish high-value operations (such as gradually concretizing functions).
[0135] CD ave It is the average difficulty value of all operation gestures, regardless of whether the operation gesture is in the corresponding instruction set. It does not use the average difficulty value of all touch screen instructions in the instruction set.
[0136] As the number of touch screen commands increases, the number of key fields in the corresponding keyword library gradually decreases, which means that the user's intention is becoming clearer and clearer, and a higher weight coefficient is given to avoid interference.
[0137] S400: Set the false touch duration according to the weight index, and filter out the execution instructions based on the weight index. Specifically including:
[0138] S401, dynamically calculate the weight index of each interactive object, set the limit time f, and calculate the false touch time of each interactive object according to the weight index k :
[0139]
[0140] Where, QZ sum is the sum of the weight indices of all interacting objects.
[0141] The accidental touch duration formula (combining the total weight and the maximum duration) enables highly active users to obtain a longer operation protection window and reduce accidental touch coverage conflicts.
[0142] S402: When the interactive object generates an operation instruction, the operation time TC1 is recorded, and the false touch duration time is time starting from the operation time TC1. k The system enters the anti-accidental touch mode and automatically refuses to execute the operation instructions of other interactive objects.
[0143] Block competing instructions with lower weights to ensure smooth operations for core users (such as navigation without interruptions).
[0144] S403: After the system cancels the accidental touch prevention mode, if it receives operation instructions from two or more interactive objects at the same time, it automatically executes the operation instruction of the interactive object with the largest weight index.
[0145] After the anti-accidental touch mode is released, instantaneous multi-command conflicts are executed with the highest weight to avoid system congestion (such as accidental operation by others in the exhibition area).
[0146] The present invention also provides a multimodal data analysis system based on intelligent interaction, which includes a data acquisition module, an object recognition module, an operation analysis module and an intelligent interaction module.
[0147] When the navigation device is activated, the data acquisition module receives touch screen instructions and collects surrounding video images.
[0148] The data acquisition module uses touchscreen commands on the navigation device to capture user gestures (including taps, slides, and pinches), command location coordinates, and a keyword library. It also activates the device's camera to capture surrounding video images. Touchscreen interactions trigger device startup, capturing time-stamped operational data and visual information.
[0149] Achieve precise and synchronous collection of multimodal data (touch commands and video streams), hierarchically map system functions through keyword libraries, provide structured input for subsequent analysis, and ensure data integrity and cross-modal alignment capabilities.
[0150] The object recognition module identifies the target object based on the video image, analyzes the touch screen instructions and matches the interactive object.
[0151] The object recognition module uses a target recognition algorithm and OpenPose joint detection technology to locate the target object, and combines it with a deep neural network to predict the physical distance of the joints. The Gaze360 model analyzes the user's 3D gaze vector, calculates the deviation between the command position and the gaze intersection position, and associates touchscreen commands with interactive objects using a matching index formula (including parameters such as arm length and operating distance).
[0152] Accurately identify valid users (through distance threshold and gaze direction), solve the problem of command attribution in multi-person scenarios, use depth information and human posture modeling to improve the mapping accuracy between physical space and screen operations, and reduce the mismatch rate.
[0153] The operation analysis module is used to establish an instruction set for each interactive object and calculate the weight index based on the instruction set.
[0154] The runtime analysis module constructs a time-sequential set of commands for each interactive object. It dynamically calculates a weight index based on the difficulty of the gesture (e.g., double-tap > light tap) and the number of key fields in the keyword library (reflecting the functional hierarchy). The calculation formula combines the ratio of the current command complexity to the average operation difficulty to quantify the user's operation priority.
[0155] By dynamically adjusting the processing logic through difficulty coefficient and function weight, high-value instructions (such as those being focused on operations or the purpose of operations becoming clearer) are identified, system resource allocation is optimized, and personalized service weight allocation is achieved.
[0156] The intelligent interaction module sets the false touch duration according to the weight index and selects the execution instructions based on the weight index.
[0157] The intelligent interaction module dynamically generates the false touch duration based on the weight index (a formula combining the total weight and the maximum duration), triggering the false touch prevention mode to block competing commands. In the event of a priority conflict, the command with the highest weight index is selected for execution, and a time window mechanism is used to reduce interference.
[0158] It effectively suppresses interference from multiple users' accidental touches (such as children's misoperation), improves response accuracy through adaptive delay control and competition arbitration mechanism, ensures smooth operation of core users, and enhances system robustness.
[0159] Example 1: Assume that when the touch screen command I5 is generated, there are three interactive objects, A1, A2, and A3. The average operation distances of all shoulder joints of these interactive objects are 30cm, 50cm, and 40cm respectively. The arm lengths are 65cm, 55cm, and 65cm respectively. The remaining parameters are:
[0160] A1 interaction object: intersection angle: 60°; deviation distance: 5cm;
[0161] A2 interaction object: intersection angle: 45°; deviation distance: 20cm;
[0162] A3 interactive object: intersection angle: 30°; deviation distance: 30cm;
[0163] When the constant α is 3.2 and the diagonal length of the navigation device screen is 100 cm, substitute the formula to calculate the matching index of each interactive object:
[0164] A1 Interaction Object Matching Index:
[0165] A2 Interaction Object Matching Index:
[0166] A3 Interaction Object Matching Index:
[0167] Since the matching index is 58.60>7.20>6.84, A1 is selected as the interaction object of the touch screen instruction I5.
[0168] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0169] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A multimodal data analysis method based on intelligent interaction, characterized by: The method includes: S100, when the navigation device is activated, it receives touch screen instructions and captures surrounding video images; S200, identifying a target object based on the video image, analyzing the touch screen command and matching the interactive object; S300, establishing an instruction set for each interactive object, and calculating a weight index according to the instruction set; S400: Setting the false touch duration according to the weight index, and filtering out execution instructions in combination with the weight index.
2. The multimodal data analysis method based on intelligent interaction according to claim 1, characterized in that: In S100, the touch screen instruction refers to an operation instruction generated by the user by touching the screen of the navigation device, including an operation gesture, an instruction position, and a keyword library; Gestures refer to the various actions a user performs on the touch screen with their fingers, used to interact with the navigation device. Command locations refer to the specific coordinates of the user's touch on the touch screen. Keyword libraries are a collection of different key fields that describe the functions that can be achieved by the system after executing the corresponding command. The number of key fields in each keyword library is different. The more key fields there are in the keyword library, the higher the function implemented by the corresponding operation instruction is in the system framework; each key field corresponds to a description of the function of a basic unit, and together constitute the functional component of the corresponding keyword library.
3. The multimodal data analysis method based on intelligent interaction according to claim 2, characterized in that: S200 includes: S201, using the target recognition algorithm to analyze the people in the video image, setting the people whose distance is less than the threshold Q and whose face is facing the navigation device as the target object, and recording the setting time TD x ; S202, intercepting the TD of each target object in the video image x The image frames at that time are used to identify the human joints of the target objects in each image frame, and the arm length of each target object is analyzed and calculated; S203: When receiving the touch screen instruction I0, record the operation time TC z , capture the TC in the video image z Image frame IMG z ; Analyze and calculate the matching index of each target object based on the image frame, so as to match the interactive object for the touch screen instruction I0.
4. The multimodal data analysis method based on intelligent interaction according to claim 3, characterized in that: S203 includes: S2031, in the image frame IMG z The joint points of each target object are marked in the image, and a deep neural network is used in combination with reference objects within the image frame to analyze the operating distance between each joint point and the navigation device screen; S2032, establishing three-dimensional models of the target object and the navigation device respectively, and constructing the relative position of the target object and the navigation device according to the operation distance; S2033, marking a target object whose operating distance of at least one shoulder joint is no greater than the corresponding arm length, and analyzing the three-dimensional gaze vector of each marked target object using the Gaze360 model; S2034, analyzing the intersection angle angle between the three-dimensional gaze vector and the navigation device screen, as well as the intersection position on the navigation device screen, and calculating the deviation distance JL between the command position of the touch screen command I0 and the intersection position; S2035. Calculate the matching index P of each marked target object according to the formula, and select the marked target object with the largest matching index as the interactive object of the touch screen instruction I0; the formula is as follows: Where α is a constant greater than 1, JL max JD is the diagonal length of the navigation device screen. ave is the average operating distance of all shoulder joints, and JD is the arm length.
5. The multimodal data analysis method based on intelligent interaction according to claim 3, characterized in that: S300 includes: S301, matching each received touch screen instruction with an interactive object in turn, establishing an instruction set for each interactive object, and sequentially placing each touch screen instruction into the instruction set of the corresponding interactive object; S302: Analyze the reception time, operation gesture, and keyword library of each touch screen instruction in the instruction set, and calculate the weight index of each interactive object respectively.
6. The multimodal data analysis method based on intelligent interaction according to claim 5, characterized in that: The steps for calculating the weight index in S302 are as follows: S3021. Obtain all touch screen commands in the command set and arrange them in order of receipt time; set different difficulty coefficients for different operation gestures, and count the number of key fields in the keyword library for each touch screen command; S3022. Obtain the difficulty coefficient corresponding to the operation gesture of each touch screen instruction and the number of key fields in the keyword library in order of arrangement, and substitute them into the formula to calculate the weight index QZ of the interactive object corresponding to the instruction set: Where m is the number of all touch screen instructions in the instruction set, CD n-1 CD is the difficulty coefficient corresponding to the n-1th touch screen instruction. ave is the average difficulty coefficient of all operation gestures, GS n-1 is the number of key fields in the n-1 touch screen instruction keyword library, GS n The number of key fields in the keyword library of the nth touch screen instruction.
7. The multimodal data analysis method based on intelligent interaction according to claim 5, characterized in that: S400 includes: S401, dynamically calculate the weight index of each interactive object, and calculate the false touch time of each interactive object according to the weight index k ; S402: When the interactive object generates an operation instruction, the operation time TC1 is recorded, and the false touch duration time is time starting from the operation time TC1. k The system enters the anti-accidental touch mode and automatically refuses to execute the operation instructions of other interactive objects; S403: After the system cancels the accidental touch prevention mode, if it receives operation instructions from two or more interactive objects at the same time, it automatically executes the operation instruction of the interactive object with the largest weight index.
8. The multimodal data analysis method based on intelligent interaction according to claim 7, characterized in that: In S401, the limit duration f is set to calculate the mis-touch duration of each interactive object: Where, QZ sum is the sum of the weight indices of all interacting objects.
9. A multimodal data analysis system based on intelligent interaction, characterized by: The system includes data acquisition module, object recognition module, operation analysis module and intelligent interaction module; When the navigation device is activated, the data acquisition module receives touch screen instructions and collects surrounding video images; The object recognition module identifies the target object based on the video image, analyzes the touch screen instructions and matches the interactive object; The operation analysis module is used to establish an instruction set for each interactive object and calculate the weight index based on the instruction set; The intelligent interaction module sets the false touch duration according to the weight index and selects the execution instructions based on the weight index.