Intelligent interaction control method and system based on visual integration
By integrating multi-camera vision and implicit context modeling, the problems of single perception and passive response in existing writing interaction systems are solved, realizing an intelligent and personalized interactive experience and improving robustness and user adaptability in complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI YOUWO TECH CO LTD
- Filing Date
- 2026-02-11
- Publication Date
- 2026-05-12
AI Technical Summary
Existing writing interaction systems have a single perception dimension, lack contextual understanding capabilities, cannot adapt to different levels of proficiency, emotional states, and usage scenarios, and have passive and mechanical response mechanisms with insufficient robustness, making it difficult to provide an intelligent and personalized user experience.
By acquiring visual features through a multi-camera collaborative architecture, performing multi-scale fusion, and combining implicit contextual state probabilistic modeling and intent-confidence pair inference, visual-action mapping and stability optimization are achieved, and the final interaction command is output.
It significantly improves perception robustness in complex environments, can automatically adapt to different users, provide timely and appropriate interactive assistance, and ensure the naturalness and consistency of interactive feedback.
Smart Images

Figure CN122018696A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of visual analysis technology, specifically relating to an intelligent interactive control method and system based on visual integration. Background Technology
[0002] In the development of human-computer interaction technology, from the early keyboard and mouse to touch screens, and then to voice and gesture recognition in recent years, the interaction methods have been constantly evolving towards a more natural and intuitive direction. However, in the fundamental field of writing interaction, existing technologies still have obvious limitations: traditional dot matrix pens mainly rely on infrared or electromagnetic sensing technology to capture the position of the pen tip, and their functions are limited to trajectory recording and simple recognition; although commercial smart pens have achieved pressure sensing and tilt detection, they lack a deep understanding of user intentions; computer vision-based interaction systems can recognize the written content, but cannot understand the implicit state in the writing process; although the academic community has tried to combine biosignals (such as electromyography and eye movement) to enhance interaction, these solutions are often complex, expensive and difficult to popularize. The fundamental flaws of existing technologies are: first, they rely on a single perceptual dimension, excessively depending on surface features such as position and pressure, while neglecting deeper information such as spatial layout, motion dynamics, and micro-texture; second, they lack contextual understanding capabilities, failing to adapt to users with varying levels of proficiency, emotional states, and usage scenarios; third, their response mechanisms are passive and mechanical, only recording completed actions and unable to predict or guide user behavior; and fourth, their systems lack robustness, exhibiting a sharp decline in performance under complex real-world environments (such as changes in lighting and paper material). These problems prevent existing writing interaction systems from providing a truly intelligent and personalized user experience, severely hindering the deep application of digital writing in education, creation, remote collaboration, and other fields. Summary of the Invention
[0003] To address the aforementioned problems in the existing technology, this invention provides an intelligent interactive control method and system based on visual integration.
[0004] The objective of this invention can be achieved through the following technical solutions: A vision-integrated intelligent interactive control method, the implementation of which includes the following steps: Step S1: Collect visual features during the user's writing process using smart devices, and perform multi-scale visual feature fusion to obtain a spatiotemporal feature tensor; Step S2: Perform intent inference based on the spatiotemporal feature tensor and output intent-confidence pairs; Step S3: Perform visual-action mapping based on the intent-confidence pair and output a preliminary set of interaction instructions; Step S4: Optimize the stability of the preliminary interactive instruction set to obtain the final instruction set.
[0005] Preferably, the multi-scale visual feature fusion in step S1 specifically involves: The visual features are collected through smart devices, and the visual features include spatial layout features, motion trajectory features, and detail texture features. The spatiotemporal feature tensor obtained based on the visual features is mathematically described as follows: ,in, For the spatiotemporal feature tensor, This is a nonlinear fusion operator, where K is the number of cameras. For spatial attention mask, This is the convolution operator, representing the fusion of local features. For texture feature weights, For normalized Laplace features, For the Laplacian operator to act on the image, Let be the image intensity captured by the k-th camera at position (x, y) and time t. To prevent small constants from being divided by zero, For motion feature weights, The normalized time variation characteristics, The rate of change of image intensity over time. For reference time, For scale-specific transformation functions, For the transformation parameters, This is element-wise multiplication.
[0006] Preferably, the intent inference in step S2 specifically includes: A preset prediction time window Δt is obtained, and the historical spatiotemporal feature tensor [0,t] is obtained simultaneously; Preset implicit context state h; Obtain the initial probability density of the implicit context state, and output the probability density of the implicit context state given the historical spatiotemporal feature tensor through the context extraction function; The system transitions to the predicted future state by outputting the state transition kernel function, given the implicit context state and the historical spatiotemporal feature tensor. The conditional probability density; Get the predicted future state The conditional probability density given the historical spatiotemporal feature tensor is mathematically described as follows: ,in, For the predicted future state, Let [0, t] be the historical spatiotemporal feature tensor. For the predicted future state Conditional probability density given a historical spatiotemporal feature tensor For implicit context space, For context extraction functions, This is the state transition kernel function. This represents the initial probability density of the implicit context state; Based on the predicted future state Given the conditional probability density function of the historical spatiotemporal feature tensor, the confidence level is obtained, mathematically described as follows: Where C is the confidence level and Var(·) is the variance. The square root of the maximum permissible variance; Matching predicted future states Given the conditional probability density under the given historical spatiotemporal feature tensor and the confidence level, the intention-confidence pair is obtained.
[0007] Preferably, the vision-action mapping in step S3 specifically includes: By mapping writing intentions to actions through action manifolds, we obtain interactive action instruction vectors, mathematically described as follows: ,in, For interactive action command vectors, As a projection operator, it maps actions onto the action manifold M. As a standard guide weight, For the purpose of writing Geometric representation, Weights for uncertain responses, For the gradient operator with respect to intention, Here, C is the uncertainty sensitivity parameter, and C is the confidence level. For adaptive noise, σ(C) is the noise intensity; The initial set of interactive instructions is output based on the interactive action instruction vector.
[0008] Preferably, the stability optimization in step S4 specifically involves: The consistency between historical interaction commands and visual feedback is evaluated using a visual-motor consistency function; The final instruction is obtained based on the vision-action consistency function, mathematically described as follows: , For final instructions, For reference time, For exponentially decaying weights, This represents the historical feedback attenuation coefficient. For visual-motor consistency function, For smoothness adjustment parameters, For reference smoothing coefficient, For command acceleration; The final instruction set is obtained based on the final instruction.
[0009] A vision-integrated intelligent interactive control system for executing the vision-integrated intelligent interactive control method described above includes a visual feature fusion module, an intent inference module, a vision-action mapping module, and an optimization module. The visual feature fusion module is used to collect visual features during the user's writing process through smart devices and perform multi-scale visual feature fusion to obtain a spatiotemporal feature tensor. The intent inference module is used to infer intent based on the spatiotemporal feature tensor and output intent-confidence pairs. The visual-action mapping module is used to perform visual-action mapping based on the intent-confidence pair and output a preliminary set of interaction instructions. The optimization module is used to optimize the stability of the initial interactive instruction set to obtain the final instruction set.
[0010] The beneficial effects of this invention are as follows: (1) By using a multi-camera collaborative architecture (wide-angle, standard, and macro), the limitations of a single sensor are overcome, and spatial layout, motion trajectory and micro texture are captured simultaneously, which significantly improves the perception robustness in complex environments.
[0011] (2) By using implicit context state probability modeling, the system can automatically adapt to users with different proficiency and emotional states, thus solving the "one-size-fits-all" interaction defect of traditional systems.
[0012] (3) By using the intent-confidence inference mechanism, the system can shift from passive recording to active guidance, providing timely and appropriate assistance during the user's writing process. This is especially valuable for improving efficiency in educational scenarios and for novice learning in professional scenarios. (4) The naturalness and coherence of interactive feedback are ensured by action manifold projection and stability optimization. Attached Figure Description
[0013] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings.
[0014] Figure 1 This is a flowchart of a vision-integrated intelligent interactive control method according to the present invention. Detailed Implementation
[0015] To better understand the invention, various aspects of the invention will be described in more detail with reference to the accompanying drawings. It should be understood that these detailed descriptions are merely illustrative of exemplary embodiments of the invention and are not intended to limit the scope of the invention in any way. Throughout the specification, the expression "and / or" includes any and all combinations of one or more of the associated listed items. As used herein, the terms "approximately," "about," and similar terms are used as expressions of approximation, not as expressions of degree, and are intended to describe inherent deviations in measured or calculated values that will be recognized by those skilled in the art. Furthermore, the order in which the steps are described in this invention does not necessarily indicate the order in which these steps occur in actual operation, unless otherwise expressly defined or deduced from the context.
[0016] It should also be understood that expressions such as "comprising," "including," "having," "containing," and / or "comprising" are open-ended rather than closed-ended expressions in this specification, indicating the presence of the stated features, elements, and / or components, but not excluding the presence of one or more other features, elements, components, and / or combinations thereof. Furthermore, when expressions such as "at least one of..." appear after a list of listed features, they modify the entire list of features, not just individual elements in the list. Additionally, when describing embodiments of the invention, the word "may" is used to mean "one or more embodiments of the invention." And the term "exemplary" is intended to refer to examples or illustrations.
[0017] Unless otherwise specified, all terms used herein (including engineering and technical terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that, unless expressly stated herein, terms defined in common dictionaries shall be interpreted as having the meaning consistent with their meaning in the context of the relevant art, and not in an idealized or overly formalized sense.
[0018] It should be noted that, unless otherwise specified, the embodiments and features described in this invention can be combined with each other. The invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0019] Example 1: Please see Figure 1 A vision-integrated intelligent interactive control method includes: Step S1: Collect visual features of the user during the writing process through smart devices (such as dot matrix pens, cameras, smart pens, etc.) and perform multi-scale visual feature fusion to obtain spatiotemporal feature tensors; Step S2: Perform intent inference based on the spatiotemporal feature tensor, output intent-confidence pairs, and also identify user-written content, etc. Step S3: Perform visual-action mapping based on the intent-confidence pair and output a preliminary set of interaction instructions; Step S4: Optimize the stability of the preliminary interactive instruction set to obtain the final instruction set.
[0020] In this embodiment, the multi-scale visual feature fusion specifically refers to: S101: The visual features are acquired through a built-in or external camera (wide-angle camera, standard camera, macro camera, etc.) of the smart device. The visual features include, but are not limited to, spatial layout features, motion trajectory features, and detail texture features. The spatial layout features are derived from images captured by the wide-angle camera, and information such as paper boundaries, lines, and the position of existing text are extracted using an edge detection algorithm. The motion trajectory features are derived from a continuous image sequence from the standard camera, and the speed and direction of pen tip movement are extracted by calculating pixel changes between adjacent frames. The detail texture features are derived from high-resolution images from the macro camera, and the pen pressure and its variation are extracted by analyzing the ink diffusion pattern between paper fibers. S102: Based on the visual features, obtain the spatiotemporal feature tensor, mathematically described as follows: ,in, For spatiotemporal feature tensors, it represents a comprehensive understanding of writing characteristics. This is a nonlinear fusion operator implemented through a neural network, which intelligently fuses features at different scales. K represents the number of cameras. This is a spatial attention mask, a matrix of the same size as the image, with values ranging from [0,1]. Values closer to 1 indicate that the system considers the region more important. For example, when the system focuses on the pen tip, the area around the pen tip... The value is relatively high. This is the convolution operator, representing the fusion of local features. Texture feature weights control the importance of Laplacian features. For normalized Laplace features, The Laplacian operator is applied to an image to detect edge and detail variations. Let be the image intensity captured by the k-th camera at position (x, y) and time t. To prevent small constants from being divided by zero, For motion feature weights, The normalized time variation characteristics, The rate of change of image intensity over time. For reference time, it is usually set to 1 second. This is a scale-specific transform function that adjusts features based on camera characteristics. The transformation parameters are obtained through system training, for example, for macro cameras. It will enhance texture details; for wide-angle cameras, It will emphasize spatial layout. This is element-wise multiplication.
[0021] In this embodiment, the intent inference specifically refers to: S201: Preset the prediction time window Δt, and at the same time obtain the historical spatiotemporal feature tensor of [0,t] (i.e., the spatiotemporal feature tensor sequence of [0,t]). S202: Preset implicit context state h, for example h={h1,h2,h3}, where h1 represents proficiency (0-1), h2 represents tension (0-1), and h3 represents fatigue (0-1); S203: Obtain the initial probability density of the implicit context state, and output the probability density of the implicit context state given the historical spatiotemporal feature tensor through the context extraction function. For example, based on previously observed features, if the user's handwriting is shaky and the speed is uneven, then the user may be a novice and in a state of tension. S204: Given the implicit context state and the historical spatiotemporal feature tensor, the system will transition to the predicted future state through the state transition kernel function. The conditional probability density, for example, when the user is a novice, might predict that the user is prone to making mistakes when drawing complex strokes; S205: Obtaining the predicted future state The conditional probability density given the historical spatiotemporal feature tensor is mathematically described as follows: ,in, The predicted future state (such as pen tip position, speed, and pressure) can be a multi-dimensional vector. Let [0, t] be the historical spatiotemporal feature tensor. For the predicted future state Conditional probability density given a historical spatiotemporal feature tensor This is an implicit context space, containing various factors that cannot be directly observed but affect writing (such as writing habits, emotional state, writing environment, etc.). For context extraction functions, This is the state transition kernel function. This represents the initial probability density of the implicit context state; S206: By predicting future states Given the conditional probability density function of the historical spatiotemporal feature tensor, the confidence level is obtained, mathematically described as follows: Where C is the confidence level and Var(·) is the variance. The square root of the maximum permissible variance; S207: Matching predicted future states Given the conditional probability density under the given historical spatiotemporal feature tensor and the confidence level, the intention-confidence pair is obtained.
[0022] In this embodiment, the vision-action mapping specifically refers to: S301: Mapping writing intent to actions through action manifolds yields interactive action instruction vectors, mathematically described as follows: ,in, For interactive action command vectors, such as A=(t prompt ,v voice ,c content These represent the prompt time, volume, and content encoding, respectively. The projection operator maps actions onto the action manifold M, where M is the action manifold, ensuring that the actions are natural and reasonable. For example, in M, the time interval between consecutive prompts is no less than 0.5 seconds. As a standard guide weight, For the purpose of writing The geometric representation, for example, for the intent to "write a vertical stroke", G outputs a path description of the standard vertical stroke and the timing of prompts. Weights for uncertain responses, This is a gradient operator with respect to intention, measuring the effect of small changes in intention on the action. Here, C is the uncertainty sensitivity parameter, and C is the confidence level. For adaptive noise, σ(C) is the noise intensity; the lower the confidence level, the greater the noise. S302: Output the preliminary interaction instruction set based on the interaction action instruction vector.
[0023] In this embodiment, the stability optimization specifically includes: S401: Evaluate the consistency between historical interaction commands and visual feedback through the visual-action consistency function. For example, if the system prompts "write to the right" at time t, but the visual feedback shows that the user writes to the left, then the consistency will be very low. S402: The final instruction is obtained based on the vision-action consistency function, mathematically described as follows: , For final instructions, For reference time, it is usually set to 1 second. The exponentially decaying weights ensure that near-term feedback is more important than long-term feedback. This represents the historical feedback attenuation coefficient. For visual-motor consistency function, For smoothness adjustment parameters, For reference smoothing coefficient, This refers to the acceleration of the command, indicating the degree of drastic change in the command. S403: Obtain the final instruction set based on the final instruction.
[0024] Example 2: A vision-integrated intelligent interactive control system includes a visual feature fusion module, an intent inference module, a vision-action mapping module, and an optimization module; The visual feature fusion module is used to collect visual features during the user's writing process through smart devices (such as dot matrix pens, cameras, smart pens, etc.) and perform multi-scale visual feature fusion to obtain a spatiotemporal feature tensor. The intent inference module is used to infer intent based on the spatiotemporal feature tensor, output intent-confidence pairs, and can also recognize user-written content, etc. The visual-action mapping module is used to perform visual-action mapping based on the intent-confidence pair and output a preliminary set of interaction instructions. The optimization module is used to optimize the stability of the initial interactive instruction set to obtain the final instruction set.
[0025] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A visual integration-based intelligent interactive control method, characterized in that, Includes the following steps: Step S1: Collect visual features during the user's writing process using smart devices, and perform multi-scale visual feature fusion to obtain a spatiotemporal feature tensor; Step S2: Perform intent inference based on the spatiotemporal feature tensor and output intent-confidence pairs; Step S3: Perform visual-action mapping based on the intent-confidence pair and output a preliminary set of interaction instructions; Step S4: Optimize the stability of the preliminary interactive instruction set to obtain the final instruction set.
2. The intelligent interactive control method based on vision integration according to claim 1, characterized in that, The multi-scale visual feature fusion in step S1 specifically involves: The visual features are collected through smart devices, and the visual features include spatial layout features, motion trajectory features, and detail texture features. The spatiotemporal feature tensor obtained based on the visual features is mathematically described as follows: ,in, For the spatiotemporal feature tensor, This is a nonlinear fusion operator, where K is the number of cameras. For spatial attention mask, This is the convolution operator, representing the fusion of local features. For texture feature weights, For normalized Laplace features, For the Laplacian operator to act on the image, Let be the image intensity captured by the k-th camera at position (x, y) and time t. To prevent small constants from being divided by zero, For motion feature weights, The normalized time variation characteristics, The rate of change of image intensity over time. For reference time, For scale-specific transformation functions, For the transformation parameters, This is element-wise multiplication.
3. The intelligent interactive control method based on vision integration according to claim 1, characterized in that, The intent inference in step S2 specifically refers to: A preset prediction time window Δt is obtained, and the historical spatiotemporal feature tensor [0,t] is obtained simultaneously; Preset implicit context state h; Obtain the initial probability density of the implicit context state, and output the probability density of the implicit context state given the historical spatiotemporal feature tensor through the context extraction function; The system transitions to the predicted future state by outputting the state transition kernel function, given the implicit context state and the historical spatiotemporal feature tensor. The conditional probability density; Get the predicted future state The conditional probability density given the historical spatiotemporal feature tensor is mathematically described as follows: ,in, For the predicted future state, Let [0, t] be the historical spatiotemporal feature tensor. For the predicted future state Conditional probability density given a historical spatiotemporal feature tensor For implicit context space, For context extraction functions, This is the state transition kernel function. This represents the initial probability density of the implicit context state; Based on the predicted future state Given the conditional probability density function of the historical spatiotemporal feature tensor, the confidence level is obtained, mathematically described as follows: Where C is the confidence level and Var(·) is the variance. The square root of the maximum permissible variance; Matching predicted future states Given the conditional probability density under the given historical spatiotemporal feature tensor and the confidence level, the intention-confidence pair is obtained.
4. The intelligent interactive control method based on vision integration according to claim 1, characterized in that, The vision-action mapping in step S3 specifically refers to: By mapping writing intentions to actions through action manifolds, we obtain interactive action instruction vectors, mathematically described as follows: ,in, For interactive action command vectors, As a projection operator, it maps actions onto the action manifold M. As a standard guide weight, For the purpose of writing Geometric representation, Weights for uncertain responses, For the gradient operator with respect to intention, Here, C is the uncertainty sensitivity parameter, and C is the confidence level. For adaptive noise, σ(C) is the noise intensity; The initial set of interactive instructions is output based on the interactive action instruction vector.
5. The intelligent interactive control method based on vision integration according to claim 4, characterized in that, The stability optimization in step S4 specifically involves: The consistency between historical interaction commands and visual feedback is evaluated using a visual-motor consistency function; The final instruction is obtained based on the vision-action consistency function, mathematically described as follows: , For final instructions, For reference time, For exponentially decaying weights, This is the historical feedback attenuation coefficient. For visual-action consistency function, For smoothness adjustment parameters, For reference smoothing coefficient, For command acceleration; The final instruction set is obtained based on the final instruction.
6. A vision-integrated intelligent interactive control system, characterized in that, The system is applied to the intelligent interactive control method based on visual integration as described in any one of claims 1-5, including a visual feature fusion module, an intent inference module, a visual-action mapping module, and an optimization module; The visual feature fusion module is used to collect visual features during the user's writing process through smart devices and perform multi-scale visual feature fusion to obtain a spatiotemporal feature tensor. The intent inference module is used to infer intent based on the spatiotemporal feature tensor and output intent-confidence pairs. The visual-action mapping module is used to perform visual-action mapping based on the intent-confidence pair and output a preliminary set of interaction instructions. The optimization module is used to optimize the stability of the initial interactive instruction set to obtain the final instruction set.