Robot visual control method and system based on spatio-temporal context perception
By using spatiotemporal context feature encoding and dynamic resource allocation of attention potential fields, the problem of lack of understanding of scene spatiotemporal context and resource waste in robot vision control is solved, achieving more efficient computing power allocation and forward-looking vision control, and improving the robot's autonomy and tracking performance in complex environments.
Patent Information
- Application Number
- CN202511522206.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-23
AI Technical Summary
Existing robot vision control methods lack a deep understanding of the spatiotemporal context of the scene, cannot adaptively adjust resource allocation strategies, and lack predictive attention mechanisms, resulting in wasted computational resources or insufficient processing of key areas in complex environments, leading to unstable robot tracking.
By combining spatiotemporal context feature encoding, attention potential field, and resource allocation matrix with lightweight convolutional networks and recurrent neural networks, dynamic feature fusion and forward-looking visual control are achieved, optimizing computational power allocation and attention prediction.
It improves the robot's autonomy and robustness in dynamic environments, reduces false detections and missed detections, optimizes computing power usage, and enhances tracking performance.
Smart Images

Figure CN120985680A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of visual control, and particularly relates to a robot visual control method and system based on spatiotemporal context perception. BACKGROUND
[0002] With the wide application of robot technology in automatic driving, industrial automation and intelligent monitoring, the performance of the visual control system is crucial. The existing robot visual control method mainly relies on static image feature extraction or simple time sequence processing, such as single-frame target detection and classification based on convolutional neural network (CNN) and motion detection based on optical flow or frame difference method. These methods are effective in some scenarios, but have obvious shortcomings: first, they often lack a deep understanding of the spatiotemporal context of the scene, and cannot dynamically integrate static features (such as texture, color) and dynamic features (such as motion patterns), resulting in attention dispersion or loss of key targets in complex environments; second, the resource allocation strategy often uses uniform or heuristic fixed allocation, which cannot adaptively adjust the computing power according to the importance of the scene, causing waste of computing resources or insufficient processing of key areas; in addition, traditional attention mechanisms are mostly limited to the current moment, lacking modeling and prediction of attention shift patterns, resulting in delayed or jittered control instructions, especially unstable robot tracking in high-speed motion scenarios. These shortcomings seriously limit the autonomy, real-time performance and robustness of robots in dynamic environments. SUMMARY
[0003] To solve the above problems in the prior art, the application provides a robot visual control method and system based on spatiotemporal context perception.
[0004] The object of the application can be achieved by the following technical solutions: A robot visual control method based on spatiotemporal context perception, the implementation of the robot visual control method based on spatiotemporal context perception includes the following steps: Step S1: acquiring an original image sequence and performing spatiotemporal context feature encoding to output a spatiotemporal context feature map; Step S2: mapping the spatiotemporal context feature map to an initial attention map through a lightweight convolutional network, and introducing an attention diffusion process to output an attention potential field; Step S3: obtaining a resource allocation matrix based on the attention potential field and performing computing power allocation; Step S4: obtaining a historical attention potential field based on the attention potential field, performing attention dynamics prediction according to the historical attention potential field to obtain a predicted attention potential field; Step S5: realizing prospective visual control based on the attention potential field and the predicted attention potential field.
[0005] Preferably, the spatio-temporal context feature encoding in the step S1 is specifically: obtaining the original image sequence; extracting features of the original image sequence by a deep neural network encoder to obtain a spatial feature map and a temporal feature map, the deep neural network encoder comprising a spatial encoder and a temporal encoder; fusing the spatial feature map and the temporal feature map by a spatio-temporal context gated fusion formula to obtain the spatio-temporal context feature map, the mathematical description of the spatio-temporal context gated fusion formula being wherein, is the spatio-temporal context feature map at time t, is a Sigmoid activation function, is a gating weight matrix, is the spatial feature map, is the temporal feature map, represents concatenation along the feature channel dimension, is a gating bias, and is a transformation weight matrix, and is a transformation bias, is a hyperbolic tangent function, is a gating weight map, is an element-wise multiplication.
[0006] Preferably, the attention diffusion process in the step S2 is specifically: mapping the spatio-temporal context feature map to the initial attention map, the mathematical description being wherein, is the initial attention map at time t, is an attention weight, is an attention bias, is the spatio-temporal context feature map at time t; converting the initial attention map to a probability distribution map by spatial Softmax normalization; obtaining the attention potential field based on the initial attention map and the probability distribution map, the mathematical description being wherein, is the attention potential field at time t, is the probability distribution map, is spatial Softmax normalization, is a diffusion coefficient, is a Plancharel operator, is a spatial discretization step size.
[0007] Preferably, the step S3 specifically comprises: obtaining total computing resource and presetting minimum computing resource guaranteed by each position; obtaining a normalization factor based on the attention potential field, which is mathematically described as wherein, is the normalization factor, and M·N is the total number of positions, is a sharpness coefficient, is the attention potential value of the position, is the attention mean value, is the attention standard deviation; obtaining the resource allocation matrix based on the total computing resource, the minimum computing resource and the normalization factor, and performing computing resource allocation, which is mathematically described as wherein, is the computing resource allocated to the position, is the minimum computing resource, is the total computing resource, is the attention potential value of the position.
[0008] Preferably, the attention dynamics prediction in the step S4 specifically comprises: learning the attention transfer rule through a recurrent neural network (RNN), which is mathematically described as wherein, is the predicted attention potential field at t+1, is the RNN hidden state at t+1, is an attention feature extraction function, is the RNN hidden state at t, is an RNN parameter, is the attention potential field at t.
[0009] Preferably, the prospective visual control in the step S5 specifically comprises: obtaining the current position coordinates of the robot; obtaining the current attention centroid world coordinates based on the attention potential field, which is mathematically described as wherein, is the current attention centroid world coordinates, is the attention potential value of the position, and M·N is the total number of positions, is the world coordinates corresponding to the position; obtaining the predicted attention centroid world coordinates based on the predicted attention potential field, which is mathematically described as wherein, predicting an attention mass world coordinate, for a predicted attention potential value of the position; outputting a control instruction of the robot based on the current position coordinate of the robot, the current attention mass world coordinate and the predicted attention mass world coordinate, mathematically described as wherein, is the control instruction, , and are proportional gain, derivative gain and integral gain respectively, is the current position coordinate of the robot, is an integral variable.
[0010] A robot vision control system based on spatiotemporal context perception is used to execute the robot vision control method based on spatiotemporal context perception described above, comprising a spatiotemporal context perception module, an attention potential field output module, a computing power allocation module, a prediction module and a vision control module. The spatiotemporal context perception module is used to obtain an original image sequence and perform spatiotemporal context feature coding, and output a spatiotemporal context feature map. The attention potential field output module is used to map the spatiotemporal context feature map into an initial attention map through a lightweight convolutional network, introduce an attention diffusion process and output an attention potential field. The computing power allocation module is used to obtain a resource allocation matrix based on the attention potential field and perform computing power allocation. The prediction module is used to obtain a historical attention potential field based on the attention potential field, perform attention dynamics prediction according to the historical attention potential field and obtain a predicted attention potential field. The vision control module is used to realize prospective vision control based on the attention potential field and the predicted attention potential field.
[0011] The present application has the following beneficial effects: (1) Through spatiotemporal context gating fusion, dynamic trade-off between static and dynamic features is realized to reduce false positives and false negatives.
[0012] (2) Through a nonlinear resource allocation strategy, high attention areas are preferentially processed to optimize computing power usage while ensuring basic monitoring.
[0013] (3) The RNN-based attention dynamics prediction enables the robot to adjust the posture in advance to improve tracking performance. BRIEF DESCRIPTION OF DRAWINGS
[0014] For the convenience of those skilled in the art, the present application will be further described below with reference to the accompanying drawings.
[0015] Figure 1 A flow chart of steps of a robot vision control method based on spatio-temporal context perception. DETAILED DESCRIPTION
[0016] For a better understanding of the present application, various aspects of the present application will be described in more detail below with reference to the accompanying drawings. It is to be noted that these detailed description is merely descriptive of exemplary embodiments of the present application and is not intended in any way to limit the scope of the present application. Throughout this specification, the expression "and / or" includes any and all combinations of one or more of the associated listed items. As used herein, the terms "substantially", "approximately", and similar terms are used as terms of approximation and not as terms of degree, and are intended to account for the inherent deviations in a measuring or computing process that would be recognized by those of ordinary skill in the art. Additionally, in the present application, the order of the steps of the processes described does not necessarily indicate the order in which the processes occur in actual operation, unless explicitly stated otherwise or derivable from context.
[0017] It should also be understood that expressions such as "include", "including", "have", "has", "contain" and / or "containing", and the like, are open-ended terms that are intended to mean one or more of the stated elements or components are present, but not excluding the presence of one or more other elements or components. Further, when such expressions as "at least one of" appear in a list of elements, and the like, it is intended to mean one or more of the elements but not excluding others not specifically listed. Further, when describing embodiments of the present application, the use of "can" means "one or more embodiments of the present application". Also, the use of the term "exemplary" is intended to present an example or an illustration.
[0018] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and not be interpreted in an overly formal or overly literal sense unless expressly so defined herein.
[0019] It should be noted that the embodiments and features of the present application can be combined with each other, if not in conflict. The present application will be described in more detail with reference to the accompanying drawings and embodiments.
[0020] Example 1: Please refer to Figure 1 A robot vision control method based on spatio-temporal context perception, comprising: Step S1: obtaining an original image sequence and performing spatio-temporal context feature coding to output a spatio-temporal context feature map; Step S2: mapping the spatio-temporal context feature map into an initial attention map of a single channel through a lightweight convolutional network, introducing an attention diffusion process for smoothing noise and enhancing spatial consistency, and outputting an attention potential field; Step S3: obtaining a resource allocation matrix based on the attention potential field and performing computing power allocation; Step S4: obtaining a historical attention potential field based on the attention potential field, performing attention dynamics prediction according to the historical attention potential field, and obtaining a predicted attention potential field; Step S5: realizing prospective visual control based on the attention potential field and the predicted attention potential field.
[0021] In the embodiment, the spatio-temporal context feature coding is specifically as follows: S101: obtaining the original image sequence by a robot; S102: performing feature extraction on the original image sequence through two parallel deep neural network encoders to obtain a spatial feature map and a temporal feature map, the deep neural network encoders including a spatial encoder and a temporal encoder, the spatial encoder being used for processing a current frame (i.e. ), extracting multi-scale spatial features through a convolutional neural network, and the feature responses thereof covering static clues such as color, texture and contrast; and the temporal encoder being used for processing a continuous frame stack (i.e. , , …), extracting motion features through a 3D convolution or an optical flow network, and the feature responses thereof showing motion regions in a scene; S103: fusing the spatial feature map and the temporal feature map through a spatio-temporal context gating fusion formula to obtain the spatio-temporal context feature map, the mathematical description of the spatio-temporal context gating fusion formula being as follows: , wherein, is the spatio-temporal context feature map at time t, is a Sigmoid activation function, is a gating weight matrix, is the spatial feature map, is the temporal feature map, represents splicing along a feature channel dimension, is a gating bias, and are transformation weight matrices, and are transformation biases, is a hyperbolic tangent function, and For the gating weight graph, Dynamically learning at each spatial location whether to trust spatial or temporal features more, for example, on a stationary but high-contrast object. Approaching 1, dependent on spatial features, on low-texture but moving objects, Approaching 0, dependent on time characteristics, This is element-wise multiplication.
[0022] In this embodiment, the attention diffusion process is specifically as follows: S201: Map the spatiotemporal context feature map to the initial attention map, mathematically described as follows: ,in, Here is the initial attention map at time t. For attention weights, For attention bias, This is the spatiotemporal context feature map at time t; S202: The initial attention map is converted into a probability distribution map by spatial Softmax normalization, and the sum of the distributions is 1, so that the attention between different images is comparable; S203: Based on the initial attention map and the probability distribution map, the attention potential field is obtained, mathematically described as follows: ,in, Let t be the attention potential field (the high potential energy region corresponds to the attention focus). This is a probability distribution diagram. For spatial Softmax normalization, The diffusion coefficient is used to control the smoothness. For the Pilates operator, The step size for spatial discretization.
[0023] In this embodiment, step S3 can be implemented through the following steps: S301: Obtain the total computing power resources and preset the minimum computing power resources guaranteed for each location (i.e., the minimum computing power resources that can be obtained by each area to maintain basic scene monitoring). S302: Obtain the normalization factor based on the attention potential field, mathematically described as follows: ,in, M is the normalization factor, and M·N is the total number of positions. This is the sharpness coefficient. The larger the area, the more resources are allocated to the high-attention area. for The attentional potential value of a location. The mean of attention ( ), The standard deviation of attention ( ); S303: The resource allocation matrix is obtained based on the total computing power resources, the minimum computing power resources, and the normalization factor, mathematically described as follows: ,in, for Location-allocated computing resources For minimum computing power resources, For total computing power resources, for The attention potential value of a location. Therefore, for regions with significantly higher than average attention, the resources they acquire can increase exponentially, ultimately achieving nonlinear adaptive allocation of resources.
[0024] In this embodiment, the attention dynamics prediction specifically refers to: The attention transfer pattern is learned through a recurrent neural network (RNN), mathematically described as follows: ,in, The attention potential field at time t+1 is predicted. This represents the hidden state of the RNN at time t+1. For attention feature extraction function, The hidden state of the RNN at time t encodes the historical information of the attention evolution. These are the parameters of the RNN.
[0025] In this embodiment, the forward-looking visual control specifically refers to: S501: Obtain the current position coordinates of the robot; S502: Based on the attention potential field, obtain the current world coordinates of the attention centroid (i.e., the system will use a virtual attention point as the tracking target, and the position of this attention point is the weighted center of the attention potential field), mathematically described as follows: ,in, As the current world coordinates of the center of mass of attention, for The attentional potential energy value of a location, where M·N is the total number of locations. for The world coordinates corresponding to the location; S503: Based on the predicted attention potential field, obtain the predicted attention centroid world coordinates, mathematically described as follows: ,in, To predict the world coordinates of the attention centroid, for Predicted attention potential value for location; S504: Output control commands for the robot based on the robot's current position coordinates, the current attention centroid world coordinates, and the predicted attention centroid world coordinates. Mathematically, this can be described as follows: wherein, is a control instruction, , and are proportional gain, derivative gain and integral gain, respectively, is a current position coordinate of the robot, is an integral variable.
[0026] Embodiment 2: A robot vision control system based on spatio-temporal context perception, comprising a spatio-temporal context perception module, an attention potential field output module, a computing power allocation module, a prediction module and a vision control module; The spatio-temporal context perception module is configured to obtain an original image sequence and perform spatio-temporal context feature encoding, and output a spatio-temporal context feature map; The attention potential field output module is configured to map the spatio-temporal context feature map into an initial attention map of a single channel through a lightweight convolutional network, introduce an attention diffusion process for smoothing noise and enhancing spatial consistency, and output an attention potential field; The computing power allocation module is configured to obtain a resource allocation matrix based on the attention potential field, and perform computing power allocation; The prediction module is configured to obtain a historical attention potential field based on the attention potential field, perform attention dynamics prediction according to the historical attention potential field, and obtain a predicted attention potential field; The vision control module is configured to implement prospective vision control based on the attention potential field and the predicted attention potential field.
[0027] The above is only a preferred embodiment of the present application, and does not limit the present application in any form. Although the present application has been disclosed as above, it is not intended to limit the present application. Any person skilled in the art can make some changes or modifications to the above disclosed technical content without departing from the scope of the present application, and any simple modification, equivalent change and modification of the above embodiment according to the technical essence of the present application are still within the scope of the present application.
Claims
1. A spatio-temporal context-aware based robot vision control method, characterized in that, The method comprises the following steps: Step S1: obtaining an original image sequence and performing spatio-temporal context feature coding to output a spatio-temporal context feature map; Step S2: mapping the spatio-temporal context feature map into an initial attention map through a lightweight convolutional network, introducing an attention diffusion process to output an attention potential field; Step S3: obtaining a resource allocation matrix based on the attention potential field and performing computing power allocation; Step S4: obtaining a historical attention potential field based on the attention potential field, performing attention dynamics prediction according to the historical attention potential field to obtain a predicted attention potential field; Step S5: realizing prospective visual control based on the attention potential field and the predicted attention potential field.
2. The spatio-temporal context-aware based robot vision control method of claim 1, wherein, The spatio-temporal context feature coding in the step S1 is specifically: obtaining the original image sequence; performing feature extraction on the original image sequence through a deep neural network encoder to obtain a spatial feature map and a temporal feature map, wherein the deep neural network encoder comprises a spatial encoder and a temporal encoder; The spatial feature map and the time feature map are fused by a spatiotemporal context gating fusion formula to obtain the spatiotemporal context feature map, and a mathematical description of the spatiotemporal context gating fusion formula is , wherein, is the spatiotemporal context feature map at t moment, is a Sigmoid activation function, is a gating weight matrix, is a spatial feature map, is a time feature map, represents splicing along the feature channel dimension, is a gating bias, and is a transformation weight matrix, and is a transformation bias, is a hyperbolic tangent function, is a gating weight map, is an element-wise multiplication. 3.The spatio-temporal context-aware based robot vision control method according to claim 1, wherein, The attention diffusion process in the step S2 is specifically: mapping the spatiotemporal context feature map to the initial attention map, mathematically described as wherein, is the initial attention map at time t, is the attention weight, is the attention bias, is the spatiotemporal context feature map at time t; converting the initial attention map into a probability distribution map through spatial Softmax normalization; obtaining the attention potential field based on the initial attention map and the probability distribution map, and a mathematical description is wherein, is the attention potential field at time t, is the probability distribution map, is the spatial Softmax normalization, is the diffusion coefficient, is the Plancharel operator, is the spatial discretization step.
4. The spatio-temporal context-aware based robot vision control method of claim 1, wherein, The step S3 specifically comprises: obtaining total computing power resources and pre-setting a minimum computing power resource guaranteed for each position; A normalization factor is obtained based on the attention potential field, which is mathematically described as wherein, is a normalization factor, and M·N is the total number of positions, is a sharpness coefficient, is is the attention potential value of the position, is the attention mean value, is the attention standard deviation; The resource allocation matrix is obtained based on the total computing resource, the minimum computing resource and the normalization factor, and computing resource allocation is performed, which is mathematically described as wherein, is the computing resource allocated to the position, is the minimum computing resource, is the total computing resource, is the attention potential value of the position.
5. The spatio-temporal context-aware based robot vision control method of claim 1, wherein, The attention dynamics prediction in the step S4 is specifically: The attention shift rule is learned by a recurrent neural network (RNN), and a mathematical description is wherein, is a predicted attention potential field at time t+1, is an RNN hidden state at time t+1, is an attention feature extraction function, is an RNN hidden state at time t, is an RNN parameter, is an attention potential field at time t.
6. The spatio-temporal context-aware based robot vision control method of claim 1, wherein, The prospective visual control in the step S5 is specifically: obtaining a current position coordinate of the robot; The current attention mass world coordinates are obtained based on the attention potential field, and a mathematical description is as follows wherein, the current attention mass world coordinates, the current attention mass world coordinates, the attention potential value of the position, and M*N is the total number of positions, the attention potential value of the position, and M*N is the total number of positions, the world coordinates corresponding to the position; Based on the predicted attention potential field, a predicted attention centroid world coordinate is obtained, which is mathematically described as wherein, is the predicted attention centroid world coordinate, is a predicted attention potential energy value of the position; outputting a control instruction of the robot based on the current position coordinate of the robot, the current attention mass center world coordinate and the predicted attention mass center world coordinate, and the mathematical description is wherein, is a control command, , and are proportional, derivative and integral gains, respectively, is a current position coordinate of the robot, is an integral variable.
7. A spatio-temporal context-aware based robot vision control system, characterized by, The system is applied to the robot visual control method based on spatio-temporal context perception as claimed in any one of claims 1-6, and comprises a spatio-temporal context perception module, an attention potential field output module, a computing power allocation module, a prediction module and a visual control module; The spatio-temporal context perception module is used to obtain an original image sequence and perform spatio-temporal context feature coding to output a spatio-temporal context feature map; The attention potential field output module is used to map the spatio-temporal context feature map into an initial attention map through a lightweight convolutional network, introduce an attention diffusion process to output an attention potential field; The computing power allocation module is used to obtain a resource allocation matrix based on the attention potential field and perform computing power allocation; The prediction module is used to obtain a historical attention potential field based on the attention potential field, perform attention dynamics prediction according to the historical attention potential field to obtain a predicted attention potential field; The visual control module is used to realize prospective visual control based on the attention potential field and the predicted attention potential field.
Citation Information
Patent Citations
Video target detection method based on deep learning
CN109583340A
Method for realizing multi-dimensional image processing by using robot
CN117671626A
Method and system for guiding three-dimensional point cloud robot based on natural language
CN119927932A
Wind power generation power prediction method based on space-time diagram convolution and gating attention
CN120597214A
Controlling robots using multi-modal language models
US20250144795A1