Multimodal Interaction and Image Processing Methods, Electronic Devices, and Storage Media for AI Glasses

Through the multimodal interaction and image processing methods of AI glasses, the spatiotemporal attention mechanism, Kalman filtering algorithm and twin network are used to solve the problem of unable to stably track the target in the existing technology, and the precise positioning and stable tracking of the target are achieved, and high-quality tracking of the target image sequence is output.

CN119785165BActive Publication Date: 2025-06-17SHENZHEN BEIBO INTELLIGENT TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510284292.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-17
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The prior art cannot track the target stably, especially when the target suddenly accelerates, turns, or is blocked, and the subsequent appearance location of the target cannot be predicted.

Method used

The multimodal interaction and image processing method of AI glasses are adopted to preprocess the original image through the spatiotemporal attention mechanism, combined with speech recognition and image processing instruction template library, dynamic tracking is used for Kalman filtering algorithm and twin network to learn and update the appearance characteristics of the target in real time.

Benefits of technology

It achieves stable tracking and precise positioning of the target, and can continuously track the target image sequence with improved clarity and texture richness in environments with complex target changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785165B_ABST
    Figure CN119785165B_ABST
Patent Text Reader

Abstract

The present invention provides a multimodal interaction and image processing method, an electronic device, and a storage medium for an AI glasses, including: preprocessing the original image collected by the AI glasses in real time based on a spatio-temporal attention mechanism to obtain a preprocessed image; performing speech recognition and semantic parsing on the voice command input by the user through the AI glasses to obtain a text command; performing keyword matching between the text command and a preset image processing command template library to determine the corresponding image processing strategy; if the image processing strategy is dynamic tracking, using the Kalman filter algorithm to predict the position of the target in the next frame in the preprocessed image, and using a Siamese network to search around the predicted position to locate the target; based on the Siamese network, learning and updating the appearance features of the target in real time to achieve stable tracking of the target and output a sequence of tracked target images. In the present invention, the defects of being unable to predict the subsequent appearance position of the target and being unable to stably track the target at present are overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a multimodal interaction and image processing method, an electronic device, and a storage medium for an AI glasses. Background Art

[0002] AI glasses have been integrated into daily life and many professional fields. In daily life, users hope to use it to conveniently identify objects of interest, track moving signs, and optimize captured images during social interactions.

[0003] However, traditional image tracking technologies have defects. They often focus on single modality or only rely on static features of images, without considering the dynamic changes of actual scenes and the advantages of multimodal fusion. In current image tracking, when the target suddenly accelerates, turns, or is blocked, it is easy to lose the target and unable to predict the subsequent appearance position of the target. Therefore, the target cannot be stably tracked. Summary of the Invention

[0004] The main objective of the present invention is to provide a multimodal interaction and image processing method, an electronic device, and a storage medium for an AI glasses, aiming to overcome the defects of being unable to predict the subsequent appearance position of the target and unable to stably track the target currently.

[0005] To achieve the above objective, the present invention provides a multimodal interaction and image processing method for an AI glasses, including the following steps:

[0006] Preprocess the original image collected in real time by the AI glasses based on a spatio-temporal attention mechanism to obtain a preprocessed image; perform speech recognition and semantic parsing on the voice command input by the user through the AI glasses to obtain a text command;

[0007] Match the keywords of the text command with a preset image processing instruction template library to determine the corresponding image processing strategy;

[0008] If the image processing strategy is dynamic tracking, use the Kalman filter algorithm to predict the position of the target in the next frame in the preprocessed image, and use a Siamese network to search around the predicted position to accurately locate the target;

[0009] Based on the Siamese network, continuously learn and update the appearance features of the target in real time to achieve stable tracking of the target and output a tracking target image sequence for display to the user on the display screen of the AI glasses.

[0010] Further, in the Kalman filter algorithm, the noise covariance matrix of the prediction model in the Kalman filter algorithm is dynamically adjusted according to the historical acceleration and angular velocity of the target movement to improve the prediction accuracy.

[0011] Further, after achieving stable tracking of the target and outputting a tracking target image sequence, it further includes:

[0012] Enhance the details of the tracked target image sequence based on the generator to obtain an enhanced image;

[0013] Based on the discriminator, determine the authenticity of the enhanced image. The generator and the discriminator are optimized against iteration to improve the clarity and texture richness of the enhanced image, which is used to be displayed to the user on the display screen of the AI glasses.

[0014] Further, preprocess the original image captured in real time by the AI glasses based on the spatio-temporal attention mechanism to obtain a preprocessed image, including:

[0015] Use a three-dimensional convolutional neural network to perform convolution on the original image in the spatial dimension, extract local detail features and global semantic features in the image to obtain a spatial feature map;

[0016] In the time dimension, for the original image, based on the long short-term memory network model, learn the temporal correlation between adjacent frames, capture the dynamic change features of the image sequence over time, and output a temporal feature sequence;

[0017] Based on the attention mechanism, fuse the spatial feature map and the temporal feature sequence to obtain a spatio-temporal coherent feature map, and through normalization processing, obtain the preprocessed image.

[0018] Further, based on the Siamese network, continuously learn and update the appearance features of the target in real time to achieve stable tracking of the target and output a tracked target image sequence, including:

[0019] Crop the image region of the target tracked in the current frame and input it into the Siamese network. Use a convolutional neural network based on the residual structure to extract features from it to obtain a template feature vector containing the key features of the target;

[0020] For the entire image of the current frame, use a densely connected convolutional network to extract features to obtain a search feature map containing potential target regions;

[0021] Calculate the similarity between each potential target region in the search feature map and the template feature vector, and compare it with the similarity threshold; among them, construct a dynamic environment evaluation model, and adjust the similarity threshold in real time and dynamically according to the light intensity change rate, the moving speed of background objects, and the acceleration of the target movement in the current image;

[0022] Obtain the region with the similarity greater than the similarity threshold and the maximum similarity, and fuse the corresponding feature information with the template feature vector with different weights, so that the template feature vector is continuously updated, achieve stable tracking of the target and output a tracked target image sequence.

[0023] Further, the method further includes:

[0024] Collect the multi-modal interaction information of the user with the AI glasses within a specified time period; wherein, the multi-modal interaction information includes the voice information issued by the user, the touch instructions of the user on the AI glasses, and the image information collected by the user controlling the AI glasses;

[0025] Send the multi-modal interaction information to the server;

[0026] The server generates permission information based on the multi-modal interaction information and sends it to the AI glasses to configure the usage permissions of the AI glasses.

[0027] Further, generating permission information based on the multi-modal interaction information includes:

[0028] Perform speech recognition on the voice information to recognize keywords; encode the keywords as numbers to obtain key numbers;

[0029] Recognize the touch instructions to obtain a touch pattern; along the touch direction of the touch pattern, sequentially add the numerical characters in the key numbers one by one to the touch pattern to obtain a numerical pattern;

[0030] Encode the image information as a data stream string, and discretely distribute the data stream string to multiple adjacent regions of the numerical pattern;

[0031] In each of the adjacent regions, obtain the characters that satisfy the preset association relationship with the numerical characters on the numerical pattern as target characters, and combine them to obtain the permission information.

[0032] The present invention also provides a multi-modal interaction and image processing device for an AI glasses, including:

[0033] A processing unit for preprocessing the original image collected by the AI glasses in real time based on the spatio-temporal attention mechanism to obtain a preprocessed image; performing speech recognition and semantic parsing on the voice instructions input by the user through the AI glasses to obtain text instructions;

[0034] A matching unit for performing keyword matching on the text instructions with a preset image processing instruction template library to determine the corresponding image processing strategy;

[0035] A positioning unit for, if the image processing strategy is dynamic tracking, predicting the position of the target in the next frame in the preprocessed image using the Kalman filter algorithm, and searching around the predicted position using a Siamese network to accurately locate the target;

[0036] A tracking unit, which is used to learn and update the appearance features of the target in real time based on the Siamese network, achieve stable tracking of the target, and output a sequence of tracking target images for display to the user on the display screen of the AI glasses.

[0037] The present invention also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the method described in any one of the above are implemented.

[0038] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0039] The multi-modal interaction and image processing method, electronic device and storage medium of the AI glasses provided by the present invention include: preprocessing the original image collected by the AI glasses in real time based on the spatio-temporal attention mechanism to obtain a preprocessed image; performing speech recognition and semantic parsing on the voice command input by the user through the AI glasses to obtain a text command; performing keyword matching on the text command with a preset image processing command template library to determine the corresponding image processing strategy; if the image processing strategy is dynamic tracking, using the Kalman filter algorithm to predict the position of the target in the preprocessed image in the next frame, and using the Siamese network to search around the predicted position to accurately locate the target; based on the Siamese network, learning and updating the appearance features of the target in real time, achieving stable tracking of the target, and outputting a sequence of tracking target images for display to the user on the display screen of the AI glasses. In the present invention, the position of the target in the next frame of the image can be predicted, and then the target can be accurately located, and the appearance features of the target can be learned and updated in real time to achieve stable tracking of the target. It overcomes the defects that the subsequent appearance position of the target cannot be predicted and the target cannot be stably tracked currently. Description of the Drawings

[0040] Figure 1 is a schematic diagram of the steps of the multi-modal interaction and image processing method of the AI glasses in an embodiment of the present invention;

[0041] Figure 2 is a block diagram of the structure of the multi-modal interaction and image processing device of the AI glasses in an embodiment of the present invention;

[0042] Figure 3 is a schematic block diagram of the structure of the electronic device in an embodiment of the present invention.

[0043] The implementation, functional features and advantages of the present invention will be further described with reference to the embodiments and the drawings. Detailed Embodiments

[0044] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0045] Referring Figure 1 , in an embodiment of the present invention, a multimodal interaction and image processing method for an AI glasses is provided, including the following steps:

[0046] Step S1, preprocess the original image collected in real time by the AI glasses based on the spatio-temporal attention mechanism to obtain a preprocessed image; perform speech recognition and semantic parsing on the voice command input by the user through the AI glasses to obtain a text command;

[0047] Step S2, perform keyword matching on the text command with a preset image processing instruction template library to determine the corresponding image processing strategy;

[0048] Step S3, if the image processing strategy is dynamic tracking, use the Kalman filter algorithm to predict the position of the target in the next frame in the preprocessed image, and use the Siamese network to search around the predicted position to accurately locate the target;

[0049] Step S4, based on the Siamese network, continuously learn and update the appearance features of the target in real time, realize stable tracking of the target, and output a sequence of tracked target images for display to the user on the display screen of the AI glasses.

[0050] In this embodiment, as described in step S1 above, the spatio-temporal attention mechanism aims to perform targeted processing on the original image collected in real time by the AI glasses from two dimensions of time and space. In the spatial dimension, it can focus on the importance degree of different regions in the image, such as highlighting the part of the image that may contain key targets or subjects, and weakening those relatively unimportant background regions, etc.; in the time dimension, considering that the image collection by the AI glasses is a continuous dynamic process, this mechanism can utilize the correlation between images at different moments, and by analyzing the changes between the front and back frame images, etc., filter out the noise, interference information, etc. in the image, so as to obtain a clearer preprocessed image with more prominent key information, laying a good foundation for subsequent image processing operations.

[0051] After the user inputs a voice command through the AI glasses, speech recognition needs to be performed first. In this step, the voice signal is converted into a text form that can be understood and processed by a computer. Subsequently, semantic analysis is carried out, which is to deeply analyze the content of the converted text to understand what the user really wants to express. Finally, an accurate text command is obtained. For example, when the user says "Track the bird in the picture", after this series of processes, a text command containing key semantic information such as "track" and "bird" will be generated for further operations in the subsequent steps.

[0052] As described in step S2 above, this step is to perform keyword matching between the text command obtained in step S1 and the preset image processing instruction template library. A large number of standard templates corresponding to different image processing functions and strategies are pre-stored in the image processing instruction template library, and each template is associated with a specific set of keywords. For example, there is a template corresponding to "dynamic tracking", and its keywords can include "track", "follow", "target movement", etc.; there is a template corresponding to "image enhancement", and the relevant keywords can be "clear", "sharpen", "color optimization", etc. By extracting the keywords in the text command and matching them with the keywords in the template library, the corresponding image processing strategy can be determined, so as to clarify what specific image processing method should be adopted next to meet the user's needs.

[0053] As described in step S3 above, when the image processing strategy is determined to be dynamic tracking, first use the Kalman filter algorithm to predict the position of the target in the preprocessed image in the next frame. The Kalman filter algorithm has the ability to optimally estimate the state of a dynamic system. In this scenario, it is based on the motion trajectory, speed and other state information of the target in the previous few frames of images, combined with a certain system model (such as considering common motion model assumptions such as uniform motion and uniform acceleration of the target), to predict the position where the target is likely to appear in the next frame of image. It can effectively filter out the noise interference in the measurement process, improve the accuracy and reliability of position prediction, and provide a general search range for more accurate target positioning in the subsequent steps.

[0054] Around the position predicted by the Kalman filter algorithm, use a Siamese network to further search for the target. The Siamese network is a special neural network architecture, which usually consists of two sub-networks with the same structure. It can perform feature extraction and comparative analysis on different input image data (such as the image area around the predicted position in the current frame image and the reference image features of the target), and accurately determine the specific position of the target in the current frame image by calculating indicators such as the similarity between the two. Even if the appearance of the target in the image changes or there is partial occlusion, it can still better identify and locate it.

[0055] As described in step S4 above, the real-time learning and updating of the appearance features of the target based on the Siamese network is a crucial point in the entire dynamic tracking process. As time goes by and the target appears under different environments, angles, etc., its appearance may change. For example, the color change caused by lighting, the pose change, etc. The Siamese network continuously learns the features of newly emerging target images and updates its understanding of the target appearance features in real time, so as to continuously and stably track the target. Even if there are many changes in the external performance of the target, the tracking clues will not be easily lost. Finally, a sequence of tracked target images is output, and these image sequences can be displayed to the user on the display screen of the AI glasses according to requirements such as appropriate frame rate, enabling the user to clearly observe the dynamic situation of the target of interest in real time.

[0056] In one embodiment, the Kalman filter algorithm dynamically adjusts the noise covariance matrix of the prediction model in the Kalman filter algorithm according to the historical acceleration and angular velocity of the target movement to improve the prediction accuracy.

[0057] In this embodiment, the Kalman filter algorithm is essentially an optimal estimation method based on the linear system state space model. It continuously corrects the estimation of the system state by predicting the system state and combining the actual observation values, so as to accurately grasp the state of the dynamic system. In the target tracking scenario of the AI glasses, the system state usually includes key information such as the position and speed of the target. Its working process is generally divided into a prediction step and an update step. The prediction step infers the current state based on the state estimation at the previous moment and the state transition equation of the system, and the update step corrects the predicted state by using the observation values obtained at the current moment (such as the information related to the target detected in the image).

[0058] In the Kalman filter algorithm, the noise covariance matrix plays a crucial role. It mainly involves two aspects of noise description: one is the process noise covariance matrix, which reflects the influence of various uncertainty factors during the state transition process itself. For example, the uncertainty of the motion state change caused by the unmodeled interference force existing in the actual motion of the target and the irregularity of its own motion, etc.; the other is the measurement noise covariance matrix, which reflects the uncertainty of the observation value due to the sensor accuracy limitation (in the case of AI glasses, such as the resolution of the image acquisition device, the noise during the acquisition process, etc.). Reasonably setting and adjusting the values of these noise covariance matrices can enable the Kalman filter algorithm to better balance the weights of the predicted value and the observation value, and then improve the accuracy of the target state estimation.

[0059] The historical acceleration of the target's motion can reflect information such as the speed change trend and the force acting on the target during the motion. For example, if the target shows a gradually increasing acceleration over a certain period of time in the past, it indicates that the target may be undergoing an accelerating motion, and this accelerating trend may continue to the next moment. At this time, by analyzing the magnitude, change direction, etc. of the historical acceleration, the process noise covariance matrix can be adjusted accordingly. Because accelerating motion means that there is greater uncertainty in the target's motion state, the prediction model based on assumptions such as conventional uniform motion may not be very accurate. Appropriately increasing the values of the corresponding dimensions (such as the dimensions related to speed and position) in the process noise covariance matrix can enable the Kalman filter algorithm to consider more of the uncertainty brought by this dynamic change when predicting the target's position and other states at the next moment, and not rely too much on the relatively stable prediction model before, so as to more accurately infer the target's position according to the actual situation and improve the prediction accuracy.

[0060] For a target with rotational characteristics or a situation where the target's own attitude changes during the motion, the historical angular velocity is particularly important. The angular velocity reflects the speed and direction change of the target's rotation. When a large change in the historical angular velocity of the target is detected, it means that there are obvious dynamic changes in the target's attitude, rotation angle, etc., which will also bring greater uncertainty to the prediction of the target's position. Based on this, the noise covariance matrix can also be dynamically adjusted. For example, appropriately increasing the value of the process noise covariance matrix for the dimensions related to the target's angle and rotation, so that the Kalman filter algorithm can fully consider the position uncertainty caused by rotation and attitude changes when predicting the target's position, making the prediction result more accurately reflect the actual position of the target at the next moment and avoiding large prediction errors caused by ignoring these dynamic change factors.

[0061] By the above method of dynamically adjusting the noise covariance matrix of the prediction model in the Kalman filter algorithm according to the historical acceleration and angular velocity of the target's motion, the Kalman filter algorithm can more adaptively handle the state changes of the target in complex and variable motion scenarios. It no longer uses a fixed noise covariance setting, but closely combines the actual motion characteristics of the target to optimize its prediction mechanism in real time. In this way, in the dynamic tracking application of the AI glasses, it can more accurately predict the position of the target in the next frame of the image, provide a more reliable basis for further accurately positioning the target using a twin network, etc., and ultimately improve the accuracy and stability of the entire AI glasses in tracking the target, and provide a better visual experience of target tracking for users.

[0062] In one embodiment, after achieving stable tracking of the target and outputting a sequence of tracking target images, it further includes:

[0063] Based on the generator, perform detail enhancement on the tracked target image sequence to obtain an enhanced image;

[0064] Based on the discriminator, determine the authenticity of the enhanced image. The generator and the discriminator are optimized against each other iteratively to improve the clarity and texture richness of the enhanced image, which is used to be displayed to the user on the AI glasses display screen.

[0065] In this embodiment, the above-mentioned generator is usually a part of the generative adversarial network (GAN) architecture in deep learning. Its core function is to generate new images based on the input tracked target image sequence, aiming to enhance the details of the original image. For example, it can learn the texture features, edge features, and color levels and other detail information of different objects in the image, and then through its own network structure (usually composed of multiple convolutional layers, activation functions, and other components to form a complex architecture), perform feature transformation and reconstruction, and refine the areas of the target image that may be relatively blurred and lack rich texture. For example, for a distant object in the tracked target image sequence, whose original outline is not clear and the surface texture is difficult to distinguish, the generator can use the general rules it has learned about the appearance characteristics of the object to add clearer contour lines and enrich its surface texture, so that the object shows more details in the image, thereby improving the visual effect of the entire tracked target image sequence.

[0066] The quality of the tracked target image sequence displayed to the user by the AI glasses directly affects the user's viewing experience. Through detail enhancement by the generator, users can more clearly see various features of the target. Especially in some complex environments or when the target itself is difficult to distinguish, clear and detail-rich images help users better understand key information such as the state and form of the target, thus meeting the user's demand for high-quality image display and making the function of the AI glasses more practical.

[0067] The above-mentioned discriminator is also an important part of the generative adversarial network (GAN) architecture. Its main task is to receive the enhanced images output by the generator and determine the authenticity of these images. Specifically, the discriminator measures the similarity between the received enhanced images and real images based on the statistical features, distribution laws, and other knowledge about real images that it has learned (which are learned through a large amount of real image data during the training phase). For example, it will judge from multiple dimensions such as whether the color distribution of the image is reasonable, whether the texture conforms to natural laws, and whether the structure of the object conforms to common sense. The discriminator is generally also a neural network composed of multiple convolutional layers, fully connected layers, etc. Its output is usually a numerical value representing the probability of the image's authenticity. For example, an output value close to 1 indicates that the discriminator believes that the image is very likely to be a real image, and close to 0 indicates that it is considered a fake image.

[0068] The discriminator exists to provide feedback to the generator, guiding the generator to generate enhanced images that are more in line with the true visual perception. In the entire technical solution, only when the generated enhanced images can deceive the discriminator, that is, when the discriminator has difficulty distinguishing them from real images, does it indicate that the images generated by the generator meet high-quality standards in terms of details, overall visual effects, etc., and meet the requirements of improving image clarity and texture richness. Therefore, the link of the discriminator judging the authenticity of enhanced images plays a crucial role in controlling the quality of the images finally presented to users.

[0069] The generator and the discriminator are in a relationship of mutual confrontation and collaborative progress in the entire solution. In the initial stage, the enhanced images generated by the generator may have many deficiencies in terms of details and authenticity, and the discriminator can easily identify them as fake images. Then, based on the discriminator's judgment results, the generator will adjust its own network parameters (for example, through the backpropagation algorithm, updating parameters such as the weights of the convolutional layer according to the error information feedback by the discriminator), and try to generate more realistic and detailed enhanced images. And the discriminator will also continuously optimize its own discriminative ability for image authenticity and improve the accuracy of judgment according to the newly received enhanced images output by the generator and the newly supplemented real image data (if any). This process will be repeated. In multiple iterations, the generator continuously improves its ability to generate high-quality enhanced images, and the discriminator also continuously improves its level of identifying image authenticity. The two are like a game, gradually reaching a relatively balanced state.

[0070] Through such an anti-iterative optimization process, the generator can finally output enhanced images with significantly improved clarity and texture richness. When these images are presented to users on the display screen of the AI glasses, they not only look more realistic but also enable users to more clearly observe the subtle details of the target, such as the texture patterns on the surface of the target object and tiny structural features, greatly optimizing the visual experience of users when using the AI glasses to observe and track the target, making the entire image processing function of the AI glasses more advantageous and better meeting the expectations of users for high-quality image display.

[0071] In one embodiment, preprocessing the original images collected in real time by the AI glasses based on the spatio-temporal attention mechanism to obtain preprocessed images, including:

[0072] Performing convolution on the original images in the spatial dimension using a three-dimensional convolutional neural network to extract local detail features and global semantic features in the images, obtaining spatial feature maps;

[0073] In the time dimension, for the original images, based on the long short-term memory network model, learning the temporal correlation between adjacent frames, capturing the dynamic change features of the image sequence over time, and outputting a temporal feature sequence;

[0074] Based on the attention mechanism, the spatial feature map and the temporal feature sequence are fused to obtain a spatio-temporal coherent feature map, and through normalization processing, the preprocessed image is obtained.

[0075] In this embodiment, traditional two-dimensional convolutional neural networks are mainly applied to process single images, focusing on mining the spatial information of images, such as extracting features like textures, edges, and shapes in images. While three-dimensional convolutional neural networks add considerations of the time dimension on this basis, however, in this step, the focus is on leveraging their convolutional capabilities in the spatial dimension. It slides a convolutional kernel of a specific size (such as a 3×3×3 convolutional kernel, where the first two dimensions correspond to the planar spatial dimensions of the image, and the third dimension can be used to handle different channels of the image, etc.) over the original image to perform convolutional operations. Such convolutional operations can extract features from local regions of the image. For example, for a local region where an object is located in the image, the convolutional kernel can capture the local texture details, such as whether the surface of the object has a smooth or rough texture; at the same time, as the convolutional kernel continuously slides over the entire image plane, it can gradually integrate the information of each local region, and then extract the global semantic features of the image, that is, understand what the image as a whole represents, such as determining whether the image shows a street scene or an indoor scene, etc. In this way, the feature information of the original image in the spatial dimension can be comprehensively and deeply mined to obtain a spatial feature map.

[0076] The spatial feature map carries the key feature information refined from the original image after spatial dimension convolutional processing, which abstracts and generalizes the complex pixel information in the original image at the feature level. These features are very important for accurately identifying targets in the image, distinguishing different objects, and understanding the image scene in the subsequent process.

[0077] The original images collected by the AI glasses are a continuous image sequence, and there is a temporal correlation between adjacent frames. The long short-term memory network model is very suitable for processing data with such time series characteristics. It can effectively solve the problems of gradient disappearance or gradient explosion that traditional neural networks are prone to when processing long sequence data through special gate structures (including input gates, forget gates, output gates, etc.), so that it can better learn the long-distance temporal dependence relationships in the image sequence. For example, for a moving target, its position, pose, etc. will change in different image frames at different times. The LSTM can remember the changing trend of the target's motion state, such as whether the target is moving at a constant speed or accelerating, based on the performance of the target in multiple previous frames, and then capture the dynamic change characteristics of the image sequence over time.

[0078] The output temporal feature sequence records the key change information of the image sequence in the temporal dimension, which reflects the evolution of the target and the entire image scene over time. For example, in a dynamic tracking scenario, the temporal feature sequence can help subsequent processes better predict the state of the target at the next moment, understand the changing trend of the target's motion trajectory, etc., providing an important basis in the temporal dimension for achieving stable and accurate tracking of the target, making the entire image processing process not only limited to the spatial information processing of a single image, but comprehensively considering the temporal dynamic changes.

[0079] The attention mechanism can simulate the difference in attention to different information in the human visual system. When fusing the spatial feature map and the temporal feature sequence, it can allocate different weights according to the importance of each of them for targeted fusion. For example, for a target that is moving rapidly, the information in the temporal feature sequence may be relatively more important because the position change of the target is relatively significant in a short period of time. At this time, the attention mechanism will assign a higher weight to the temporal feature sequence, making it play a more dominant role in the fused spatio-temporal coherent feature map; while for a relatively static background area, the information about its spatial features such as texture and layout in the spatial feature map is more worthy of attention, and appropriate weights will be given for fusion accordingly. In this way, the key features in the spatial and temporal dimensions can be organically combined to form a spatio-temporal coherent feature map, which not only contains rich details and semantic information of the image in space, but also incorporates the dynamic change features in the temporal dimension, more comprehensively reflecting the actual situation of the original image.

[0080] After obtaining the spatio-temporal coherent feature map, normalization processing is performed to normalize the feature values in the feature map to meet certain numerical range requirements, facilitating subsequent calculations and processing, and also helping to improve the stability and convergence speed of the model.

[0081] In one embodiment, based on the Siamese network, the appearance features of the target are learned and updated in real time to achieve stable tracking of the target and output a tracking target image sequence, including:

[0082] Crop the image region of the target tracked in the current frame and input it into the Siamese network, and use a convolutional neural network based on the residual structure to extract features from it to obtain a template feature vector containing the key features of the target;

[0083] For the entire image of the current frame, use a densely connected convolutional network to extract features to obtain a search feature map containing potential target regions;

[0084] Calculate the similarity between each potential target region in the search feature map and the template feature vector, and compare it with the similarity threshold; among them, construct a dynamic environment evaluation model, and adjust the similarity threshold in real time dynamically according to the light intensity change rate, the moving speed of background objects and the acceleration of the target movement in the current image;

[0085] Obtain the region with similarity greater than the similarity threshold and the maximum similarity, fuse the corresponding feature information with the template feature vector by assigning different weights, continuously update the template feature vector, realize stable tracking of the target and output the tracking target image sequence.

[0086] In this embodiment, during the process of tracking the target, every time the target is tracked in the current frame, the corresponding image region is cropped out. This can focus on the target itself and exclude a large amount of interference from irrelevant background information, enabling subsequent feature extraction operations to more intensively mine the key features of the target itself. For example, in a complex scene, if no cropping is performed, the entire image may contain many background objects, complex light and shadow changes, etc. factors. After cropping out the target region, it is possible to specifically process information closely related to target recognition and tracking, such as the shape, texture, color, etc. of the target, improving the efficiency and accuracy of feature extraction.

[0087] The convolutional neural network with residual structure has significant advantages in feature extraction. In the process of deepening the network layers to obtain more abstract and richer features, traditional convolutional neural networks are prone to problems such as gradient disappearance or gradient explosion, resulting in difficulties in network training and performance degradation. The residual structure, by introducing shortcut connections (i.e., skip connections that can directly pass the input information to the subsequent layers), enables the network to more easily learn the residual mapping between the input and output, effectively alleviating the above problems. When using such a network to extract features from the cropped target image region, it can extract the detailed features and abstract features of the target from different levels. For example, extract the basic features such as the edges and textures of the target from the shallow network layers, and extract the more semantic overall shape, category-related and other key features of the target from the deep network layers. Finally, integrate these features to form a template feature vector containing the key features of the target, providing an accurate target feature representation for subsequent similarity comparison and other operations.

[0088] The core feature of the DenseNet (Dense Convolutional Network) is that each layer is directly connected to all the previous layers, enabling more sufficient information transmission in the network and maximizing the reuse of features. When processing the entire image of the current frame, since the image may contain multiple potential target regions (for example, the target may not be easily located accurately due to occlusion, interference from similar objects, etc.), using this network structure can comprehensively extract various feature information in the image without missing any local region features that may contain the target. It can extract features from different angles and different local ranges of the image, and synthesize these features to form a search feature map containing potential target regions. This search feature map is actually a representation of the entire image at the feature level, providing rich candidate region information for subsequent searching for regions that match the target template feature vector.

[0089] The search feature map carries rich feature information of the current frame image, covering the feature clues that may exist for the target in various positions and local regions of the image. By performing operations such as calculating the similarity with the template feature vector in the subsequent process, the region most similar to the target can be located in this search feature map, and then the actual position of the target in the current frame image can be determined. Even if there are changes in the appearance of the target, the environment it is in, etc., effective searching and positioning can be carried out based on the feature information in the search feature map.

[0090] Calculating the similarity between each potential target region in the search feature map and the template feature vector is to find the region most likely to be the target in the current frame image. Commonly used similarity calculation methods, such as cosine similarity, Euclidean distance, etc., can measure the similarity degree between two feature vectors. By comparing the feature vector corresponding to each potential target region with the template feature vector, the corresponding similarity value is obtained, and then compared with the similarity threshold. If the similarity is greater than the threshold, it is considered that this region may be the target region; otherwise, it is probably not the target region. Such comparison operations can screen out the target candidate regions that meet the requirements, narrow the scope of target positioning, and improve the accuracy of target tracking.

[0091] In an actual tracking scenario, environmental factors are complex and variable. For example, when the change rate of light intensity is large, features such as the appearance color and light-dark contrast of the target will change, which will cause areas that were originally judged to be similar according to a fixed threshold to become dissimilar; when the moving speed of background objects is relatively fast, interference areas similar to the target are likely to be generated, which will also affect the similarity judgment; the change in the acceleration of the target's movement will also cause the appearance features of the target itself to show different degrees of change at different times. Constructing a dynamic environment assessment model is to consider in real time the impact of these environmental factors on the appearance features of the target and the similarity judgment. According to indicators such as the change rate of light intensity, the moving speed of background objects, and the acceleration of the target's movement, the similarity threshold is dynamically adjusted. For example, when the light intensity changes violently, the similarity threshold is appropriately relaxed to avoid the situation that the target features cannot be recognized due to slight changes in the light; when the moving speed of background objects is fast and there is more interference, the similarity threshold may be appropriately increased to enhance the strictness of screening and ensure that the real target area is found, so that the similarity judgment is more in line with the actual changing environment and improves the stability and accuracy of target tracking.

[0092] Obtaining the region with a similarity greater than the similarity threshold and the maximum similarity means finding the region in the current frame image that best conforms to the target features, that is, the position where the target is most likely to be located. Fusing the feature information corresponding to this region with the template feature vector takes into account that the appearance features of the target will continuously change during movement, such as angle changes and appearance changes caused by partial occlusion. By fusing with different weights, for example, a higher weight is given to the part of the newly obtained region features that is more reliable and related to the stable features of the target, while a lower weight is given to the part that may be affected by noise and temporary interference, so as to update the template feature vector, enabling it to keep up with the changes in the appearance features of the target in a timely manner and always maintaining an accurate description of the target features, thereby achieving stable tracking of the target.

[0093] Finally, an image sequence of the tracked target is output, enabling the user to intuitively see the state of the target at different times through the AI glasses. The continuous image sequence can clearly show the movement trajectory, appearance changes, etc. of the target, meeting the user's need for dynamic observation of the target, and also reflecting the integrity and practicality of the entire AI glasses in the target tracking function, providing a good user experience for the user.

[0094] In one embodiment, the method further includes:

[0095] Collecting the multi-modal interaction information of the user with the AI glasses within a specified time period; wherein, the multi-modal interaction information includes the voice information issued by the user, the touch instructions of the user to the AI glasses, and the image information collected by the user controlling the AI glasses;

[0096] Send the multi-modal interaction information to the server;

[0097] The server generates permission information based on the multi-modal interaction information and sends it to the AI glasses to configure the usage permissions of the AI glasses.

[0098] In this embodiment, the multi-modal interaction information covers data of the user's interaction with the AI glasses in multiple dimensions, specifically including the voice information sent by the user, the touch instructions of the user on the AI glasses, and the image information collected by the user controlling the AI glasses.

[0099] The voice information reflects the intentions and requirements conveyed by the user to the AI glasses in a voice manner. For example, voice commands such as "set permissions" and "magnify the object in the picture" are issued. It reflects the specific functional operations that the user expects the AI glasses to perform. It is an intuitive and convenient interaction method. Collecting this voice information helps to comprehensively understand various demand expressions of the user during the use process.

[0100] The touch instructions are the instructions generated by the user's touch operations on the AI glasses and are also important interaction data. For example, operations such as switching the function interface and adjusting the display brightness are achieved by touching specific areas. These touch instructions can reflect the user's habits in manual operation and the usage frequency of different functions, etc., and have important value for analyzing user behavior and personalized needs.

[0101] The image information collected by the user controlling the AI glasses records the visual content that the user is interested in, which may be a specific scene, a specific object, etc. This can reflect the user's interest points. For example, if the user frequently collects images of natural scenery, it means that the user may have higher requirements for functions related to such images (such as image enhancement, sharing, etc.).

[0102] Collecting the multi-modal interaction information within the specified time period can comprehensively master the interaction situation between the user and the AI glasses from multiple perspectives, providing a rich data basis for subsequent operations such as further analyzing user behavior and configuring corresponding permissions based on this information.

[0103] Sending the collected multi-modal interaction information to the server is to utilize the powerful computing, storage, and analysis capabilities of the server to process this data. Due to the limitations of the hardware resources of the AI glasses themselves (such as processor performance, storage space, etc.), it is difficult to independently complete operations such as in-depth analysis of a large amount of complex interaction information and permission configuration. Through network communication technology, the data is stably transmitted to the server to ensure that the server can receive complete and accurate multi-modal interaction information, enabling it to be used as input data for subsequent generation of permission information, thereby better overall management and allocation of the usage permissions of the AI glasses.

[0104] After receiving the multimodal interaction information, the server will use corresponding data analysis algorithms and permission configuration policies to generate permission information. For example, based on the functions corresponding to the voice commands frequently used by the user, the server can determine the core function modules commonly used by the user, and give higher usage permissions to these core functions to ensure that the user can continue to use them smoothly. For those functions that the user rarely touches, are relatively complex or have certain security risks (such as advanced setting functions involving access to private data), their permissions can be appropriately restricted based on comprehensive judgments such as their usage frequency and the user's operation proficiency.

[0105] Sending the generated permission information back to the AI glasses to configure their usage permissions has several important significances. On the one hand, it can enhance the user experience. By ensuring the high-permission and smooth use of the user's frequently used functions, the user can operate the AI glasses more proficiently and reduce the situation of operation obstruction caused by insufficient permissions. On the other hand, from the perspective of security and privacy protection, it can avoid problems such as data leakage and infringement of others' rights caused by the user's misoperation or inappropriate use, standardize the usage scope and method of the AI glasses, ensure the legality, security and rationality of the entire usage process, and also help to better manage and maintain the stable operation of the AI glasses system. In some embodiments, the above permission information can also be a permission verification code, and the user of the AI glasses can use all functions of the AI glasses by using this permission verification code.

[0106] In one embodiment, generating permission information based on the multimodal interaction information includes:

[0107] Performing speech recognition on the speech information to recognize keywords; encoding the keywords into numbers to obtain key numbers;

[0108] Identifying the touch command to obtain a touch pattern; sequentially adding the numeric characters in the key numbers along the touch direction of the touch pattern to obtain a numeric pattern;

[0109] Encoding the image information into a data stream string, and discretely distributing the data stream string to multiple adjacent regions of the numeric pattern;

[0110] In each of the adjacent regions, obtaining the characters that satisfy the preset association relationship with the numeric characters on the numeric pattern as target characters, and combining them to obtain the permission information.

[0111] In this embodiment, first, the collected voice information is subjected to speech recognition. This process uses speech recognition technology to convert the content expressed by the user in voice into a text form that can be understood by a computer. Then, keywords are extracted from these texts. Keywords are usually the core words that can reflect the user's intention and the functional requirements involved. For example, when the user says "I want to set permissions", after speech recognition, the extracted keywords are "set" and "permissions", etc. These keywords carry the key information of the user's operation expectations and are an important basis for subsequent further processing.

[0112] Furthermore, encoding the recognized keywords into numbers is to convert the information at the semantic level into a digital form that is convenient for subsequent calculation and processing. A pre-set encoding rule can be adopted. For example, a mapping table of keywords and numbers is established to map different keywords to specific numbers. Through such an encoding method, key numbers are obtained, making the voice information exist in a simple and standardized digital sequence form, which is convenient for subsequent operations such as fusing with other modal interaction information, and is also conducive to unified rule-based processing during the generation of permission information.

[0113] Regarding the touch instructions of the user on the AI glasses, it is necessary to use corresponding touch sensing and recognition technologies to clarify the specific touch operation situation, and then obtain a touch pattern. The touch pattern can be understood as an abstract representation on a plane of the user's touch behavior such as the touch path and touch range. For example, when the user slides a finger in the touch area of the AI glasses to form a circular touch trajectory, the corresponding circular touch pattern will be obtained after recognition; or when the user clicks several specific points in sequence, these points connected together also form a touch pattern, which intuitively reflects the trajectory and method of the user's manual operation.

[0114] After obtaining the touch pattern, along its touch direction, the numerical characters in the key numbers encoded from the voice information before are sequentially added to the touch pattern one by one. The purpose of doing this is to fuse the functional intention represented by the voice information with the operation behavior reflected by the touch instruction. By adding the numerical characters to the touch pattern, a new digital pattern is formed. This digital pattern not only contains the spatial information of the touch operation but also incorporates the digital encoding information corresponding to the voice instruction, making it a unique representation form that combines the interaction information of two modalities.

[0115] The image information itself is complex data composed of the color, brightness, and other information of numerous pixels. To facilitate the combined processing with other modal information, it needs to be encoded and converted into a data stream string. This can be achieved through an image encoding algorithm. For example, the image pixel data is arranged in a certain order (such as from left to right, from top to bottom), and an appropriate encoding method (such as binary encoding, etc.) is used to convert it into a continuous data stream string, so that the image information is presented in a linear and easy-to-process data form.

[0116] Dispersing the encoded data stream string into multiple adjacent regions of the digital graphic means integrating the image information into the digital graphic that has already incorporated voice and touch instruction information in a dispersed manner. Each adjacent region can be regarded as a unit carrying a part of the image data stream. Through this dispersed distribution method, the image information can be closely associated with the comprehensive interaction information represented by the digital graphic, further enriching the data content used to generate the permission information, so that the finally generated permission information can comprehensively consider all aspects of the multi-modal interaction information.

[0117] In each adjacent region, search for the characters associated with the digital characters on the digital graphic according to the pre-set association relationship. This pre-set association relationship can be set based on pre-set algorithm rules or mapping logic. For example, it is stipulated that in the adjacent region corresponding to the region where the number "1" is located in the digital graphic, search for the characters that conform to a certain encoding rule (such as the characters that meet the conditions after performing a certain operation on the ASCII code value of the character and the number, etc.). The characters found are the target characters. Or the positional relationship (distance, angle, etc.) between the characters in the adjacent region and the digital characters in the digital graphic can be calculated. When the positional relationship meets the pre-set conditions, it is used as the target character. In this way, the characters that meet the requirements are screened out from the complex data structure integrating multi-modal interaction information, and these characters carry the key information related to permission configuration extracted from different modal interaction information.

[0118] Combining the target characters obtained from each adjacent region gives the final permission information. The permission information is presented in the form of a character combination that has undergone multiple steps of processing and integrates the key information extracted from multi-modal interaction information. It can accurately reflect the usage permissions of the AI glasses configured based on the user's multi-faceted interaction behaviors such as voice, touch, and image acquisition. Furthermore, it can be sent back to the AI glasses for use as the permission code for AI verification. The AI glasses user can fully use the permissions of the AI glasses with this permission information, while other users without the permission information cannot use it.

[0119] In one embodiment, generating the permission information based on the multi-modal interaction information includes:

[0120] Perform speech recognition on the speech information to identify keywords; encode the keywords to obtain encoded characters;

[0121] Encode the image information into a data stream string, and the data stream string;

[0122] Based on the quantity characteristics and character type characteristics of the encoded characters, rearrange the data stream string to obtain a rearranged string;

[0123] Discretely distribute the rearranged string into a preset graph; wherein, the preset graph is composed of multiple circles with the same center and different radii;

[0124] Identify the touch command to obtain a touch graph; superimpose the touch graph onto the preset graph in a preset manner;

[0125] Obtain the characters at the positions where the touch graph intersects with the preset graph as target characters, and combine the target characters in the order from the outside to the inside to obtain a character combination as the permission information.

[0126] In this embodiment, the speech information input by the user through the AI glasses is converted into text form. Subsequently, keywords are screened out from the converted text. These keywords are usually words closely related to the operations the user expects to perform and the objects of concern. Perform an encoding operation on the identified keywords to convert the keywords at the semantic level into an encoded character form that is convenient for subsequent calculation and processing. A flexible and extensible encoding rule can be preset, such as assigning different character encodings according to factors such as the function category, occurrence frequency, and importance degree of the keywords.

[0127] Encoding the image information into a data stream string is to convert the complex image data at the visual level into a linear and computer - processable data form. By using algorithms such as lossless compression encoding, while retaining the key information of the image, the multi - dimensional information such as the color, brightness, and position of the image pixels can be arranged and combined in a specific order to generate a compact data stream string, enabling the image information to cooperate with information such as speech and touch commands in the same data processing framework.

[0128] Based on the quantity characteristics and character type characteristics of the encoded characters, rearrange the data stream string. For example, if there are more numeric characters in the encoded characters, then when rearranging the data stream string, the numeric character part is placed in the front; if there are more English characters, then when rearranging the data stream string, the English character part is placed in the front.

[0129] The preset graph consisting of multiple circles with the same center and different radii is selected because it has natural hierarchical and regional division characteristics. The rearranged character string is discretely distributed to each area of ​​the preset graph according to certain rules. For example, according to the order of the characters in the character string, the characters at the beginning of the test are placed in the inner circle area, and the characters at the end are placed in the outer circle area, so as to achieve a structured layout of the information.

[0130] Then, with the help of high-precision touch sensing technology, the user's touch operation trajectory, touch point distribution, touch duration and other information on the AI ​​glasses are identified, and a touch graph is abstracted. This touch graph intuitively reflects the user's manual control intention, such as the user drawing a right-pointing arrow or a closed circle.

[0131] According to the preset method, such as taking the center of the preset graphic as the reference point, the touch graphic is aligned and superimposed according to the geometric center position to ensure that the touch graphic and the preset graphic are accurately integrated in space. After the touch graphic and the preset graphic are superimposed, the intersection area of ​​the two contains characters that combine the key information of the three modes of voice, image, and touch. The characters at these intersections are extracted as target characters because they are information carriers after multi-modal interaction and fusion.

[0132] Combine each target character in order from the outside to the inside, and the final character combination is the permission information, which accurately reflects the usage permissions (permission code) that should be granted to AI glasses based on the user's multimodal interaction behavior, and can be used for subsequent permission configuration management.

[0133] In one embodiment, generating permission information based on the multimodal interaction information includes:

[0134] Analyze the intonation information of speech information and construct intonation curve;

[0135] Analyze the touch duration of touch instructions and construct a touch curve;

[0136] Analyze the color characteristics and color characteristic curves of each specified area of ​​the graphic information;

[0137] The intonation curve, the touch curve, and the color characteristic curve are superimposed in the same coordinate system to generate a combined curve; the combined curve is encoded to generate an authority code as the authority information.

[0138] In this embodiment, first, the intonation information of the voice information is analyzed to construct an intonation curve. When the user speaks to the AI ​​glasses, the voice analysis method is used to dig out the emotions and emphasis hidden in the intonation. By extracting fundamental frequency and other technologies, the intonation curve is drawn with time as the horizontal axis and intonation quantitative indicators as the vertical axis.

[0139] Next, analyze the touch duration of the touch command and construct a touch curve. The AI ​​glasses' high-sensitivity touch sensing hardware and time recording module can accurately measure the duration of each touch from start to end. Short clicks are often used to quickly select basic functions, such as switching display modes; long presses are often used to go into deep settings and activate complex functions, such as calling up image detail adjustments. Draw a curve with the touch time sequence as the horizontal axis and the duration as the vertical axis to show the user's manual control rhythm and gain insight into their operating habits and current needs.

[0140] Then, the color characteristics of each designated area of ​​the graphic information are analyzed to construct a color characteristic curve. When processing images collected by users, the areas are divided according to preset rules, such as the center of the picture, the target hot spot area, etc. For each area, the image color analysis algorithm is used to extract features such as average color value, contrast, saturation, etc., and a curve is constructed with the area identifier as the horizontal axis and the characteristic value as the vertical axis. The color dimension reflects the image dynamics or regional differences, adding visual perception characteristics to permission generation.

[0141] After that, the intonation curve, touch curve, and color feature curve are superimposed on the same coordinate system to generate a combined curve. With time or operation sequence as the horizontal axis, the three curves are merged to intuitively present the coordinated changes of multimodal information. For example, when the peak of intonation, long touch, and high contrast of color in the key area of ​​the image are synchronized, it may indicate that the user has made a fine operation on a specific area of ​​the screen. The combined curve fully captures the cross-modal linkage information.

[0142] Finally, the combined curve is encoded and a permission code is generated as permission information. The coding system is designed based on the geometric characteristics of the combined curve, such as the peak value, valley value, slope, area under the curve, and its distribution in different time periods and regions. The complex curve information is converted into a simple permission code to obtain the AI ​​glasses usage permission that the user deserves for multimodal interaction, which is used for subsequent permission configuration to ensure that it meets the needs and is safe and controllable.

[0143] Reference Figure 2 In another embodiment of the present invention, a multimodal interaction and image processing device for AI glasses is provided, comprising:

[0144] A processing unit is used to pre-process the original image collected by the AI ​​glasses in real time based on the spatiotemporal attention mechanism to obtain a pre-processed image; perform speech recognition and semantic analysis on the voice command input by the user through the AI ​​glasses to obtain a text command;

[0145] A matching unit, used for matching the text instruction with a preset image processing instruction template library by keywords to determine a corresponding image processing strategy;

[0146] A positioning unit, which is used to, if the image processing strategy is dynamic tracking, predict the position of the target in the preprocessed image in the next frame by using the Kalman filtering algorithm, and search around the predicted position by using a Siamese network to accurately locate the target;

[0147] A tracking unit, which is used to continuously learn and update the appearance features of the target based on the Siamese network, realize stable tracking of the target, and output a sequence of tracking target images for display to the user on the display screen of the AI glasses.

[0148] In this embodiment, for the specific implementation of each unit in the above device embodiment, please refer to the description in the above method embodiment, and details will not be repeated here.

[0149] Refer to Figure 3 , in an embodiment of the present invention, an electronic device is further provided. The internal structure of the electronic device may be as Figure 3 shown. The electronic device includes a processor, a memory, a display screen, an input device, a network interface, and a database connected through a system bus. Among them, the processor of this computer design is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the electronic device is used to store the corresponding data in this embodiment. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.

[0150] Those skilled in the art can understand that Figure 3 the structure shown in

[0151] is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the electronic device to which the solution of the present invention is applied.

[0152] In summary, the multimodal interaction and image processing method, electronic device, and storage medium of the AI glasses provided in the embodiments of the present invention include: preprocessing the original image collected by the AI glasses in real time based on a spatio-temporal attention mechanism to obtain a preprocessed image; performing speech recognition and semantic parsing on the voice command input by the user through the AI glasses to obtain a text command; performing keyword matching between the text command and a preset image processing command template library to determine the corresponding image processing strategy; if the image processing strategy is dynamic tracking, using the Kalman filter algorithm to predict the position of the target in the next frame of the preprocessed image, and using a Siamese network to search around the predicted position to accurately locate the target; based on the Siamese network, continuously learning and updating the appearance features of the target in real time to achieve stable tracking of the target and output a sequence of tracked target images for display to the user on the display screen of the AI glasses. In the present invention, the position of the target in the next frame of the image can be predicted, and then the target can be accurately located, and the appearance features of the target can be continuously learned and updated in real time to achieve stable tracking of the target. It overcomes the defects that the subsequent appearance position of the target cannot be predicted and the target cannot be stably tracked currently.

[0153] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in the present invention and in the embodiments can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be obtained in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.

[0154] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, apparatus, article or method comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, apparatus, article or method. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, apparatus, article or method comprising such element.

[0155] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present invention by the same token.

Claims

1. A multimodal interaction and image processing method for AI glasses, characterized in that: The following steps are involved: Based on the spatiotemporal attention mechanism, the original images collected by the AI ​​glasses in real time are preprocessed to obtain preprocessed images; the voice commands input by the user through the AI ​​glasses are subjected to speech recognition and semantic analysis to obtain text commands; Perform keyword matching between the text instruction and a preset image processing instruction template library to determine a corresponding image processing strategy; If the image processing strategy is dynamic tracking, the Kalman filter algorithm is used to predict the position of the target in the preprocessed image in the next frame, and the twin network is used to search around the predicted position to accurately locate the target; Based on the twin network, the appearance features of the target are updated in real time to achieve stable tracking of the target and output a tracking target image sequence for display to the user on the AI ​​glasses display screen; The method further comprises: Collecting multimodal interaction information between the user and the AI ​​glasses within a specified time period; wherein the multimodal interaction information includes voice information sent by the user, touch instructions of the user to the AI ​​glasses, and image information collected by the user controlling the AI ​​glasses; Sending the multimodal interaction information to a server; The server generates permission information based on the multimodal interaction information, and sends the permission information to the AI ​​glasses to configure the usage permission of the AI ​​glasses; Generating permission information based on the multimodal interaction information includes: Performing speech recognition on the speech information to identify keywords; encoding the keywords into numbers to obtain key numbers; The touch instruction is identified to obtain a touch pattern; along the touch direction of the touch pattern, the digital characters in the key number are added to the touch pattern one by one in sequence to obtain a digital pattern; Encoding the image information into a data stream character string, and discretely dividing the data stream character string into a plurality of adjacent areas of the digital pattern; In each of the adjacent areas, characters that satisfy a preset association relationship with the digital characters on the digital graphic are obtained as target characters, and the authority information is obtained by combining them.

2. The multimodal interaction and image processing method of AI glasses according to claim 1, characterized in that: in, The Kalman filter algorithm dynamically adjusts the noise covariance matrix of the prediction model in the Kalman filter algorithm according to the historical acceleration and angular velocity of the target motion to improve the prediction accuracy.

3. The multimodal interaction and image processing method of AI glasses according to claim 1, characterized in that: After achieving stable tracking of the target and outputting the tracking target image sequence, it also includes: Performing detail enhancement on the tracking target image sequence based on the generator to obtain an enhanced image; Based on the discriminator determining the authenticity of the enhanced image, the generator and the discriminator are anti-iteratively optimized to improve the clarity and texture richness of the enhanced image for display to the user on the AI ​​glasses display.

4. The multimodal interaction and image processing method of AI glasses according to claim 1, characterized in that: The original images collected by AI glasses in real time are preprocessed based on the spatiotemporal attention mechanism to obtain preprocessed images, including: Use a three-dimensional convolutional neural network to convolve the original image in the spatial dimension, extract local detail features and global semantic features in the image, and obtain a spatial feature map; In the time dimension, for the original image, based on the long short-term memory network model, the temporal correlation between adjacent frames is learned, the dynamic change characteristics of the image sequence over time are captured, and the temporal feature sequence is output; Based on the attention mechanism, the spatial feature map is fused with the temporal feature sequence to obtain a spatiotemporal coherent feature map, and the preprocessed image is obtained through normalization.

5. The multimodal interaction and image processing method of AI glasses according to claim 1, characterized in that: Based on the twin network, the appearance features of the target are updated in real time through learning, the stable tracking of the target is achieved, and a tracking target image sequence is output, including: The image area of ​​the target tracked in the current frame is cropped and input into the twin network. The convolutional neural network based on the residual structure is used to extract its features to obtain the template feature vector containing the key features of the target. For the entire image of the current frame, a densely connected convolutional network is used to extract features to obtain a search feature map containing potential target areas; Calculate the similarity between each potential target area in the search feature map and the template feature vector, and compare it with the similarity threshold; construct a dynamic environment assessment model to dynamically adjust the similarity threshold in real time according to the rate of change of light intensity in the current image, the moving speed of background objects and the acceleration of target motion; The area with the greatest similarity and a similarity greater than the similarity threshold is obtained, and its corresponding feature information is fused with the template feature vector by assigning different weights so that the template feature vector is continuously updated to achieve stable tracking of the target and output the tracking target image sequence.

6. A multimodal interaction and image processing device for AI glasses, characterized in that: include: A processing unit, used for preprocessing the original image collected by the AI ​​glasses in real time based on the spatiotemporal attention mechanism to obtain a preprocessed image; Perform speech recognition and semantic analysis on the voice commands input by the user through AI glasses to obtain text commands; A matching unit, used for matching the text instruction with a preset image processing instruction template library by keywords to determine a corresponding image processing strategy; A positioning unit, which is used to predict the position of the target in the preprocessed image in the next frame by using a Kalman filter algorithm if the image processing strategy is dynamic tracking, and to search around the predicted position by using a twin network to accurately locate the target; A tracking unit, used to learn and update the appearance features of the target in real time based on the twin network, to achieve stable tracking of the target and output a tracking target image sequence for display to the user on the AI ​​glasses display; Also includes: Collecting multimodal interaction information between the user and the AI ​​glasses within a specified time period; wherein the multimodal interaction information includes voice information sent by the user, touch instructions of the user to the AI ​​glasses, and image information collected by the user controlling the AI ​​glasses; Sending the multimodal interaction information to a server; The server generates permission information based on the multimodal interaction information, and sends the permission information to the AI ​​glasses to configure the usage permission of the AI ​​glasses; Generating permission information based on the multimodal interaction information includes: Performing speech recognition on the speech information to identify keywords; encoding the keywords into numbers to obtain key numbers; The touch instruction is identified to obtain a touch pattern; along the touch direction of the touch pattern, the digital characters in the key number are added to the touch pattern one by one in sequence to obtain a digital pattern; Encoding the image information into a data stream character string, and discretely dividing the data stream character string into a plurality of adjacent areas of the digital pattern; In each of the adjacent areas, characters that satisfy a preset association relationship with the digital characters on the digital graphic are obtained as target characters, and the authority information is obtained by combining them.

7. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Interactive remote expert cooperation maintenance system and method based on augmented reality technology

    CN106339094A

  • Anti-occlusion target tracking method fusing Kalman filtering and particle filtering

    CN119399243A