Visual interactive control methods and devices, electronic devices and storage media
By automatically identifying visual dependency types and collecting images through smart glasses, and combining them with a multimodal intent analysis model, the problem of modal fragmentation in human-computer interaction of smart glasses devices has been solved, realizing the fusion of semantic understanding and visual perception, and improving the ability to handle complex real-world scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN EVERBEST MACHINERY IND
- Filing Date
- 2026-01-20
- Publication Date
- 2026-05-26
AI Technical Summary
Existing smart glasses devices suffer from a fragmented modal perception problem in human-computer interaction. Voice interaction and visual acquisition capabilities cannot be effectively combined, resulting in an inability to establish effective perception of the physical world in real-time interaction and limiting the ability to handle complex real-world scenarios.
The system acquires target interaction commands through smart glasses, determines the interaction type, and automatically controls image acquisition when the interaction type depends on the visual input. It then uses a pre-trained multimodal intent analysis model to fuse voice commands and visual images, identifies the user's visual intent, calls the corresponding business processing interface for professional processing, and finally generates feedback information.
It enables smart devices to proactively perceive visual information in the early stages of interaction, accurately identify complex user interaction needs, build a closed loop from intent recognition to business execution, improve the ability to handle complex real-world scenarios, and ensure a natural and smooth interaction process.
Smart Images

Figure CN122086237A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a visual interaction control method and device, electronic device and storage medium. Background Technology
[0002] With the development of electronic information technology, wearable smart devices, especially smart glasses, are gradually becoming the next generation of mobile computing platforms after smartphones. These devices typically integrate optical display modules, audio acquisition units, and image sensors, aiming to free users' hands and provide them with an immersive first-person perspective experience and convenient information interaction services. Through these devices, users can obtain information related to their surrounding environment anytime, anywhere.
[0003] In related technologies, the human-computer interaction logic of smart glasses often suffers from the problem of modal perception fragmentation. Although current devices have both voice interaction and visual acquisition capabilities, these two modalities are often independent of each other. When environmental information is involved in natural dialogue, existing systems lack the ability to capture the logical connection between voice commands and environmental vision, which makes it impossible for devices to establish effective perception of the physical world in real-time interaction, thus limiting the ability of smart devices to handle complex real-world scenarios. Summary of the Invention
[0004] The main objective of this application is to propose a visual interaction control method, device, electronic device, and storage medium that can proactively identify the user's visual interaction needs based on the dialogue context and automatically mobilize visual perception capabilities to achieve the fusion of semantic understanding and visual perception, thereby improving the ability of intelligent devices to handle complex real-world scenarios.
[0005] To achieve the above objectives, a first aspect of this application proposes a visual interaction control method, the method comprising: In response to a target interaction command issued by the smart glasses, the target interaction type of the target interaction command is determined; When the target interaction type is characterized as a visually dependent type, an image acquisition command is sent to the smart glasses, and the target image acquired by the smart glasses according to the image acquisition command is received. The target interaction command and the target image input are processed by a pre-trained multimodal intent analysis model to determine the target visual intent; The corresponding target business processing interface is determined based on the target visual intent, and the target business processing interface is called to process the target image to obtain the business processing result; Target feedback information is generated based on the business processing results.
[0006] In some embodiments, determining the target interaction type of the target interaction instruction includes: Keyword extraction is performed on the target interaction command to obtain the target keywords; If the target keyword is a preset spatial indicator pronoun, then the target interaction type is determined to be the visual dependency type; or, The target interaction instructions are semantically parsed to determine the target task category; The target task category is matched with a preset scene logic library. If the target task category belongs to a predefined visual entity category in the scene logic library, the target interaction type is determined to be the visual dependency type.
[0007] In some embodiments, the method further includes: If the target interaction type is characterized as a non-visual dependent type, then a pre-set general dialogue model is invoked to perform text semantic processing on the target interaction instruction and generate text response information. The text reply information is sent to the smart glasses.
[0008] In some embodiments, before processing the target interaction instruction and the target image input into a pre-trained multimodal intent analysis model to determine the target visual intent, the method further includes: The sharpness and exposure of the target image are obtained, and it is determined whether the sharpness and exposure meet a preset image quality threshold. The multimodal intent analysis model is used to perform image-text relevance verification on the target image and the target interaction command, and to determine whether the visual entities extracted from the target image match the semantic entities in the target interaction command. If the clarity or exposure does not meet the preset image quality threshold, or if the image-text correlation check fails, a guidance prompt message is generated and used as the target feedback message.
[0009] In some embodiments, if the target visual intent is a visual navigation intent, the step of determining the corresponding target service processing interface based on the target visual intent and calling the target service processing interface to process the target image to obtain a service processing result includes: The target business processing interface is determined to be a navigation service interface; Visual landmark information is extracted from the target image, and real-time positioning data is obtained; Perform a spatial topology consistency check on the visual landmark information and the real-time positioning data; If the verification passes, the navigation service interface is invoked to generate path planning data based on the visual landmark information, which is then used as the business processing result.
[0010] In some embodiments, if the target visual intent is a product price comparison intent, the step of determining the corresponding target business processing interface based on the target visual intent and calling the target business processing interface to process the target image to obtain a business processing result includes: The target business processing interface is identified as an e-commerce price comparison interface; Identify the brand characteristics, model characteristics, and specifications of the target product in the target image; Based on the identified features and parameters, the e-commerce price comparison interface is called to obtain a real-time price list, which is then used as the result of the business processing.
[0011] In some embodiments, if the target visual intent is a general visual question-and-answer intent, the step of determining the corresponding target business processing interface based on the target visual intent and calling the target business processing interface to process the target image to obtain a business processing result includes: The target business processing interface is determined to be a general visual large model interface; The general visual large model interface is invoked to perform image content understanding on the target image, and natural language answer information for the target image is generated in combination with the target interaction instructions. The natural language answer information is used as the business processing result.
[0012] To achieve the above objectives, a second aspect of this application provides a visual interaction control device, comprising: A visual interaction control system, the visual interaction control system being used to execute the visual interaction control method as described in the first aspect; The smart glasses are used to acquire the target interaction command and send the target interaction command to the visual interaction control system, perform visual image acquisition in response to the image acquisition command issued by the visual interaction control system, send the acquired target image to the visual interaction control system, and receive and output the target feedback information issued by the visual interaction control system.
[0013] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect of the embodiment.
[0014] To achieve the above objectives, a fourth aspect of the present application provides a storage medium, which is a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect of the present application.
[0015] The visual interaction control method, device, electronic device, and storage medium proposed in this application have the following beneficial effects: First, by acquiring the target interaction command and determining its target interaction type, it is possible to proactively determine whether there is visual dependency at the semantic level in the early stages of interaction. When it is determined to be a visual dependency type, the smart glasses are automatically controlled to acquire images. This mechanism changes the passive mode of relying on manual triggering or explicit commands to wake up visual hardware in traditional technologies, enabling smart devices to proactively perceive and acquire the required visual information based on the dialogue context, realizing on-demand automatic mobilization of perception capabilities, and ensuring the naturalness and smoothness of the interaction process. Second, this application inputs the target interaction command and the acquired target image into a pre-trained multimodal intent analysis model for joint analysis. Through this deep fusion processing of multimodal data, speech and visual information are no longer processed in isolation, but rather the semantic background of the command and the visual features of the image can be combined to accurately lock the user's target visual intent. This processing method effectively solves the problem of intent understanding bias under a single modality, ensuring accurate identification of the user's complex interaction needs. Finally, this application matches the corresponding target business processing interface according to the determined target visual intent, and directly generates feedback information for output based on the business processing results returned by the interface. Through this intent-based targeted service distribution mechanism, the system can automatically invoke appropriate backend service resources to perform professional image processing based on different visual intents, and directly feed the processing results back to the user. This constructs a closed loop from intent recognition to service execution and result feedback, avoiding the problem that visual recognition only stays at a superficial labeling level and cannot solve actual tasks. In summary, this application can proactively identify the user's visual interaction needs based on the dialogue context and automatically mobilize visual perception capabilities, achieving the fusion of semantic understanding and visual perception, thereby improving the ability of intelligent devices to handle complex real-world scenarios. Attached Figure Description
[0016] Figure 1 This is a flowchart of the visual interaction control method provided in the embodiments of this application; Figure 2 This is another flowchart of the visual interaction control method provided in the embodiments of this application; Figure 3 This is another flowchart of the visual interaction control method provided in the embodiments of this application; Figure 4 This is another flowchart of the visual interaction control method provided in the embodiments of this application; Figure 5This is another flowchart of the visual interaction control method provided in the embodiments of this application; Figure 6 This is another flowchart of the visual interaction control method provided in the embodiments of this application; Figure 7 This is a schematic diagram of the visual interaction control device provided in the embodiments of this application; Figure 8 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0018] It should be noted that although functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0020] With the development of electronic information technology, wearable smart devices, especially smart glasses, are gradually becoming the next generation of mobile computing platforms after smartphones. These devices typically integrate optical display modules, audio acquisition units, and image sensors, aiming to free users' hands and provide them with an immersive first-person perspective experience and convenient information interaction services. Through these devices, users can obtain information related to their surrounding environment anytime, anywhere.
[0021] In related technologies, the human-computer interaction logic of smart glasses often suffers from the problem of modal perception fragmentation. Although current devices have both voice interaction and visual acquisition capabilities, these two modalities are often independent of each other. When environmental information is involved in natural dialogue, existing systems lack the ability to capture the logical connection between voice commands and environmental vision, which makes it impossible for devices to establish effective perception of the physical world in real-time interaction, thus limiting the ability of smart devices to handle complex real-world scenarios.
[0022] Based on this, embodiments of this application provide a visual interaction control method and device, electronic device and storage medium, which can actively identify the user's visual interaction needs based on the dialogue context and automatically mobilize visual perception capabilities to achieve the fusion of semantic understanding and visual perception, thereby improving the ability of intelligent devices to handle complex real-world scenarios.
[0023] The visual interaction control method, device, electronic device, and storage medium provided in the embodiments of this application are specifically described through the following embodiments. First, the visual interaction control method in the embodiments of this application is described.
[0024] The visual interaction control method in this application can be illustrated through the following embodiments.
[0025] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. For example, when obtaining user-stored data and user cached data access requests, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards of the relevant countries and regions. In addition, when embodiments of this application need to obtain sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to a confirmation page. Only after obtaining the user's separate permission or consent is the necessary user-related data for the normal operation of the embodiments of this application obtained.
[0026] Figure 1 This is an optional flowchart of the visual interaction control method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps 101 to 105. It is also understood that this embodiment... Figure 1 The order of steps 101 to 105 is not specifically limited. The order of steps can be adjusted or some steps can be reduced or added according to actual needs.
[0027] Step 101: In response to the target interaction command issued by the smart glasses, determine the target interaction type of the target interaction command.
[0028] Step 102: When the target interaction type is characterized as a visually dependent type, send an image acquisition command to the smart glasses and receive the target image acquired by the smart glasses according to the image acquisition command.
[0029] Step 103: Process the target interaction command and the target image input pre-trained multimodal intent analysis model to determine the target visual intent.
[0030] Step 104: Determine the corresponding target business processing interface based on the target visual intent, and call the target business processing interface to process the target image to obtain the business processing result.
[0031] Step 105: Generate target feedback information based on the business processing results.
[0032] This application first acquires the target interaction command and determines its target interaction type, enabling proactive semantic-level judgment of visual dependency at the initial stage of interaction. When a visual dependency is determined, the smart glasses are automatically controlled to acquire images. This mechanism changes the passive mode of traditional technology, which relies on manual triggering or explicit commands to wake up visual hardware. It allows smart devices to proactively perceive and acquire the required visual information based on the dialogue context, achieving on-demand automatic mobilization of perception capabilities and ensuring a natural and smooth interaction process. Second, this application inputs the target interaction command and the acquired target image into a pre-trained multimodal intent analysis model for joint analysis. Through this deep fusion processing of multimodal data, speech and visual information are no longer processed in isolation, but rather the semantic background of the command and the visual features of the image are combined to accurately pinpoint the user's target visual intent. This processing method effectively solves the problem of intent understanding bias under a single modality, ensuring accurate identification of complex user interaction needs. Finally, this application matches the corresponding target business processing interface according to the determined target visual intent and directly generates feedback information for output based on the business processing results returned by the interface. Through this intent-based targeted service distribution mechanism, the system can automatically invoke appropriate backend service resources to perform professional image processing based on different visual intents, and directly feed the processing results back to the user. This constructs a closed loop from intent recognition to service execution and result feedback, avoiding the problem that visual recognition only stays at a superficial labeling level and cannot solve actual tasks. In summary, this application can proactively identify the user's visual interaction needs based on the dialogue context and automatically mobilize visual perception capabilities, achieving the fusion of semantic understanding and visual perception, thereby improving the ability of intelligent devices to handle complex real-world scenarios.
[0033] In step 101 of some embodiments, the target interaction instruction refers to voice data input by the user through the audio acquisition module (such as a microphone array) of the smart glasses, or text or instruction code generated through touch interaction or gesture interaction. This instruction carries the user's current immediate needs. Determining the target interaction type refers to performing semantic dimension parsing on the above-mentioned unstructured interaction instruction to identify the task attribute label corresponding to the instruction. Specifically, after the server or local control center receives the instruction data packet uploaded by the smart glasses, it uses natural language processing (NLP) algorithms to extract features from the verbs, nouns, and context of the instruction, thereby classifying the instruction into a visually dependent type that requires visual information intervention (e.g., instructions involving "identifying objects" or "finding routes") or a non-visually dependent type that only requires plain text processing (e.g., instructions involving "checking the weather" or "chatting"). This serves as the logical basis for subsequent judgment on whether to activate the visual hardware.
[0034] Please see Figure 2 In some embodiments, the step of determining the target interaction type of the target interaction instruction in step 101 may include, but is not limited to, steps 201 to 204.
[0035] Step 201: Extract keywords from the target interaction command to obtain the target keywords.
[0036] Step 202: If the target keyword is a preset spatial indicator pronoun, then the target interaction type is determined to be a visual dependency type.
[0037] Step 203: Perform semantic parsing on the target interaction instructions to determine the target task category.
[0038] Step 204: Match the target task category with the preset scene logic library. If the target task category belongs to the predefined visual entity category in the scene logic library, then determine the target interaction type as a visual dependency type.
[0039] In step 201 of some embodiments, target keywords refer to lexical units with significant semantic features extracted from unstructured speech-transcribed text or input commands input by the user. Keyword extraction from target interactive commands involves using word segmentation algorithms, part-of-speech tagging, or key phrase extraction algorithms in natural language processing to decompose continuous natural language commands into a sequence, and then filtering out semantically meaningless function words using a pre-set stop word list to select nouns, pronouns, or verbs that represent the core components of the sentence as target keywords. For example, in the command "Look at what this is," the extracted target keywords include "look," "this," and "what."
[0040] In step 202 of some embodiments, the preset spatial indicator pronouns refer to a set of pronouns used in linguistics to indicate spatial location or specific objects, such as "this," "that," "here," "over there," "before one's eyes," and "ahead." These words have a contextual relevance in human-computer interaction, implying that the object of the user's attention is located in the current physical space. When the system detects that the target keyword matches the preset indicator pronoun vocabulary, it immediately determines that the intent of the instruction is strongly related to the current environment, thereby directly identifying the target interaction type as a visually dependent type. This rule-based rapid determination mechanism can respond to explicit visual interaction needs with extremely low computational latency.
[0041] In step 203 of some embodiments, the target task category is a functional label that summarizes the underlying intent behind the user's instruction. Semantic parsing of the target interaction instruction refers to using an intent recognition model or a deep semantic understanding algorithm based on dependency parsing to analyze the logical relationship between the predicate verb and object of the instruction, thereby inferring the specific task type the user wants to perform. For example, for the instruction "Help me compare prices," although it does not contain demonstrative pronouns such as "this," semantic parsing can determine its target task category as "shopping / price comparison"; for the instruction "How do I get there?", its target task category is parsed as "navigation / wayfinding."
[0042] In step 204 of some embodiments, the preset scene logic library is a configuration table that defines the mapping relationship between different task categories and hardware modal requirements. The visual entity category refers to the set of tasks that, in business logic, must rely on image data to complete. This step uses the target task category determined in step 203 as the query key to perform a search and match in the scene logic library. If the task category (e.g., "text translation") is marked as a visual entity category, it indicates that the task relies on visual input, thus determining the target interaction type as a visually dependent type. For example, when the parsed task category is object recognition, even if no spatial indicator pronoun appears in the instruction, the logic library matching result will still trigger a visual dependency determination.
[0043] Through steps 201 to 204 above, the embodiments of this application, on the one hand, can quickly respond to explicit referential commands from users by recognizing spatial indicator pronouns, which aligns with human intuition; on the other hand, through semantic parsing and scene logic matching, they can deeply understand complex intentions that do not contain indicator pronouns but are essentially dependent on visual information, effectively avoiding missed triggers. This complementary judgment logic not only ensures the trigger accuracy of visual perception and maximizes the coverage of users' multimodal interaction scenarios, but also avoids privacy violations and power consumption waste caused by mistakenly activating the camera in non-visual scenarios.
[0044] In some embodiments, the method provided in this application may further include, but is not limited to, the following steps: If the target interaction type is represented as a non-visual dependent type, a pre-built general dialogue model is invoked to perform text semantic processing on the target interaction command and generate text response information. Send a text reply to the smart glasses.
[0045] In some embodiments, when the target interaction type is determined to be a non-visually dependent type, a plain text task routing strategy is executed. Non-visually dependent types refer to interaction requests that can be completed independently without the aid of visual information, such as encyclopedic Q&A ("Who is Newton?"), logical calculations ("What does 1 plus 1 equal?"), or casual conversation. In this case, the system does not send a wake-up signal to the smart glasses' camera module, ensuring the camera is in a switched-off or low-power standby state, and routes the target interaction command to a pre-built general dialogue model. This general dialogue model typically refers to a generative large language model trained on a large-scale plain text corpus, which focuses on natural language understanding, reasoning, and generation, without including a visual encoder component, thus offering higher inference speed and lower computational cost when processing plain text tasks.
[0046] In the subsequent processing, the system invokes a general dialogue model to perform text semantic processing on the target interaction command. Specifically, the model first converts the command text into high-dimensional word embedding vectors and performs attention calculations based on the historical context of the current session to capture semantic dependencies within the command. Subsequently, the model predicts and generates a natural language text sequence that conforms to grammatical rules and logical facts, i.e., the text response information. For example, for the command "What's the weather like today?", the model invokes a weather plugin or retrieves information from a knowledge base to generate the text response "Today it's cloudy turning sunny, temperature 25 degrees Celsius".
[0047] The generated text response is then sent to the smart glasses. This process involves encapsulating the text data into lightweight communication protocol packets and transmitting them to the smart glasses' communication module via Bluetooth Low Energy (BLE) or Wi-Fi. Upon receiving the information, the smart glasses, based on the user's output preferences, either invoke a local text-to-speech engine to convert it into an audio stream for playback, or invoke an optical display driver engine to render the text information onto the lenses, thus providing real-time feedback to the user.
[0048] Through the above steps, the embodiments of this application accurately identify non-visual tasks in the early stages of interaction and divert them to the plain text processing link with lower computational cost. More importantly, this mechanism eliminates the need to turn on the camera in unnecessary scenarios, reduces the hardware power consumption of smart glasses, and also eliminates users' privacy concerns about the device being filmed at all times, ensuring the performance, power consumption and privacy security of the multimodal interaction system.
[0049] In step 102 of some embodiments, the visual dependency type indicates that the current user's interaction intent cannot be achieved solely through textual semantics and must be combined with visual information from the physical environment. The image acquisition command is a control signal containing a specific hardware control protocol, sent to the underlying firmware of the smart glasses via Bluetooth, WiFi, or other near-field communication links. Upon receiving this signal, the smart glasses activate the camera module in the background to capture the scene within the user's current field of view, generating digital image data containing environmental information as the target image. This process achieves automated control logic that automatically acquires visual input based solely on semantic triggering without requiring manual operation of the physical shutter, and then transmits the acquired image data back to the control terminal.
[0050] In step 103 of some embodiments, the pre-trained multimodal intent analysis model refers to a deep neural network model (such as the Transformer architecture) trained on a large-scale image and text dataset, capable of simultaneously processing textual and visual modal data. Inputting the target interaction command and target image into the model means aligning and interacting the text vector of the command with the visual feature vector of the image at the model's feature fusion layer. Determining the target visual intent means that the model infers the user's specific task objective in the current visual scene based on the fused multimodal features. For example, when the command is "How much is this?" and the image contains a product, the model outputs the intent of comparing product prices; when the command is "What is that building in front?" and the image contains a landmark building, the model outputs the intent of visual landmark recognition or navigation positioning.
[0051] Please see Figure 3 In some embodiments, steps 301 to 303 may be included before step 103.
[0052] Step 301: Obtain the sharpness and exposure of the target image, and determine whether the sharpness and exposure meet the preset image quality threshold.
[0053] Step 302: Use a multimodal intent analysis model to perform image-text correlation verification on the target image and the target interaction command, and determine whether the visual entities extracted from the target image match the semantic entities in the target interaction command.
[0054] Step 303: If the clarity or exposure does not meet the preset image quality threshold, or the image-text correlation verification fails, a guidance prompt message is generated and used as the target feedback message.
[0055] In step 301 of some embodiments, obtaining the sharpness and exposure of the target image and determining whether the sharpness and exposure meet a preset image quality threshold involves using image processing operators in the field of computer vision to perform statistical feature calculations on the original pixel data. Specifically, for sharpness detection, the Laplacian operator can be used to perform convolution operations on the target image to calculate the second derivative response of the image pixel grayscale values in the spatial domain and to calculate the variance of this response value. This variance value characterizes the richness of high-frequency components at the image edges; the lower the variance value, the more severe the loss of high-frequency information due to focus failure or motion blur. For exposure detection, the pixel brightness distribution is statistically analyzed by calculating the average pixel value in the brightness channel of the image or by constructing a grayscale histogram. If the average brightness value is lower than a preset dark threshold (e.g., pixel value less than 40), or the proportion of pixels in the dark area of the histogram exceeds a preset ratio, it is determined to be underexposed; otherwise, it is determined to be overexposed. Only when the calculated sharpness variance value is higher than the preset threshold and the exposure index is within the valid range is the current image quality determined to meet the basic requirements of subsequent algorithm processing.
[0056] In step 302 of some embodiments, a multimodal intent analysis model is used to perform image-text relevance verification on the target image and the target interaction command, which involves calculating a semantic consistency measure in a cross-modal feature space. Specifically, the model's internal text encoder can extract the text semantic vector of the target interaction command, while a visual encoder can extract the global visual feature vector of the target image. These two vectors are then mapped to a shared semantic space of the same dimension. Within this space, the cosine similarity or Euclidean distance between the text vector and the visual vector is calculated. This similarity value quantifies the degree of overlap between the semantic entity in the command (such as "this road sign") and the visual content in the image at the feature level. For example, when the command mentions a specific object but the corresponding visual features cannot be extracted from the image, the similarity score will be lower than a preset relevance threshold, thus determining that the image-text relevance verification fails, indicating that the visual input at this time cannot support the user's interaction request.
[0057] In step 303 of some embodiments, if the sharpness or exposure does not meet the threshold, the system will generate guidance information prompting the user to adjust the lighting or maintain stability; if the image-text correlation verification fails, the system will generate guidance information prompting the user to adjust the shooting angle or aim at the target. This guidance information is directly output to the user, replacing subsequent business processing feedback.
[0058] Through steps 301 to 303 described above, this embodiment introduces an image validity verification mechanism before performing intent analysis. By conducting dual checks on image quality indicators and image-text relevance, it can effectively intercept low-quality or invalid image data, preventing invalid data from entering subsequent calculation processes, thereby saving system computing resources and network bandwidth. Simultaneously, by generating guidance prompts, it can promptly instruct users to correct shooting conditions, improving the success rate of visual interaction.
[0059] In step 104 of some embodiments, the target business processing interface refers to a pre-configured application programming interface (API) or microservice entry point associated with a specific domain service, such as a navigation map API, e-commerce search API, or translation engine API. Determining the interface based on the target visual intent establishes a mapping relationship from abstract intent labels to specific execution services. Calling the interface to process the target image means sending image data or key visual feature parameters extracted from the image (such as text extracted by OCR, or recognized object category IDs) to the corresponding third-party service or internal business module. The business processing result refers to the structured data returned by the aforementioned interface, such as the "product price list and purchase link" returned by the e-commerce interface, or the "path planning coordinates from the current location to the destination" returned by the navigation interface.
[0060] Please see Figure 4 In some embodiments, if the target visual intent is a visual navigation intent, step 104 may include, but is not limited to, steps 401 to 404.
[0061] Step 401: Determine the target business processing interface as the navigation service interface.
[0062] Step 402: Extract visual landmark information based on the target image and obtain real-time positioning data.
[0063] Step 403: Perform spatial topology consistency verification between visual landmark information and real-time positioning data.
[0064] Step 404: If the verification passes, call the navigation service interface to generate route planning data based on visual landmark information, which will be used as the business processing result.
[0065] In step 401 of some embodiments, determining that the target business processing interface is a navigation service interface means indexing the corresponding map service entry from a pre-set interface configuration table based on the visual navigation intent tag identified in the previous steps. Specifically, the system internally maintains a mapping relationship between intent types and service APIs. When the intent tag is confirmed to be a visual navigation intent tag, the logical routing module automatically locks the backend service responsible for spatial calculation and path planning (such as an AR navigation engine or a high-precision map API) and initializes the call parameter template of the interface, preparing to receive subsequent input image feature data and positioning coordinate data.
[0066] In step 402 of some embodiments, visual landmark information is extracted based on the target image, and real-time positioning data is acquired, involving the concurrent acquisition of visual feature extraction and sensor data. Extracting visual landmark information refers to using OCR (Optical Character Recognition) to identify entities with significant spatial location attributes from the target image, such as shop sign text, road signs, the exterior of specific buildings, or traffic lights. Simultaneously, acquiring real-time positioning data involves calling the global navigation satellite system module built into the smart glasses to read the current latitude and longitude coordinates, and combining this with IMU (Inertial Measurement Unit) data to obtain the device's orientation, thereby constructing the user's initial spatial state in the physical world.
[0067] In step 403 of some embodiments, the spatial topology consistency check between the visual landmark information and the real-time positioning data refers to retrieving surrounding POI (Point of Interest) data centered on the acquired real-time positioning coordinates, determining whether the visual landmark identified in the image exists within a reasonable line-of-sight range of the coordinates, and whether the relative orientation of the landmark matches the device's orientation data. This step aims to solve the drift problem that may exist in single GPS positioning (e.g., GPS shows a landmark on Street A, but the image shows a landmark on Street B), or the false detection problem that may occur in visual recognition (e.g., misidentifying a bus advertisement as a physical store), ensuring the physical authenticity of the navigation starting point.
[0068] In step 404 of some embodiments, if the verification passes, the navigation service interface is invoked to generate route planning data based on visual landmark information as the business processing result. This means initiating a navigation request using confirmed and accurate visual landmarks as high-precision spatial anchor points, setting the precise coordinates of the visual landmarks as the navigation start point or calibration point, and sending a request to the navigation service interface in conjunction with the destination information in the user's instructions. The route planning data returned by the interface not only includes traditional latitude and longitude trajectories but also includes 3D guidance information based on visual perspective (such as AR vector data for "turn left at the currently seen red building"). This data constitutes the final business processing result.
[0069] Through steps 401 to 404 described above, this embodiment of the application introduces spatial topology consistency verification and utilizes visual landmarks to effectively correct the drift error common in satellite positioning in urban environments, ensuring the accuracy of the navigation system's understanding of the user's current location. This mechanism enables smart glasses to provide high-precision augmented reality navigation services based on the real field of view, avoiding navigation guidance errors caused by positioning deviations and significantly improving the user's route-finding efficiency in complex road conditions.
[0070] Please see Figure 5 In some embodiments, if the target visual intent is to compare prices of goods, step 104 may include, but is not limited to, steps 501 to 503.
[0071] Step 501: Determine the target business processing interface as the e-commerce price comparison interface.
[0072] Step 502: Identify the brand characteristics, model characteristics, and specifications of the target product in the target image.
[0073] Step 503: Based on the identified features and parameters, call the e-commerce price comparison interface to obtain a real-time price list as the business processing result.
[0074] In step 501 of some embodiments, determining that the target business processing interface is an e-commerce price comparison interface means that the system searches and matches in a pre-set interface routing table based on the product price comparison intent tag output in step 103. Specifically, during execution, the system reads the category field from the intent tag, the logic control module automatically retrieves the access address of a pre-integrated third-party e-commerce platform API or aggregated price comparison service, and initializes the interface call request, preparing to receive product feature parameters extracted in subsequent steps as query input.
[0075] In step 502 of some embodiments, identifying the brand characteristics, model characteristics, and specification parameters of the target product in the target image involves fine-grained feature extraction of the main body of the product in the image. Specifically, firstly, an object detection model (such as YOLO) is used to crop out the main body region of the product in the image, and background noise is removed. Subsequently, within the main body region, an image matching algorithm is used to extract the logo texture to determine the brand characteristics, optical character recognition (OCR) technology is used to extract the alphanumeric sequence on the packaging to determine the model characteristics, and the color histogram or contour structure of the product is extracted as specification parameters. This step aims to transform unstructured image information into a structured keyword combination that can be retrieved by search engines.
[0076] In step 503 of some embodiments, calling the e-commerce price comparison interface to obtain a real-time price list based on the identified features and parameters, as a business processing result, refers to performing a standardized network data retrieval operation. The system encapsulates the brand, model, and specification data extracted in step 502 into a request message in JSON or XML format and sends it to the e-commerce price comparison interface. This interface performs concurrent queries on multi-source e-commerce databases on the server side, filters out matching products, and returns a structured list of data containing real-time prices, inventory status, and discount information from each platform. The system receives this list of data, caches it, and marks it as the business processing result of this interaction.
[0077] Through steps 501 to 503 described above, this embodiment of the application implements an automated price comparison mechanism based on visual features. The system extracts multi-dimensional features (brand, model, specifications) from product images, transforms visual information into precise search keywords, and automatically connects to e-commerce interfaces in vertical industries to obtain real-time data. This processing method directly maps physical product images to internet price data, ensuring the accuracy of information retrieval and improving the efficiency of users obtaining product information.
[0078] Please see Figure 6 In some embodiments, if the target visual intent is a general visual question-and-answer intent, step 104 may include, but is not limited to, steps 601 to 602.
[0079] Step 601: Determine the target business processing interface as the general visual large model interface.
[0080] Step 602: Call the general visual large model interface to perform image content understanding on the target image, and combine the target interaction instructions to generate natural language answer information for the target image, and use the natural language answer information as the business processing result.
[0081] In step 601 of some embodiments, determining that the target business processing interface is a general visual large model interface means switching the service call path to the access point of the generative artificial intelligence service based on the general visual question-answering intent tags identified in the previous steps. Specifically, when the logical routing module detects that the intent tag belongs to a non-vertical task such as "image description," "general object recognition," or "open domain question answering," it locks the API gateway of the pre-trained multimodal large model (VLM). This interface typically connects to a high-performance inference cluster deployed in the cloud, supporting the processing of unstructured image input and natural language commands, establishing a dedicated data channel to address the problem of openness without fixed patterns.
[0082] In step 602 of some embodiments, a general visual large model interface is invoked to perform image content understanding on the target image, and natural language answer information for the target image is generated in conjunction with the target interaction command, involving a multimodal generative reasoning process. Specifically, the system combines the Base64 encoded data or feature vector of the target image with the text Prompt of the target interaction command to construct a multimodal input sequence and send it to the VLM interface. Internally, the model uses a visual encoder to extract semantic features of the image and aligns them with the command text through a cross-attention mechanism. Then, a decoder uses word-by-word prediction based on visual context to generate a logical text response (e.g., explaining physical phenomena in the image, translating foreign language menus, or describing scene details). This generated text sequence is directly encapsulated as natural language answer information, serving as the final business processing result of this interaction.
[0083] Through steps 601 to 602 described above, this embodiment of the application leverages the strong generalization and reasoning capabilities of a general visual large model to solve interactive scenario problems that traditional rule engines cannot cover. By directly entrusting open-ended visual questions to a generative large model, it achieves deep understanding and free question-and-answer capabilities for arbitrary image content, without the need for predefined templates for each object or scene. This mechanism expands the knowledge boundaries of smart glasses, enabling them to handle everything from encyclopedic popular science to complex scene reasoning, providing stronger general visual assistance capabilities.
[0084] In step 105 of some embodiments, the target feedback information refers to converting structured business processing results into interactive content that conforms to human cognitive habits, including text-to-speech (TTS), visual graphic cards, or augmented reality (AR) overlay information. Generating feedback information based on business processing results involves cleaning, summarizing, and formatting the raw data. For example, a "price list" might be transformed into a short voice announcement such as "The lowest price on the entire network is XXX," while simultaneously generating a card data stream containing detailed parameters. This information is ultimately sent to smart glasses to play the audio or to the optical display module to present visual content, completing the final response to the user's command.
[0085] Please see Figure 7 This application also provides a visual interaction control device that can implement the above-described visual interaction control method. The device includes: A visual interaction control system is used to execute the aforementioned visual interaction control method; The smart glasses are used to acquire target interaction commands and send them to the visual interaction control system. They also perform visual image acquisition in response to image acquisition commands issued by the visual interaction control system, send the acquired target images to the visual interaction control system, and receive target feedback information from the visual interaction control system for output.
[0086] In some embodiments, smart glasses, serving as the system's sensing and interaction terminal, integrate a high-precision multimodal sensor array and a low-power processor. Specifically, the smart glasses are equipped with an audio acquisition unit and touch or gesture sensors to assist in acquiring non-voice-based target interaction commands. The smart glasses also feature a visual acquisition module (such as a wide-angle camera or depth camera), which is normally in standby or low-power mode and is only activated upon receiving specific commands from the visual interaction control system to perform single-frame capture or continuous video stream acquisition, thereby acquiring target images containing environmental information. Furthermore, the smart glasses include a feedback output unit, such as bone conduction headphones or a near-eye display optical module, for receiving and playing target feedback information from the system to realize the presentation of audio or augmented reality (AR) visual content.
[0087] In some embodiments, the visual interaction control system serves as the processing center of the entire device, typically deployed on a cloud server cluster or high-performance edge computing node, to execute the visual interaction control methods described in the aforementioned embodiments. Logically, the system includes an interaction intent analysis module, a device control module, a multimodal reasoning module, and a business service scheduling module. The interaction intent analysis module performs semantic analysis on the received target interaction instructions to determine whether visual intervention is required (i.e., distinguishing between visually dependent and non-visually dependent types). The device control module generates hardware-level image acquisition instructions based on the analysis results and sends them to the smart glasses. The multimodal reasoning module, equipped with a pre-trained large-scale image-text model, performs joint analysis of the returned target image and instructions to determine the user's specific visual intent. The business service scheduling module automatically calls external navigation, e-commerce, or question-and-answer API interfaces based on the intent and generates the final target feedback information based on the returned data.
[0088] In some embodiments, the smart glasses and the visual interaction control system work in a cloud-edge collaborative manner. First, the smart glasses acquire the user's target interaction command and send it to the visual interaction control system via a wireless network (such as Wi-Fi, 5G, or Bluetooth connection to a mobile gateway). Second, the visual interaction control system performs rapid semantic analysis on the command, confirming it as a visual dependency type, and then sends an image acquisition command back to the smart glasses. Next, in response to this command, the smart glasses automatically activate their camera to focus and capture an image, sending the acquired target image back to the visual interaction control system. Finally, the visual interaction control system completes intent analysis and business processing, sending the generated target feedback information to the smart glasses, which then completes the final output.
[0089] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described visual interaction control method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0090] Please refer to the figure. Figure 8 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 801 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 802 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 802 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 802 and is called and executed by the processor 801 using the visual interaction control method of the embodiments of this application. The 803 input / output interface is used to implement information input and output. The communication interface 804 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 805 transmits information between various components of the device (e.g., processor 801, memory 802, input / output interface 803, and communication interface 804); The processor 801, memory 802, input / output interface 803, and communication interface 804 are connected to each other within the device via bus 805.
[0091] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described visual interaction control method.
[0092] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0093] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0094] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0095] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0096] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0097] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0098] It should be understood that in this application, "at least one" and "several" refer to one or more, and "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0099] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. The couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0100] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0101] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0102] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0103] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A visual interaction control method, characterized in that, The method includes: In response to a target interaction command issued by the smart glasses, the target interaction type of the target interaction command is determined; When the target interaction type is characterized as a visually dependent type, an image acquisition command is sent to the smart glasses, and the target image acquired by the smart glasses according to the image acquisition command is received. The target interaction command and the target image input are processed by a pre-trained multimodal intent analysis model to determine the target visual intent; The corresponding target business processing interface is determined based on the target visual intent, and the target business processing interface is called to process the target image to obtain the business processing result; Target feedback information is generated based on the business processing results.
2. The visual interaction control method according to claim 1, characterized in that, Determining the target interaction type of the target interaction instruction includes: Keyword extraction is performed on the target interaction command to obtain the target keywords; If the target keyword is a preset spatial indicator pronoun, then the target interaction type is determined to be the visual dependency type; or, The target interaction instructions are semantically parsed to determine the target task category; The target task category is matched with a preset scene logic library. If the target task category belongs to a predefined visual entity category in the scene logic library, the target interaction type is determined to be the visual dependency type.
3. The visual interaction control method according to claim 2, characterized in that, The method further includes: If the target interaction type is characterized as a non-visual dependent type, then a pre-set general dialogue model is invoked to perform text semantic processing on the target interaction instruction and generate text response information. The text reply information is sent to the smart glasses.
4. The visual interaction control method according to claim 1, characterized in that, Before processing the target interaction command and the target image input into a pre-trained multimodal intent analysis model to determine the target visual intent, the method further includes: The sharpness and exposure of the target image are obtained, and it is determined whether the sharpness and exposure meet a preset image quality threshold. The multimodal intent analysis model is used to perform image-text relevance verification on the target image and the target interaction command, and to determine whether the visual entities extracted from the target image match the semantic entities in the target interaction command. If the clarity or exposure does not meet the preset image quality threshold, or if the image-text correlation verification fails, a guidance prompt message is generated and used as the target feedback message.
5. The visual interaction control method according to claim 1, characterized in that, If the target visual intent is a visual navigation intent, the step of determining the corresponding target business processing interface based on the target visual intent and calling the target business processing interface to process the target image to obtain a business processing result includes: The target business processing interface is determined to be a navigation service interface; Visual landmark information is extracted from the target image, and real-time positioning data is obtained; Perform a spatial topology consistency check on the visual landmark information and the real-time positioning data; If the verification passes, the navigation service interface is invoked to generate path planning data based on the visual landmark information, which is then used as the business processing result.
6. The visual interaction control method according to claim 1, characterized in that, If the target visual intent is a product price comparison intent, the step of determining the corresponding target business processing interface based on the target visual intent and calling the target business processing interface to process the target image to obtain a business processing result includes: The target business processing interface is identified as an e-commerce price comparison interface; Identify the brand characteristics, model characteristics, and specifications of the target product in the target image; Based on the identified features and parameters, the e-commerce price comparison interface is called to obtain a real-time price list, which is then used as the result of the business processing.
7. The visual interaction control method according to claim 1, characterized in that, If the target visual intent is a general visual question-and-answer intent, the step of determining the corresponding target business processing interface based on the target visual intent and calling the target business processing interface to process the target image to obtain a business processing result includes: The target business processing interface is determined to be a general visual large model interface; The general visual large model interface is invoked to perform image content understanding on the target image, and natural language answer information for the target image is generated in combination with the target interaction command. The natural language answer information is used as the business processing result.
8. A visual interaction control device, characterized in that, The device includes: A visual interaction control system, wherein the visual interaction control system is used to execute the visual interaction control method as described in any one of claims 1 to 7; The smart glasses are used to acquire the target interaction command and send the target interaction command to the visual interaction control system, perform visual image acquisition in response to the image acquisition command issued by the visual interaction control system, send the acquired target image to the visual interaction control system, and receive and output the target feedback information issued by the visual interaction control system.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the visual interaction control method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the visual interaction control method according to any one of claims 1 to 7.