Object query method, device and electronic equipment

By displaying voice dialogue controls on the electronic device interface and combining them with a multimodal visual perception model, the problem of inaccurate object recognition in environmental perception is solved, enabling precise object positioning and spatial relationship understanding, thereby improving the accuracy and intelligence of environmental perception.

CN122285718APending Publication Date: 2026-06-26VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
VIVO MOBILE COMM CO LTD
Filing Date
2026-03-18
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

In existing technologies, electronic devices cannot accurately identify objects in images when sensing the environment, resulting in low accuracy in environmental perception.

Method used

By displaying a voice dialogue control in the interface, the user enters the voice interaction mode. The system combines a multimodal visual perception model to find objects, outputs the relative position information of the objects, and improves the spatial perception capability of the model through training data, including the annotation of two-dimensional and three-dimensional positional relationships.

Benefits of technology

It improves the accuracy and intelligence of environmental perception, enabling precise location of objects and understanding of their spatial relationships with other objects in the environment, thereby enhancing the user's ability to perceive scene structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122285718A_ABST
    Figure CN122285718A_ABST
Patent Text Reader

Abstract

This application discloses an object query method, apparatus, and electronic device, relating to the field of artificial intelligence. The specific technical solution is as follows: A first interface displays first dialogue information, which is used to indicate the search for a first object. The first interface includes a preview image captured by a camera. Based on the first dialogue information and the preview image, an object search is performed, and first information is output. The first information includes at least one of the following: the relative position information of the first object relative to the user; and the relative position information between the first object and a second object in the preview image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence, specifically relating to an object query method, device, and electronic device. Background Technology

[0002] With the development of electronic information technology, electronic devices can use cameras to capture environmental images for environmental perception, thereby assisting users in obtaining visual information about the environment.

[0003] In related technologies, when a user needs to know the objects in their environment, they can capture real-time or still images through the camera of an electronic device. The electronic device uses object recognition algorithms to analyze the captured image content, and after identifying objects in the image, it outputs the names of these objects. For example, the electronic device outputs object names such as "computer" and "water cup." However, existing technologies identify image content, and their output results are labels for each object detected in the image. Therefore, the environmental perception capability of existing technologies is limited, making it impossible for users to accurately obtain visual information in the image, resulting in low accuracy of environmental perception. Summary of the Invention

[0004] The purpose of this application is to provide an object query method, apparatus, and electronic device that can greatly improve the accuracy and intelligence of environmental perception.

[0005] In a first aspect, embodiments of this application provide an object query method, the method comprising: displaying first dialogue information on a first interface, the first dialogue information being used to indicate the search for a first object, the first interface including a preview image captured by a camera; performing an object search based on the first dialogue information and the preview image, and outputting first information; wherein the first information includes at least one of the following: relative position information of the first object relative to the user; relative position information between the first object and a second object in the preview image.

[0006] In some embodiments of this application, the first interface further includes a voice dialogue control; displaying first dialogue information on the first interface includes: receiving a first input to the voice dialogue control; responding to the first input, activating a voice interaction mode, and displaying the first dialogue information collected by the microphone on the first interface.

[0007] In this embodiment, by displaying a voice dialogue control on the first interface, users can intuitively see the voice dialogue control and conveniently activate the auxiliary query function. Secondly, after the voice dialogue control is triggered, it enters voice interaction mode and performs voice acquisition, avoiding misidentification caused by user error and improving the reliability of the interaction.

[0008] In some embodiments of this application, after performing object search based on the first dialogue information and the preview image, the method further includes: displaying a first identifier in a first interface; wherein the first identifier is used to mark the display area of ​​the first object in the preview image.

[0009] In this embodiment, by displaying a first identifier on a first interface, a real-time visual marking operation is performed on the first object after successful recognition, thereby realizing a feedback mechanism that combines voice description and visual indication. Specifically, the voice feedback provides spatial location relationships that can be obtained through hearing, while the display of the first identifier provides visual object indication. In this way, by combining voice feedback and visual indication, the user's overall perception of the scene structure is enhanced, enabling the user to quickly and unambiguously understand the location of the first object by combining the two types of information.

[0010] In some embodiments of this application, object lookup is performed based on first dialogue information and a preview image, and first information is output, including: inputting the first dialogue information and the preview image into a multimodal visual perception model, performing multimodal reasoning on the first dialogue information and the preview image through the multimodal visual perception model to obtain prediction information; the prediction information includes at least one of the following: coordinate information of the first object, relative position information between the first object and at least one reference object in the preview image; and outputting the first information based on the prediction information.

[0011] In this embodiment, an end-to-end processing framework employing a multimodal perception model improves the accuracy and efficiency of object querying. Specifically, through joint inference, the multimodal visual perception model can locate specific objects in an image based on the user's natural language commands, avoiding the error problems caused by first identifying all objects in the entire image and then performing text matching, thereby improving positioning accuracy. Secondly, the multimodal perception model not only outputs the coordinates of the object but also understands and outputs the spatial relationship between the first object and other reference objects in the environment, thus obtaining a multi-level spatial description of the scene context, improving the comprehensiveness and accuracy of object querying.

[0012] In some embodiments of this application, the multimodal visual perception model includes a language encoding module, a visual encoding module, and a cross-modal processing module. The multimodal visual perception model performs multimodal reasoning on the first dialogue information and the preview image to obtain predicted information, including: using the language encoding module to perform text encoding processing on the first dialogue information to obtain text feature information; using the visual encoding module to perform image encoding processing on the preview image to obtain image feature information; and using the cross-modal processing module to perform fusion reasoning on the text feature information and the image feature information to obtain predicted information.

[0013] In this embodiment, by performing cross-modal fusion inference on preview images and dialogue information, the semantic understanding and visual perception in object queries can be deeply unified, achieving precise fine-grained alignment from semantics to vision. Specifically, the language encoding module transforms ambiguous natural language instructions into precise query vectors, the visual encoding module provides rich environmental visual information, and the cross-modal processing module achieves matching between the two through an attention mechanism. This enables accurate execution of queries containing complex attributes, relationships, and referential meanings, thereby significantly improving the accuracy of object queries in complex scenarios.

[0014] In some embodiments of this application, before performing object lookup processing based on the first dialogue information and the preview image and outputting the first information, the method further includes: inputting the preview image into a multimodal visual perception model, locating the object in the preview image through the multimodal visual perception model, and obtaining the coordinate information of the first object; and outputting the second information when it is determined based on the coordinate information that the first object is located in the edge region of the preview image; wherein the second information is used to prompt the user to move the electronic device so that the first object is located in the center region of the preview image.

[0015] In this embodiment, by actively judging and guiding the user to move the electronic device, the first object is moved to a non-edge area of ​​the image, thereby improving the accuracy of subsequent multimodal inference and avoiding errors or incompleteness in the location description caused by partial occlusion of the target or poor viewing angle. Secondly, by providing the user with real-time and clear spatial guidance through the second information, the user can accurately move the electronic device so that the first object can be placed in the center area of ​​the preview image, thereby accurately locating the first object.

[0016] In some embodiments of this application, before performing multimodal inference on the first dialogue information and the preview image using a multimodal visual perception model, the method further includes: acquiring first training data, second training data, and third training data; training an initial model based on the first training data, second training data, and third training data to obtain a multimodal visual perception model; wherein, the first training data includes at least one first sample image and first annotation information corresponding to each first sample image; the first annotation information corresponding to each first sample image includes object description information and coordinate information of at least one object in the first sample image; the second training data includes at least one second sample image and second annotation information corresponding to each second sample image; the second annotation information corresponding to each second sample image includes relative position information between any two objects in the second sample image; the third training data includes at least one third sample image and first dialogue information corresponding to each third sample image.

[0017] In this embodiment, training the initial model with first and second training data enhances its fine-grained localization capabilities and accurately answers spatial orientation information. Specifically, the new data labeling method introduces two new training data formats to characterize the spatial relationships and semantic constraints between targets. Through joint training with multiple data types, the model can not only achieve precise target localization but also perform fine-grained matching of complex semantic descriptions and accurately answer spatial orientation information related to the target, thereby significantly improving the expressive power and application scope of visual localization tasks. Secondly, a spatial perception capability model is constructed through mixed training with multiple data types. Specifically, at the data level, a mixed training strategy using multiple data types is adopted. By using reasonable data ratios to jointly optimize the model, the model not only possesses the ability to semantically describe the content of the image but also accurately understands the spatial position and orientation relationships of target objects, thereby supporting coordinate-level localization and spatial position information output, significantly improving the model's spatial perception capability.

[0018] In some embodiments of this application, obtaining first training data includes: obtaining at least one first sample image and description information corresponding to each first sample image; the description information includes object description information of at least one object in the first sample image; performing semantic parsing processing on the description information corresponding to each first sample image to determine at least one object in each first sample image; performing visual positioning on each first sample image using a visual positioning model to determine the coordinate information of at least one object in each first sample image in the first sample image; adding the coordinate information of each object in the at least one object to the first text position corresponding to each object to obtain first training data; the first text position is: the text position where the object description information of the object is located in the first description information, and the first description information is the description information corresponding to the first sample image.

[0019] In this embodiment, by utilizing image-descriptive text and an efficient visual localization model to construct the first training data, a massive amount of image-text pairs with precise coordinate annotations can be quickly generated, ensuring the accuracy and consistency of the annotations. By strictly corresponding the text parsing objects with the visual localization results and embedding them into the original descriptive information, the alignment between text descriptions, visual entities, and spatial locations in the generated training samples is ensured, providing reliable information for model learning and thereby improving the model's ability to understand fine-grained semantics and complex scenes.

[0020] In some embodiments of this application, obtaining second training data includes: obtaining at least one second sample image; determining the spatial positional relationship between at least two objects in each second sample image; the spatial positional relationship includes two-dimensional spatial positional relationship and three-dimensional spatial positional relationship; and constructing second training data based on the spatial positional relationship between at least two objects in each second sample image.

[0021] In this embodiment, a second training data comprising annotations of two-dimensional and three-dimensional spatial positions is constructed and used for model training. This data primarily characterizes the spatial orientation relationships between target objects, specifically including six directions: up, down, left, right, front, and back. By introducing this second training data, the overall perception capability of the multimodal visual perception model in both on-screen and general scenarios is significantly improved. Specifically, by introducing spatial perception data containing annotations of spatial positions such as up / down, left / right, and front / back, the model can not only accurately return the two-dimensional coordinates of the target but also understand and infer the relative spatial orientation relationships between targets, significantly improving spatial perception and spatial reasoning capabilities.

[0022] In some embodiments of this application, the spatial positional relationship includes a two-dimensional spatial positional relationship; determining the spatial positional relationship between at least two objects in each second sample image includes: performing content analysis on each second sample image using a multimodal large model to obtain descriptive information corresponding to each object in each second sample image; performing visual positioning on the descriptive information using a visual positioning model to obtain coordinate information of each object in the second sample image; and determining the two-dimensional spatial positional relationship between at least two adjacent objects in each second sample image based on the coordinate information of at least one object in each second sample image.

[0023] In this embodiment, by integrating the content parsing capability of a multimodal large model with the precise positioning capability of a visual positioning model, and designing automatic relationship derivation rules based on geometric coordinates, the positional relationship between objects in the image is accurately determined. Based on this positional relationship, high-quality training data is constructed, enabling the model to learn the orientation logic of "up, down, left, right" between objects, thus enabling it to accurately answer user queries about relative positions during reasoning.

[0024] In some embodiments of this application, the spatial positional relationship includes a three-dimensional spatial positional relationship; determining the spatial positional relationship between at least two objects in each second sample image includes: performing depth estimation processing on each second sample image using a depth estimation model to obtain depth information of each second sample image; and determining the three-dimensional spatial positional relationship between at least two objects in each second sample image based on the depth information of each second sample image and the coordinate information of at least one object in each second sample image in the second sample image.

[0025] In this embodiment of the application, by utilizing the image depth estimation capability of the depth estimation model and based on the inference rules of front-back relationships, the front-back positional relationship between each object in the image is accurately determined. Based on this positional relationship, high-quality training data is constructed, enabling the model to learn the orientation logic of the spatial relative position of objects, thereby enabling it to output accurate spatial relative positions during inference.

[0026] In some embodiments of this application, an initial model is trained based on training data to obtain a multimodal visual perception model, including: using the initial model, performing N inference samplings on each first sample image to obtain N first prediction results, each first prediction result including prediction description information of at least one predicted object in the first sample image; calculating N first reward values ​​based on the N first prediction results using a first reward function; and using the initial model, performing M inference samplings on each second sample image to obtain M second prediction results, each second prediction result including prediction relative position information between at least two predicted objects in the second sample image. Based on M second prediction results, M second reward values ​​are calculated using a second reward function. Based on the first and second differences, the parameters of the initial model are updated to obtain a multimodal visual perception model. The first difference is the difference between N first reward values; the second difference is the difference between M second reward values. The first reward function evaluates the consistency between the predicted description information of the predicted object and the actual visual content of the predicted object. The predicted description information of the predicted object is obtained by inference sampling of the first sample image using the initial model. The second reward function evaluates the matching degree between the predicted relative position information between at least two predicted objects and the actual relative position information between at least two predicted objects. The predicted relative position information between two predicted objects is obtained by inference sampling of the second sample image using the initial model. N is a positive integer, and M is a positive integer.

[0027] In this embodiment, during the model training phase, a DAM model is introduced as a constraint mechanism in the reinforcement learning process to generate a structured description of the target localization region. This description is then combined with a multimodal large model for consistency judgment and illusion detection. This effectively suppresses semantic illusions generated by the model that do not conform to visual reality while optimizing localization and description capabilities, thereby improving the accuracy and reliability of the output results. Specifically, a novel establishment function is constructed, introducing the model as a visual consistency constraint. Furthermore, a differentiated reward mechanism is designed for different data. In the reinforcement learning phase of the edge-side 3B multimodal visual model, the DAM model is introduced as a visual consistency constraint module to perform local visual description verification of the target region inferred by the large model, and the verification result is directly fed back to the reward function. When the description generated by the large model is inconsistent with the model's description of the corresponding visual region, the reward value is penalized, thereby explicitly suppressing the generation of descriptive illusions during training. Thus, by designing differentiated reward evaluation and penalty mechanisms based on the characteristics of different types of training data, the model's textual description capability and spatial perception capability are improved.

[0028] In some embodiments of this application, N first reward values ​​are calculated based on N first prediction results using a first reward function, including: cropping the image region corresponding to the prediction bounding box from the first sample image based on the prediction bounding box corresponding to each predicted object in each first prediction result; performing image description on each image region using an image description model to obtain reference description information for each predicted object; performing a consistency comparison between the prediction description information of each predicted object in the first prediction results and the reference description information of each predicted object to obtain consistency comparison result information; and calculating the first reward value corresponding to the first prediction result using the first reward function based on the consistency comparison result information; wherein the first reward value is used to characterize the degree of consistency between the prediction description information generated by the initial model and the real visual content.

[0029] In this embodiment, by comparing the model's output with the model's objective description of the same visual region, errors in the model-generated description, such as color and shape errors, can be accurately located and quantified, enabling the model to learn the ability to accurately generate descriptive information. This improves the accuracy and reliability of the final model output description.

[0030] In some embodiments of this application, M second reward values ​​are calculated based on M second prediction results using a second reward function, including: comparing the predicted relative position information and reference relative position information in each second prediction result to determine a first quantity; the first quantity is the number of predicted relative position information that matches the reference relative position information; and calculating a second reward value based on the first quantity and a second quantity using a second reward function; wherein the second quantity is the total number of predicted relative position information; and the second reward value is used to characterize the degree of consistency between the predicted relative position information of the initial model and the actual relative position.

[0031] In this embodiment, during the reinforcement learning phase, a policy gradient-based optimization method is employed to sample the same input sample multiple times, generating different inference results. Each sample independently calculates its corresponding reward value. By comparing the results of multiple samplings and using the differences in reward values ​​to guide model parameter updates, inference results with high consistency and low illusion are given higher weights, thereby gradually improving the stability and reliability of the model's output during training. Through this reinforcement learning mechanism combining multiple sampling and constraints, the model can effectively suppress illusion problems in visual description and spatial reasoning while maintaining its localization capabilities.

[0032] Secondly, embodiments of this application provide an object query device, which includes: a processing module, configured to display first dialogue information on a first interface, the first dialogue information being used to indicate the search for a first object, the first interface including a preview image captured by a camera; the processing module is further configured to perform object search based on the first dialogue information and the preview image, and output first information; wherein the first information includes at least one of the following: relative position information of the first object relative to the user; relative position information between the first object and a second object in the preview image.

[0033] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0034] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0035] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0036] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0037] In this embodiment, the electronic device displays first dialogue information on a first interface, which is used to indicate the search for a first object. The first interface includes a preview image captured by a camera. Based on the first dialogue information and the preview image, the device performs an object search and outputs first information. The first information includes at least one of the following: the relative position information of the first object relative to the user; and the relative position information between the first object and a second object in the preview image. Through this scheme, the electronic device can output information including the relative position information between the first object and the user, or the relative position information between the first object and other objects, based on the first dialogue information and the preview image input by the user. This allows the user to quickly and accurately determine the location of the first object based on its relative position to itself, or based on its relative position to other known objects, thus greatly improving the accuracy and intelligence of environmental perception. Attached Figure Description

[0038] Figure 1 A flowchart illustrating an object query method provided for some embodiments of this application;

[0039] Figure 2 A schematic diagram of a first interface provided for some embodiments of this application;

[0040] Figure 3A A schematic diagram of a settings interface provided for some embodiments of this application;

[0041] Figure 3B A schematic diagram illustrating shortcut and auxiliary interfaces provided for some embodiments of this application;

[0042] Figure 3C A schematic diagram of an accessibility interface provided for some embodiments of this application;

[0043] Figure 4 A schematic diagram of a first interface provided for some embodiments of this application;

[0044] Figure 5 A schematic diagram of a first interface provided for some embodiments of this application;

[0045] Figure 6 A schematic diagram of a first interface provided for some embodiments of this application;

[0046] Figure 7 A schematic diagram of a first interface provided for some embodiments of this application;

[0047] Figure 8 A schematic diagram of a first interface provided for some embodiments of this application;

[0048] Figure 9 A schematic diagram illustrating the visualization effect of the inference results provided in the embodiments of this application;

[0049] Figure 10 A flowchart illustrating the process of constructing second training data provided for some embodiments of this application;

[0050] Figure 11 A flowchart illustrating an object query method provided for some embodiments of this application;

[0051] Figure 12 Schematic diagrams of the structure of an object query device provided for some embodiments of this application;

[0052] Figure 13 Schematic diagrams of the structure of electronic devices provided for some embodiments of this application;

[0053] Figure 14 A schematic diagram of the hardware structure of an electronic device provided for some embodiments of this application. Detailed Implementation

[0054] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0055] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0056] The terms "at least one," "at least one of," etc., used in the specification and claims of this application refer to any one, any two, or a combination of two or more of the included items. For example, at least one of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more items, and its meaning is similar to that of "at least one."

[0057] The terminology used in the implementation section of this application is only for explaining specific embodiments of this application and is not intended to limit this application. The terminology involved in the embodiments of this application is explained below.

[0058] Terminal side: mainly refers to screen display devices, including but not limited to portable terminal devices such as mobile phones and computers.

[0059] 3B Large Language Model on the Device Side: This mainly refers to a large language model with 3B parameters, which can be directly built into a mobile phone to handle various scenario tasks on the mobile device.

[0060] Multimodal large models: refer to large-scale neural network models that can simultaneously process and understand multiple types of data such as text and images, and are used to realize the perception, understanding and generation of cross-modal information.

[0061] Visual perception refers to the process of analyzing images or screen content through visual models to obtain semantic information, spatial location and relative relationships of target objects, providing basic perceptual capabilities for target recognition, localization, understanding and interaction.

[0062] Target-assisted query: refers to assisting users in querying and responding to the location or attributes of a target object by perceiving and recognizing the target object in a natural language manner.

[0063] Reinforcement learning: refers to machine learning methods that enable models to learn optimal decision-making strategies autonomously through environmental interaction, trial and error, and reward feedback mechanisms.

[0064] ViT-Lora training: refers to the method of using LoRA low-rank adaptation technique to efficiently fine-tune a visual Transformer model, such as the ViT model, in order to achieve customized training of a large model for a specific task.

[0065] With the development of multimodal large-scale models and computer vision technology, smart terminals are gradually gaining the ability to comprehensively understand image content. Visual perception technology not only focuses on the existence of target objects but also emphasizes the overall modeling of the semantic description, spatial location, and relative relationships of the targets. Current visual perception technologies are mostly based on visual Transformers (ViT) or convolutional neural networks to extract multi-scale visual features, and combine them with language models or multimodal encoders to perform semantic description and understanding of targets. On the other hand, they locate target objects in images based on text descriptions. The core idea is to predict the location of the object referred to by the text in the image through image-text feature alignment and similarity calculation. Existing methods are mostly based on deep neural network structures, such as using visual Transformers (ViT) to extract image feature information, combining them with text encoders (such as CLIP, BERT, etc.) to generate text feature information, and achieving cross-modal localization through feature fusion and similarity measurement.

[0066] The object query method provided in this application can be executed by an object query device, which can be an electronic device, or a functional module or functional entity within an electronic device. The following description uses an electronic device executing the object query method as an example to illustrate the object query method provided in this application.

[0067] The object query method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0068] The object query method provided in this application can be applied to various scenarios where it is necessary to locate specific objects in the environment.

[0069] One specific scenario is when a user needs to find frequently used items such as a mobile phone, keys, or a water glass at home or in the office. For example, when a user asks "Where is my phone?" via voice, the electronic device not only needs to recognize the phone's presence but also needs to further understand the user's query intent, accurately locate the phone, determine its relative position to the user, and its positional relationship with surrounding objects. It then provides clear spatial location information to the user, such as indicating that the phone is on the table in front of them, to the right of the water glass, next to a lamp and a vase, thus helping the user quickly find their phone.

[0070] Another specific scenario is when users need to locate large or specific personal items, such as suitcases, while traveling or in unfamiliar environments. For example, in a waiting room, a user might ask, "Where is my blue suitcase?" The electronic device needs to recognize the suitcase and guide the user to rotate the device until the suitcase is fully in the frame when it's at the edge of the screen. For instance, if the device detects the suitcase at the edge, it prompts the user to move the device towards the suitcase, with a voice prompt like, "Try moving the camera to the right." When the device continues to move until the suitcase appears in the frame, a voice reply states, "The suitcase is in front of you to your right, behind the table," thus providing the user with information about the suitcase's location and its relative position to other objects, helping the user quickly find their suitcase.

[0071] Figure 1 This is a flowchart illustrating the object query method provided in an embodiment of this application, such as... Figure 1 As shown, the object query method may include the following steps 201 and 202:

[0072] Step 201: The electronic device displays the first dialogue information on the first interface.

[0073] The first dialogue information is used to indicate the search for a first object, and the first interface includes a preview image captured by the camera.

[0074] In some embodiments of this application, the first interface described above is an interactive interface for implementing object query function, which is mainly used to provide users with real-time environmental visual feedback and convenient function operation entry.

[0075] In some embodiments of this application, the first interface includes an image display area and a function control area. The image display area is used to display images captured by the camera in real time, allowing the user to know the current view. Exemplarily, the function control area includes multiple function controls, such as an "explore" control, a "find" control, a "take a picture" control, a voice dialogue control, a "capture" control, and a "camera flip" control.

[0076] Specifically, the aforementioned "Explore" control is used to enter the auxiliary query mode, start or stay on the first interface. When on the first interface, the electronic device can automatically identify and announce the object the user is looking for based on the first dialogue information entered by the user. The aforementioned "Search" control is used to start the "Search" mode. When the user clicks the "Search" control, the electronic device displays an object selection interface or enters a specific search mode. The object selection interface includes an object selection list from which the user can select or specify the object to be searched for via voice. Subsequently, the electronic device performs targeted search and location based on the camera image. The aforementioned "Take Photo" control is used to trigger entry into the shooting mode. The aforementioned voice dialogue control is used to enable the voice interaction mode and obtain the first dialogue information. The aforementioned "Take Photo" control is used to trigger taking a photo. The aforementioned "Camera Flip" control is used to switch between the front and rear cameras of the electronic device.

[0077] For example, when a user opens the application, the first interface, namely the "Explore" mode interface, is entered by default. The electronic device displays the image captured by the camera in real time on the first interface. If the user wants to actively find the key, he / she can click the voice dialogue control and say the search intention according to the prompt, such as "Help me find my phone". The electronic device searches for the phone on the screen and reports its location.

[0078] In some embodiments of this application, the preview image described above is an image captured by an electronic device through a camera and displayed on a first interface.

[0079] For example, the preview image mentioned above is an image captured by the camera when or after acquiring the first dialogue information.

[0080] In some embodiments of this application, the preview images are captured by a camera in real time or periodically.

[0081] The first interface provided in the embodiments of this application is illustrated below with reference to the accompanying drawings.

[0082] For example, such as Figure 2 As shown, the electronic device displays a first interface 21, which includes an image 22 captured by the camera. The first interface 21 also includes an "Explore" control 23, a "Find" control 24, a "Take Photo" control 25, a voice dialogue control 26, a "Shoot" control 27, and a "Flip Camera" control 28. When the user clicks the voice dialogue control 26, the electronic device displays text or graphic prompts such as "Ask me a question" on the first interface 21, indicating that the electronic device has entered a voice inquiry state and is waiting for the user to issue a query command, such as "Where is my phone?".

[0083] In some embodiments of this application, when the electronic device has the auxiliary query function enabled, it displays the first interface and executes the object query method described in the embodiments of this application.

[0084] In some embodiments of this application, the auxiliary query function can be found and activated from the settings interface, the shortcut and assistance interface, and the accessibility interface.

[0085] For example, such as Figure 3A As shown, the electronic device displays a settings interface 31, which includes options such as "Shortcuts and Accessibility," "Privacy," "RAM and Storage," and "Battery." After the user clicks the "Shortcuts and Accessibility" option, as shown... Figure 3B As shown, the electronic device displays a shortcut and accessibility interface 32, which includes options such as "Accessibility," "Senior Mode," "Easy Mode," and "Remote Assistance." After the user clicks the "Accessibility" option, as shown... Figure 3C As shown, an accessibility interface 33 is displayed, which includes options for "assisted query", "screen reading", "read aloud", and "input method". After the user clicks the "assisted query" option, they enter the assisted query mode.

[0086] It is understood that the illustrated interface, such as the settings interface, is only for illustration and does not constitute a limitation on the actual interface. In practice, the interface may include more content; the above function entry path is only an exemplary implementation.

[0087] In some embodiments of this application, this function can also be enabled via desktop shortcuts, negative one screen cards, or voice assistants.

[0088] It should be noted that the auxiliary query function described in the embodiments of this application may also be referred to as the "target auxiliary query" function or the "see" function in specific products. Regardless of its specific name, as long as the technical solution it implements is substantially the same as or equivalent to the solution disclosed in this application, it falls within the protection scope of this application.

[0089] In some embodiments of this application, the first dialogue information may be voice command information input by the user, or it may be text information converted from the voice command information input by the user, or it may be text command information directly input by the user.

[0090] In some embodiments of this application, when the first interface is displayed, the user can input first dialogue information through the first interface.

[0091] In some examples, users can speak directly into the microphone of an electronic device, such as saying "Where is my phone?" The electronic device captures this speech to obtain the first dialogue information.

[0092] In some examples, users can speak directly into the microphone of an electronic device, such as saying "Where is my phone?" The electronic device captures this speech and converts it into text information, i.e., the first dialogue information.

[0093] In some embodiments of this application, the aforementioned first dialogue information is used to indicate the user's query intent. For example, the first dialogue information may include the object name, object description information, etc., of the object to be searched.

[0094] In some embodiments of this application, a user can click on the voice dialogue control in the first interface to trigger the electronic device to enter the voice listening state, and then state the query intent.

[0095] In some embodiments of this application, the first object described above can be any entity with a visual form that the user needs to locate in the physical environment.

[0096] For example, the first object mentioned above may include, but is not limited to, objects of the following categories:

[0097] Everyday items: such as mobile phones, keys, suitcases, wallets, water bottles, remote controls, books, glasses, and other personal belongings or household items.

[0098] People or other animals: such as family members, friends, pets, etc.

[0099] Specific objects: such as furniture like tables, chairs, and sofas; home appliances like televisions and refrigerators; electronic devices like laptops and tablets; and vehicles like bicycles and buses.

[0100] Environmental elements and landmarks: such as doors, windows, stairs, shop signs, and exit signs in shopping malls.

[0101] In some embodiments of this application, the object description information includes descriptions such as object attributes of the object the user wants to find. For example, the object description information may include the object category, object attributes, and relationship with other objects of the first object. For instance, the object category included in the object description information may be "cell phone," "keys," "suitcase," etc.; the object attributes included in the object description information may be "black" or "large"; and the relationship with other objects included in the object description information may be "on the table," "in the corner," etc.

[0102] In some examples, the user inputs the voice: "Help me find my phone?", and the electronic device converts the voice into text through the voice recognition module. The first dialogue information obtained is "Help me find my phone?", which includes the object name of the first object, namely "phone".

[0103] In some examples, the user inputs the voice: "Help me find my white water glass?" The electronic device converts the voice into text through the voice recognition module, and the first dialogue information obtained is "Help me find my white water glass?", which includes the object description information of the first object, namely "white water glass".

[0104] In some examples, the user inputs the voice: "Help me find my suitcase?", and the electronic device converts the voice into text through the voice recognition module. The first dialogue information obtained is "Help me find my suitcase?", which includes the object name of the first object, namely "suitcase".

[0105] The object query method provided in this application allows users to interactively ask questions to an electronic device through dialogue. Based on visual perception and semantic understanding, the electronic device can accurately distinguish the first object to be searched and perform precise matching and spatial positioning of that object, achieving target-assisted querying and spatial positioning. For example, the electronic device can accurately distinguish that the user wants to find a white water glass instead of a green water glass and accurately answer the location of the white water glass, thereby improving the accuracy of object querying.

[0106] Step 202: The electronic device performs object search based on the first dialogue information and the preview image, and outputs the first information.

[0107] In some embodiments of this application, the first information mentioned above includes at least one of the following:

[0108] The relative position information of the first object relative to the user;

[0109] The preview image shows the relative position information between the first and second objects.

[0110] In some embodiments of this application, the second object can be an object different from the first object in the preview image. For example, the second object can be an object adjacent to the first object. For instance, if the preview image includes a table with a water glass, a mobile phone, and a vase placed on it, with the water glass to the left of the mobile phone and the vase to the right, then the first object can be the mobile phone, and the second object can be the table, or the second object can be the table and the water glass, or the second object can be the table and the vase, or the second object can be the table, the water glass, and the vase.

[0111] In some embodiments of this application, the first information may include a relative orientation description based on the user's perspective, that is, the relative position information of the first object relative to the user. For example, if a user asks "Where is my phone?", the electronic device may respond to the user with "Your phone is located to your right front".

[0112] In some embodiments of this application, the first information may include a description of the relative spatial positions of various objects within the preview image, that is, the relative position information of the first object with respect to other objects in the preview image. For example, if a user asks "Where is my phone?", the electronic device may tell the user "It's on the table to the right of the stool, in front of the green plants and the lamp."

[0113] In some embodiments of this application, the first information may include the relative position information of the first object relative to the user and the relative position information of the first object with other objects in the preview image. For example, if a user asks "Where is my phone?", the electronic device may tell the user "The phone is on the table in front of you, to the right of the water glass, next to a lamp and a vase."

[0114] In this embodiment of the application, by integrating the user's relative coordinate description with the description of the spatial position relationship of multiple targets within the screen, the electronic device can indicate the spatial position of the first object to the user in a multi-level manner. Specifically, by using the spatial position relationship between the user and the first object and the spatial position relationship between the first object and other objects in the environment, the user can accurately know the relative position of the first object relative to himself and the relative position of the first object relative to other surrounding objects, thereby quickly and accurately locating the position of the first object and improving the accuracy of object position determination.

[0115] In some embodiments of this application, the electronic device may generate first information based on first dialogue information and a preview image, which comprehensively describes the location of the first object in the environment so that the user can understand it.

[0116] In some embodiments of this application, the electronic device can use an object recognition algorithm to identify one or more objects included in a preview image and find the first object therefrom.

[0117] In some embodiments of this application, the electronic device can jointly process the first dialogue information and the acquired preview image through a multimodal visual perception model deployed on the device side. Exemplarily, the electronic device uses this multimodal visual perception model to simultaneously perform deep encoding, alignment, and reasoning on information from both linguistic and visual modalities. Exemplarily, during processing, the multimodal visual perception model searches and matches in the image feature information space based on the semantics of the first dialogue information to determine the position of the first object, such as its coordinates, and simultaneously understands its spatial positional relationship with other objects in the scene. Then, based on the reasoned spatial positional relationship, structured positional information is generated. Based on this positional information, the electronic device generates positional description information, i.e., the first information, thereby achieving more accurate and efficient intent understanding and object localization.

[0118] In some embodiments of this application, the electronic device can use a built-in computer vision algorithm to perform visual understanding and analysis on the acquired images, identify objects in the images and their approximate locations; and, combined with semantic parsing of the first dialogue information, determine the first object that the user needs to query; then, based on the visual analysis results and semantic understanding results, generate location description information for the first object, i.e., first information, thereby achieving rapid identification and positioning of the object.

[0119] In some embodiments of this application, after generating the first information, the electronic device can use various methods to feed the information back to the user in order to adapt to different usage scenarios and user needs.

[0120] In some embodiments of this application, the electronic device may broadcast the first information through a speech synthesis module; or, the first information may be displayed in text or graphic form on the display interface of the electronic device; or, the electronic device may combine speech broadcasting and visual display to provide multiple types of comprehensive feedback.

[0121] In some examples, the user inputs the voice command, "Help me find my phone?" After recognizing the current environment, the electronic device broadcasts the first piece of information via voice: "The phone is on the table in front of you, to the right of the water glass, next to a lamp and a vase." This feedback simultaneously includes information about the phone's relative position to the user, as well as its relative position to other objects in the scene, such as the table, water glass, lamp, and vase.

[0122] In some examples, the user inputs a voice message: "Help me find my white water glass?" After the electronic device locates a white water glass on the screen, it displays the text message "The white water glass is located slightly to the right of the center of the screen" overlaid at the top of the screen.

[0123] In some examples, the user inputs a voice message: "Help me find my suitcase?" After recognizing the suitcase, the electronic device announces, "The suitcase is about three meters directly in front of you, behind the red chair."

[0124] In this embodiment, the electronic device can locate a specific target in the real-time environment based on the user's dialogue information and provide feedback with spatial location information that is easy for the user to understand. This enables a more comprehensive and accurate understanding and interactive feedback of the screen content. The user can use the terminal device to query the location of the target object, and the terminal device can use voice to reply with the location of the target object to assist the user in perceiving the screen content, thereby providing convenience for users, especially visually impaired people.

[0125] The object query method provided in this application embodiment involves an electronic device displaying first dialogue information on a first interface. This first dialogue information is used to indicate the search for a first object. The first interface includes a preview image captured by a camera. Based on the first dialogue information and the preview image, an object search is performed, and first information is output. The first information includes at least one of the following: the relative position information of the first object relative to the user, and the relative position information of the first object with other objects in the preview image. Through this scheme, the electronic device can output information including the relative position information between the first object and the user, or the relative position information between the first object and other objects, based on the first dialogue information and the preview image input by the user. This allows the user to quickly and accurately determine the location of the first object based on its relative position with itself, or based on its relative position with other known objects, thus greatly improving the accuracy and intelligence of environmental perception.

[0126] In some embodiments of this application, the first interface includes a voice dialogue control; exemplarily, step 201, the process of obtaining the first dialogue information, may include steps 201a and 201b:

[0127] Step 201a: The electronic device receives the first input to the voice dialogue control.

[0128] Step 201b: The electronic device responds to the first input, activates the voice interaction mode, and displays the first dialogue information collected by the microphone on the first interface.

[0129] In some embodiments of this application, the first input is used to trigger the activation of the voice interaction mode, trigger the collection of first dialogue information through the microphone, and display the first dialogue information on the first interface.

[0130] In some embodiments of this application, the first input may include any of the following: user touch input, voice input, gesture input, folding or unfolding operation of the foldable screen, or other feasible inputs, which are not limited in this application embodiment. Further, the touch input may be: user click input, swipe input, press input, etc. Further, the click operation may be any number of clicks. The swipe operation may be a swipe operation in any direction, such as swiping up, swiping down, swiping left, or swiping right, which are not limited in this application embodiment.

[0131] In some embodiments of this application, the aforementioned voice dialogue control can be an interactive element in a graphical user interface, and its form includes, but is not limited to, buttons, icons, or floating operation buttons.

[0132] In some embodiments of this application, when in voice interaction mode, the electronic device begins to collect ambient sound through its microphone array and waits to receive specific instructions spoken by the user. Subsequently, the electronic device continues to collect the user's voice stream until the end of the voice is detected, thereby obtaining complete first dialogue information. This first dialogue information can be used directly for subsequent processing, or converted into text form in real time by an Automatic Speech Recognition (ASR) module.

[0133] For example, combined Figure 2 With the first interface 21 displayed, the user clicks the voice dialogue control 26 in the lower left corner. The electronic device responds to this click by displaying the text "Ask me a question" on the first interface 21, indicating that voice interaction mode has been entered. The user then says, "Where is my phone?" The electronic device captures this voice message through its microphone and uses it as the first dialogue information to be processed.

[0134] In this embodiment, by displaying a voice dialogue control on the first interface, users can intuitively see the voice dialogue control and conveniently activate the auxiliary query function. Secondly, after the voice dialogue control is triggered, it enters voice interaction mode and performs voice acquisition, avoiding misrecognition caused by user error and improving the reliability of the interaction.

[0135] In some embodiments of this application, exemplarily speaking, after performing object lookup based on the first dialogue information and the preview image in step 202 above, the object query method provided in this application embodiment may further include the following step 204:

[0136] Step 204: The electronic device displays the first identifier on the first interface.

[0137] The first identifier is used to mark the display area of ​​the first object in the preview image.

[0138] In some embodiments of this application, after the electronic device identifies a first object based on a preview image and first dialogue information, the electronic device displays a first identifier on a first interface to mark the first object in the preview image.

[0139] In some embodiments of this application, the electronic device displays a first identifier on a preview image displayed on a first interface to highlight the first object that has been successfully located by enhancing visual cues.

[0140] In some embodiments of this application, the above-mentioned display of the first identifier may include, but is not limited to: drawing a bounding box surrounding the first object, outlining the object outline with a highlight color, making the area where the first object is located flash, and overlaying a specific icon or indicator arrow above the first object.

[0141] It should be noted that the purpose of displaying the first identifier is to highlight the first object, and its specific visual style can be set according to actual needs. This application embodiment does not limit this.

[0142] For example, combined Figure 2 The user inputs a voice message, "Help me find my phone." After recognizing the phone in the current frame, the electronic device first announces its location via voice: "The phone is on the table to the right of the stool, in front of the green plants and the lamp." Simultaneously, a rectangular frame is displayed on the real-time video feed on the first interface 21, selecting the phone to visually mark it in the frame. This creates consistent spatial guidance between the voice feedback and the visual result, enhancing the user's overall perception of the scene structure and providing intuitive location guidance.

[0143] In this embodiment, by displaying a first identifier on a first interface, a real-time visual marking operation is performed on the first object after successful recognition, thereby realizing a feedback mechanism that combines voice description and visual indication. Specifically, the voice feedback provides spatial location relationships that can be obtained through hearing, while the display of the first identifier provides visual object indication. In this way, by combining voice feedback and visual indication, the user's overall perception of the scene structure is enhanced, enabling the user to quickly and unambiguously understand the location of the first object by combining the two types of information.

[0144] In some embodiments of this application, step 202 may include steps 202a and 202b:

[0145] Step 202a: The electronic device inputs the first dialogue information and the preview image into the multimodal visual perception model. Through the multimodal visual perception model, the first dialogue information and the preview image are processed in a multimodal manner to obtain prediction information.

[0146] The aforementioned prediction information includes at least one of the following: coordinate information of the first object, and relative position information between the first object and at least one reference object in the preview image.

[0147] Step 202b: The electronic device outputs the first information based on the predicted information.

[0148] In some embodiments of this application, the electronic device may preprocess the first dialogue information and the preview image, and use the preprocessed first dialogue information and the preview image as joint input to a locally deployed multimodal visual perception model.

[0149] In some embodiments of this application, the multimodal visual perception model is a specially trained neural network model capable of simultaneously processing and understanding linguistic and visual information. For example, the multimodal visual perception model is a 3B multimodal visual perception model.

[0150] In this embodiment, a multimodal 3B visual big model deployed on the device side is used to realize the perception, understanding, and localization reasoning of screen content without relying on cloud computing or external requests. This reduces network dependence and improves real-time response while ensuring user privacy and data security, making it suitable for terminal application scenarios with high privacy and latency requirements. Furthermore, this application uses a device-side 3B model, eliminating the need to upload data to the cloud, thus providing excellent protection for user privacy. The model is directly installed on the user's electronic device, such as a mobile phone, eliminating the need for the user to download the model after purchasing the phone. When the user activates accessibility features, the system should request permissions and prompt the user to enable the offline big model's visual perception function for camera and screen content.

[0151] In some embodiments of this application, after receiving input, the multimodal visual perception model performs end-to-end multimodal inference.

[0152] In some embodiments of this application, a multimodal visual perception model is used to perform deep cross-modal interaction and joint reasoning on the input first dialogue information and preview image, understand the semantics of the dialogue instructions, and selectively search and match visual objects that match the description of the dialogue information in the image feature information space, and analyze the geometric and semantic relationship between the object and reference objects in the scene to obtain structured prediction information.

[0153] In some embodiments of this application, the aforementioned reference object can be an object different from the first object in the preview image. For example, the reference object can be an object adjacent to the first object in position, or an object relatively close to the first object. For instance, if the first object is a mobile phone on a table, the reference object can be the table on which the phone is placed. As another example, if the first object is a mobile phone on a table, the reference objects can be the table on which the phone is placed, and a vase placed on the table to the right of the phone.

[0154] In some embodiments of this application, the structured prediction information output by the above inference process may include at least one of the following:

[0155] The coordinate information of the first object: for example, in the form of bounding box coordinates [x1, y1, x2, y2], accurately characterizing the position and range of the first object in the pixel coordinate system of the preview image.

[0156] The relative positional information between the first object and other objects in the preview image: for example, a structured or vectorized relational representation used to describe the spatial positional relationship of the first object, such as "to the left of A" or "above B".

[0157] In some embodiments of this application, after obtaining structured prediction information, the electronic device converts it into an output form that the user can naturally perceive, such as outputting audio information or displaying text information.

[0158] For example, the electronic device maps coordinate information into a location description based on the user's perspective, such as "in front of your right", and organizes the relative positional relationships between objects into a coherent statement that conforms to everyday expression habits, that is, generates first information, which is a complete natural language description that integrates at least one of the above-mentioned positional dimensions, and broadcasts it through the device's speech synthesis module, and / or displays it on the display interface.

[0159] For example, the user inputs the voice message: "Where is my phone?" The preview image is a view of the current living room scene. After performing multimodal reasoning through a multimodal visual perception model, the electronic device outputs predicted information: the remote control's coordinates are [200, 150, 280, 200], and the relationship information: the phone is on the table. Then, based on this, the electronic device generates the first message: "The phone is on the table to your left and in front of you," and announces this message via voice, while simultaneously drawing a highlighted box at the corresponding coordinates on the screen to emphasize the phone.

[0160] The object query method provided in this application is illustrated by specific examples below.

[0161] For example, the electronic device performs location information perception and assisted interaction based on visual reasoning, calling upon a trained 3B multimodal visual perception model to achieve semantic-level understanding and location information perception of the screen content. This function is specifically designed to meet the actual needs of visually impaired users when using mobile devices.

[0162] For example, in actual use, users can input target query commands via voice through the device, such as "Where is my phone?". The electronic device first converts the voice signal into a text command and inputs this command along with the current screen image into a pre-trained 3b edge-side multimodal visual perception model. After receiving the image and text command, the model performs multi-layer semantic understanding and spatial analysis of the screen content based on a multimodal feature alignment mechanism. Specifically, the model extracts target region features from the image through a visual encoding module and parses object information in the user command through a language encoding module, including category, color, appearance features, and spatial constraints. Then, the model matches the target attributes with the image region in the joint feature space to achieve precise localization of the target object.

[0163] It should be noted that, because the model incorporates descriptive data containing location information and spatial positional relationship annotation data during the training phase, and strengthens visual consistency through model constraints, its inference results can not only distinguish multiple objects with highly similar appearances in the picture, such as "red stockings" and "black short socks", but also accurately determine the spatial positional relationship and relative orientation between objects.

[0164] For example, while generating visual understanding results, the electronic device can return the location information of the first object to the user in natural language. This location information can include a relative orientation description based on the user's perspective, such as "your phone is located to your right front," and can also include a description of the relative spatial relationship within the screen, such as "located on the table to the right of the stool, in front of the green plants and the lamp." By integrating the user's relative coordinate description with the description of the spatial relationship between multiple objects within the screen, the electronic device can express the spatial location of the target object in a multi-layered manner, significantly improving the accuracy and understandability of the location feedback.

[0165] For example, combined Figure 2 ,like Figure 4 As shown, the user clicks the voice dialogue control 26 in the lower left corner of the screen to enter voice interaction mode, such as... Figure 5 As shown, the electronic device enters a query-reading state and displays the message "Ask me a question" on the interface; the user then gives a targeted query command via voice, such as "Where is my phone?". Figure 6As shown, after receiving a voice command, the electronic device performs multimodal understanding and spatial analysis on the current screen image and generates query results including the target location and spatial relationships. The electronic device then broadcasts detailed spatial location information of the target object to the user via voice, such as "The phone is on the table in front of you, to the right of the water glass, next to a lamp and a vase." During the voice broadcast, interface 21 displays a "Click to interrupt" prompt, allowing the user to interrupt the voice feedback at any time, thus improving the flexibility and controllability of the interaction. Simultaneously with the output of the verbal description, the electronic device visually highlights the corresponding area of ​​the target object on the screen, creating consistent spatial guidance between the voice feedback and the visual results, thereby enhancing the user's overall perception of the scene structure. This multimodal feedback method effectively reduces ambiguity caused by single voice or single visual prompts, and is particularly suitable for the spatial cognitive needs of visually impaired users regarding complex screen content. Thus, through the above interaction mechanism, visually impaired users can query and locate target objects on the screen without relying on visual operations, enabling them to independently complete daily information acquisition and interactive operations, significantly improving the practicality and user experience of smart terminals in accessible usage scenarios.

[0166] In this embodiment, an end-to-end processing framework employing a multimodal perception model improves the accuracy and efficiency of object querying. Specifically, through joint inference, the multimodal visual perception model can locate specific objects in an image based on the user's natural language commands, avoiding the error problems caused by first identifying all objects in the entire image and then performing text matching, thereby improving positioning accuracy. Secondly, the multimodal perception model not only outputs the coordinates of the object but also understands and outputs the spatial relationship between the first object and other reference objects in the environment, thus obtaining a multi-level spatial description of the scene context, improving the comprehensiveness and accuracy of object querying.

[0167] In some embodiments of this application, the multimodal visual perception model includes a language encoding module, a visual encoding module, and a cross-modal processing module; exemplarily, step 202a may include steps 202a1 to 202a3:

[0168] Step 202a1: The electronic device performs text encoding processing on the first dialogue information through the language encoding module to obtain text feature information.

[0169] Step 202a2: The electronic device performs image encoding processing on the preview image through the visual encoding module to obtain image feature information.

[0170] Step 202a3: The electronic device performs fusion reasoning processing on text feature information and image feature information through the cross-modal processing module to obtain prediction information.

[0171] In some embodiments of this application, the language encoding module described above can be built based on a pre-trained language model, such as based on a Transformer encoder architecture. This language encoding module receives first dialogue information in text form, such as "Where is my phone?", and performs lexical parsing, semantic understanding, and contextual representation on it. Its output is text feature information, specifically encoding the complete semantic intent of the user query, key attributes such as "phone", and the implicit query target.

[0172] For example, for the first dialogue information "find a blue suitcase made of metal", the language encoding module will parse the first object "suitcase", whose attributes are "metal" and "blue". The output text feature information is a mathematical vector representing this complex query intent, thereby providing clear semantic query conditions for subsequent visual search.

[0173] In some embodiments of this application, the visual encoding module is typically based on a deep convolutional neural network or a visual Transformer architecture. Exemplarily, this visual encoding module receives a preview image and extracts visual information from the preview image, ranging from low-level edges and textures to high-level semantic objects, through multiple layers of nonlinear transformation. Its output is image feature information, which can be a global image feature information vector or a set of region feature vectors related to the spatial location of the image. This feature vector encodes the visual appearance and contextual information of each region in the image.

[0174] In some embodiments of this application, the cross-modal processing module receives text feature information and image feature information from the first two modules. In the image feature information space, it determines the specific region that best matches the text query intent and decodes it into the precise coordinate information of the first object. At the same time, by analyzing the relationship between the features of the target region where the first object is located and the features of other regions in the image in the joint space, it infers and outputs in a structured manner the relative position information between the target and other significant objects in the scene, such as spatial orientation.

[0175] For example, in conjunction with the above example, the textual feature information representing a blue suitcase made of metal is input into the cross-modal processing module. This cross-modal processing module guides the model to focus on areas in the image feature information that conform to the shapes of "metallic material", "blue", and "suitcase". Finally, it outputs structured prediction information containing the coordinates of the suitcase and its position relative to the wall, such as "in front of the wall" or "the blue suitcase made of metal is in front of the wall".

[0176] In this embodiment, by performing cross-modal fusion inference on preview images and dialogue information, the semantic understanding and visual perception in object queries can be deeply unified, achieving precise fine-grained alignment from semantics to vision. Specifically, the language encoding module transforms ambiguous natural language instructions into precise query vectors, the visual encoding module provides rich environmental visual information, and the cross-modal processing module achieves matching between the two through an attention mechanism. This enables accurate execution of queries containing complex attributes, relationships, and referential meanings, thereby significantly improving the accuracy of object queries in complex scenarios.

[0177] In some embodiments of this application, prior to step 202 above, the object query method provided in this application may further include steps 205 and 206:

[0178] Step 205: The electronic device inputs the preview image into the multimodal visual perception model, and uses the multimodal visual perception model to locate objects in the preview image to obtain the coordinate information of the first object.

[0179] Step 206: Based on the coordinate information, the electronic device determines that the first object is located in the edge area of ​​the preview image and outputs the second information.

[0180] The second information is used to prompt the user to move the electronic device so that the first object is located in the center area of ​​the preview image.

[0181] In some embodiments of this application, the electronic device can analyze the camera image in real time, determine the composition position of the first object, and guide the user to move the electronic device when the first object is located at the edge of the image so that the first object is in the ideal field of view, thereby enabling subsequent accurate identification.

[0182] In some embodiments of this application, after receiving the user's initial dialogue information, the electronic device begins to continuously input image frames from the video stream into a multimodal visual perception model. The multimodal visual perception model processes each input image frame to quickly locate objects that match the user's command description.

[0183] In some embodiments of this application, after the electronic device obtains the coordinate information of the first object, it determines whether the first object is located in the edge region of the current image by calculating the distance between the center point of the bounding box of the first object and the center point of the image, or by determining whether the bounding box intersects with the image boundary by more than a certain proportion, based on a preset threshold.

[0184] In some embodiments of this application, the aforementioned second information is used to prompt the user to move the electronic device toward the direction of the first object, so as to place the first object in the center area of ​​the subsequently acquired image. For example, if the first object is off-center to the left in the image, the user is prompted to "move left"; if it is off-center to the right, the user is prompted to "move right". When the user moves the electronic device accordingly, the camera orientation changes, and the first object moves toward the center of the image in subsequent image frames.

[0185] For example, when a user commands "Where is my suitcase?", the electronic device identifies a suitcase in the initial camera view, but its bounding box is close to the right edge of the image. The electronic device calculates its center point and determines that it is located in the edge region of the preview image. Then, the electronic device announces a second message via voice: "The suitcase is on your right. Please slowly turn your phone to the right." The user follows the prompt and turns the phone to the right, and the suitcase moves towards the center in the subsequent image. When the electronic device determines that the suitcase is no longer in the edge region, it uses the image captured at this time, where the suitcase is located in the center of the image, as the preview image, and generates and outputs a first message based on this preview image, such as "The suitcase is about two meters directly in front of you."

[0186] In some embodiments of this application, the electronic device may display second information or broadcast second information via voice on a first interface. For example, it may display "The suitcase is about two meters in front of you" on the first interface, or broadcast "The suitcase is about two meters in front of you" via voice.

[0187] The object query method provided in this application is illustrated by specific examples below.

[0188] For example, the edge-side 3B large model also has the function of recognizing objects at the edge of the screen and can output very detailed descriptions. When the queried object is at the edge of the screen, if the model detects that the object is at the edge of the screen, it will guide the user to move the screen in the direction of the object. As the object slowly appears on the screen, it will provide a voice response. For example, if the user asks, "Can you help me find my suitcase?", combined with... Figure 5 If the device detects a suitcase at the edge of the frame, it will prompt the user to move the device towards the suitcase, saying, "Try moving the camera to the right." After the user moves the electronic device to the right, as shown... Figure 7 As shown, most of the suitcase appears in the picture, such as Figure 8 As shown, when the device moves continuously until the suitcase is fully visible in the frame, it outputs the voice message, "The suitcase is in front of you to your right, behind the table," while simultaneously highlighting the suitcase on the screen.

[0189] For example, if the object being queried is neither in the frame nor at the edge, a message will be displayed indicating that there is no first object in the current frame, suggesting that the user move the camera and make another judgment. For instance, when there is no "car key" in the frame, the user can ask, "Can you find my car key?" If the model does not detect "car key" in the frame, it will reply with a voice message, "I don't see the car key right now. Can you move the camera?" The user can then move the device and ask again.

[0190] In this embodiment, by actively judging and guiding the user to move the electronic device, the first object is moved to a non-edge area of ​​the image, thereby improving the accuracy of subsequent multimodal inference and avoiding errors or incompleteness in the location description caused by partial occlusion of the target or poor viewing angle. Secondly, by providing the user with real-time and clear spatial guidance through the second information, the user can accurately move the electronic device so that the first object can be placed in the center area of ​​the preview image, thereby accurately locating the first object.

[0191] In some embodiments of this application, prior to step 202a above, the object query method provided in this application may further include steps 208 and 209:

[0192] Step 208: The electronic device acquires the first training data, the second training data, and the third training data.

[0193] Step 209: The electronic device trains the initial model based on the first training data, the second training data, and the third training data to obtain a multimodal visual perception model.

[0194] The first training data includes at least one first sample image and first annotation information corresponding to each first sample image; the first annotation information corresponding to each first sample image includes object description information and coordinate information of at least one object in the first sample image; the second training data includes at least one second sample image and second annotation information corresponding to each second sample image; the second annotation information corresponding to each second sample image includes relative position information between any two objects in the second sample image; the third training data includes at least one third sample image and first dialogue information corresponding to each third sample image.

[0195] In some embodiments of this application, the aforementioned first training data is used to establish fine-grained semantic, visual, and positional associations. Exemplarily, the first training data includes at least one first sample image and first annotation information corresponding to each first sample image. Each first annotation information contains natural language description information of at least one object in the first sample image, and also associates the object's coordinate information in the image, such as a bounding box [x1, y1, x2, y2]. Through the first training data, the model can learn to correspond specific textual descriptions with specific pixel regions and their spatial ranges in the image, thereby achieving accurate localization.

[0196] For example, the first training data mentioned above is the original image caption, with the corresponding coordinate information of each object inserted.

[0197] In some embodiments of this application, the aforementioned second training data is used to provide prior knowledge of spatial relationships. It includes at least one second sample image and second annotation information corresponding to each second sample image. Each second annotation information describes a preset spatial relative positional relationship between multiple objects in the second sample image, such as "up, down, left, right, front, back". Through structured or natural language relationship annotations, the model is able to understand and infer the layout relationships of different objects in two-dimensional and three-dimensional space.

[0198] For example, the second training data mentioned above is spatial perception data, such as spatial position relationship data pairs. This second training data is mainly used to characterize the spatial orientation relationship between objects, specifically including six directions: up, down, left, right, front, and back.

[0199] In some embodiments of this application, the aforementioned third training data is used to align the model with the user's first dialogue information. It includes at least one third sample image and the first dialogue information corresponding to each third sample image. Exemplarily, this first dialogue information is used to simulate natural language queries that the user might make, such as "Where is my phone?" or "Find all the cups in the picture." This type of data is used to enable the model to learn and understand the various ways, intentions, and referential habits of user queries, thereby improving its ability to understand the user's complex query intentions.

[0200] For example, the third training data may include multi-object localization data and referential data. The multi-object localization data includes multiple objects, with each object corresponding to a coordinate. For example, the multi-object localization data could be: Please find the mobile phone and the water cup. The referential data includes a description of an object, which corresponds to a coordinate. For example, the referential data could be: Please find the white water cup.

[0201] It should be noted that in visually impaired scenarios, the model needs to have spatial perception capabilities, that is, it needs to know the specific location information of the target object [x1, y1, x2, y2], as well as the relative positional relationship, for example, the "phone" is to the right of the "water cup" or the "phone" is on top of the "table". Therefore, it is necessary to construct training data that includes object position information to train the model.

[0202] In some embodiments of this application, the parameters of the initial model are iteratively optimized using the training data described above through machine learning paradigms such as supervised learning and pre-training fine-tuning, so that the initial model learns the mapping ability from complex instructions to precise positioning and relational description.

[0203] In this embodiment, training the initial model with first and second training data enhances its fine-grained localization capabilities and accurately answers spatial orientation information. Specifically, the new data labeling method introduces two new training data formats to characterize the spatial relationships and semantic constraints between targets. Through joint training with multiple data types, the model can not only achieve precise target localization but also perform fine-grained matching of complex semantic descriptions and accurately answer spatial orientation information related to the target, thereby significantly improving the expressive power and application scope of visual localization tasks. Secondly, a spatial perception capability model is constructed through mixed training with multiple data types. Specifically, at the data level, a mixed training strategy using multiple data types is adopted. By using reasonable data ratios to jointly optimize the model, the model not only possesses the ability to semantically describe the content of the image but also accurately understands the spatial position and orientation relationships of target objects, thereby supporting coordinate-level localization and spatial position information output, significantly improving the model's spatial perception capability.

[0204] In some embodiments of this application, the process of obtaining the first training data in step 208 above may include the following steps A1 to A4:

[0205] Step A1: The electronic device acquires at least one first sample image and descriptive information corresponding to each first sample image.

[0206] The aforementioned descriptive information includes object description information of at least one object in the first sample image.

[0207] Step A2: The electronic device performs semantic parsing processing on the description information corresponding to each first sample image to determine at least one object in each first sample image.

[0208] Step A3: The electronic device performs visual positioning on each first sample image using a visual positioning model to determine the coordinate information of at least one object in each first sample image.

[0209] Step A4: The electronic device adds the coordinate information of each object in at least one object to the first text position corresponding to each object to obtain the first training data.

[0210] The first text position mentioned above refers to the text position where the object description information of the object is located in the first description information, and the first description information is the description information corresponding to the first sample image.

[0211] In some embodiments of this application, the first sample image is the original image to be labeled. Exemplarily, the first sample image may be a scene photograph from a public dataset or collected in-house. Exemplarily, each first sample image is associated with a descriptive message, which may be a natural language text description of the image's content, including a description of at least one object in the image. For example, the descriptive message for an indoor scene image might be: "There is a table in the living room, on the table are a mobile phone, a vase, and a table lamp; there is a chair to the left of the table, and a suitcase to the right of the table."

[0212] In some embodiments of this application, the electronic device can utilize natural language processing technology to perform text analysis on descriptive information. For example, by using methods such as part-of-speech tagging, dependency parsing, or named entity recognition, noun phrases representing specific visual entities can be extracted from the descriptive text, thereby identifying at least one object. For instance, objects such as "sofa," "chair," and "suitcase" can be parsed from the above description.

[0213] In some embodiments of this application, the aforementioned visual localization model is a pre-trained model capable of locating image regions based on text descriptions. Exemplarily, the electronic device inputs the parsed text description of each object along with the corresponding first sample image into the model. The model outputs the coordinate information of the object referred to by each text description in the image.

[0214] In some embodiments of this application, the aforementioned first text position refers to the position of a word or phrase describing a specific object in the original image description information text. The electronic device will obtain coordinate information, such as [x1, y1, x2, y2], and embed it in a structured manner, for example, as a suffix or inserted into a specific mark, after or at the corresponding position of the object description information in the first description information.

[0215] For example, the specific data labeling method is as follows: First, the Qwen3-VL-Plus model is used to analyze the original description and extract the various main objects involved in the description; then, a self-developed visual positioning model is called to process the image and label the precise coordinate information of each subject; finally, the coordinate information of each subject is embedded into the original caption to form a correspondence between the subject and the coordinates, thereby constructing training samples that support fine-grained positioning and spatial relationship understanding. This method allows the model to acquire semantic information and coordinate position information simultaneously during training, enhancing the ability to accurately identify targets and perform visual reasoning in complex scenes.

[0216] In some examples, the image description is: "A cat is lying on a windowsill, looking outside." The electronic device resolves the objects "a cat" and "windowsill." A visual localization model outlines the cat's position in the image, such as [100, 200, 250, 350], and the windowsill's position, such as [50, 180, 300, 400]. Then, the coordinate information is inserted into the corresponding text positions, and the annotation in the generated first training data becomes: "A cat [100, 200, 250, 350] is lying on a windowsill [50, 180, 300, 400], looking outside."

[0217] In some embodiments of this application, the object description information may include the object's location information, attribute information, etc. For example, assuming the object is a person, the object description information may include the object's age, clothing, height, actions, facial features, expression, etc.

[0218] In some examples, the description information corresponding to the sample image and the coordinate information of at least one object in the sample image are as follows: Image description: "This is a real photograph taken in an open lawn area of ​​a city park or cultural square. Slightly to the left of center in the image is a boy approximately 3–5 years old." <box> [279,307,825,928]< / box> Wearing a beige printed jacket <box> [279,445,597,699]< / box> (Cartoon print), pink trousers <box> [332,652,697,868]< / box> (Pants with Mickey Mouse pattern on the cuffs) and black sneakers <box> [499,749,659,928]< / box> Wearing a bright blue fisherman's hat <box> [314,307,535,466]< / box> Wearing round sunglasses with pink frames <box> [358,370,501,413]< / box> Sit relaxed and hold the blue plastic rotator with both hands. <box> [534,558,675,677]< / box> With a bright smile and sparkling eyes, she leaned slightly forward, displaying a lively demeanor. Two adults can be seen in the background: one in a black suit. <box> [576,251,612,340]< / box> Standing at a distance, another person wearing a long brown coat <box> [739,250,782,344]< / box> Holding colorful kites <box> [0,165,244,635]< / box>The child is running, the movements natural. In the distance, the buildings are modern in style, with large glass curtain walls and light gray stone facades, without any signs or markings. A corner of a red garment is visible on the left edge, possibly belonging to a child. <box> [279,307,825,928]< / box> Standing. The overall environment is winter or early spring, with bare trees, withered grass, a gray sky, soft lighting, and no strong shadows, consistent with overcast or foggy weather. Although there is no textual information, combined with the character interactions, props (kites, spinning tops), and architectural background, it can be inferred that the event is a family outdoor recreational activity, with a relaxed and joyful atmosphere, full of childlike fun and a sense of freedom. The benefits of this training data, such as... Figure 9 As shown, the left side is a visualization of the reasoning results, which will select all objects in detail. Corresponding to the image description information above, you can see that each box has a very detailed description, which has a strong fine-grained perception and a more complete and detailed description.

[0219] It should be noted that [x1, y1, x2, y2], where x1 and y1 are the x and y coordinates of the top left corner of the rectangular area where the object is located, respectively, and x2 and y2 are the x and y coordinates of the bottom right corner of the rectangular area where the object is located, respectively.

[0220] In this embodiment, by utilizing image-descriptive text and an efficient visual localization model to construct the first training data, a massive amount of image-text pairs with precise coordinate annotations can be quickly generated, ensuring the accuracy and consistency of the annotations. By strictly corresponding the text parsing objects with the visual localization results and embedding them into the original descriptive information, the alignment between text descriptions, visual entities, and spatial locations in the generated training samples is ensured, providing reliable information for model learning and thereby improving the model's ability to understand fine-grained semantics and complex scenes.

[0221] In some embodiments of this application, the process of obtaining the second training data in step 208 above may include the following steps B1 to B3:

[0222] Step B1: The electronic device acquires at least one second sample image.

[0223] Step B2: The electronic device determines the spatial relationship between at least two objects in each second sample image.

[0224] The aforementioned spatial positional relationships include two-dimensional spatial positional relationships and three-dimensional spatial positional relationships.

[0225] Step B3: The electronic device constructs the second training data based on the spatial relationship between at least two objects in each second sample image.

[0226] In some embodiments of this application, the aforementioned second sample image is the original image used to construct a spatial location relationship dataset, which can cover various scenes such as indoor, outdoor, natural, and artificial environments, to ensure that the spatial location relationships learned by the model have broad generalization.

[0227] In some embodiments of this application, the aforementioned spatial positional relationships include two-dimensional spatial positional relationships, such as up, down, left, and right, and may also include three-dimensional spatial positional relationships, such as front and back, constituting a complete six-degree-of-freedom description of the orientation between objects. It should be noted that the two-dimensional relationship is determined based on the image plane coordinates, while the three-dimensional relationship requires the introduction of depth information for reasoning.

[0228] In some embodiments of this application, the electronic device associates the spatial positional relationship of the second sample image with the corresponding second sample image in a structured format to form a second training data sample.

[0229] It should be noted that the second training data includes spatial location relationship data pairs corresponding to the images, that is, the second training sample is a dataset containing "image-spatial location relationship pairs".

[0230] For example, for a second sample image containing "water glass", "table", and "person", the constructed second training data can be: [

[0232] {"subject": "water cup", "object": "table", "relation": "on"},

[0233] {"subject": "person", "object": "table", "relation": "after"},

[0234] {"subject": "person", "object": "water cup", "relation": "after"}

[0235] ].

[0236] In this embodiment, a second training data comprising annotations of two-dimensional and three-dimensional spatial positions is constructed and used for model training. This data primarily characterizes the spatial orientation relationships between target objects, specifically including six directions: up, down, left, right, front, and back. By introducing this second training data, the overall perception capability of the multimodal visual perception model in both on-screen and general scenarios is significantly improved. Specifically, by introducing spatial perception data containing annotations of spatial positions such as up / down, left / right, and front / back, the model can not only accurately return the two-dimensional coordinates of the target but also understand and infer the relative spatial orientation relationships between targets, significantly improving spatial perception and spatial reasoning capabilities.

[0237] In some embodiments of this application, the above spatial positional relationship includes a two-dimensional spatial positional relationship; exemplarily, step B2 may include steps C1 to C3:

[0238] Step C1: The electronic device performs content analysis on each second sample image using a multimodal large model to obtain descriptive information corresponding to each object in each second sample image.

[0239] Step C2: The electronic device uses a visual positioning model to visually locate the description information and obtain the coordinate information of each object in the second sample image.

[0240] Step C3: The electronic device determines the two-dimensional spatial positional relationship between at least two adjacent objects in each second sample image based on the coordinate information of at least one object in each second sample image.

[0241] In some embodiments of this application, the aforementioned multimodal large model is a pre-trained large model capable of simultaneously processing and understanding images and text. For example, a second sample image is input into the multimodal large model, which performs semantic parsing on the image content, identifies objects present in the image, and generates a corresponding, accurate text description for each identified object. For instance, for an office image, the model might output: [“a black laptop,” “a white coffee cup,” “a book”].

[0242] In some embodiments of this application, after obtaining the text description of the object, the text description and the original second sample image are input into a visual localization model. Exemplarily, this visual localization model can be an open-vocabulary object detection model, used to locate the corresponding object region in the image based on text prompts. For each text description, the model outputs its corresponding bounding box coordinates in the image, typically represented in the format [x1, y1, x2, y2], thereby assigning a precise two-dimensional spatial location to each semantic object.

[0243] In some embodiments of this application, after obtaining the coordinates of all objects, the electronic device calculates their relative positions by comparing the coordinate values ​​of the bounding boxes of different objects. For example, the X-axis coordinates of the center points of two bounding boxes are compared. The object with the smaller X-coordinate value is determined to be "to the left" of the other object, and vice versa. The Y-axis coordinates of the center points of two bounding boxes are compared. In an image coordinate system, the Y-axis is generally positive downwards, so the object with the smaller Y-coordinate value is determined to be "above" the other object, and vice versa. In this way, by comparing all objects pairwise, a set of structured two-dimensional spatial positional relationships can be generated.

[0244] For example, in the data construction process, the input image is first parsed using the multimodal large model Qwen-VL-Plus to extract multiple main targets and their corresponding text descriptions. Then, the text information of the main targets is input into the visual localization model GroundingDINO to obtain the corresponding coordinate positions of each main target in the image. After obtaining the coordinate information, by comparing and analyzing the coordinate values ​​of different targets, the two-dimensional spatial positional relationship between the target objects can be determined, including relative orientations such as up, down, left, and right.

[0245] In this embodiment, by integrating the content parsing capability of a multimodal large model with the precise positioning capability of a visual positioning model, and designing automatic relationship derivation rules based on geometric coordinates, the positional relationship between objects in the image is accurately determined. Based on this positional relationship, high-quality training data is constructed, enabling the model to learn the orientation logic of "up, down, left, right" between objects, thus enabling it to accurately answer user queries about relative positions during reasoning.

[0246] In some embodiments of this application, the above-mentioned spatial positional relationship includes a three-dimensional spatial positional relationship; exemplarily, step B2 may include steps D1 and D2:

[0247] Step D1: The electronic device performs depth estimation processing on each second sample image through a depth estimation model to obtain the depth information of each second sample image.

[0248] Step D2: The electronic device determines the three-dimensional spatial relationship between at least two objects in each second sample image based on the depth information of each second sample image and the coordinate information of at least one object in each second sample image.

[0249] In some embodiments of this application, for determining the spatial relationship between objects, a depth estimation model MiDaS 3.1 is introduced to generate a depth map from the image and, combined with the spatial coordinate information of each target, infer the spatial relationship between different targets in three-dimensional space. Finally, based on the spatial relationship information between target objects, corresponding spatial relationship data pairs are constructed to train a multimodal model with spatial perception and spatial reasoning capabilities.

[0250] In some embodiments of this application, the depth estimation model described above is a pre-trained neural network model capable of predicting scene depth maps from monocular RGB images.

[0251] In some embodiments of this application, after the second sample image is input into the depth estimation model, the depth estimation model outputs a depth map with the same size as the input image. The value of each pixel in this depth map represents the relative distance of that point to the camera in three-dimensional space; a smaller value generally indicates that the point is closer to the camera, and a larger value indicates that the point is farther away.

[0252] In some embodiments of this application, the electronic device can perform 3D reasoning using the coordinate and depth information of objects in a second sample image. For example, firstly, based on the coordinate information of each object in the image, such as a bounding box, its pixel region in the image is determined. Then, on the depth map corresponding to the second sample image, the depth values ​​of all pixels within that pixel region are extracted. By calculating the average or median depth value of that region, the representative depth of the object is obtained. Finally, by comparing the representative depth values ​​of different objects, the relative positions of the objects can be determined: objects with smaller representative depth values ​​are determined to be closer to the camera, and objects with larger depth values ​​are determined to be farther from the camera.

[0253] It should be noted that the determination of the coordinate information of the objects in the second sample image can be found in the descriptions of steps C1 and C2 above, and will not be repeated here.

[0254] For example, an image contains a "chair" and a "table". After obtaining the two-dimensional coordinate information of both and generating a depth map of the image, the electronic device calculates the average depth within the coordinate bounding boxes of the "chair" and "table" on the depth map, respectively. Assume the average depth of the "chair" region is 0.3 and the average depth of the "table" region is 0.7. Since 0.3 < 0.7, it is determined that the "chair" is closer to the camera than the "table," meaning the three-dimensional spatial relationship is "the chair is in front of the table."

[0255] In this embodiment of the application, by utilizing the image depth estimation capability of the depth estimation model and based on the inference rules of front-back relationships, the front-back positional relationship between each object in the image is accurately determined. Based on this positional relationship, high-quality training data is constructed, enabling the model to learn the orientation logic of the spatial relative position of objects, thereby enabling it to output accurate spatial relative positions during inference.

[0256] The process of constructing the second training data described above will be illustrated by specific examples below.

[0257] For example, such as Figure 10 As shown, the process of constructing the second training data may include the following steps:

[0258] Step 11: The electronic device acquires the original image.

[0259] For example, the original image shows a water glass on a table with a person sitting behind it. Input this original image.

[0260] Step 12: The electronic device inputs the raw image into the depth estimation model MiDaS 3.1.

[0261] Step 13: The electronic device outputs a black-and-white depth map through the model.

[0262] Step 14: The electronic device inputs the raw image into the multimodal large model Qwen-VL-Plus.

[0263] Step 15: The electronic device outputs the objects in the image and their text descriptions through the model.

[0264] It should be noted that the objects in the image are the main subjects in the image.

[0265] For example, the objects in the image include: "cup", "table", and "person".

[0266] Step 16: The electronic device inputs the original image and the text description of the object into the visual positioning model GroundingDINO.

[0267] Step 17: The electronic device outputs the coordinate positions of each object through the visual positioning model.

[0268] For example: cup: (x1, y1, x2, y2); table: (x3, y3, x4, y4); person: (x5, y5, x6, y6).

[0269] Step 18: The electronic device analyzes and determines the positional relationship of each object in the two-dimensional image space based on its coordinate position.

[0270] Examples include top, bottom, left, and right relationships.

[0271] Step 19: The electronic device combines the output depth map and the coordinate information of each main target to infer the front-back position relationship of each object in three-dimensional space.

[0272] Step 20: The electronic device constructs spatial position relationship description data pairs based on the two-dimensional positional relationship and the front-back positional relationship.

[0273] Examples include: "the cup is above the table", "the person is behind the table", and "the cup is in front and the person is behind".

[0274] Step 21: The electronic device uses the above spatial location relationship description data pairs as the second training data.

[0275] In this embodiment, by constructing refined labeled data, the overall perception capability of the multimodal visual perception model in both screen and general scenarios is significantly improved. By explicitly embedding multi-subject coordinate information into the image description, the model can simultaneously learn the correspondence between semantics, objects, and positions during the training phase. This enhances the coverage and completeness of descriptions of multiple targets in complex scenes, effectively avoiding the problem of only being able to locate a single target or missing key subjects. Furthermore, the introduction of spatial perception data containing annotations of spatial positional relationships such as up / down, left / right, and front / back enables the model not only to accurately return the two-dimensional coordinates of targets but also to understand and infer the relative spatial orientation relationships between targets, significantly improving spatial perception and spatial reasoning capabilities. Thus, the model's accuracy in target localization, its ability to distinguish semantic references, and the reliability of spatial orientation determination are improved.

[0276] In some embodiments of this application, step 209 may include steps 209a to 209e:

[0277] Step 209a: The electronic device performs N inference samplings on each first sample image using the initial model to obtain N first prediction results.

[0278] Each first prediction result includes prediction description information for at least one predicted object in the first sample image.

[0279] Step 209b: Based on the above N first prediction results, the electronic device calculates N first reward values ​​through the first reward function.

[0280] Step 209c: The electronic device performs M inference samplings on each second sample image using the initial model to obtain M second prediction results.

[0281] Each second prediction result includes the predicted relative position information between at least two predicted objects in the second sample image.

[0282] Step 209d: Based on the above M second prediction results, the electronic device calculates M second reward values ​​through the second reward function.

[0283] Step 209e: Based on the first difference and the second difference, update the parameters of the initial model to obtain the multimodal visual perception model.

[0284] The first difference is the difference between N first reward values; the second difference is the difference between M second reward values.

[0285] The first reward function is used to evaluate the consistency between the predicted description information of the predicted object and the actual visual content of the predicted object. The predicted description information of the predicted object is obtained by inference sampling of the first sample image through the initial model. The second reward function is used to evaluate the matching degree between the predicted relative position information between at least two predicted objects and the actual relative position information between at least two predicted objects. The predicted relative position information between two predicted objects is obtained by inference sampling of the second sample image through the initial model. N is a positive integer and M is a positive integer.

[0286] In some embodiments of this application, for each first sample image in the first training data, N independent inference samplings are performed through the initial model to obtain N independent first prediction results, each first prediction result containing descriptive information generated by the model for that first sample image. For example, N and M can be 4.

[0287] It should be noted that due to the inherent randomness within the initial model, each sampling may produce slightly different outputs. The N first predictions obtained through N inferences constitute the distribution of the model's ability to describe the image under the current parameters.

[0288] In some embodiments of this application, after obtaining multiple first prediction results, the electronic device inputs each first prediction result into the first reward function of the model for calculation. This first reward function is used to evaluate the degree of consistency between the predicted descriptive information generated by the model and the real visual content, i.e., to suppress descriptive illusion. For example, the electronic device compares the predicted descriptive information of a first sample image with the image content of the first sample image. For instance, it uses the predicted descriptive information and the objective descriptive information generated by the DAM model for the first sample image as a reference, and calculates the consistency between the two through semantic similarity calculation or other methods. The higher the consistency, the higher the first reward value.

[0289] In some embodiments of this application, for each second sample image in the second training data, M independent inference samplings are performed using an initial model. Each second prediction result contains a set of predicted relative position information about multiple predicted objects in that image, as predicted by the model, for example, a set of data pairs including subject, relationship, and object. Exemplarily, M can be 4.

[0290] In some embodiments of this application, each second prediction result is input into a second reward function, which is used to evaluate the degree of matching between the relative position information between objects predicted by the model and the actual relative position information labeled. For example, the electronic device can directly compare the predicted relative position information with the actual labeled relative position information; the higher the degree of matching, the higher the second reward value.

[0291] In some embodiments of this application, the electronic device uses the differences between all the calculated first reward values ​​and the differences between all the second reward values ​​to backpropagate gradients through a policy gradient algorithm, and updates the parameters of the initial model so that the model parameters are adjusted in a direction that can more stably produce high reward value outputs. After multiple iterations, the model learns to generate accurate descriptive information and positional relationships.

[0292] For example, a first sample image is sampled four times, resulting in four descriptions. Two of these descriptions are correctly identified as "a blue suitcase," earning higher reward values; the other two are incorrectly identified as "a gray chair," earning lower reward values. The model can be adjusted to more accurate descriptions of attributes such as "blue" and "suitcase" based on higher reward values. A second sample image is sampled four times, resulting in four sets of predicted relative position information. One set of predicted relative position information, "the phone is on the table, and there is a vase to the right of the phone," has a reward value of 1, while the incorrect prediction, "the phone is to the right of the mouse pad," has a reward value of 2. Since reward value 1 is greater than reward value 2, the model can be adjusted to more accurate predictions of attributes such as "the phone is on the table" and "the vase is to the right of the phone" based on higher reward values, thereby enabling the model to learn correct spatial relationship reasoning abilities.

[0293] It should be noted that in the visual reasoning process of multimodal large models, especially in visual localization and image description tasks, the model is prone to the description illusion problem. That is, although the model provides a formally reasonable textual answer, its descriptive details are inconsistent with the actual visual content of the target area, or it makes incorrect inferences in scenarios such as spatial position relationships and multi-object differentiation. This problem is particularly prominent in edge-side models, and it is difficult to effectively constrain the above-mentioned illusionary behavior by relying solely on language consistency or localization accuracy indicators.

[0294] In this embodiment, during the model training phase, a DAM model is introduced as a constraint mechanism in the reinforcement learning process to generate a structured description of the target localization region. This description is then combined with a multimodal large model for consistency judgment and illusion detection. This effectively suppresses semantic illusions generated by the model that do not conform to visual reality while optimizing localization and description capabilities, thereby improving the accuracy and reliability of the output results. Specifically, a novel establishment function is constructed, introducing the model as a visual consistency constraint. Furthermore, a differentiated reward mechanism is designed for different data. In the reinforcement learning phase of the edge-side 3B multimodal visual model, the DAM model is introduced as a visual consistency constraint module to perform local visual description verification of the target region inferred by the large model, and the verification result is directly fed back to the reward function. When the description generated by the large model is inconsistent with the model's description of the corresponding visual region, the reward value is penalized, thereby explicitly suppressing the generation of descriptive illusions during training. Thus, by designing differentiated reward evaluation and penalty mechanisms based on the characteristics of different types of training data, the model's textual description capability and spatial perception capability are improved.

[0295] In some embodiments of this application, step 209b may include steps 209b1 to 209b4:

[0296] Step 209b1: Based on the prediction bounding box corresponding to each prediction object in the first prediction result, crop the image region corresponding to the prediction bounding box from the first sample image.

[0297] Step 209b2: The electronic device uses an image description model to perform image description on each image region, thereby obtaining reference description information for each predicted object.

[0298] Step 209b3: The electronic device performs a consistency comparison between the prediction description information of each prediction object in the first prediction result and the reference description information of each prediction object to obtain the consistency comparison result information.

[0299] Step 209b4: Based on the consistency comparison result information, the electronic device calculates the first reward value corresponding to the first prediction result through the first reward function.

[0300] The first reward value is used to characterize the degree of consistency between the predicted descriptive information generated by the initial model and the real visual content.

[0301] In some embodiments of this application, each first prediction result includes a text description of the first sample image, and also includes coordinate information corresponding to each predicted object in the first sample image, such as the coordinates of the bounding box of the predicted object. The electronic device uses the coordinate information of each predicted object to crop out a local image region corresponding to each predicted object from the first sample image. For example, the coordinate information corresponding to the predicted object can be box coordinate information.

[0302] In some embodiments of this application, an image description model (DAM) is used to describe each cropped image region to obtain reference description information for each predicted object. For example, each cropped local image region is input into the image description model. For each local image block, the DAM generates a reference description, which is an objective description based on the visual content of that region.

[0303] It should be noted that DAM is a model specifically trained to generate objective and detailed text descriptions for any given image region.

[0304] In some embodiments of this application, the predicted description information of each predicted object in the first prediction result of the initial model is compared with the reference description information generated by the DAM for the corresponding image region to determine the predicted description information that matches the reference description information. For example, a similarity score is obtained by calculating the semantic similarity between the two, thereby quantifying the degree of consistency between them.

[0305] In some embodiments of this application, the electronic device calculates a first reward value based on the determined number of predicted description information that matches the reference description information using a first reward function.

[0306] It should be noted that this first reward value represents the degree of consistency between the overall predicted descriptive information generated by the initial model and the real visual content.

[0307] The following is an illustrative explanation of the calculation of the first reward value in the embodiments of this application, using formulas.

[0308] For example, during the reinforcement learning training of the initial model, for a sample in the first training data, namely a first sample image and its annotation, after performing a forward inference sampling through the initial model, a first prediction result is output. This first prediction result contains object description information of multiple objects predicted by the model and the coordinates of their corresponding predicted bounding boxes, i.e., box coordinate information.

[0309] It should be noted that the purpose of calculating this first reward value is to evaluate the accuracy of the model's localization, i.e., whether the predicted bounding box matches the ground truth bounding box, and to evaluate the accuracy of its description, i.e. whether the predicted description is consistent with the actual image content.

[0310] For example, the first reward value R is jointly determined by the F1 score, a basic indicator reflecting positioning accuracy, and the illusion penalty term, which reflects the consistency of description. The formula for calculating the first reward value R is shown in formula (1):

[0311] (1)

[0312] F1 is a score calculated based on precision and recall in positioning. Precision is the ratio of precision to recall.

[0313] E represents the proportion of illusion errors in the descriptive information used in this prediction. Let the set of descriptive information involved in the current sample be: Here, each descriptive information rkr, krk represents the predicted descriptive information of an object, such as "blue suitcase" or "car key". The set of reference descriptive information obtained through the model is R*. The error of a single descriptive information is non-zero, i.e., 1. The illusion ratio E is defined as: Where K is the total number of samples, and rk indicates whether the k-th sample experiences hallucinations. This is an indicator function; `rk` has a value of 1 when a specific condition is met, and 0 otherwise. For example, if `rk` represents whether a hallucination exists, then... The value is 1 when hallucinations are present, and 0 otherwise. The normalized hallucination error ratio used when calculating hallucination punishment has a value range of [0,1].

[0314] It is the smoothing function of sigmoid, used to smooth the penalty term. Defined as: , where x is the input value, which can be any real number; specifically, by adding a hallucination penalty to the F1 algorithm, the sigmoid smoothing function in the above formula is used to smooth the penalty term, and the first reward value is calculated;

[0315] This represents the maximum impact weight of the hallucination penalty on the final reward, i.e., the maximum penalty magnitude. It is used to limit the upper bound of the penalty on the reward, preventing the reward from being over-compressed. The larger α is, the lower the hallucination tolerance. Experiments... It can be set to 0.5.

[0316] This is the hallucination tolerance threshold, for example, set to 0.2;

[0317] k is the penalty slope, and k can be set to 10.

[0318] It should be noted that, in order to avoid the instability of the training process caused by linear penalties, a non-linear smoothing mapping is introduced into the hallucination penalty term, so that the model maintains training robustness when there is slight semantic bias, while significantly reducing the reward value under severe semantic inconsistency, thereby effectively suppressing inference hallucinations.

[0319] The following is an exemplary description of the calculation of the first reward value in the embodiments of this application, using the formula.

[0320] For example, all bounding boxes predicted by the initial model are matched with the ground truth bounding boxes. For instance, a prediction is considered correct when the LU (Local Area) between the predicted box and a ground truth bounding box is greater than 0.5. Assuming image 1 contains 10 ground truth objects, and the initial model predicts the object descriptions of 8 objects and the coordinates of their corresponding 8 bounding boxes, with 6 of them correctly matching the ground truth boxes, then the precision P = 6 / 8 = 0.75, and the recall R = 6 / 10 = 0.6. Therefore, the F1 score is calculated as: F1 = 2(P*R) / (P + R) = 2(0.75*0.6) / (0.75 + 0.6), approximately 0.667. For the six bounding boxes correctly predicted by the model, six corresponding local regions are cropped from Image 1. These six regions are then input into the image description model to generate six reference descriptions based on visual facts. The predicted descriptions generated by the model for these six objects are semantically consistent with the six reference descriptions generated by the DAM (Digital Awareness Model). For example, a large language model such as Qwen-VL-Plus is used for the judgment. The hallucination ratio E is calculated. Assuming that the comparison finds that the predicted descriptions of three of the six objects have key details inconsistent with the reference descriptions, then the relation number Kerror = 3 for hallucination is generated. Substituting into the above hallucination ratio calculation formula, E = 3 / 6 = 0.5. With F1 = 0.667, E = 0.5, and preset parameters α = 0.5, k = 10, =0.2 Calculation formula (1) to obtain the final reward value R.

[0321] In this embodiment, by comparing the model's output with the model's objective description of the same visual region, errors in the model-generated description, such as color and shape errors, can be accurately located and quantified, enabling the model to learn the ability to accurately generate descriptive information. This improves the accuracy and reliability of the final model output description.

[0322] In some embodiments of this application, step 209d may include steps 209d1 and 209d2:

[0323] Step 209d1: The electronic device compares the predicted relative position information and the reference relative position information in each second prediction result to determine the first quantity.

[0324] The first quantity mentioned above refers to the number of predicted relative position information that matches the reference relative position information.

[0325] Step 209d2: The electronic device calculates the second reward value based on the first quantity and the second quantity using the second reward function.

[0326] The second quantity mentioned above refers to the total number of predicted relative position information; the second reward value mentioned above is used to characterize the degree of consistency between the predicted relative position information of the initial model and the actual relative position.

[0327] In some embodiments of this application, each second prediction result contains a set of predicted relative position information predicted by the model. Each set of predicted relative position information can be represented by a structured set, such as [{"object": "item A", "object": "item B" "relationship": "above"}, {"object": "person", "object": "furniture C" "relationship": "behind"}].

[0328] It should be noted that this reference relative position information comes from the manual or automatic annotation of the second sample image in the second training data, and can be understood as a set of real spatial positional relationships.

[0329] In some embodiments of this application, for each second sample image, the electronic device can match the predicted relative position information with the corresponding reference relative position information item by item, and count the number of items in the predicted relative position information that can be successfully matched with the reference relative position information, which is recorded as the number of matched predicted relative position information.

[0330] In some embodiments of this application, after calculating the number of matches, the ratio between the number of matched predicted relative position information and the total number of predicted relative position information is calculated to obtain the second reward value.

[0331] In this application embodiment, for spatial perception data involving spatial relationships, such as up / down, left / right, and front / back, to improve the model's spatial perception capability in screen visual perception, this application proposes a spatial positional relationship reward function based on structured annotation. Specifically, the spatial positional relationship data in each image is represented in a structured form, and each spatial positional relationship consists of "target pair + relationship type". For example: [

[0333] {"subject": "water cup", "object": "table", "relation": "on"},

[0334] {"subject": "person", "object": "table", "relation": "after"},

[0335] {"subject": "person", "object": "water cup", "relation": "after"} ];

[0337] The relationship types are limited to six directions: up, down, left, right, front, and back. Structured annotation allows reward value calculation to directly align the prediction with the ground truth (GT), eliminating the need for natural language text similarity evaluation, avoiding ambiguity, and improving accuracy. For each data point, the correctness of the spatial relationship can be directly determined. Then, the reward value for the spatial relationship is obtained by averaging all relationship data pairs within the same image.

[0338] For example, the formula for calculating the second reward value is shown in formula (2):

[0339] …… ……(2)

[0340] in, : Indicates the accuracy or score of spatial location relationship prediction, with a value between 0 and 1;

[0341] K: Represents the total number of spatial relationships that need to be evaluated, for example, 3;

[0342] This is an indicator function or scoring function used to determine whether the prediction of the i-th spatial relationship is correct. For example, if the prediction is correct, =1; if the prediction is incorrect, =0.

[0343] For example, combining the above formula (2), assume that the reference relative position information of a second sample image is: [

[0344] {"subject": "water cup", "object": "table", "relation": "on"},

[0345] {"subject": "person", "object": "table", "relation": "after"},

[0346] {"subject": "person", "object": "water cup", "relation": "after"}

[0347] The second prediction result output by the model after one sampling is:

[0348] {"subject": "water cup", "object": "table", "relation": "on"},

[0349] {"subject": "person", "object": "table", "relation": "front"},

[0350] {"subject": "person", "object": "water cup", "relation": "after"}

[0351] From the given information, we can see that the ground truth (GT) is: "the glass is on the table," "the person is behind the table," and "the person is behind the glass." The predicted pred is: "the glass is on the table," "the person is in front of the table," and "the person is behind the glass." The first item, "the glass is on the table," is a perfect match. The second item, "the person is in front of the table," does not match the reference "the person is behind the table." The third item is also a perfect match. Therefore, the number of matches is 2, and the total number of predictions is 3. =2 / 3, approximately 0.667, meaning the reward value is 0.667.

[0352] In this embodiment, during the reinforcement learning phase, a policy gradient-based optimization method is employed to sample the same input sample multiple times, generating different inference results. Each sample independently calculates its corresponding reward value. By comparing the results of multiple samplings and using the differences in reward values ​​to guide model parameter updates, inference results with high consistency and low illusion are given higher weights, thereby gradually improving the stability and reliability of the model's output during training. Through this reinforcement learning mechanism combining multiple sampling and constraints, the model can effectively suppress illusion problems in visual description and spatial reasoning while maintaining its localization capabilities.

[0353] In some embodiments of this application, in order to train the initial model to accurately understand and respond to diverse user natural language queries, the reinforcement learning training process also needs to utilize third training data. For such data, the training process introduces a third reward function specifically designed to evaluate and optimize the model's semantic understanding and target referencing ability of text instructions.

[0354] For example, the above training process further includes the following steps: L inference samplings are performed on each sample in the third training data (i.e., the third sample image and its corresponding first dialogue information) using the initial model to obtain L third prediction results. Each third prediction result contains the target referential information output by the model after understanding the corresponding third sample image based on the input dialogue instruction. This target referential information is the model's answer to "What is the target referred to by the user instruction?", and its form can be the object category of the target, such as "water cup," or a description with attributes, such as "blue suitcase." For example, the electronic device calculates L third reward values ​​based on the L third prediction results using a third reward function. This third reward function is used to guide the model to accurately map the semantics of the text instruction to the correct object and enable it to learn the ability to distinguish detailed attributes.

[0355] For example, the calculation process of the third reward function can be as follows: First, extract the target referential text output by the model in the third prediction result. Then, compare the predicted text with the real dialogue instructions labeled in the third training data as input, for example, by using a text semantic encoder to calculate the semantic similarity score between the two. Finally, map the semantic similarity score to the corresponding reward value, i.e., the third reward value.

[0356] For example, for multi-object detection data (where multiple objects correspond to multiple boxes) and referential detection data (where text referencing corresponds to a single object box), the large model first outputs the object category or referential result. Then, a text semantic encoder, such as a reranker, is used to calculate semantic similarity between the model's predicted target text and the labeled text. The similarity result is directly mapped to a reward score. If there is a semantic mismatch or incorrect referencing, the reward is penalized. This approach effectively constrains the model's accuracy in object category recognition and referential understanding.

[0357] In some embodiments of this application, the electronic device updates the parameters of the initial model through a reinforcement learning optimization algorithm based on the differences between the N first reward values, the differences between the M second reward values, and the differences between the L third reward values, and finally obtains a multimodal visual perception model with fully optimized performance.

[0358] In this embodiment, a first reward function is introduced to suppress descriptive illusions, and a second reward function is introduced to improve the accuracy of spatial location relationships. Specifically, a third reward function is used to enhance instruction understanding and fine-grained referencing capabilities, enabling the model to learn how to accurately parse and understand constraints such as attribute descriptions in the user's natural language, accurately match the prediction of instruction details, and thus reliably perform complex queries. Training the model using the first reward function ensures that the descriptions generated by the model are consistent with the actual visual content, enabling the model to output image description information more accurately. Training the model using the second reward function ensures that the spatial location relationships between objects predicted by the model conform to the actual spatial location relationships, thereby improving the accuracy and reliability of spatial location prediction.

[0359] The following specific examples illustrate the flow of the object query method provided in the embodiments of this application.

[0360] For example, such as Figure 11 As shown, the process of this object query method may include the following steps:

[0361] Step 301: The electronic device constructs the first training data, the second training data, and the third training data.

[0362] In this step, it is necessary to process the model data for visual perception. In visually impaired scenarios, the model needs to have spatial perception capabilities, that is, it needs to know the specific location information of the target object [x1, y1, x2, y2], as well as the relative positional relationship, for example, "the phone" is to the right of the "water cup" or "the phone" is on top of the "table".

[0363] It should be noted that, compared to existing visual perception work, this application adds two types of self-developed labeled data, which can enhance the model's fine-grained localization capability and accurately answer spatial orientation information. Furthermore, it introduces new data labeling methods, building upon traditional multi-target localization data and referential localization data, by introducing two new training data formats to characterize the spatial relationships and semantic constraints between targets. Through joint training with multiple types of data, the model can not only achieve precise target localization but also perform fine-grained matching of complex semantic descriptions and accurately answer spatial orientation information related to targets, thereby significantly improving the expressive power and application scope of visual localization tasks.

[0364] The first type of training data added in this application involves inserting the corresponding coordinate information of each subject object into the original image description. The specific data labeling method is as follows: First, the Qwen3-VL-Plus model is used to analyze the original description and extract the various subject objects involved in the description; then, a self-developed visual localization model is used to process the image and label the precise coordinate information of each subject; finally, the coordinate information of each subject is embedded into the original caption, forming a subject-coordinate correspondence, thereby constructing training samples that support fine-grained localization and spatial relationship understanding. This method allows the model to simultaneously acquire semantic and coordinate information during training, enhancing the accurate recognition and visual reasoning capabilities of targets in complex scenes.

[0365] The second type of data added in this application is spatial perception data, which is mainly used to characterize the spatial orientation relationship between target objects, specifically including six directions: up, down, left, right, front, and back.

[0366] For example, in the data construction process, the input image is first parsed using the multimodal large model Qwen-VL-Plus to extract multiple main targets and their corresponding text descriptions. Then, the text information of the main targets is input into the visual localization model GroundingDINO to obtain the corresponding coordinate positions of each main target in the image. After obtaining the coordinate information, the two-dimensional spatial positional relationship between the target objects can be determined by comparing and analyzing the coordinate values ​​of different targets, including relative orientations such as up, down, left, and right. For determining the front-back spatial positional relationship, this invention further introduces the depth estimation model MiDaS 3.1 to generate a depth map of the image and, combined with the spatial coordinate information of each target, infers the front-back relationship of different targets in three-dimensional space. Finally, based on the spatial positional relationship information between the target objects, corresponding spatial positional relationship data pairs are constructed to train a multimodal model with spatial perception and spatial reasoning capabilities.

[0367] With the combined effect of the two types of data, the model has been systematically improved in terms of the accuracy of target localization, the ability to distinguish semantic references, and the reliability of spatial orientation expression, thus better supporting complex query and assisted perception application scenarios.

[0368] Step 302: Train the initial model to obtain the multimodal visual perception model.

[0369] The purpose of this invention is to improve spatial perception and reduce visual illusions. It constructs a novel establishment function by introducing a model as a visual consistency constraint. Furthermore, it designs a differentiated reward mechanism for different types of data. In the reinforcement learning stage of the edge-side 3B multimodal visual model, this application introduces a model DAM as a visual consistency constraint module to perform local visual description verification on the target region obtained by the large model inference, and directly feeds the verification result back to the reward function. When the description generated by the large model is inconsistent with the model's description of the corresponding visual region, the reward metric is penalized, thereby explicitly suppressing the generation of descriptive illusions during training. This reinforcement learning scheme designs differentiated reward evaluation and penalty mechanisms based on the characteristics of different types of training data.

[0370] The design and processing methods for reward functions for different data types are as follows:

[0371] First: Design of reward function for multi-target localization data and referential detection data;

[0372] For multi-object detection data (multiple objects corresponding to multiple boxes) and referential detection data (textual referencing corresponding to a single object box): First, the large model outputs the object category or referential result; then, a text semantic encoder (reranker) calculates the semantic similarity between the model's predicted target text and the labeled text; the similarity result is directly mapped to a reward score; if there is a semantic mismatch or referential error, the reward is penalized. This method effectively constrains the model's accuracy in object category recognition and referential understanding. The specific calculation method is expressed by the following formula.

[0373] Second: Description-Location Corresponding Data Reward Function Design (Fine-grained Description Constraints);

[0374] For data whose descriptions explicitly include box coordinate information: Based on the box region predicted by the large model, the corresponding local region is cropped from the original image; the cropped region is input into the DAM model to generate a visual description of the local region; the description generated by the DAM and the description output by the large model are compared using the qwen-vl-plus text large model for consistency judgment; if the description has illusions or inconsistencies in key details, a further penalty is applied on top of the precision / recall (PR) metric; the final reward for the sample is calculated comprehensively. This mechanism can effectively constrain the model's ability in terms of fine-grained visual description and localization consistency, significantly reducing the problem of "correct localization but incorrect description".

[0375] Third: Design of reward function for spatial perception data (relational reasoning constraints);

[0376] For spatial perception data involving spatial relationships (up / down, left / right, front / back), to improve the model's spatial perception capability in screen visual perception, this invention proposes a spatial relationship reward function based on structured annotation. Specifically, the spatial relationship data in each image is represented in a structured form, and each spatial relationship consists of "target pair + relationship type". For example: [

[0378] {"subject": "water cup", "object": "table", "relation": "on"},

[0379] {"subject": "person", "object": "table", "relation": "after"},

[0380] {"subject": "person", "object": "water cup", "relation": "after"} ]

[0382] The relationship types are limited to six directions: up, down, left, right, front, and back. Structured annotation allows reward calculation to directly align the prediction with the ground truth (GT), eliminating the need for natural language text similarity evaluation, avoiding ambiguity, and improving accuracy. For each data point, the correctness of the spatial relationship can be directly determined. Then, the sample average is taken for all relationship data pairs within the same graph to obtain the reward result for the spatial relationship.

[0383] Step 303: Deploy the multimodal visual perception model on the electronic device.

[0384] For example, this application uses an on-device 3B model, which eliminates the need to upload data to the cloud, thus effectively protecting user privacy. The model is directly installed on the user's phone, eliminating the need for the user to download it again after purchasing the phone. When the user activates accessibility features, the system should request permissions from the user, prompting them to enable visual perception of the offline large model for camera and screen content.

[0385] Step 304: The electronic device performs multimodal reasoning on the first dialogue information and the preview image through a multimodal visual perception model to obtain prediction information.

[0386] For example, based on visual reasoning-based location information perception and assisted interaction, the electronic device invokes a trained 3B multimodal visual perception model to achieve semantic-level understanding and location information perception of the screen content. This function is specifically designed to meet the actual needs of visually impaired users when using mobile devices. When the user uses this function, real-time video stream information is displayed on the screen, a voice conversation can be initiated by clicking the icon in the lower left corner, a photo can be taken by clicking the circular button in the middle, and the front and rear cameras can be switched by clicking the button in the lower right corner.

[0387] In practical use, users can input target query commands via voice through the edge device, such as "Where is the red suitcase?". The edge device first converts the voice signal into a text command and inputs this command along with the current screen image into the pre-trained 3b edge multimodal visual perception model. After receiving the image and text command, the model performs multi-layer semantic understanding and spatial analysis of the screen content based on a multimodal feature alignment mechanism. Specifically, the model extracts target region features from the image through a visual encoding module and parses target attribute information from the user command through a language encoding module, including target category, color, appearance features, and spatial constraints. Subsequently, the model matches the target attributes with the image region in the joint feature space to accurately locate the target object. Because the model incorporates descriptive data containing positioning information and spatial positional relationship annotation data during the training phase, and strengthens visual consistency through model constraints, its inference results can not only distinguish multiple targets with highly similar appearances in the screen, such as "red stockings" and "black socks," but also accurately determine the spatial positional relationship and relative orientation between targets.

[0388] For example, while generating visual understanding results, the electronic device can simultaneously return the location information of the target object to the user in natural language. This location information includes not only a relative orientation description based on the user's perspective, such as "your phone is located to your right front," but also a description of the relative spatial relationships within the screen, such as "located on the table to the right of the stool, in front of the green plants and the lamp." By integrating the user's relative coordinate description with the description of the spatial relationships of multiple targets within the screen, the system can express the spatial location of the target object in a multi-layered manner, significantly improving the accuracy and understandability of the location feedback.

[0389] Step 305: The electronic device plays the first information through a speaker based on the predicted information.

[0390] For example, the interaction process for the aforementioned target-assisted query includes: the user clicks the function button in the lower left corner of the screen to enter voice interaction mode; the system enters a state awaiting inquiry and displays the prompt "Ask me a question" on the interface; in this step, the user submits a target query command via voice, such as "Help me find my phone?"; after receiving the voice command, the system performs multimodal understanding and spatial analysis on the current screen image and generates query results including the target location and spatial relationships; finally, the system broadcasts detailed spatial location information of the target object to the user via voice, such as "The phone is on the table in front of you, to the right of the water glass, next to a lamp and a vase." During the voice broadcast, the interface displays a "Click to interrupt" prompt, which the user can click to interrupt the voice feedback at any time, thereby improving the flexibility and controllability of the interaction. While outputting the verbal description, the system visually highlights the corresponding area of ​​the target object on the screen, making the voice feedback and visual results form consistent spatial guidance, thereby enhancing the user's overall perception of the scene structure. This multimodal feedback method can effectively reduce the ambiguity of understanding caused by a single voice or single visual cues, and is especially suitable for the spatial cognition needs of visually impaired users for complex screen content.

[0391] Through the above-mentioned interaction mechanism, visually impaired users can query and locate target objects on the screen without relying on visual operations, and can independently complete daily information acquisition and interaction operations, significantly improving the practicality and user experience of smart terminals in accessible use scenarios.

[0392] In this embodiment, based on the enhanced 3B visual perception model, the electronic device possesses semantic reasoning and spatial perception capabilities, enabling it to understand natural language questions and return precise location information. For example, when a user asks "Where are my keys?", the system can accurately distinguish different key categories and output their relative positions to the user, as well as the relative positions of objects in the image. It will also highlight specific areas in the image to help visually impaired users gain a realistic spatial perception experience, greatly improving the practicality and intelligence of the interaction.

[0393] It should be noted that the current solution can be used not only for screen recognition and assistive scenarios for visually impaired people, but also for intelligent composition scenarios in cameras and photo albums.

[0394] It should be noted that each of the above method embodiments, or various possible implementations of each method embodiment, can be executed individually or in combination of any two or more. The specific implementation can be determined according to actual usage requirements, and this application embodiment does not impose any restrictions on this.

[0395] The object query method provided in this application can be executed by an object query device. This application uses an object query device executing the object query method as an example to illustrate the object query device provided in this application.

[0396] Figure 12 This is a schematic diagram of the structure of an object query device provided in an embodiment of this application. Figure 12 As shown, the object query device includes: a processing module 601, used to display first dialogue information on a first interface, the first dialogue information being used to indicate the search for a first object, the first interface including a preview image captured by a camera; the processing module 601 is also used to perform object search based on the first dialogue information and the preview image, and output first information; wherein, the first information includes at least one of the following: the relative position information of the first object relative to the user; the relative position information between the first object and a second object in the preview image.

[0397] In some embodiments of this application, the first interface includes a voice dialogue control;

[0398] The receiving module is used to receive the first input to the voice dialogue control;

[0399] The processing module is specifically used to respond to the first input received by the receiving module and enable the voice interaction mode.

[0400] The display module is used to display the first dialogue information captured by the microphone on the first interface.

[0401] In some embodiments of this application, the display module is configured to display a first identifier on the first interface;

[0402] The first identifier is used to mark the display area of ​​the first object in the preview image.

[0403] In some embodiments of this application, the processing module is specifically used for:

[0404] The first dialogue information and the preview image are input into the multimodal visual perception model. The multimodal visual perception model performs multimodal reasoning on the first dialogue information and the preview image to obtain prediction information. The prediction information includes at least one of the following: the coordinate information of the first object and the relative position information between the first object and at least one reference object in the preview image.

[0405] Based on the predicted information, the first piece of information is output.

[0406] In some embodiments of this application, the multimodal visual perception model includes a language encoding module, a visual encoding module, and a cross-modal processing module; the processing module is specifically used for:

[0407] The language encoding module performs text encoding processing on the first dialogue information to obtain text feature information;

[0408] The preview image is processed by the visual encoding module to obtain image feature information;

[0409] The cross-modal processing module fuses and infers textual and image features to obtain prediction information.

[0410] In some embodiments of this application, the processing module is further configured to:

[0411] The preview image is input into the multimodal visual perception model. The multimodal visual perception model is used to locate objects in the preview image and obtain the coordinate information of the first object.

[0412] Based on the coordinate information, if it is determined that the first object is located in the edge region of the preview image, the second information is output;

[0413] The second information is used to prompt the user to move the electronic device in the first direction so that the first object is located in the center area of ​​the preview image.

[0414] In some embodiments of this application, the processing module is further configured to acquire first training data, second training data, and third training data before performing multimodal reasoning on the first dialogue information and the preview image through a multimodal visual perception model;

[0415] The processing module is also used to train the initial model based on the first training data, the second training data and the third training data to obtain a multimodal visual perception model.

[0416] The first training data includes at least one first sample image and first annotation information corresponding to each first sample image. The first annotation information corresponding to each first sample image includes object description information and coordinate information of at least one object in the first sample image. The second training data includes at least one second sample image and second annotation information corresponding to each second sample image. The second annotation information corresponding to each second sample image includes relative position information between any two objects in the second sample image. The third training data includes at least one third sample image and first dialogue information corresponding to each third sample image.

[0417] In some embodiments of this application, the processing module is specifically used for:

[0418] Obtain at least one first sample image and corresponding descriptive information for each first sample image; the descriptive information includes object description information for at least one object in the first sample image.

[0419] Semantic parsing is performed on the description information corresponding to each first sample image to determine at least one object in each first sample image;

[0420] By using a visual positioning model, visual positioning is performed on each first sample image to determine the coordinate information of at least one object in each first sample image.

[0421] Add the coordinate information of each object in at least one object to the first text position corresponding to each object to obtain the first training data; the first text position is: the text position where the object description information of the object is located in the first description information, and the first description information is the description information corresponding to the first sample image.

[0422] In some embodiments of this application, the processing module is specifically used for:

[0423] Obtain at least one second sample image;

[0424] Determine the spatial relationship between at least two objects in each second sample image, including two-dimensional and three-dimensional spatial relationships;

[0425] Second training data is constructed based on the spatial relationship between at least two objects in each second sample image.

[0426] In some embodiments of this application, the spatial positional relationship includes a two-dimensional spatial positional relationship; the processing module is specifically used for:

[0427] By performing content analysis on each second sample image using a multimodal large model, descriptive information corresponding to each object in each second sample image is obtained;

[0428] The description information is visually located using a visual positioning model to obtain the coordinate information of each object in the second sample image.

[0429] Based on the coordinate information of at least one object in each second sample image, determine the two-dimensional spatial positional relationship between at least two adjacent objects in each second sample image.

[0430] In some embodiments of this application, spatial positional relationships include three-dimensional spatial positional relationships; the processing module is specifically used for:

[0431] The depth estimation model is used to perform depth estimation processing on each second sample image to obtain the depth information of each second sample image;

[0432] Based on the depth information of each second sample image and the coordinate information of at least one object in each second sample image, the three-dimensional spatial positional relationship between at least two objects in each second sample image is determined.

[0433] In some embodiments of this application, the processing module is specifically used for:

[0434] Using the initial model, N inference samplings are performed on each first sample image to obtain N first prediction results. Each first prediction result includes prediction description information of at least one predicted object in the first sample image.

[0435] Based on N first prediction results, N first reward values ​​are calculated using the first reward function;

[0436] Using the initial model, M inference samplings are performed on each second sample image to obtain M second prediction results; each second prediction result includes the predicted relative position information between at least two predicted objects in the second sample image.

[0437] Based on the M second prediction results, M second reward values ​​are calculated using the second reward function;

[0438] Based on the first and second differences, the parameters of the initial model are updated to obtain the multimodal visual perception model; the first difference is the difference between N first reward values; the second difference is the difference between M second reward values;

[0439] The first reward function is used to evaluate the consistency between the predicted description information of the predicted object and the actual visual content of the predicted object. The predicted description information of the predicted object is obtained by inference sampling of the first sample image through the initial model. The second reward function is used to evaluate the matching degree between the predicted relative position information between at least two predicted objects and the actual relative position information between at least two predicted objects. The predicted relative position information between two predicted objects is obtained by inference sampling of the second sample image through the initial model. N is a positive integer and M is a positive integer.

[0440] In some embodiments of this application, the processing module is specifically used for:

[0441] Based on the predicted bounding box corresponding to each predicted object in each first prediction result, the image region corresponding to the predicted bounding box is cropped from the first sample image;

[0442] By using an image description model, each image region is described to obtain reference description information for each predicted object;

[0443] The consistency comparison is performed between the prediction description information of each prediction object in the first prediction result and the reference description information of each prediction object to obtain the consistency comparison result information.

[0444] Based on the consistency comparison results, the first reward value corresponding to the first prediction result is calculated using the first reward function.

[0445] The first reward value is used to characterize the degree of consistency between the predicted descriptive information generated by the initial model and the real visual content.

[0446] In some embodiments of this application, the processing module is specifically used for:

[0447] The predicted relative position information in each second prediction result is compared with the reference relative position information to determine the first quantity; the first quantity is the number of predicted relative position information that matches the reference relative position information.

[0448] Based on the first quantity and the second quantity, a second reward value is calculated through a second reward function; wherein, the second quantity is the total quantity of predicted relative position information; and the second reward value is used to characterize the degree of consistency between the predicted relative position information of the initial model and the actual relative position.

[0449] The object query device provided in this application embodiment displays first dialogue information on a first interface, which is used to indicate the search for a first object. The first interface includes a preview image captured by a camera. Based on the first dialogue information and the preview image, the device performs an object search and outputs first information. The first information includes at least one of the following: the relative position information of the first object relative to the user; and the relative position information between the first object and a second object in the preview image. Through this scheme, the object query device can output information including the relative position information between the first object and the user, or the relative position information between the first object and other objects, based on the first dialogue information and the preview image input by the user. This allows the user to quickly and accurately determine the location of the first object based on its relative position to itself, or based on its relative position to other known objects, thus greatly improving the accuracy and intelligence of environmental perception.

[0450] The object query device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0451] The object query device in this application embodiment can be a device with an operating system. The operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.

[0452] The object query device provided in this application embodiment can implement the various processes implemented in the object query method embodiment, and will not be described again here to avoid repetition.

[0453] Optionally, such as Figure 13 As shown, this application embodiment also provides an electronic device 700, including a processor 701 and a memory 702. The memory 702 stores a program or instructions that can run on the processor 701. When the program or instructions are executed by the processor 701, they implement the various steps of the above-described object query method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0454] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0455] Figure 14 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0456] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, and processor 110.

[0457] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 14 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0458] The processor 110 is configured to display first dialogue information on a first interface, the first dialogue information being used to indicate the search for a first object, the first interface including a preview image captured by a camera; the processor 110 is also configured to perform object search based on the first dialogue information and the preview image, and output first information; wherein the first information includes at least one of the following: the relative position information of the first object relative to the user; the relative position information between the first object and a second object in the preview image.

[0459] In some embodiments of this application, the first interface includes a voice dialogue control;

[0460] User input unit 107 is used to receive the first input to the voice dialogue control;

[0461] The processor is specifically configured to respond to the first input received by the user input unit 107 and enable the voice interaction mode;

[0462] Display unit 106 is used to display the first dialogue information collected by the microphone on the first interface.

[0463] In some embodiments of this application, the display unit 106 is configured to display a first identifier on the first interface;

[0464] The first identifier is used to mark the display area of ​​the first object in the preview image.

[0465] In some embodiments of this application, the processor is specifically used for:

[0466] The first dialogue information and the preview image are input into the multimodal visual perception model. The multimodal visual perception model performs multimodal reasoning on the first dialogue information and the preview image to obtain prediction information. The prediction information includes at least one of the following: the coordinate information of the first object and the relative position information between the first object and at least one reference object in the preview image.

[0467] Based on the predicted information, the first piece of information is output.

[0468] In some embodiments of this application, the multimodal visual perception model includes a language encoding module, a visual encoding module, and a cross-modal processor; the processor is specifically used for:

[0469] The language encoding module performs text encoding processing on the first dialogue information to obtain text feature information;

[0470] The preview image is processed by the visual encoding module to obtain image feature information;

[0471] By using a cross-modal processor, textual and image feature information are fused and inferred to obtain prediction information.

[0472] In some embodiments of this application, the processor is further configured to:

[0473] The preview image is input into the multimodal visual perception model. The multimodal visual perception model is used to locate objects in the preview image and obtain the coordinate information of the first object.

[0474] Based on the coordinate information, if it is determined that the first object is located in the edge region of the preview image, the second information is output;

[0475] The second information is used to prompt the user to move the electronic device in the first direction so that the first object is located in the center area of ​​the preview image.

[0476] In some embodiments of this application, the processor is further configured to acquire first training data, second training data, and third training data before performing multimodal inference on the first dialogue information and the preview image through a multimodal visual perception model;

[0477] The processor is also used to train the initial model based on the first training data, the second training data, and the third training data to obtain a multimodal visual perception model.

[0478] The first training data includes at least one first sample image and first annotation information corresponding to each first sample image. The first annotation information corresponding to each first sample image includes object description information and coordinate information of at least one object in the first sample image. The second training data includes at least one second sample image and second annotation information corresponding to each second sample image. The second annotation information corresponding to each second sample image includes relative position information between any two objects in the second sample image. The third training data includes at least one third sample image and first dialogue information corresponding to each third sample image.

[0479] In some embodiments of this application, the processor is specifically used for:

[0480] Obtain at least one first sample image and corresponding descriptive information for each first sample image; the descriptive information includes object description information for at least one object in the first sample image.

[0481] Semantic parsing is performed on the description information corresponding to each first sample image to determine at least one object in each first sample image;

[0482] By using a visual positioning model, visual positioning is performed on each first sample image to determine the coordinate information of at least one object in each first sample image.

[0483] Add the coordinate information of each object in at least one object to the first text position corresponding to each object to obtain the first training data; the first text position is: the text position where the object description information of the object is located in the first description information, and the first description information is the description information corresponding to the first sample image.

[0484] In some embodiments of this application, the processor is specifically used for:

[0485] Obtain at least one second sample image;

[0486] Determine the spatial relationship between at least two objects in each second sample image, including two-dimensional and three-dimensional spatial relationships;

[0487] Second training data is constructed based on the spatial relationship between at least two objects in each second sample image.

[0488] In some embodiments of this application, the spatial positional relationship includes a two-dimensional spatial positional relationship; the processor is specifically used for:

[0489] By performing content analysis on each second sample image using a multimodal large model, descriptive information corresponding to each object in each second sample image is obtained;

[0490] The description information is visually located using a visual positioning model to obtain the coordinate information of each object in the second sample image.

[0491] Based on the coordinate information of at least one object in each second sample image, determine the two-dimensional spatial positional relationship between at least two adjacent objects in each second sample image.

[0492] In some embodiments of this application, spatial positional relationships include three-dimensional spatial positional relationships; the processor is specifically used for:

[0493] The depth estimation model is used to perform depth estimation processing on each second sample image to obtain the depth information of each second sample image;

[0494] Based on the depth information of each second sample image and the coordinate information of at least one object in each second sample image, the three-dimensional spatial positional relationship between at least two objects in each second sample image is determined.

[0495] In some embodiments of this application, the processor is specifically used for:

[0496] Using the initial model, N inference samplings are performed on each first sample image to obtain N first prediction results. Each first prediction result includes prediction description information of at least one predicted object in the first sample image.

[0497] Based on N first prediction results, N first reward values ​​are calculated using the first reward function;

[0498] Using the initial model, M inference samplings are performed on each second sample image to obtain M second prediction results; each second prediction result includes the predicted relative position information between at least two predicted objects in the second sample image.

[0499] Based on the M second prediction results, M second reward values ​​are calculated using the second reward function;

[0500] Based on the first and second differences, the parameters of the initial model are updated to obtain the multimodal visual perception model; the first difference is the difference between N first reward values; the second difference is the difference between M second reward values;

[0501] The first reward function is used to evaluate the consistency between the predicted description information of the predicted object and the actual visual content of the predicted object. The predicted description information of the predicted object is obtained by inference sampling of the first sample image through the initial model. The second reward function is used to evaluate the matching degree between the predicted relative position information between at least two predicted objects and the actual relative position information between at least two predicted objects. The predicted relative position information between two predicted objects is obtained by inference sampling of the second sample image through the initial model. N is a positive integer and M is a positive integer.

[0502] In some embodiments of this application, the processor is specifically used for:

[0503] Based on the predicted bounding box corresponding to each predicted object in each first prediction result, the image region corresponding to the predicted bounding box is cropped from the first sample image;

[0504] By using an image description model, each image region is described to obtain reference description information for each predicted object;

[0505] The consistency comparison is performed between the prediction description information of each prediction object in the first prediction result and the reference description information of each prediction object to obtain the consistency comparison result information.

[0506] Based on the consistency comparison results, the first reward value corresponding to the first prediction result is calculated using the first reward function.

[0507] The first reward value is used to characterize the degree of consistency between the predicted descriptive information generated by the initial model and the real visual content.

[0508] In some embodiments of this application, the processor is specifically used for:

[0509] The predicted relative position information in each second prediction result is compared with the reference relative position information to determine the first quantity; the first quantity is the number of predicted relative position information that matches the reference relative position information.

[0510] Based on the first quantity and the second quantity, a second reward value is calculated through a second reward function; wherein, the second quantity is the total quantity of predicted relative position information; and the second reward value is used to characterize the degree of consistency between the predicted relative position information of the initial model and the actual relative position.

[0511] The electronic device provided in this application embodiment displays first dialogue information on a first interface, which is used to indicate the search for a first object. The first interface includes a preview image captured by a camera. Based on the first dialogue information and the preview image, the device performs object search and outputs first information. The first information includes at least one of the following: the relative position information of the first object relative to the user; and the relative position information between the first object and a second object in the preview image. Through this scheme, the electronic device can output information including the relative position information between the first object and the user or the relative position information between the first object and other objects based on the first dialogue information and the preview image input by the user. This allows the user to quickly and accurately determine the location of the first object based on its relative position to itself or its relative position to other known objects, thus greatly improving the accuracy and intelligence of environmental perception.

[0512] It should be understood that, in this embodiment, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0513] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 109 may include volatile memory or non-volatile memory, or it may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0514] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.

[0515] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described object query method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0516] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0517] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described object query method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0518] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0519] This application provides a computer program product that is stored in a storage medium and executed by at least one processor to implement the various processes of the object query method embodiment described above, and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0520] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0521] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0522] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An object query method, characterized in that, The method includes: On the first interface, first dialogue information is displayed, which is used to indicate the search for a first object. The first interface includes a preview image captured by a camera. Based on the first dialogue information and the preview image, an object search is performed, and the first information is output. The first information includes at least one of the following: The relative position information of the first object with respect to the user; The relative position information between the first object and the second object in the preview image.

2. The method according to claim 1, characterized in that, The first interface also includes a voice dialogue control; The first interface displays first dialogue information, including: Receive the first input to the voice dialogue control; In response to the first input, the voice interaction mode is activated, and the first dialogue information captured by the microphone is displayed on the first interface.

3. The method according to claim 1, characterized in that, After performing object lookup based on the first dialogue information and the preview image, the method further includes: The first identifier is displayed on the first interface; The first identifier is used to mark the display area of ​​the first object in the preview image.

4. The method according to claim 1, characterized in that, The step of performing object lookup based on the first dialogue information and the preview image, and outputting first information, includes: The first dialogue information and the preview image are input into a multimodal visual perception model. The multimodal visual perception model is used to perform multimodal reasoning on the first dialogue information and the preview image to obtain prediction information. The prediction information includes at least one of the following: coordinate information of the first object and relative position information between the first object and at least one reference object in the preview image. Based on the predicted information, the first information is output.

5. The method according to claim 4, characterized in that, The multimodal visual perception model includes a language encoding module, a visual encoding module, and a cross-modal processing module; the step of performing multimodal inference on the first dialogue information and the preview image through the multimodal visual perception model to obtain prediction information includes: The language encoding module performs text encoding processing on the first dialogue information to obtain text feature information; The visual encoding module performs image encoding processing on the preview image to obtain image feature information; The cross-modal processing module fuses and infers the text feature information and the image feature information to obtain prediction information.

6. The method according to claim 1, characterized in that, Before performing object lookup processing based on the first dialogue information and the preview image, and outputting the first information, the method further includes: The preview image is input into a multimodal visual perception model, and the object is located in the preview image through the multimodal visual perception model to obtain the coordinate information of the first object; If, based on the coordinate information, it is determined that the first object is located in the edge region of the preview image, the second information is output; The second information is used to prompt the user to move the electronic device so that the first object is located in the center area of ​​the preview image.

7. The method according to claim 4, characterized in that, Before performing multimodal reasoning on the first dialogue information and the preview image using the multimodal visual perception model, the method further includes: Acquire the first training data, the second training data, and the third training data; Based on the first training data, the second training data, and the third training data, the initial model is trained to obtain a multimodal visual perception model. The first training data includes at least one first sample image and first annotation information corresponding to each first sample image; the first annotation information corresponding to each first sample image includes object description information and coordinate information of at least one object in the first sample image; the second training data includes at least one second sample image and second annotation information corresponding to each second sample image; the second annotation information corresponding to each second sample image includes relative position information between any two objects in the second sample image; the third training data includes at least one third sample image and first dialogue information corresponding to each third sample image.

8. The method according to claim 7, characterized in that, The acquisition of the first training data includes: Obtain at least one first sample image and corresponding descriptive information for each first sample image; the descriptive information includes object description information of at least one object in the first sample image. Semantic parsing is performed on the description information corresponding to each first sample image to determine at least one object in each first sample image; Using a visual positioning model, visual positioning is performed on each first sample image to determine the coordinate information of at least one object in each first sample image. The coordinate information of each object in the at least one object is added to the first text position corresponding to each object to obtain the first training data; the first text position is: the text position where the object description information of the object is located in the first description information, and the first description information is the description information corresponding to the first sample image.

9. The method according to claim 7, characterized in that, The acquisition of the second training data includes: Obtain at least one second sample image; Determine the spatial relationship between at least two objects in each second sample image; the spatial relationship includes two-dimensional spatial relationship and three-dimensional spatial relationship. Second training data is constructed based on the spatial relationship between at least two objects in each second sample image.

10. The method according to claim 9, characterized in that, The spatial positional relationship includes a two-dimensional spatial positional relationship; Determining the spatial relationship between at least two objects in each second sample image includes: By using a multimodal large model, content analysis is performed on each second sample image to obtain descriptive information for each object in each second sample image; The description information is visually located using a visual positioning model to obtain the coordinate information of each object in the second sample image. Based on the coordinate information of at least one object in each second sample image, determine the two-dimensional spatial positional relationship between at least two adjacent objects in each second sample image.

11. The method according to claim 9, characterized in that, The spatial positional relationship includes a three-dimensional spatial positional relationship; Determining the spatial relationship between at least two objects in each second sample image includes: The depth estimation model is used to perform depth estimation processing on each second sample image to obtain the depth information of each second sample image; Based on the depth information of each second sample image and the coordinate information of at least one object in each second sample image, the three-dimensional spatial positional relationship between at least two objects in each second sample image is determined.

12. The method according to any one of claims 7 to 11, characterized in that, The step of training the initial model based on the training data to obtain the multimodal visual perception model includes: Using the initial model, N inference samplings are performed on each first sample image to obtain N first prediction results. Each first prediction result includes prediction description information of at least one predicted object in the first sample image. Based on the N first prediction results, N first reward values ​​are calculated using the first reward function; Using the initial model, each second sample image is subjected to M inference samplings to obtain M second prediction results; each second prediction result includes the predicted relative position information between at least two predicted objects in the second sample image. Based on the M second prediction results, M second reward values ​​are calculated using the second reward function; Based on the first difference and the second difference, the parameters of the initial model are updated to obtain the multimodal visual perception model; the first difference is the difference between the N first reward values; the second difference is the difference between the M second reward values; Wherein, the first reward function is used to evaluate the degree of consistency between the predicted description information of the predicted object and the real visual content of the predicted object, and the predicted description information of the predicted object is obtained by inference sampling of the first sample image through the initial model; the second reward function is used to evaluate the degree of matching between the predicted relative position information between at least two predicted objects and the real relative position information between the at least two predicted objects, and the predicted relative position information between the two predicted objects is obtained by inference sampling of the second sample image through the initial model; N is a positive integer, and M is a positive integer.

13. The method according to claim 12, characterized in that, The step of calculating N first reward values ​​based on the N first prediction results using a first reward function includes: Based on the prediction bounding box corresponding to each predicted object in each first prediction result, the image region corresponding to the prediction bounding box is cropped from the first sample image; By using an image description model, each image region is described to obtain reference description information for each predicted object; The prediction description information of each predicted object in the first prediction result is compared with the reference description information of each predicted object to obtain the consistency comparison result information. Based on the consistency comparison result information, the first reward value corresponding to the first prediction result is calculated using the first reward function; The first reward value is used to characterize the degree of consistency between the predicted description information generated by the initial model and the real visual content.

14. The method according to claim 12, characterized in that, The step of calculating M second reward values ​​based on the M second prediction results using the second reward function includes: The predicted relative position information and the reference relative position information in each of the second prediction results are compared to determine a first quantity; the first quantity is the number of predicted relative position information that matches the reference relative position information. Based on the first quantity and the second quantity, a second reward value is calculated using the second reward function; wherein, the second quantity is the total quantity of the predicted relative position information; and the second reward value is used to characterize the degree of consistency between the predicted relative position information of the initial model and the actual relative position.

15. An object query device, characterized in that, The device includes: The processing module is used to display first dialogue information on a first interface, the first dialogue information being used to indicate the search for a first object, and the first interface including a preview image captured by a camera. The processing module is further configured to perform object lookup based on the first dialogue information and the preview image, and output the first information; The first information includes at least one of the following: The relative position information of the first object with respect to the user; The relative position information between the first object and the second object in the preview image.

16. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the object query method as described in any one of claims 1-14.