Target positioning method, device, electronic device and storage medium based on large model

By receiving object information in voice commands and using large models for target detection and visual retrieval, the problem of inaccurate target positioning in existing technologies is solved, and precise positioning in robot navigation and self-driving cars is achieved.

CN119469150BActive Publication Date: 2025-09-26BEIJING BAIDU NETCOM SCI & TECH CO LTD

Patent Information

Application Number
CN202411578022.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-06
Publication Date
2025-09-26
Estimated Expiration
2044-11-06

AI Technical Summary

Technical Problem

Existing technologies have difficulty in accurately identifying targets in images or video sequences and determining their precise locations in target positioning, especially in fields such as robot navigation and self-driving cars, where there is a problem of inaccurate positioning.

Method used

By receiving the positioning request from the target terminal, extracting the object information in the voice command, using the large model for target detection and visual retrieval, determining the location information of the target object, and sending it to the target terminal to achieve precise positioning.

Benefits of technology

It improves the accuracy and efficiency of target positioning, especially in robot navigation and self-driving cars, achieving precise positioning of target objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119469150B_ABST
    Figure CN119469150B_ABST
Patent Text Reader

Abstract

The present application discloses a target positioning method, device, electronic device, and storage medium based on a large model, and relates to the fields of computer technology, particularly to large models, voice technology, computer vision, deep learning, and the like. The solution is as follows: receiving a positioning request sent by a target terminal, the positioning request including a target image and a voice command; extracting first object information of the object to be positioned from the voice command, performing target detection on the target image based on the first object information, and obtaining a detection result; intercepting an object image of the candidate object from the target image based on the position information of the candidate object in the target image; determining the target object from the candidate objects using the large model based on the object image, the position information of the candidate object in the target image, and the first object information; and sending the position information of the target object in the target image to the terminal, so that the target terminal determines the position information of the target object relative to the target terminal based on the position information of the target object in the target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, in particular to artificial intelligence fields such as large models, speech technology, computer vision, and deep learning, and specifically to a target positioning method, device, electronic device, and storage medium based on a large model. Background Art

[0002] Target localization technology is a key technology in computer vision. It can identify targets in images or video sequences and determine their precise positions in the images or scenes. Target localization technology is widely used in many fields, such as robot navigation, self-driving cars, video surveillance, etc. Summary of the Invention

[0003] The present application provides a target positioning method, device, electronic device and storage medium based on a large model.

[0004] According to one aspect of the present application, a target positioning method based on a large model is provided, comprising:

[0005] Receiving a positioning request sent by a target terminal; wherein the positioning request includes a target image and a voice command;

[0006] Extracting first object information of the object to be located from the voice command, and performing object detection on the target image based on the first object information to obtain a detection result; wherein the detection result includes position information of the candidate object in the target image;

[0007] intercepting an object image of the candidate object from the target image according to position information of the candidate object in the target image;

[0008] Determine the target object from the candidate objects using a large model based on the object image, the position information of the candidate object in the target image, and the first object information;

[0009] The position information of the target object in the target picture is sent to the terminal, so that the target terminal determines the position information of the target object relative to the target terminal according to the position information of the target object in the target picture.

[0010] According to another aspect of the present application, a target positioning device based on a large model is provided, comprising:

[0011] A receiving module, configured to receive a positioning request sent by a target terminal; wherein the positioning request includes a target image and a voice command;

[0012] an extraction module, configured to extract first object information of an object to be located from the voice command;

[0013] a detection module, configured to perform target detection on the target image based on the first object information to obtain a detection result; wherein the detection result includes position information of the candidate object in the target image;

[0014] a screenshot module, configured to capture an object image of the candidate object from the target image based on position information of the candidate object in the target image;

[0015] a retrieval module, configured to determine a target object from the candidate objects using a large model based on the object image, position information of the candidate object in the target image, and the first object information;

[0016] The sending module is used to send the position information of the target object in the target picture to the terminal, so that the target terminal determines the position information of the target object relative to the target terminal according to the position information of the target object in the target picture.

[0017] According to another aspect of the present application, an electronic device is provided, including:

[0018] at least one processor; and

[0019] a memory communicatively connected to the at least one processor; wherein,

[0020] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the above embodiment.

[0021] According to another aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method according to the above embodiment.

[0022] According to another aspect of the present application, a computer program product is provided, including a computer program, which implements the steps of the method described in the above embodiment when executed by a processor.

[0023] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present application.

[0025] Figure 1A flowchart of a target positioning method based on a large model provided in one embodiment of the present application;

[0026] Figure 2 A schematic diagram of a flow chart of a target positioning method based on a large model provided in another embodiment of the present application;

[0027] Figure 3 A schematic diagram of a flow chart of a target positioning method based on a large model provided in another embodiment of the present application;

[0028] Figure 4 A schematic diagram of a process of a target positioning method based on a large model provided in an embodiment of the present application;

[0029] Figure 5 A schematic diagram of the structure of a target positioning device based on a large model provided in one embodiment of the present application;

[0030] Figure 6 It is a block diagram of an electronic device used to implement the large model-based target positioning method of an embodiment of the present application. DETAILED DESCRIPTION

[0031] The following description of exemplary embodiments of the present application is made in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0032] The following describes the target positioning method, device, electronic device and storage medium based on a large model according to the embodiments of the present application with reference to the accompanying drawings.

[0033] Figure 1 A flowchart of a large model-based target positioning method provided in one embodiment of the present application.

[0034] The target positioning method based on a large model in the embodiment of the present application can be executed by the target positioning device based on a large model in the embodiment of the present application. The device can be configured in an electronic device to realize the target positioning function.

[0035] Among them, the electronic device can be any device with computing capabilities, such as a personal computer, mobile terminal, server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, and other hardware devices with various operating systems, touch screens and / or display screens.

[0036] like Figure 1 As shown, the target positioning method based on the large model includes:

[0037] Step 101: Receive a positioning request sent by a target terminal.

[0038] In this application, the target terminal may include a camera device that can be used to shoot scenes within the shooting range. In addition, the target terminal also has the function of recording and playing sounds, and the target terminal can interact with the user through voice. After the target terminal collects the user's voice, it can perform intent recognition on the collected voice to determine the user's intention and determine whether the user's intention includes target positioning requirements. If the user's intention includes target positioning requirements, a positioning request can be sent to the server, and the server can receive the positioning request sent by the target terminal.

[0039] The positioning request may include a target image or voice command captured by the target terminal. For example, the target image may be captured by the target terminal while simultaneously capturing the user's voice, or captured after determining that the user's intention includes a target positioning request. For example, the voice command may be a voice message containing a target positioning request.

[0040] For example, the user inputs the voice command "Give me the apple on the left". After the target terminal recognizes the intention of the voice command and determines that the user's intention includes the need to locate the apple, it sends a positioning request to the server.

[0041] For example, the microphone of the target terminal may be always on, and the user may directly input voice, or there may be a voice input control on the target terminal, and the user may input voice while triggering the voice input control.

[0042] Step 102 : extracting first object information of the object to be located from the voice command, and performing target detection on the target image based on the first object information to obtain a detection result.

[0043] The first object information may include the name of the object to be located, and may also include features such as color, shape, size, and position of the object to be located.

[0044] Exemplarily, voice recognition can be performed on the voice instruction to convert the voice instruction into text, and the text can be segmented to obtain word segments, and part-of-speech tagging can be performed on the word segments. Then, named entity recognition can be performed on the part-of-speech tagging results to obtain the first object information.

[0045] In this application, the detection result of the target image may include the position information of the candidate object in the target image, the detection score, etc.

[0046] Exemplarily, the position information of the candidate object in the target image may refer to the position information of the detection box of the candidate object in the target image, and the position information may include the coordinates of four vertices of the detection box in the target image.

[0047] The detection score can be used to represent the confidence level of the position information of the detected candidate object. In addition, if there are multiple candidate objects, each candidate object can correspond to a detection score.

[0048] As one possible implementation, if the first object information includes an object name, feature extraction can be performed on the object name to obtain text features, and feature extraction can be performed on the target image to obtain image features. Detection results can then be obtained based on the text features and image features. Thus, by combining the image features extracted from the target image with the text features extracted from the object name for target detection, the accuracy and efficiency of detection results can be improved.

[0049] For example, if the name of the object to be located is extracted from the voice command as apple, then target detection can be performed on the target image based on the object name to detect the apple from the target image and obtain the location information, detection score, etc. of the apple in the target image.

[0050] It should be noted that, in the present application, the number of detected candidate objects can be one or more, and there is no limitation on this.

[0051] Step 103 : capturing an object image of the candidate object from the target image based on the position information of the candidate object in the target image.

[0052] In this application, the target image can be captured according to the detection frame of the candidate object to obtain the object image of the candidate object.

[0053] Step 104 : Determine the target object from the candidate objects using the large model based on the object image, the position information of the candidate objects in the target image, and the first object information.

[0054] In this application, the object image of the candidate object, the position information of the candidate object in the target image, the first object information, etc. can be input into the large model, and the large model can be used to retrieve the target object to be finally located from the detected candidate objects.

[0055] Step 105 : Sending the position information of the target object in the target picture to the terminal, so that the target terminal determines the position information of the target object relative to the target terminal based on the position information of the target object in the target picture.

[0056] In this application, the position information of the target object in the target image can be sent to the target terminal. The target terminal determines the distance from the target object to the target terminal based on the position information of the target object in the target image. Based on the distance, the position information of the target object relative to the target terminal can be determined.

[0057] In an embodiment of the present application, the object information of the object to be located is extracted from the voice command sent by the target terminal, and target detection is performed on the target image sent by the target terminal based on the object information. Based on the detection results and the extracted object information, the target object to be located is determined by using a large model, and then the position information of the target object in the target image is sent to the target terminal, so that the target terminal determines the position of the target object relative to the target terminal based on the position information, thereby achieving the positioning of the target object. Thus, by extracting the information of the specified object to be located from natural language, and then performing target detection, and using the large model to retrieve the object to be located from the detection results, the accuracy of target positioning can be improved. In addition, combining the object information extracted from natural language with target detection on the target image can improve the accuracy of the detection results.

[0058] Figure 2 A flowchart of a large model-based target positioning method provided in another embodiment of the present application.

[0059] like Figure 2 As shown, the target positioning method based on the large model includes:

[0060] Step 201: Receive a positioning request sent by a target terminal.

[0061] In the present application, step 201 can be implemented in any of the embodiments of the present application, so it will not be described in detail here.

[0062] Step 202 : extracting first object information of the object to be located from the voice command, and performing target detection on the target image based on the first object information to obtain a detection result.

[0063] Exemplarily, feature extraction can be performed on the first object information to obtain a first text feature, and feature extraction can be performed on the target image to obtain a first image feature. The self-attention mechanism, image-to-text cross-attention, and text-to-image cross-attention can be used to fuse the first text feature and the first image feature to obtain a second text feature and a second image feature. The target image feature with the highest matching degree with the second text feature is determined in the second image feature, and decoding is performed based on the target image feature to obtain a detection result.

[0064] For example, the target image features may be used as queries for a decoder to decode the target image features and obtain a detection result, wherein the decoder query may include a content query and a location query.

[0065] Therefore, using text features to guide the selection of image features and decoding based on the selected target image features can improve the accuracy of the detection results.

[0066] Step 203 : capturing an object image of the candidate object from the target image based on the position information of the candidate object in the target image.

[0067] In the present application, step 203 can be implemented in any of the embodiments of the present application, so it will not be described in detail here.

[0068] Step 204 : Obtain prompt information based on the object image, the position information of the candidate object in the target image, and the first object information.

[0069] In the present application, the preset prompt template may be filled according to the object image of the candidate object, the position information of the candidate object in the target image, the first object information, etc., to obtain prompt information.

[0070] Among them, the prompt information can be used to prompt the large model to perform visual retrieval tasks.

[0071] Exemplarily, if the number of candidate objects exceeds a preset number, an object picture containing these candidate objects can be captured from the target picture, and prompt information can be generated based on the object picture, position information of each candidate object in the target picture, first object information, etc.

[0072] Exemplarily, if the number of candidates does not exceed a preset number, a picture of each candidate object can be captured from the target picture, and for each candidate object, prompt information corresponding to each candidate object is generated based on the object picture of each candidate object, the position information of the candidate object in the target picture, the first object information, etc.

[0073] Step 205: Input the prompt information into the large model to determine whether the candidate object meets the positioning requirements.

[0074] In this application, a prompt can be input into the large model, and the large model determines whether the candidate object meets the positioning requirements based on the prompt information, that is, whether the candidate object is the object that the user wants to locate. Exemplarily, the positioning requirements can include the first object information extracted from the voice command.

[0075] Exemplarily, in addition to the object image of the candidate object, the position information of the candidate object in the target image, the first object information, etc., the prompt information may also include large model output requirements, such as output format requirements. For example, outputting "1" indicates that the candidate object is the object that the user wants to locate, and outputting "0" indicates that the candidate object is not the object that the user wants to locate.

[0076] For example, the large model can be used to extract object information from the object image and the location information of the candidate object in the target image to obtain second object information. Based on the match between the second object information and the first object information, it can be determined whether the candidate object meets the positioning requirements. For example, if the second object information matches the first object information, it can be determined that the candidate object meets the positioning requirements. If the second object information does not match the first object information, it can be determined that the candidate object does not meet the positioning requirements.

[0077] The second object information may include but is not limited to the name, color, shape, position, etc. of the candidate object. The position here may refer to the position of the candidate object in the target image.

[0078] Exemplarily, the first object information includes the position information of the object to be located. Based on the picture of the candidate object, it can be determined whether the position information of the candidate object in the detection result matches the position information of the object to be located in the first object information. If they match, it can be determined that the candidate object meets the positioning requirements, that is, the candidate object is the object that the user wants to locate. If they do not match, it can be determined that the candidate object is not the object that the user wants to locate.

[0079] Therefore, the large model is used to extract object information from the object image and the position information of the candidate object in the target image. Based on the matching between the extracted second object information and the first object information extracted from the voice command, it is determined whether the candidate object is the object the user wants to locate, thereby using the large model for visual retrieval to improve the accuracy of the retrieval.

[0080] Step 206: Determine the target object based on the judgment result of the large model.

[0081] In this application, the candidate objects that meet the positioning requirements and the number of candidate objects that meet the positioning requirements are determined based on the judgment results of the large model. If there is a candidate object that meets the positioning requirements, the candidate object can be determined as the target object. If there are multiple candidate objects that meet the positioning requirements, the candidate object with the highest detection score among the multiple candidate objects can be determined as the target object.

[0082] Therefore, when the large model determines that there are multiple candidate objects that are the objects that the user wants to locate, the candidate object with the highest detection score can be used as the final object to be located, thereby improving the accuracy of object retrieval.

[0083] Step 207 : Send the position information of the target object in the target picture to the terminal, so that the target terminal determines the position information of the target object relative to the target terminal based on the position information of the target object in the target picture.

[0084] In the present application, step 207 can be implemented in any of the embodiments of the present application, so it will not be described in detail here.

[0085] In an embodiment of the present application, prompt information is constructed based on the object image, the position information of the candidate object in the target image and the first object information, and the prompt information is input into the large model. The large model is used to perform visual retrieval to retrieve the object to be ultimately located from the candidate objects, thereby improving the accuracy of the retrieval.

[0086] Figure 3 A flowchart of a large model-based target positioning method provided in another embodiment of the present application.

[0087] like Figure 3 As shown, the target positioning method based on the large model includes:

[0088] Step 301: Receive a positioning request sent by a target terminal.

[0089] In this application, step 301 can be implemented in any of the embodiments of this application, so it will not be described in detail here.

[0090] Step 302: Perform voice recognition on the voice instruction to convert the voice instruction into text.

[0091] In this application, a pre-trained speech recognition model can be used to perform speech recognition on the voice instructions to convert the voice instructions into text.

[0092] Step 303: Use the large model to extract object information from the text to obtain first object information.

[0093] In this application, the corresponding prompt template can be filled in according to the text converted from the voice command to obtain the corresponding prompt information. The prompt information can be used to prompt the large model to perform the object information extraction task. The prompt information is input into the large model, and the large model is used to extract object information from the text to obtain the first object information.

[0094] Step 304: Perform target detection on the target image according to the first object information to obtain a detection result.

[0095] Step 305 : capturing an object image of the candidate object from the target image based on the position information of the candidate object in the target image.

[0096] Step 306 : Obtain prompt information based on the object image, the position information of the candidate object in the target image, and the first object information.

[0097] Step 307: Input the prompt information into the large model to determine whether the candidate object meets the positioning requirements.

[0098] Step 308: Determine the target object based on the judgment result of the large model.

[0099] In the present application, steps 304 to 308 can be implemented in any of the embodiments of the present application, and therefore will not be described in detail here.

[0100] Step 309 : Send the position information of the target object in the target picture to the terminal, so that the target terminal determines the position information of the target object relative to the target terminal based on the position information of the target object in the target picture.

[0101] Exemplarily, the target terminal may include a distance measuring device, and the target terminal may use the distance measuring device to measure the distance between the target object and the target terminal based on the position information of the target object in the target image, and determine the position information of the target object relative to the target terminal based on the distance and the position information of the target object in the target image.

[0102] Thus, the target terminal can measure the distance between the target terminal and the target object through the distance measuring device, and then determine the position information of the target object relative to the target terminal, thereby achieving accurate positioning of the target object in the scene.

[0103] Illustratively, the distance measuring device may be an infrared sensor, an ultrasonic device, or other device that can be used to measure the distance between objects, which is not limited.

[0104] For example, the target terminal may be provided with an infrared sensor, and the horizontal straight-line distance between the target object and the target terminal may be measured by the infrared sensor to ultimately obtain the coordinates of the target object relative to the target terminal.

[0105] Exemplarily, the target terminal may be a robot having a robotic arm. The target terminal may also control the robotic arm to pick up the target object based on position information of the target object relative to the target terminal.

[0106] Therefore, the target terminal can be a robot. During the human-computer interaction process, the target can be located using a large model based on the voice commands input by the user. The robot can pick up objects based on the positioning results, thereby completing complex human-computer interaction behaviors and enriching the human-computer interaction experience.

[0107] In an embodiment of the present application, a large model can be used to extract the object to be located specified by the user from natural language, thereby improving the accuracy of information extraction, and the large model can be used to retrieve the final object to be located from the candidate objects through visual retrieval, thereby improving the accuracy of retrieval and thus improving the accuracy of target positioning.

[0108] For ease of understanding, the following Figure 4 To explain, Figure 4 A schematic diagram of the process of a target positioning method based on a large model provided in an embodiment of the present application.

[0109] Taking the target terminal as a smart terminal as an example, the smart terminal can include a camera, which can capture the front image and transmit it to the cloud service. In addition, the camera also contains an infrared sensor that can measure the horizontal straight-line distance of objects in the camera image. At the same time, the smart terminal has the ability to record and play sound, and it can transmit the user's voice commands to the cloud service.

[0110] In the process of target positioning based on large models, the services used can be as follows:

[0111] Access service: Calling various other services based on the user's location request;

[0112] Speech conversion service: a service that converts speech into text;

[0113] Extraction service: This service extracts the name of the object (such as apple or clothing) and its features (such as color, shape, and location) to be retrieved in the user's voice command by calling the model service;

[0114] Object Detection Service: The input is the name of the object to be retrieved and the target image, and the output is the location of the object in the target image and a screenshot;

[0115] Retrieval service: Retrieve the object images identified by the target detection service and use a large model (such as the image itself, location, etc.) to determine whether it is the object the user wants to retrieve. If multiple conditions are met, the object with the highest score from the target detection service is returned, that is, the object with the highest detection score.

[0116] Model services: services provided by large models.

[0117] like Figure 4 As shown in the figure, the target positioning process based on the large model is as follows:

[0118] The user transmits a voice command to the [smart terminal], which sends a positioning request via the HTTP protocol. The service is triggered by an HTTP trigger and the positioning request is transmitted to the [access service]. The [access service] calls the [voice conversion service] to convert the voice command in the positioning request into text, then calls the [extraction service] to extract the object name and features such as color, shape, and position from the text. The [target detection service] is then called to pass in the object name to obtain one or more detected object images, locations, detection scores, and other information. The [retrieval service] is then called for retrieval. The [retrieval service] can construct a prompt word based on the object image, location information in the target image, object features extracted by the extraction service, and object name, and send it to the [model service] to determine whether it is the object the user wants to locate. If there are multiple objects, the position of the object with the highest score from the [target detection service] is returned in the target image. The [smart terminal] then measures the horizontal straight-line distance of the object in the target image by calling the infrared sensor, ultimately obtaining a coordinate relative to the [smart terminal].

[0119] In order to implement the above embodiment, the embodiment of the present application also proposes a target positioning device based on a large model. Figure 5 A schematic structural diagram of a large model-based target positioning device provided in one embodiment of the present application.

[0120] like Figure 5 As shown, the target positioning device 500 based on the large model includes:

[0121] The receiving module 510 is configured to receive a positioning request sent by a target terminal; wherein the positioning request includes a target image and a voice command;

[0122] An extraction module 520 is configured to extract first object information of the object to be located from the voice command;

[0123] a detection module 530 configured to perform object detection on the target image based on the first object information to obtain a detection result; wherein the detection result includes position information of the candidate object in the target image;

[0124] A screenshot module 540 is configured to capture an object image of the candidate object from the target image based on position information of the candidate object in the target image;

[0125] a retrieval module 550 for determining a target object from the candidate objects using a large model based on the object image, the position information of the candidate objects in the target image, and the first object information;

[0126] The sending module 560 is configured to send the position information of the target object in the target image to the terminal, so that the target terminal determines the position information of the target object relative to the target terminal based on the position information of the target object in the target image.

[0127] Optionally, the retrieval module 550 is configured to:

[0128] Acquire prompt information according to the object picture, the position information of the candidate object in the target picture, and the first object information;

[0129] Inputting the prompt information into the large model to determine whether the candidate object meets the positioning requirements;

[0130] The target object is determined according to the judgment result of the large model.

[0131] Optionally, the detection result further includes a detection score, and the retrieval module 550 is configured to:

[0132] In response to determining, according to the judgment result, that there are multiple candidate objects that meet the positioning requirement, a candidate object with the highest detection score among the multiple candidate objects is determined as the target object.

[0133] Optionally, the retrieval module 550 is configured to:

[0134] Using the large model, extracting object information from the object image and the position information of the candidate object in the target image to obtain second object information;

[0135] According to the matching between the second object information and the first object information, it is determined whether the candidate object meets the positioning requirement.

[0136] Optionally, the extraction module 520 is configured to:

[0137] performing speech recognition on the voice instruction to convert the voice instruction into text;

[0138] The large model is used to extract object information from the text to obtain the first object information.

[0139] Optionally, the detection module 530 is configured to:

[0140] performing feature extraction on the first object information to obtain a first text feature;

[0141] Performing feature extraction on the target image to obtain a first image feature;

[0142] fusing the first text feature and the first image feature to obtain a second text feature and a second image feature;

[0143] Determining a target image feature having the highest matching degree with the second text feature among the second image features;

[0144] Decoding is performed according to the target image features to obtain the detection result.

[0145] Optionally, the target terminal includes a distance measuring device, and the target terminal uses the distance measuring device to measure the distance between the target object and the target terminal according to the position information of the target object in the target image; and determines the position information of the target object relative to the target terminal based on the distance.

[0146] Optionally, the target terminal is a robot having a robotic arm, and the target terminal is further configured to control the robotic arm to pick up the target object according to position information of the target object relative to the target terminal.

[0147] It should be noted that the explanation of the aforementioned embodiment of the target positioning method based on a large model is also applicable to the target positioning device based on a large model of this embodiment, so it will not be repeated here.

[0148] In an embodiment of the present application, the object information of the object to be located is extracted from the voice command sent by the target terminal, and target detection is performed on the target image sent by the target terminal based on the object information. Based on the detection results and the extracted object information, the target object to be located is determined by using a large model, and then the position information of the target object in the target image is sent to the target terminal, so that the target terminal determines the position of the target object relative to the target terminal based on the position information, thereby achieving the positioning of the target object. Thus, by extracting the information of the specified object to be located from natural language, and then performing target detection, and using the large model to retrieve the object to be located from the detection results, the accuracy of target positioning can be improved. In addition, combining the object information extracted from natural language with target detection on the target image can improve the accuracy of the detection results.

[0149] According to an embodiment of the present application, the present application also provides an electronic device, a readable storage medium and a computer program product.

[0150] Figure 6A schematic block diagram of an example electronic device 600 that can be used to implement an embodiment of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0151] like Figure 6 As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 602 or a computer program loaded from a storage unit 608 into a RAM (Random Access Memory) 603. Various programs and data required for the operation of the device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An I / O (Input / Output) interface 605 is also connected to the bus 604.

[0152] Various components in device 600 are connected to I / O interface 605, including an input unit 606, such as a keyboard, mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, optical disk, etc.; and a communication unit 609, such as a network card, modem, wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0153] The computing unit 601 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various specialized AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the large-model-based target localization method. For example, in some embodiments, the large-model-based target localization method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the large-model-based target localization method described above can be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute the large model-based target positioning method in any other appropriate manner (for example, by means of firmware).

[0154] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0155] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow charts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0156] In the context of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0157] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0158] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.

[0159] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The client-server relationship is established by computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and poor scalability of traditional physical hosts and VPS services. The server may also be a server in a distributed system or a server integrated with blockchain.

[0160] According to an embodiment of the present application, the present application also provides a computer program product, which, when an instruction processor in the computer program product executes, executes the large model-based target positioning method proposed in the above embodiment of the present application.

[0161] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved. This is not a limitation herein.

[0162] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.

Claims

1. A target positioning method based on a large model, comprising: Receiving a positioning request sent by a target terminal; wherein the positioning request includes a target image and a voice command; Extracting first object information of the object to be located from the voice command, and performing object detection on the target image based on the first object information to obtain a detection result; wherein the detection result includes position information of the candidate object in the target image; intercepting an object image of the candidate object from the target image according to position information of the candidate object in the target image; Acquire prompt information according to the object picture, the position information of the candidate object in the target picture, and the first object information; Using the large model, extracting object information from the object image and the position information of the candidate object in the target image to obtain second object information; determining whether the candidate object meets positioning requirements based on a match between the second object information and the first object information; Determining the target object according to the judgment result of the large model; The position information of the target object in the target picture is sent to the terminal, so that the target terminal determines the position information of the target object relative to the target terminal according to the position information of the target object in the target picture.

2. The method according to claim 1, wherein The detection result also includes a detection score. The step of determining the target object based on the judgment result of the large model includes: In response to determining, according to the judgment result, that there are multiple candidate objects that meet the positioning requirement, a candidate object with the highest detection score among the multiple candidate objects is determined as the target object.

3. The method according to claim 1, wherein The extracting first object information of the object to be located from the voice command includes: performing speech recognition on the voice instruction to convert the voice instruction into text; The large model is used to extract object information from the text to obtain the first object information.

4. The method according to claim 1, wherein The performing target detection on the target image according to the first object information to obtain a detection result includes: performing feature extraction on the first object information to obtain a first text feature; Performing feature extraction on the target image to obtain a first image feature; fusing the first text feature and the first image feature to obtain a second text feature and a second image feature; Determining a target image feature having the highest matching degree with the second text feature among the second image features; Decoding is performed according to the target image features to obtain the detection result.

5. The method according to claim 1, wherein The target terminal includes a distance measuring device, and the target terminal determines the position information of the target object relative to the target terminal according to the position information of the target object in the target image, including: The target terminal measures the distance between the target object and the target terminal using the distance measuring device according to the position information of the target object in the target image; Determine the position information of the target object relative to the target terminal according to the distance.

6. The method according to any one of claims 1 to 5, wherein The target terminal is a robot having a robotic arm. The target terminal is further configured to control the robotic arm to pick up the target object according to position information of the target object relative to the target terminal.

7. A target positioning device based on a large model, comprising: A receiving module, configured to receive a positioning request sent by a target terminal; wherein the positioning request includes a target image and a voice command; an extraction module, configured to extract first object information of an object to be located from the voice command; a detection module, configured to perform target detection on the target image based on the first object information to obtain a detection result; wherein the detection result includes position information of the candidate object in the target image; a screenshot module, configured to capture an object image of the candidate object from the target image based on position information of the candidate object in the target image; a retrieval module configured to obtain prompt information based on the object image, the position information of the candidate object in the target image, and the first object information; and to extract object information from the object image and the position information of the candidate object in the target image using the large model to obtain second object information; determine whether the candidate object meets positioning requirements based on a match between the second object information and the first object information; and determine the target object based on the determination result of the large model; The sending module is used to send the position information of the target object in the target picture to the terminal, so that the target terminal determines the position information of the target object relative to the target terminal according to the position information of the target object in the target picture.

8. The device according to claim 7, wherein The detection result also includes a detection score, and the retrieval module is used to: In response to determining, according to the judgment result, that there are multiple candidate objects that meet the positioning requirement, a candidate object with the highest detection score among the multiple candidate objects is determined as the target object.

9. The device according to claim 7, wherein The extraction module is used to: performing speech recognition on the voice instruction to convert the voice instruction into text; The large model is used to extract object information from the text to obtain the first object information.

10. The device according to claim 7, wherein The detection module is used to: performing feature extraction on the first object information to obtain a first text feature; Performing feature extraction on the target image to obtain a first image feature; fusing the first text feature and the first image feature to obtain a second text feature and a second image feature; Determining a target image feature having the highest matching degree with the second text feature among the second image features; Decoding is performed according to the target image features to obtain the detection result.

11. The device according to claim 7, wherein The target terminal includes a distance measuring device, and the target terminal uses the distance measuring device to measure the distance between the target object and the target terminal according to the position information of the target object in the target picture; Determine the position information of the target object relative to the target terminal according to the distance.

12. The device according to any one of claims 7 to 11, wherein: The target terminal is a robot having a robotic arm. The target terminal is further configured to control the robotic arm to pick up the target object according to position information of the target object relative to the target terminal.

13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.

15. A computer program product comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image object recognition method and device and storage medium

    CN108681743A

  • Image positioning method and device based on interactive input, equipment and storage medium

    CN111400523A

Cited By

  • A large model target detection method for a belt foreign object sorting robot

    CN122500688A