Object recognition method and device and computer program product
By receiving the query request and orientation information from the mobile terminal, querying and rewriting user voice text in the object database, the problems of low recognition efficiency and low accuracy in the prior art are solved, and more efficient and accurate object recognition is achieved.
Patent Information
- Application Number
- CN202411999252.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, users need to obtain relevant information by photographing and identifying objects, resulting in low recognition efficiency and low accuracy.
By receiving the query request for the target object sent by the mobile terminal, including the user's voice text and the mobile terminal orientation, the object to be filtered is queried in the object database based on this information, and the user's voice text is rewritten through a collection of semantic tags, and the feedback text is generated and sent.
It significantly improves the accuracy and efficiency of object recognition, optimizes the user experience, and can instantly obtain users' query needs and current location information, providing a basis for subsequent object screening.
Smart Images

Figure CN120011536A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of identification technology, and in particular to an object identification method, device and computer program product. Background Art
[0002] With the rapid development of mobile Internet and smart devices, users' demands for information acquisition are becoming increasingly diverse and they pursue immediacy. Especially in mobile scenarios, users often want to know the relevant information of the currently observed object. However, in related technologies, users need to take out their mobile phones to shoot the objects to be identified before they can further identify them. The recognition accuracy depends on the shooting quality and shooting angle, which not only has low recognition efficiency, but also low recognition accuracy. Summary of the invention
[0003] In view of this, embodiments of the present application at least provide an object recognition method, device, and computer program product.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] In one aspect, an embodiment of the present application provides an object recognition method, the method comprising:
[0006] Receiving a query request for a target object sent by a mobile terminal; the query request includes a user voice text and a mobile terminal location;
[0007] Based on the mobile terminal location and the user voice text, the object to be screened is searched in the object database; the object database sets the appearance attribute of the preset object by at least one semantic tag in the semantic tag set;
[0008] Rewrite the user voice text based on the semantic tag set to obtain a query text;
[0009] Based on the query text and the appearance attributes of the object to be screened, a feedback text is generated and sent to the mobile terminal.
[0010] On the other hand, an embodiment of the present application provides an object recognition device, the device comprising:
[0011] A receiving module, configured to receive a query request for a target object sent by a mobile terminal; the query request includes a user voice text and a mobile terminal location;
[0012] A query module, configured to query an object database for an object to be screened based on the mobile terminal position and the user voice text; the object database sets the appearance attribute of a preset object by at least one semantic tag in a semantic tag set;
[0013] A rewriting module, used for rewriting the user voice text based on the semantic tag set to obtain a query text;
[0014] The feedback module is used to generate and send feedback text to the mobile terminal based on the query text and the appearance attributes of the object to be screened.
[0015] On the other hand, an embodiment of the present application provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements some or all of the steps in the above method.
[0016] In the embodiment of the present application, by receiving a query request for a target object sent by a mobile terminal, including user voice text and mobile terminal location, the server can instantly obtain the user's query needs and current location information, providing a basis for subsequent object screening; at the same time, based on the mobile terminal location and user voice text, the object to be screened is queried in the object database, and objects related to the user's current location and query intent can be preliminarily screened out, thereby improving the pertinence and efficiency of the query; the user voice text is rewritten based on a semantic tag set to obtain a more accurate and standardized query text, which can further narrow the query scope and reduce the interference of irrelevant objects; finally, based on the rewritten query text and the appearance attributes of the object to be screened, a feedback text containing relevant object information is generated and sent to the mobile terminal, which can provide users with intuitive and accurate query results to meet the actual needs of users. Based on the embodiments provided by the present application, the accuracy and efficiency of object recognition can be significantly improved, and the user experience can be optimized.
[0017] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and are used together with the specification to illustrate the technical solution of the present application.
[0019] Figure 1 A schematic diagram of the architecture of an object recognition system provided in an embodiment of the present application;
[0020] Figure 2 A schematic diagram of an implementation flow of an object recognition method provided in an embodiment of the present application;
[0021] Figure 3 A schematic diagram of an implementation flow of an object recognition method provided in an embodiment of the present application;
[0022] Figure 4 A schematic diagram of an implementation flow of an object recognition method provided in an embodiment of the present application;
[0023] Figure 5 A schematic diagram of an implementation flow of an object recognition method provided in an embodiment of the present application;
[0024] Figure 6 A schematic diagram of an implementation flow of an object recognition method provided in an embodiment of the present application;
[0025] Figure 7 A schematic diagram of an implementation flow of an object recognition method provided in an embodiment of the present application;
[0026] Figure 8 A schematic diagram of a POI data production process provided in an embodiment of the present application;
[0027] Fig. 9 A schematic diagram of a route identification system provided in an embodiment of the present application;
[0028] Fig.10 A schematic diagram of the structure of an object recognition device provided in an embodiment of the present application;
[0029] Fig.11 A hardware entity diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0030] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices consistent with some aspects of the present disclosure as detailed in the appended claims.
[0031] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions of the present application are further elaborated in detail below in conjunction with the drawings and embodiments. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.
[0032] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments, but it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict. The terms "first / second / third" involved are merely used to distinguish similar objects and do not represent a specific order for the objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing this application and are not intended to limit this application.
[0034] See also Figure 1 , Figure 1 It is an optional architecture diagram of the object recognition system 100 provided in an embodiment of the present application. To support an object recognition application, a mobile terminal (mobile terminal 400-1 and mobile terminal 400-2 are shown as examples) is connected to a server terminal 200 via a network 300. The network 300 may be a wide area network or a local area network, or a combination of the two. Figure 1 It is also shown that the server 200 can be a server cluster, which includes servers 200-1 to 200-3. Similarly, the servers 200-1 to 200-3 can be physical machines or virtual machines built using virtualization technology (such as container technology and virtual machine technology, etc.). The embodiment of the present application does not limit this. Of course, a single server can also be used to provide services in this embodiment.
[0035] As an example, the server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, and a smart watch, etc., but is not limited to this. The terminal and the server 200 can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.
[0036] Figure 2 is a schematic diagram of an implementation flow of an object recognition method provided in an embodiment of the present application. The method can be executed by a server. Figure 2 The steps shown are explained.
[0037] Step S201: receiving a query request for a target object sent by a mobile terminal; the query request includes a user voice text and a mobile terminal location.
[0038] In some embodiments, the query request refers to a retrieval requirement for a target object sent by a mobile terminal to a server terminal, and the query request includes text formed by user voice input (i.e., user voice text) and the current geographic location of the mobile terminal (i.e., mobile terminal location).
[0039] The user voice text refers to the voice information recorded by the user through a voice input device (such as a microphone), which is converted into text form after being processed by voice recognition technology. In the embodiment of the present application, the user voice text may include descriptive information of the target object that the user wants to find or identify.
[0040] Among them, the mobile terminal's location refers to the geographical location information and direction information of the mobile terminal. The geographical location information can be obtained through location information, usually through the Global Positioning System (GPS) or network-based positioning technology (such as base station positioning, Wi-Fi positioning); the direction information can be obtained through geomagnetic sensors, accelerometers, gyroscopes and other sensors. It can be understood that the direction information can characterize the orientation and movement status of the mobile terminal.
[0041] In some embodiments, the target object is an actual object in the physical world (such as a commodity, artwork, person, place, landmark, etc.), wherein the user can issue a query voice for the target object after observing the target object; the client collects the query voice, converts it into user voice text and sends it to the server. Therefore, the user voice text carries the user's description information for the target object, which is used to reflect the user's query intention.
[0042] In some implementation scenarios, the user issues a query voice for the target object, the mobile terminal software collects this voice, converts it into text in the form of user voice text using voice recognition technology, and obtains the current location and direction information of the mobile terminal as the mobile terminal position. Then, this information is packaged into a query request and sent to the server through the network.
[0043] For example, a user is visiting a museum and is interested in an ancient pottery jar, but does not know its details. The user issues a voice query through a mobile terminal (which can be a smartphone): "What is that jar with blue stripes?" At the same time, the GPS function of the mobile terminal automatically records and sends the current exhibition hall location information.
[0044] For example, a user is driving and is interested in a building on the roadside, but does not know its detailed information. The user issues a voice query through a mobile terminal (which can be a car computer): "What is that red building that looks like a bucket?" At the same time, the GPS function of the mobile terminal automatically records and sends the current vehicle location information.
[0045] The above only describes a few scenario examples of the present application. Of course, the present application can also be applied to other scenarios, such as asking for product information in a supermarket, asking for student information in a class, etc.
[0046] Step S202: searching an object database for an object to be screened based on the mobile terminal location and the user voice text; the object database sets appearance attributes of a preset object at least through at least one semantic tag in a semantic tag set.
[0047] In some embodiments, the object database is a database system that stores multiple preset objects and corresponding object information. The object information includes the name, location, appearance photo, semantic tag, etc. of the preset object, which is used to support operations such as query, retrieval and analysis. Among them, the semantic tag is a word or phrase used to describe the appearance characteristics or appearance attributes, such as color, shape, material, etc. In the object database, each preset object is marked with at least one semantic tag. Accordingly, the semantic tag set can be obtained by statistically analyzing the vocabulary set formed by the semantic tags of each preset object in the object database; the location of the preset object is the geographical location information corresponding to the preset object, which is used to indicate the specific location of the object in the real world.
[0048] In some embodiments, the above-mentioned process of querying the objects to be screened may include two screening processes, one screening process is a screening process based on the mobile terminal location, and the other screening process is a screening process based on semantic tags. It is understandable that these two screening processes can be performed simultaneously or successively, that is, the screening process based on the mobile terminal location is performed first, and then the screening process based on semantic tags is performed in the preset objects obtained after screening, or the screening process based on semantic tags is performed first, and then the screening process based on mobile terminal location is performed in the preset objects obtained after screening.
[0049] Among them, the screening process based on the mobile terminal's position may include: obtaining the location information of the preset object from the object database, calculating the distance between the mobile terminal's position and the preset object; obtaining a distance threshold (setting a reasonable distance value for screening preset objects located near the mobile terminal according to actual needs), comparing the calculated distance with the set distance threshold, and screening out preset objects whose distance is less than or equal to the threshold.
[0050] Among them, the screening process based on semantic tags may include: parsing the user's voice text, extracting keywords that reflect the user's query intention, matching the extracted keywords with the semantic tag set, and finding preset objects with relevant semantic tags.
[0051] In some embodiments, the number of the above-mentioned objects to be screened may be one or more.
[0052] Step S203: rewrite the user voice text based on the semantic tag set to obtain a query text.
[0053] The above query text is obtained after rewriting and optimization based on the user's voice text to more accurately reflect the user's query intention and improve the accuracy and efficiency of the query.
[0054] In some embodiments, taking into account the mismatch between the user voice text and the semantic tags in the object database, which leads to low query accuracy, the present application uses a semantic tag set to rewrite the user voice text, thereby eliminating ambiguity and vagueness in the user voice text, so that it more accurately reflects the user's query intention and improves the accuracy and efficiency of the query.
[0055] In some embodiments, the above rewriting process may include at least one of the following: modal particle filtering and keyword rewriting.
[0056] Among them, the server can pre-process the user's voice text, identify and filter out modal particles therein, where modal particles do not provide specific information about the query object and may increase the ambiguity of the query, and can include, for example: "ah", "um", "maybe", etc.
[0057] The server can analyze the keywords describing the object in the user's voice text and find the semantic tags corresponding to these keywords according to the semantic tag set. For example, for a query like "What is that building that looks like a bucket?", "bucket" can be rewritten as a semantic tag such as "shape: cylindrical" or "appearance feature: curved surface".
[0058] Finally, the server will combine the filtered modal particles and rewritten keywords into a new query text, which more accurately reflects the user's query intention than the original user voice text and can be used for precise retrieval in the object database.
[0059] For example, suppose the user voice text is: "What is that building that looks like a bucket?" After receiving the corresponding user voice text, the server will first filter out the modal particles "that" and "what", then analyze the keyword "bucket", and find the corresponding tag in the semantic tag set, such as "shape: cylindrical" and / or "appearance feature: curved surface". Finally, the following query text can be generated: "shape: cylindrical AND appearance feature: curved surface AND object type: building", or "find buildings with cylindrical shape and curved appearance".
[0060] For example, suppose the user voice text is: "What is that can with blue stripes?" After receiving the corresponding user voice text, the server will first filter out the modal particles "that" and "what", and then rewrite the keyword "blue stripes" into a more precise semantic tag "color: blue AND decoration: stripes". Finally, the following query text can be generated: "color: blue AND decoration: stripes AND type: can", or "find cans with blue color and stripes".
[0061] Step S204: Based on the query text and the appearance attributes of the object to be screened, generate and send feedback text to the mobile terminal.
[0062] The appearance attributes of the object to be screened are stored in the object database. In some embodiments, the appearance attributes of the object to be screened may be represented by at least one semantic tag.
[0063] In some embodiments, for the objects to be screened obtained by the query in step S202, the objects to be screened can be traversed, the appearance attributes (at least one semantic label) of each object to be screened can be compared with the query text, and the object to be screened with the highest matching degree can be determined, and the feedback text can be generated based on the object to be screened with the highest matching degree.
[0064] In some embodiments, when the object to be screened with the highest matching degree is found, other key information of the object, such as name, location, appearance features, etc., can be extracted from the object database, and feedback text can be generated based on this information.
[0065] In some embodiments, the server may send the generated feedback text to the mobile terminal and display it to the user.
[0066] For example, assuming that a user queries "What is that red round building?" by voice, after receiving the query request, the server will generate a query text: "Color: red AND shape: round AND object type: building", or "Find buildings with a round shape and red color"; traverse the objects to be screened obtained in step S202, assuming that there is a building named "Red Round Building" among them, whose appearance attributes are red and round, and which has the highest match with the query text, the information of this object will be extracted from the object database and a feedback text will be generated: "The red round building you queried is 'Red Round Building', located at No. XX, XX Road." Finally, the server sends this feedback text to the mobile terminal.
[0067] For example, the server traverses the objects to be screened and finds the collection with the highest matching degree with the query text, such as a collection called "blue-striped Han Dynasty pottery jar", whose appearance attributes fully meet the requirements in the query text. The server can extract detailed information of the collection from the database, such as name, age (Han Dynasty), material (clay), size, etc., and generate feedback text: "The blue-striped pottery jar you queried is a 'blue-striped Han Dynasty pottery jar', which is a Han Dynasty cultural relic, located in exhibition hall A, exhibit number 001." Finally, the server sends this feedback text to the user through the mobile application, and the user can see detailed collection information on the mobile phone screen.
[0068] In the embodiment of the present application, by receiving a query request for a target object sent by a mobile terminal, including user voice text and mobile terminal location, the server can instantly obtain the user's query needs and current location information, providing a basis for subsequent object screening; at the same time, based on the mobile terminal location and user voice text, the object to be screened is queried in the object database, and objects related to the user's current location and query intent can be preliminarily screened out, thereby improving the pertinence and efficiency of the query; the user voice text is rewritten based on a semantic tag set to obtain a more accurate and standardized query text, which can further narrow the query scope and reduce the interference of irrelevant objects; finally, based on the rewritten query text and the appearance attributes of the object to be screened, a feedback text containing relevant object information is generated and sent to the mobile terminal, which can provide users with intuitive and accurate query results to meet the actual needs of users. Based on the embodiments provided by the present application, the accuracy and efficiency of object recognition can be significantly improved, and the user experience can be optimized.
[0069] Figure 3 Schematic diagram of the implementation process of an object recognition method provided in an embodiment of the present application. Figure 2 , Figure 2 Step S203 in can be updated to step S301 to step S303, combining Figure 3 The steps shown are explained.
[0070] Step S301: extracting an initial description tag for the target object in the user voice text.
[0071] The initial description tag refers to a keyword extracted from the user's voice text through natural language processing technology that can characterize the abstract features of the target object. In some embodiments, the initial description tag is a colloquial description and an abstract description of the appearance features of the target object.
[0072] In some embodiments, considering that the initial description tag is a keyword included in the user's voice text, the natural language processing (NLP) technology can be used to extract nouns, adjectives, verbs and other key words related to the target object from the user's voice text as the initial description tag. The natural language processing here can be used to perform word segmentation, part-of-speech tagging and other processing on the user's voice text.
[0073] For example, when the user voice text is "What is that building that looks like a red bucket?", the user voice text "What is that building that looks like a red bucket?" is parsed through natural language processing technology, and "red" and "bucket" are extracted as initial description tags.
[0074] Step S302: searching the semantic tag set for a target semantic tag that matches the initial description tag.
[0075] In some embodiments, it is possible to first search in the semantic tag set whether there is a semantic tag that is completely identical to the initial description tag. If there is a semantic tag that is completely identical to the initial description tag, the initial description tag (or the corresponding semantic tag) is directly determined as the target semantic tag that matches it. If there is no semantic tag that is completely identical to the initial description tag, a semantic matching algorithm is used to search in the semantic tag set for at least one semantic tag that has the highest similarity to the initial description tag as the target semantic tag that matches it.
[0076] In some embodiments, at least one semantic tag with the highest similarity to the initial description tag can be found in the following manner: using machine learning methods, the initial description tag and the semantic tags in the semantic tag set are respectively converted into feature vectors; the similarity between the initial description tag and each semantic tag is respectively calculated to determine at least one semantic tag with the highest similarity.
[0077] Exemplarily, continuing based on the above example, after extracting "red" and "bucket" as initial description tags, the tags that best match "red" and "bucket" can be searched in the predefined semantic tag set. Among them, the semantic tag set may include the semantic tag of "red", so the semantic tag of "red" is directly used as the target semantic tag; because the semantic tag set may not include the semantic tag of "bucket", it is necessary to use the semantic matching algorithm to search for tags that match "bucket" in the semantic tag set, for example, to find the semantic tag of "cylindrical" (or "bucket-shaped"). All the obtained tags "red", "cylindrical" (or "bucket-shaped") are determined as target semantic tags.
[0078] Step S303: rewrite the user voice text based on the target semantic tag to obtain the query text.
[0079] In some embodiments, the information in the user's voice text can be replaced or supplemented based on the target semantic tag to generate a new query text, including but not limited to: replacing the initial description tag with a more precise target semantic tag, and adding a standardized template for the query statement (such as query XX for XX).
[0080] Exemplarily, continuing based on the above example, the user voice text "What is that building that looks like a red bucket?" can be rewritten as: "Query for buildings that are red in color and cylindrical in shape."
[0081] In the embodiment of the present application, by extracting the initial description tag for the target object in the user's voice text, the user's descriptive information about the target object can be effectively captured. At the same time, by searching for the target semantic tag that matches the initial description tag in the semantic tag set, the user's natural language description can be converted into a more standardized and structured semantic representation, which helps to improve the accuracy and efficiency of subsequent queries. In addition, by rewriting the user's voice text based on the target semantic tag, a clearer and more specific query text can be obtained, further improving the accuracy of subsequent queries.
[0082] Figure 4 This is a schematic diagram of the implementation process of an object recognition method provided in an embodiment of the present application. Figure 4 , the method can be executed by a processor of a computer device. Figure 2 The method may further include steps S401 to S403, combining Figure 4 The steps shown are explained.
[0083] Step S401: Acquire at least one appearance image of a preset object.
[0084] Here, the appearance image is used to display the external shape, color, texture and other features of the preset object. Exemplarily, the appearance image may be, but is not limited to, a digital photo, a video frame, a rendered image and the like.
[0085] In some embodiments, the appearance image of the preset object can be obtained by at least one of the following methods: capturing the appearance image of the preset object through a photographic device (such as a camera, mobile phone, etc.); pulling the image of the preset object from an existing image library or database; using web crawler technology to automatically capture images related to the preset object from the Internet.
[0086] In some embodiments, the at least one appearance image is an image of the preset object in different scenes, where different scenes include different viewing angles, different weather, and different seasons. The appearance images of the preset object in different viewing angles are images taken from various directions or angles of the object, such as the front, side, back, top, etc.; different weathers may include sunny days, cloudy days, rainy days, snowy days, etc.; different seasons are images of the preset object taken in different seasons, such as spring, summer, autumn, winter, etc.
[0087] Step S402: extracting features of the appearance of the preset object in each of the appearance images to obtain appearance feature description information corresponding to each of the appearance images.
[0088] In some embodiments, a multimodal large model (Vision-and-Language Model, VLM) may be used to extract features of the appearance of a preset object in an appearance image to obtain appearance feature description information of the preset object in the appearance image.
[0089] Among them, the multimodal large model is a deep learning model that can process and understand visual (image) and language (text) information at the same time. It combines the technology of computer vision and natural language processing, and can learn the relationship between images and text, so as to achieve the understanding and description of complex visual scenes. In the feature extraction task, the multimodal large model can use rich visual and language knowledge to more accurately extract the key information in the image.
[0090] In step S402, the multimodal large model can be used to accurately identify and encode features such as color, texture, and shape in the appearance image, thereby generating appearance feature description information that can fully describe the appearance of the preset object.
[0091] In some embodiments, for each appearance image, feature extraction may be performed on the appearance image using a multimodal large model to obtain appearance feature description information corresponding to each appearance image.
[0092] Step S403: setting at least one semantic label of the preset object based on the appearance feature description information corresponding to each of the appearance images.
[0093] In some embodiments, the appearance feature description information corresponding to all the appearance images obtained can be statistically processed to find out at least one appearance feature description information that can reflect the general appearance features of the preset object, and then the at least one appearance feature description information obtained can be set as at least one semantic label of the preset object.
[0094] In some embodiments, for all appearance feature description information describing the preset object, the number of occurrences of each type of appearance description information may be counted, and the appearance description information whose number of occurrences exceeds a preset threshold is used as the semantic label.
[0095] In other embodiments, a correlation analysis method may be used to perform cluster analysis on all appearance feature description information of the preset object, thereby classifying similar appearance feature description information into one category and extracting semantic labels corresponding to each category of appearance feature description information.
[0096] In an embodiment of the present application, by acquiring at least one appearance image of a preset object in different scenes (including different viewing angles, different weather conditions, and different seasons), the appearance features of the preset object under different conditions can be fully captured, providing a rich data basis for subsequent feature extraction and semantic label setting. At the same time, by using a large multimodal model (VLM) to extract features of the preset object in each appearance image, it is possible to deeply mine key information such as color, texture, and shape in the image, and generate accurate and comprehensive appearance feature description information. Finally, the appearance feature description information corresponding to each appearance image is summarized to obtain at least one semantic label of the preset object, which is convenient for subsequent retrieval tasks. Based on the embodiments provided in the present application, comprehensive capture and accurate description of the appearance features of the preset object can be achieved.
[0097] In some embodiments, the at least one semantic tag includes a semantic tag of at least one category; the setting of at least one semantic tag of the preset object based on the appearance feature description information corresponding to each of the appearance images includes: for each of the categories, extracting the appearance tag of the category from the appearance feature description information corresponding to each of the appearance images; and determining the semantic tag of the category in combination with the appearance tag of each of the appearance images.
[0098] In some embodiments, the semantic tags are distinguished by categories, and illustratively, the categories may include color, shape, size, style, texture, etc. Therefore, in the process of extracting the semantic tags of the preset objects, targeted extraction may be performed for each category, and a semantic tag may be present for each category as much as possible, so that the description of the preset objects in the final object database may be more comprehensive.
[0099] In some embodiments, after obtaining the appearance feature description information of each appearance image, all the appearance feature description information can be classified based on the above categories to obtain a description information set for each category. It can be understood that the description information set of a category includes all the appearance feature description information of this category. Afterwards, for each category description information set, a statistical method is used to determine the semantic label of each category.
[0100] For example, assume that all appearance feature description information of a preset object (such as a red cylindrical high-rise building) includes red (appears 6 times), light red (appears 2 times), light gray (appears 2 times), column (appears 9 times), square (appears 1 time), and circle (appears 2 times); first classify these appearance feature description information, for example, into color categories (including red, light red, light gray) and shape categories (including column, square, circle). For the color category, since red appears the most times, the semantic label of the color category can be set to red; similarly, for the shape category, since the column appears the most times, the semantic label of the shape category can be set to column.
[0101] In the embodiment of the present application, accurate and representative semantic labels can be set for each category of the preset objects. These semantic labels can intuitively describe the appearance feature categories of the preset objects and provide strong data support for subsequent object recognition. In addition, since the setting of semantic labels is based on the results of statistics and analysis, they have high stability and robustness and can cope with image data under different scenes and conditions.
[0102] In some embodiments, the target object is a target landmark object, and the preset object is a preset landmark object; the method also includes: obtaining multiple original landmark objects and appearance feature description information of each of the original landmark objects; based on the appearance feature description information of each of the original landmark objects, deleting the original landmark objects that cannot be observed by the mobile terminal to obtain the preset landmark object.
[0103] The original landmark object refers to a set of landmark objects obtained from a landmark pull interface (which may be provided by a map software). The original landmark objects may include various types of landmarks, such as buildings, scenic spots, etc., and may include some landmarks that cannot be observed by the mobile terminal (such as indoor landmarks). When the original landmark object is obtained, the appearance feature description information of the original landmark object will be obtained. The appearance feature description information includes the indoor and outdoor properties of the original landmark object. The indoor and outdoor properties are used to distinguish whether the original landmark object can be observed by the mobile terminal.
[0104] In some embodiments, the indoor and outdoor attributes can be reflected by an attribute identifier, such as "I" for indoor and "O" for outdoor; it can also be determined by the object name of the original landmark object, for example, "XX Cinema" and "XX Restaurant" can be determined as indoor, and "XX Building" and "XX Monument" can be determined as outdoor.
[0105] In some embodiments, multiple original landmark objects can be pulled from the map software, and appearance feature description information of these original landmark objects can be obtained at the same time. Based on these appearance feature description information, especially indoor and outdoor properties, those original landmark objects that cannot be observed by the mobile terminal (such as indoor landmarks) are deleted, and the remaining original landmark objects are the preset landmark objects.
[0106] In the embodiments of the present application, based on the indoor and outdoor attributes in the appearance feature description information, the original landmark objects (such as indoor landmarks) that cannot be observed by the mobile terminal can be accurately identified and deleted, thereby ensuring the validity of the preset landmark objects and reducing the interference of invalid landmark information; in addition, using attribute identifiers and object names to judge indoor and outdoor attributes not only improves the accuracy and efficiency of the judgment, but also enhances the flexibility and applicability of the method. Based on the embodiments provided by the present application, a more accurate and effective set of preset landmark objects can be constructed.
[0107] Figure 5 Schematic diagram of the implementation process of an object recognition method provided in an embodiment of the present application. Figure 2 , Figure 2 Step S204 in can be updated to step S501 to step S503, combining Figure 5 The steps shown are explained.
[0108] Step S501: input the query text and the appearance attributes of the object to be screened into a large language reasoning model to obtain a reasoning result output by the large language reasoning model.
[0109] Among them, the Large Language Model is a natural language processing model based on deep learning technology, which can perform logical reasoning, context understanding, knowledge reasoning and other tasks by understanding and analyzing text information. In the embodiment of the present application, the Large Language Model is used to receive the query text and the appearance attributes of the object to be screened, and output the reasoning result to predict the object that best matches the query text.
[0110] In some embodiments, the appearance attributes of the objects to be screened include at least one semantic tag. In the above step S501, the query text and at least one semantic tag of each object to be screened can be input into the large language reasoning model. Here, the large language reasoning model can find the objects to be screened that match the query text among the objects to be screened based on the appearance description of the target object that the user wants to query in the query text and at least one semantic tag of each object to be screened.
[0111] In some embodiments, the query text and the appearance attributes of the object to be screened can be integrated to generate the target input text. Exemplarily, the query text and the appearance attributes of the object to be screened can be integrated based on a preset integration template, for example, the integration template can be "the query scope includes: object to be screened n, whose corresponding semantic tags are XX, ...; please search for the object to be screened that is most relevant to (query text) within the above query scope"; or "search for the object to be screened that is most relevant to (query text) among the object to be screened n, whose corresponding semantic tags are XX, ...".
[0112] Assume that the objects to be screened and their corresponding semantic labels include:
[0113] Building A, white and square;
[0114] Building B, red and square;
[0115] Building C, red and round;
[0116] Building D is white and round.
[0117] The query text includes: "a building that is red in color and round in shape".
[0118] The generated target input text can be:
[0119] "Find the objects that are most relevant to the query text among the objects to be filtered:
[0120] Object to be screened 1, building A, has a corresponding semantic label of white and square; object to be screened 2, building B, has a corresponding semantic label of red and square; object to be screened 3, building C, has a corresponding semantic label of red and round; object to be screened 4, building D, has a corresponding semantic label of white and round.
[0121] Please find the most relevant objects to be filtered for 'buildings that are red in color and round in shape' within the above query range. "
[0122] The generated target input text can also be:
[0123] "Find the object most relevant to 'red circular building' among the following objects to be filtered:
[0124] Building A: white, square; Building B: red, square; Building C: red, round; Building D: white, round. ”
[0125] Accordingly, the inference result output by the large language inference model may be: "The most relevant object to be screened is building C."
[0126] Step S502: When the inference result includes a predicted object, verify the predicted object to obtain an object to be output.
[0127] Here, the predicted object is the object output by the large language reasoning model. Considering the possible error problem when the large language reasoning model outputs the predicted object, that is, the predicted object output by the large language reasoning model may not exist in the object to be screened, resulting in object recognition error. Step S502 introduces a verification mechanism to further verify the predicted object to ensure that the output object not only conforms to the result of model reasoning, but also exists in the object to be screened, thereby improving the accuracy of object recognition.
[0128] In some embodiments, the above-mentioned verification of the predicted object can be implemented through steps S5021 and S5022 to obtain the object to be output.
[0129] Step S5021: When the predicted object exists in the objects to be screened, use the predicted object as the object to be output.
[0130] Exemplarily, based on the above example, it is assumed that in the inference result output by the large language inference model, the predicted object is "Building C". Since "Building C" is among the objects to be screened, the predicted object "Building C" is directly used as the object to be output.
[0131] Step S5022: When the predicted object does not exist in the objects to be screened, the object to be screened that has the highest correlation with the predicted object is used as the object to be output.
[0132] In some embodiments, the reasons why the predicted object does not exist in the objects to be screened include but are not limited to: (1) model error of the large language reasoning model; (2) the information provided by the query text is not sufficient to accurately match an object in the objects to be screened, and the large language reasoning model may output a predicted object that does not completely match the object to be screened but is relatively close; (3) the search range of the large language reasoning model exceeds the range of the objects to be screened.
[0133] In some embodiments, the correlation between the predicted object and each object to be screened can be calculated, and the object to be screened with the highest correlation can be used as the object to be output. The correlation can be determined by at least one of the following factors: semantic similarity, object attribute similarity, etc.
[0134] Exemplarily, based on the above example, assuming that the predicted object in the inference result output by the large language inference model is "Building E", since "Building E" is not among the objects to be screened, it is necessary to recalculate the association between "Building E" and "Building A" to "Building D" respectively. Taking semantic similarity as an example, assuming that the semantic similarity between "Building A" and "Building E" is the highest, then "Building A" is used as the object to be output. It should be noted that "Building A" to "Building D" in the above example are only exemplary descriptions for easy distinction, not the actual names of the buildings. Therefore, the semantic similarity between "Building A" and "Building E" is actually the semantic similarity between the real names.
[0135] In the embodiment of the present application, while ensuring the accuracy of the output, it is possible to flexibly deal with the situation where the predicted object is missing, and by providing the alternative options with the highest correlation, the robustness of the system and the user experience are effectively improved.
[0136] Step S503: Generate the feedback text based on the object to be output.
[0137] In some embodiments, the object to be output may be organized into colloquial (object to be output) text content based on a preset feedback text template.
[0138] In some embodiments, there may be multiple preset feedback text templates. When generating feedback text, a target feedback text template matching the user voice text may be determined from multiple preset feedback text templates. Then, the feedback text is generated using the target feedback text template and the object to be output.
[0139] For example, for the above-mentioned query related to buildings, the system may preset the following templates: "The building of {building feature description} you mentioned is {building name}.", "The building that looks like {metaphor} is actually {building name}.", etc. Based on the user's voice text "What is that building that looks like a red bucket?", the preset template is searched and the target template that matches it is determined. For example, the template "The building that looks like {metaphor} is actually {building name}." can be selected because it can match the metaphorical description proposed by the user well. After that, "red bucket" can be filled in the {metaphor} position, and "building C" can be filled in the {building name} position, so that a colloquial and easy-to-understand feedback text can be generated: "The building that looks like a red bucket is actually building C."
[0140] In the embodiment of the present application, the preset feedback text template and colloquial expressions can be used to generate text content that is more natural and close to the user's daily language habits. In addition, by selecting the most suitable template according to the user input, a feedback text that is closer to the user's needs and expectations can be generated, further improving the user's satisfaction.
[0141] In some embodiments, the method may further include step S504.
[0142] Step S504: if the inference result does not include the predicted object, generate a feedback text for instructing the user to supplement the description.
[0143] In some embodiments, the reasoning result not including the predicted object may be based on at least one of the following reasons: insufficient input data, vague information, or model capability limitation, etc. The above feedback text is text content used to convey information to the user, request input, provide guidance, or explain the results.
[0144] In this embodiment, in order to achieve more accurate reasoning and meet user needs, step S504 is expected to guide the user to supplement the description by generating feedback text, thereby providing more useful information. In some embodiments, the feedback text can ask the user for more information about the target object.
[0145] For example, in the example of the red bucket, the following feedback text may be generated: “I’m sorry, but I’m not sure what the red bucket-like building you mentioned is. Can you provide more descriptions of the building? For example, its specific location, surroundings, or other notable features?”
[0146] In the embodiment of the present application, by guiding the user to provide more useful information, the user's intention and needs can be understood more accurately, thereby obtaining a more accurate reasoning result. In addition, this step also helps to improve the user's trust and satisfaction with the system, because the user can feel that the system is actively working to solve their problems and making improvements and optimizations based on their feedback.
[0147] Figure 6 Schematic diagram of the implementation process of an object recognition method provided in an embodiment of the present application. Figure 2 , Figure 2 Step S202 in can be updated to step S601 to step S603, combining Figure 6 The steps shown are explained.
[0148] Step S601: extracting position description information from the user voice text; the position description information is used to characterize the relative position of the target object with respect to the mobile terminal.
[0149] The position description information is used to indicate the direction and distance of the target object relative to the mobile terminal when describing the spatial position relationship. For example, the position description information may include "left", "right", "front", "back", "far", "near", etc.
[0150] In some embodiments, the user's voice text may be preprocessed first, including noise removal, word segmentation, part-of-speech tagging, etc.; and named entity recognition or keyword extraction technology in natural language processing may be used to identify the location description information in the text.
[0151] For example, by extracting the user voice text "What is the building on the right that looks like a red bucket?", the location description information "on the right" can be obtained.
[0152] Step S602: Determine the geographical range information of the target object based on the mobile terminal position and the position description information.
[0153] In some embodiments, the mobile terminal orientation is at least one of the following: mobile terminal position and mobile terminal movement direction.
[0154] Among them, the mobile terminal position refers to the specific position of the mobile terminal in the geographic space, which can be obtained through satellite navigation systems such as GPS and Beidou. It can be understood that the coordinate system where the mobile terminal position is located is the same as the coordinate system where the position of the preset object in the object database is located. The direction of movement of the mobile terminal is the direction or direction of travel of the mobile terminal during the movement. For example, if the mobile terminal is a vehicle, the direction of movement of the mobile terminal can be the driving direction of the vehicle; if the mobile terminal is a mobile phone held by the user, the direction of movement of the mobile terminal can be the walking direction of the user.
[0155] In some embodiments, the above-mentioned determination of the geographical range information of the target object based on the mobile terminal location and the location description information can be achieved through at least one of steps S6021 to 6023.
[0156] Step S6021: When the mobile terminal position includes the mobile terminal location and the position description information includes distance information, generate the geographic range information based on the mobile terminal location and the distance information.
[0157] The distance information is information in the position description information used to describe the relative distance between the target object and the mobile terminal. Exemplarily, the distance information can be expressed in distance units such as meters and kilometers.
[0158] In some embodiments, when the mobile terminal's orientation includes location information and the orientation description information includes distance information, the server can calculate a circular or annular area with the mobile terminal's location as the center and the distance as the radius based on the mobile terminal's current location coordinates and the distance information provided by the user as the possible geographical range where the target object may be located.
[0159] For example, if the distance information is "within 50 meters", the generated geographic range information is a circular range with the mobile terminal position as the center and a radius of 50 meters; if the distance information is "outside 50 meters", the generated geographic range information is a circular range with the mobile terminal position as the center, 50 meters as the inner ring radius, and 500 meters (the preset maximum visible distance) as the outer ring radius.
[0160] Step S6022: When the mobile terminal position includes the mobile terminal movement direction and the position description information includes direction information, generate the geographic range information based on the mobile terminal movement direction and the direction information.
[0161] The direction information is information in the position description information that is used to describe the relative direction between the target object and the mobile terminal.
[0162] In some embodiments, when the mobile terminal position includes the direction of movement of the mobile terminal, and the position description information includes direction information, the server can calculate a sector area at a certain angle to the direction of movement of the mobile terminal based on the current direction of movement of the mobile terminal and the direction information provided by the user. This sector area takes the current position of the mobile terminal as the center of the circle and the direction described by the user as the center direction of the sector. The boundary angle of the sector can be set according to the actual application scenario and user needs, and the radius distance of the sector can be a preset maximum visible distance.
[0163] For example, if the movement direction of the mobile terminal is 0 degrees due north and the direction information provided by the user is right, the generated geographic range information is centered on the mobile terminal position, with a radius of 500 meters (the preset maximum visible distance), and the fan center direction is 90 degrees due east (assuming the preset fan angle range is 90 degrees). The angle of one boundary is 45 degrees northeast, and the angle of the other boundary is a fan range of 135 degrees southeast.
[0164] Step S6023: When the mobile terminal orientation includes the mobile terminal position and the mobile terminal movement direction, and the orientation description information includes the distance information and the direction information, generate the geographic range information based on the mobile terminal position, the distance information, the mobile terminal movement direction and the direction information.
[0165] Step S6023 is actually a combination of step S6021 and step S6022. When the mobile terminal orientation includes the mobile terminal position and the mobile terminal movement direction, and the orientation description information includes the distance information and the direction information, the size of the geographic range information can be further reduced.
[0166] For example, if the distance information is "within 50 meters", the direction of movement of the mobile terminal is 0 degrees due north, and the direction information provided by the user is right, then the generated geographic range information is centered on the mobile terminal position, with a radius of 50 meters, and the fan-shaped center direction is 90 degrees due east (assuming the preset fan-shaped angle range is 90 degrees). The angle of one boundary is 45 degrees northeast, and the angle of the other boundary is a fan-shaped range of 135 degrees southeast.
[0167] In the embodiment of the present application, the geographical range information of the target object is determined by the mobile terminal position and position description information. In this way, the query range of the object to be screened can be narrowed, the transmission bandwidth for data interaction with the object database can be reduced, and the calculation amount for subsequent object recognition and matching can be reduced, thereby improving recognition efficiency and accuracy.
[0168] Step S603: query the object to be screened in the object database based on the geographic range information.
[0169] In some embodiments, the corresponding spatial query conditions are constructed based on the geographic range information (such as a circular area, a sector area, etc.) calculated in step S602; thereafter, the spatial query interface or function provided by the object database can be used to apply the geographic range information as a query condition to the preset objects in the object database, and perform a spatial matching operation, thereby filtering out all the preset objects within the geographic range information and returning them to the server as objects to be filtered.
[0170] In an embodiment of the present application, by extracting the location description information in the user's voice text, the user's intuitive description of the target object relative to the mobile terminal's location can be accurately captured; at the same time, based on the combination of the mobile terminal's location and the location description information, the geographic range information where the target object may be located can be calculated; based on this geographic range information, querying the object database for objects to be screened can effectively narrow the search scope, improve recognition efficiency, and enhance the correlation between the returned results and the target object location described by the user.
[0171] In some embodiments, the appearance attribute includes the location of the object; the above-mentioned querying of the object to be screened in the object database based on the geographic range information can be implemented through step S6031.
[0172] Step S6031: Determine the object to be screened based on the geographic range information and the object position of the preset object in the object database.
[0173] Among them, the object database also stores the object position of the preset object, which is the specific position of the preset object in the geographic space and can be obtained through satellite navigation systems such as GPS and Beidou.
[0174] In some embodiments, the spatial query interface provided by the object database can be used to match the geographic range information with the object location of the preset object as the query condition, and then the preset objects whose object locations are within the geographic range information can be screened out and returned to the server as the objects to be screened.
[0175] In some embodiments, the appearance attributes also include object height; determining the objects to be screened based on the geographic range information and the object positions of the preset objects in the object database includes: determining the objects to be pulled within the geographic range information among the preset objects based on the geographic range information and the object positions of the preset objects in the object database; removing the target objects to be pulled based on the object heights of the objects to be pulled and the orientation of the mobile terminal to obtain the objects to be screened; the target objects to be pulled are objects that cannot be observed by the mobile terminal.
[0176] The object height is the vertical distance of a specific object (such as a building, sculpture, etc.) relative to a reference plane (such as the ground) in three-dimensional space. In this embodiment, on the basis of determining the objects to be screened based on the geographic range information and the object position, the object height and the mobile terminal orientation are further considered to remove those objects that cannot be observed by the mobile terminal, thereby improving the accuracy and practicality of the screening results.
[0177] In some embodiments, the server preliminarily screens out the objects to be pulled within the specified range based on the geographic range information and the object location information in the object database. Then, the system obtains the object height information of these objects to be pulled, and calculates the visibility of each object to be pulled relative to the mobile terminal in combination with the mobile terminal position (which may also include the height of the mobile terminal) of the mobile terminal (such as a vehicle). For those objects that cannot be observed by the mobile terminal because they are too low or blocked by other objects to be pulled (i.e., the target objects to be pulled), they are removed from the objects to be screened; finally, the final objects to be screened are obtained.
[0178] For example, for a mobile terminal, building A, building B, and building C located in a straight line, if the height of the mobile terminal is 1 meter, the height of building A is 15 meters and the distance from the mobile terminal is 30 meters, the height of building B is 13 meters and the distance from the mobile terminal is 40 meters, and the height of building C is 28 meters and the distance from the mobile terminal is 60 meters. Based on the triangular relationship, it can be known that building B will be blocked by building A, and building C will not be completely blocked by building A and building B. Therefore, building B is removed as the target object to be pulled, and the objects to be screened include building A and building C.
[0179] In an embodiment of the present application, by combining the object height of each object to be pulled and the orientation of the mobile terminal, calculating the visibility of each object, and removing those objects that cannot be observed by the mobile terminal, it can be ensured that the objects in the screening results can be seen by the user in the actual scene, thereby improving the accuracy and practicality of the screening results.
[0180] In some embodiments, the query request also includes an image of the environment in which the mobile terminal is located; after removing the target object to be pulled, the method also includes: obtaining the appearance images of the remaining objects to be pulled from the object database; based on the environment image and the appearance images of the remaining objects to be pulled, removing the remaining objects to be pulled that do not exist in the environment image.
[0181] The above-mentioned environment image of the mobile terminal is a real-time image of the environment of the mobile terminal, which is used to reflect the actual scene in the current environment. For example, when the mobile terminal is a vehicle, the real-time image can be collected by a panoramic camera set on the vehicle.
[0182] In some embodiments, after removing the target object to be pulled, the server can obtain the appearance images of the remaining objects to be pulled from the object database; using the image matching algorithm, the appearance image of each object to be pulled is compared with the environment image provided by the mobile terminal. For those objects to be pulled that do not find a match in the environment image, they are regarded as objects that do not exist in the actual environment and are removed from the objects to be screened. What can be obtained in the end is the object to be screened that is both within the geographical range and exists in the actual environment.
[0183] A practical scenario for the application of this solution is that during the driving of the current mobile terminal (vehicle), if there is an obstacle on the right side that does not exist in the object database, such as a large vehicle, a temporary construction building, etc., which blocks the user's line of sight in one direction, and also blocks the vehicle's camera for collecting environmental images, then the user is actually unable to observe all the buildings on the right side. Therefore, even if there is an object to be pulled on the right side of the vehicle, it needs to be deleted to improve recognition accuracy.
[0184] In an embodiment of the present application, by introducing the environmental image of the mobile terminal as part of the query request, the actual scene in the current environment can be reflected in real time, providing richer information for the subsequent screening process; at the same time, after removing the target object to be pulled, the appearance images of the remaining objects to be pulled are obtained from the object database and compared with the environmental image, so as to further accurately screen the objects to be pulled and remove those objects that are located within the geographical scope but are invisible in the actual environment due to being blocked by obstacles or other reasons.
[0185] In some embodiments, the querying of the object to be screened in the object database based on the mobile terminal location and the user voice text further includes: when there is no location description information in the user voice text, constructing the geographical range information of the target object based on the mobile terminal location.
[0186] In the embodiment of the present application, the location of the mobile terminal and the preset query range can be directly used to determine the geographical range information of the target object.
[0187] Exemplarily, the location of the mobile terminal can be used as the center of the circle and the query range can be used as the radius, such as 100 meters, to generate a circular range with a radius of 100 meters as the geographic range information.
[0188] In an embodiment of the present application, by utilizing the mobile terminal location and a preset query range to construct geographic range information, even if the user does not provide a specific location description, a reasonable geographic range can be automatically inferred based on the user's current location and possible query intentions, thereby improving convenience and flexibility.
[0189] Figure 7 FIG. 1 is a schematic diagram of an implementation flow of an object recognition method provided in an embodiment of the present application, and the method can be executed by a mobile terminal. Figure 7 The steps shown are explained.
[0190] Step S701: Send a query request carrying the user's voice text and the location of the mobile terminal to the server.
[0191] The query request refers to the search requirement for the target object sent by the mobile terminal to the server, and the query request includes the text formed by the user's voice input (ie, the user voice text) and the current geographical location of the mobile terminal (ie, the mobile terminal location).
[0192] The user voice text refers to the voice information recorded by the user through a voice input device (such as a microphone), which is converted into text form after being processed by voice recognition technology. In the embodiment of the present application, the user voice text may include descriptive information of the target object that the user wants to find or identify.
[0193] Among them, the mobile terminal's location refers to the geographical location information and direction information of the mobile terminal. The geographical location information can be obtained through location information, usually through the Global Positioning System (GPS) or network-based positioning technology (such as base station positioning, Wi-Fi positioning); the direction information can be obtained through geomagnetic sensors, accelerometers, gyroscopes and other sensors. It can be understood that the direction information can characterize the orientation and movement status of the mobile terminal.
[0194] Step S702, receiving a feedback statement sent by the server; wherein the feedback statement is generated by the server based on the query text and the appearance attributes of the object to be screened; the object to be screened is obtained by querying the object database based on the mobile terminal location and the user voice text, and the query text is obtained by rewriting the user voice text based on the semantic tags of the semantic tag set in the object database.
[0195] Here, the implementation scheme for generating feedback statements on the server side can refer to Figures 2 to 6 The implementation methods provided are not repeated in this application.
[0196] In some embodiments, the method further includes: collecting the mobile terminal position of the mobile terminal in real time; and in response to receiving a query voice, converting the query voice into the user voice text.
[0197] Among them, the mobile terminal can collect the current location information (longitude, latitude, etc.) in real time through technical means such as GPS positioning system or network base station positioning; when the mobile terminal receives the user's query voice, it can use voice recognition technology to convert the voice signal into the user's voice text in text form. These location information and user voice text are then integrated into the query request and sent to the server for processing.
[0198] In some embodiments, the method further includes: converting the feedback text into feedback voice, and playing the feedback voice.
[0199] Among them, when the mobile terminal receives the feedback text from the server, it can use text-to-speech technology to convert the text information into sound signals. Then these sound signals are played to the user through the speaker so that the user can hear the feedback results of the system. In some embodiments, in order to provide a better user experience, the system can also adjust the volume, speed and tone of the sound according to the user's preferences and scene requirements.
[0200] In this way, by converting the feedback text into feedback voice and playing it, the user can still obtain the system feedback information even when it is inconvenient for the user to view the screen. This not only improves the usability and ease of use of the system, but also enhances the user experience and satisfaction.
[0201] The following describes the application of the object recognition method provided in the embodiments of the present application in actual scenarios, mainly involving a route recognition system applied to a vehicle.
[0202] First, with the continuous evolution of large models, their intelligence is constantly improving, and they are reshaping all walks of life at an unprecedented speed. As the focus of the current integration of informatization and intelligence, the in-vehicle smart cockpit has naturally become an important field for the application of large models.
[0203] Secondly, the navigation function in the car cabin, after a series of technological innovations from desktop map navigation to augmented reality (AR) navigation, seems to have encountered a development bottleneck. The existing navigation system lacks interactivity and intelligence, and the focus of technological development is mostly on the visualization of navigation effects, while intelligence still needs to be strengthened. In addition, the navigation function mainly focuses on route planning, road conditions, and the positioning of starting and ending points, but fails to fully perceive the surrounding buildings and environment, such as map points of interest (POI), or its perception ability is not smart enough to answer questions raised by users about the surrounding environment, such as: "What is the name of the bucket-shaped building in front?" "What is the red tower on the left?" and so on. These questions are extremely common in self-driving trips or parent-child trips, but there is currently a lack of effective technical means to answer them.
[0204] In order to solve the above problems, the present application provides a smart route identification system, which is divided into two aspects: POI data production and intelligent matching service.
[0205] First, in terms of POI data production, POI data is pre-extracted based on the map data of the map supplier, and a database of appearance photos of all POIs is established based on the extracted POI data. With the help of the capabilities of the multimodal large model (VLM), the appearance of each POI is identified and described from multiple dimensions, and finally text information is formed; at the same time, based on the powerful semantic understanding and information extraction and summarization capabilities of the large language model, the appearance data of each POI is summarized and summarized, and finally the unique appearance feature description information of each POI is formed. In addition, this application will also extract geographic information such as the latitude and longitude coordinates of each POI and the building attribute information given by the map. Finally, the geographic information and appearance feature description information of all POIs are summarized as the basic database of the intelligent route recognition system.
[0206] Secondly, in terms of intelligent matching services, this application uses the text understanding and reasoning capabilities of the large language model to combine the user's question "What is the name of the building in front that looks like a bucket?", the buildings in front of the current location and heading, the building's feature information, attribute information, etc., as well as typical examples given to the large model, into input information and give it to the large model. Using the large model's powerful text understanding and reasoning capabilities, the application selects the building that best suits the user's interest from the candidate buildings, which may be one, multiple, or none.
[0207] Therefore, through the construction of POI database and the use of large models, we can provide very extensive, very abstract and very interesting route recognition capabilities. In addition, the location service system can expand operations and services based on the precise route recognition capabilities provided, and provide users with the most timely and accurate information services. Finally, the intelligent route recognition system will be further explained from the perspective of flow charts and system block diagrams.
[0208] See also Figure 8 , which shows a schematic diagram of a POI data production process.
[0209] Step 811: Obtain POI original data.
[0210] All POIs and attribute data corresponding to each POI may be obtained from the map provider 820. The POIs and corresponding attribute data obtained here are POI original data.
[0211] Step 812: Screening of POI raw data.
[0212] Among them, POIs that are invisible while driving, such as indoor POIs, can be deleted, and the screening of POI raw data has been completed.
[0213] Step 813: Obtain a POI appearance picture.
[0214] The appearance picture of each POI can be obtained by crawling to establish an appearance photo library of each POI.
[0215] Step 814: Extract POI features based on the multimodal large model.
[0216] Among them, based on the image understanding ability of the multimodal large model, the appearance features of each POI can be extracted and summarized. The extraction here refers to saving the appearance features of the POI in the appearance image in the form of text (such as keywords), and the summary refers to summarizing all the extracted texts describing the appearance features to form a semantic label describing the POI and storing it in the object database.
[0217] Step 815: Summarize POI features and geographic information.
[0218] The POI feature is the semantic label of the POI obtained in the above steps, and the geographic information refers to the location of the POI in the real scene, such as longitude and latitude, which can be aggregated to form a POI database 830.
[0219] See also Fig. 9 , which shows a system schematic diagram of a route recognition system.
[0220] Among them, the route recognition system consists of the cockpit end (mobile end 910) of the vehicle machine and the server end 920. On the mobile end 910, relying on the perception capabilities of the cockpit (voice recognition module, body camera and positioning system, etc.), after recognizing the user's intention, the information can be summarized through the vehicle-side information summary module and the cloud service can be requested. After receiving the user's request, the smart route recognition service of the server end 920 extracts the request information and queries the POI database; then, it constructs the input information of the large model reasoning request (LLM reasoning request), selects the most suitable POI information by calling the large model service (LLM service) and returns it to the smart route recognition service. The smart route recognition service performs LLM information post-processing on the information returned by the large model, that is, it performs verification and post-processing again to avoid the occurrence of error information such as large model hallucinations. The post-processed information can be returned to the mobile end 910, and after TTS synthesis, the voice playback system is used to play the answer.
[0221] Based on the above embodiments, firstly, by establishing a POI database, POI data without visual appearance features such as stores and floor POIs in the map supplier data are screened out; secondly, by establishing a set of appearance features for each POI through POI photos, and combining the attribute information of POI, etc., more and more specific information can be provided than the original map POI data, surpassing the original map data in the information dimension of POI data; finally, by constructing a large model reasoning input through geographic information retrieval, POI feature information, and map attributes, relying on the powerful text understanding and reasoning ability of the large model, the user's questions and candidate POIs are accurately matched. Through the above technical solution, in addition to supporting POIs such as "community" supported by the map, POIs can also be accurately found according to the abstract and vivid description information of different users. On the basis of expanding the map solution, the service capability and generalization capability are greatly improved. At the same time, it provides better support for different age groups, especially children's visualization problems.
[0222] Based on the foregoing embodiments, an embodiment of the present application provides an object recognition device, which includes the units included and the modules included in the units, which can be implemented by a processor in a computer device; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.
[0223] Fig.10 A schematic diagram of the structure of an object recognition device provided in an embodiment of the present application is shown in FIG. Fig.10 As shown, the object recognition device 1000 includes: a receiving module 1010, a query module 1020, a rewriting module 1030, and a feedback module 1040, wherein:
[0224] The receiving module 1010 is used to receive a query request for a target object sent by a mobile terminal; the query request includes a user voice text and a mobile terminal location;
[0225] A query module 1020 is used to query the object to be screened in an object database based on the mobile terminal position and the user voice text; the object database sets the appearance attribute of the preset object by at least one semantic tag in the semantic tag set;
[0226] A rewriting module 1030, configured to rewrite the user speech text based on the semantic tag set to obtain a query text;
[0227] The feedback module 1040 is used to generate and send feedback text to the mobile terminal based on the query text and the appearance attributes of the object to be screened.
[0228] In some embodiments, the rewriting module 1030 is also used to: extract the initial description tag for the target object in the user voice text; search for a target semantic tag that matches the initial description tag in the semantic tag set; and rewrite the user voice text based on the target semantic tag to obtain the query text.
[0229] In some embodiments, the object recognition device 1000 includes a construction module, which is used to: obtain at least one appearance image of a preset object; perform feature extraction on the appearance of the preset object in each of the appearance images to obtain appearance feature description information corresponding to each of the appearance images; and set at least one semantic label of the preset object based on the appearance feature description information corresponding to each of the appearance images.
[0230] In some embodiments, the target object is a target landmark object, and the preset object is a preset landmark object; the construction module is used to: obtain multiple original landmark objects and appearance feature description information of each of the original landmark objects; based on the appearance feature description information of each of the original landmark objects, delete the original landmark objects that cannot be observed by the mobile terminal to obtain the preset landmark object.
[0231] In some embodiments, the feedback module 1040 is further used to: input the query text and the appearance attributes of the object to be screened into a large language reasoning model to obtain an inference result output by the large language reasoning model; when the inference result includes a predicted object, verify the predicted object to obtain the object to be output; and generate the feedback text based on the object to be output.
[0232] In some embodiments, the feedback module 1040 is also used to: when the predicted object exists in the objects to be screened, use the predicted object as the object to be output; when the predicted object does not exist in the objects to be screened, use the object to be screened with the highest correlation with the predicted object as the object to be output.
[0233] In some embodiments, the feedback module 1040 is further used to: generate feedback text for instructing the user to supplement the description when the inference result does not include the predicted object.
[0234] In some embodiments, the query module 1020 is also used to: extract the location description information in the user voice text; the location description information is used to characterize the relative location of the target object relative to the mobile terminal; based on the mobile terminal location and the location description information, determine the geographical range information of the target object; based on the geographical range information, query the object to be screened in the object database.
[0235] In some embodiments, the mobile terminal location is at least one of the following: a mobile terminal position and a mobile terminal movement direction; the query module 1020 is further used for: when the mobile terminal location includes the mobile terminal location and the location description information includes distance information, generating the geographic range information based on the mobile terminal location and the distance information; when the mobile terminal location includes the mobile terminal movement direction and the location description information includes direction information, generating the geographic range information based on the mobile terminal movement direction and the direction information; when the mobile terminal location includes the mobile terminal location and the mobile terminal movement direction and the location description information includes the distance information and the direction information, generating the geographic range information based on the mobile terminal location, the distance information, the mobile terminal movement direction and the direction information; wherein the distance information is information in the location description information used to describe the relative distance between the target object and the mobile terminal; and the direction information is information in the location description information used to describe the relative direction between the target object and the mobile terminal.
[0236] In some embodiments, the query module 1020 is further used to determine the object to be screened based on the geographic range information and the object position of the preset object in the object database.
[0237] In some embodiments, the appearance attributes also include object height; the query module 1020 is also used to: determine the objects to be pulled within the geographic range information in the preset objects based on the geographic range information and the object position of the preset objects in the object database; based on the object height of each object to be pulled and the orientation of the mobile terminal, remove the target object to be pulled to obtain the object to be screened; the target object to be pulled is an object that cannot be observed by the mobile terminal.
[0238] In some embodiments, the query request also includes an image of the environment in which the mobile terminal is located; after removing the target object to be pulled, the query module 1020 is also used to: obtain the appearance images of the remaining objects to be pulled from the object database; based on the environment image and the appearance images of the remaining objects to be pulled, remove the remaining objects to be pulled that do not exist in the environment image.
[0239] In some embodiments, the query module 1020 is further used to: construct the geographical range information of the target object based on the mobile terminal location when there is no location description information in the user voice text.
[0240] The description of the above device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. In some embodiments, the functions or modules included in the device provided in the embodiment of the present application can be used to execute the method described in the above method embodiment. For technical details not disclosed in the device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.
[0241] It should be noted that in the embodiment of the present application, if the above-mentioned object recognition method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific hardware, software or firmware, or any combination of hardware, software, and firmware.
[0242] An embodiment of the present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, some or all of the steps in the above method are implemented.
[0243] The embodiment of the present application provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, some or all of the steps in the above method are implemented. The computer-readable storage medium can be transient or non-transient.
[0244] An embodiment of the present application provides a computer program, including a computer-readable code. When the computer-readable code is run in a computer device, a processor in the computer device executes some or all of the steps for implementing the above method.
[0245] The embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product can be implemented specifically by hardware, software or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium, and in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK) and the like.
[0246] It should be noted here that the description of the various embodiments above tends to emphasize the differences between the various embodiments, and the same or similar aspects can be referenced to each other. The description of the above device, storage medium, computer program and computer program product embodiments is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the embodiments of the device, storage medium, computer program and computer program product of this application, please refer to the description of the method embodiment of this application for understanding.
[0247] Fig.11 A hardware entity diagram of a computer device provided in an embodiment of the present application is shown in FIG. Fig.11 As shown, the hardware entity of the computer device 1100 includes: a processor 1101 and a memory 1102, wherein the memory 1102 stores a computer program that can be run on the processor 1101, and the processor 1101 implements the steps in the method of any of the above embodiments when executing the program.
[0248] The memory 1102 stores computer programs that can be run on the processor. The memory 1102 is configured to store instructions and applications executable by the processor 1101. It can also cache data to be processed or processed by the processor 1101 and various modules in the computer device 1100 (for example, image data, audio data, voice communication data, and video communication data). This can be achieved through flash memory (FLASH) or random access memory (Random Access Memory, RAM).
[0249] When the processor 1101 executes the program, the steps of any of the above object recognition methods are implemented. The processor 1101 generally controls the overall operation of the computer device 1100.
[0250] An embodiment of the present application provides a computer storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the steps of the object recognition method in any of the above embodiments.
[0251] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0252] The processor may be at least one of an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Central Processing Unit (CPU), a controller, a microcontroller, and a microprocessor. It is understandable that the electronic device that implements the functions of the processor may also be other, and the embodiments of the present application are not specifically limited.
[0253] The above-mentioned computer storage medium / memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disk, or a compact disc read-only memory (CD-ROM) and the like; it can also be various terminals including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0254] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the serial numbers of the above-mentioned steps / processes does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.
[0255] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0256] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0257] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0258] In addition, all functional units in the embodiments of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the above integrated unit can be implemented in the form of hardware or in the form of hardware plus software functional units. A person of ordinary skill in the art can understand that all or part of the steps of the above method embodiments can be completed by hardware related to program instructions, and the above program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above method embodiments are executed; and the above storage medium includes: mobile storage devices, read-only memory (ROM), disks or optical disks, etc. Various media that can store program codes.
[0259] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0260] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.
[0261] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0262] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. An object recognition method, characterized in that: Applied to the server, the method includes: Receiving a query request for a target object sent by a mobile terminal; the query request includes a user voice text and a mobile terminal location; Based on the mobile terminal location and the user voice text, the object to be screened is searched in the object database; the object database sets the appearance attribute of the preset object by at least one semantic tag in the semantic tag set; Rewrite the user voice text based on the semantic tag set to obtain a query text; Based on the query text and the appearance attributes of the object to be screened, a feedback text is generated and sent to the mobile terminal.
2. The object recognition method according to claim 1, characterized in that: The rewriting of the user voice text based on the semantic tag set to obtain the query text includes: Extracting an initial description tag for the target object in the user voice text; Searching for a target semantic tag matching the initial description tag in the semantic tag set; The user voice text is rewritten based on the target semantic tag to obtain the query text.
3. The object recognition method according to claim 1, characterized in that: The method further comprises: Acquire at least one appearance image of a preset object; Extracting features of the appearance of the preset object in each of the appearance images to obtain appearance feature description information corresponding to each of the appearance images; At least one semantic label of the preset object is set based on the appearance feature description information corresponding to each of the appearance images.
4. The object recognition method according to claim 3, characterized in that: The target object is a target landmark object, and the preset object is a preset landmark object; the method further includes: Acquire multiple original landmark objects and appearance feature description information of each of the original landmark objects; Based on the appearance feature description information of each of the original landmark objects, the original landmark objects that cannot be observed by the mobile terminal are deleted to obtain the preset landmark objects.
5. The object recognition method according to any one of claims 1 to 4, characterized in that: The generating and sending feedback text to the mobile terminal based on the query text and the appearance attributes of the object to be screened includes: Inputting the query text and the appearance attributes of the object to be screened into a large language reasoning model to obtain a reasoning result output by the large language reasoning model; In the case where the inference result includes a predicted object, verifying the predicted object to obtain an object to be output; The feedback text is generated based on the object to be output.
6. The object recognition method according to claim 5, characterized in that: The step of verifying the predicted object to obtain the object to be output includes: In the case where the predicted object exists in the objects to be screened, taking the predicted object as the object to be output; When the predicted object does not exist in the objects to be screened, the object to be screened that has the highest correlation with the predicted object is used as the object to be output.
7. The object recognition method according to claim 5, characterized in that: The method further comprises: In a case where the inference result does not include the predicted object, a feedback text is generated for instructing the user to supplement the description.
8. The object recognition method according to any one of claims 1 to 4, characterized in that: The step of searching for objects to be screened in an object database based on the location of the mobile terminal and the user voice text includes: Extracting the position description information in the user voice text; the position description information is used to characterize the relative position of the target object relative to the mobile terminal; Determine the geographical range information of the target object based on the mobile terminal position and the position description information; The object to be screened is queried in the object database based on the geographic range information.
9. The object recognition method according to claim 8, characterized in that: The mobile terminal position is at least one of the following: the mobile terminal position and the mobile terminal movement direction; the determining the geographical range information of the target object based on the mobile terminal position and the position description information includes: In a case where the mobile terminal position includes a mobile terminal location, and the position description information includes distance information, generating the geographic range information based on the mobile terminal location and the distance information; In a case where the mobile terminal position includes the mobile terminal movement direction, and the position description information includes direction information, generating the geographic range information based on the mobile terminal movement direction and the direction information; In a case where the mobile terminal position includes a mobile terminal location and a mobile terminal movement direction, and the position description information includes the distance information and the direction information, generating the geographic range information based on the mobile terminal location, the distance information, the mobile terminal movement direction and the direction information; The distance information is information in the position description information used to describe the relative distance between the target object and the mobile terminal; the direction information is information in the position description information used to describe the relative direction between the target object and the mobile terminal.
10. The object recognition method according to claim 8, characterized in that: The appearance attribute includes the object location; and the querying the object to be screened in the object database based on the geographic range information includes: The object to be screened is determined based on the geographic range information and the object position of the preset object in the object database.
11. The object recognition method according to claim 10, characterized in that: The appearance attribute also includes object height; and determining the object to be screened based on the geographic range information and the object position of the preset object in the object database includes: Based on the geographic range information and the object position of the preset object in the object database, determining the object to be pulled within the geographic range information in the preset object; Based on the object height of each object to be pulled and the orientation of the mobile terminal, the target object to be pulled is removed to obtain the object to be screened; the target object to be pulled is an object that cannot be observed by the mobile terminal.
12. The object recognition method according to claim 11, characterized in that: The query request also includes an image of the environment in which the mobile terminal is located; after removing the target object to be pulled, the method further includes: Acquire the appearance images of the remaining objects to be pulled from the object database; Based on the environment image and the appearance images of the remaining objects to be pulled, the remaining objects to be pulled that do not exist in the environment image are removed.
13. The object recognition method according to claim 8, characterized in that: The step of searching the object database for the object to be screened based on the mobile terminal location and the user voice text further includes: In the case that there is no location description information in the user voice text, the geographical range information of the target object is constructed based on the location of the mobile terminal.
14. An object recognition device, characterized in that: The device comprises: A receiving module, configured to receive a query request for a target object sent by a mobile terminal; the query request includes a user voice text and a mobile terminal location; A query module, configured to query an object to be screened in an object database based on the mobile terminal position and the user voice text; the object database sets the appearance attribute of a preset object by at least one semantic tag in a semantic tag set; A rewriting module, used for rewriting the user voice text based on the semantic tag set to obtain a query text; The feedback module is used to generate and send feedback text to the mobile terminal based on the query text and the appearance attributes of the object to be screened.
15. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps in the method according to any one of claims 1 to 13 are implemented.