Image recognition method based on visual language, controller, robot and medium
By generating an auxiliary object-finding knowledge graph and combining scene features to understand object-finding instructions, the problem of inaccurate object recognition by the visual language model in specific scenarios is solved, achieving higher object search accuracy and reliability.
Patent Information
- Application Number
- CN202510856503.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-06-24
AI Technical Summary
Existing visual language models find it difficult to accurately identify items for specific user groups in specific scenarios, especially items needed by the elderly, and are unable to understand the semantic association between item search instructions and specific scene images.
By obtaining the target user's item usage record information, an auxiliary search knowledge graph is generated, the graph's structured semantic features are encoded, and attention fusion is performed in combination with the current scene features. The joint features of the graph screen are used to understand the natural language instructions for searching, and the semantic association between the target item description information and the user's personalized behavior pattern in a specific scenario is strengthened.
It improves the accuracy and reliability of item search in specific scenarios, reduces semantic understanding errors caused by lack of personalized background knowledge, and enhances the robot's ability to find items in specific user group scenarios.
Smart Images

Figure CN120726435A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and is applicable to the fields of financial technology and medical health, and in particular to an image recognition method, controller, robot, and medium based on visual language. Background Art
[0002] Vision-Language Models (VLM) are an artificial intelligence model that combines computer vision and natural language processing technologies. They are widely used in image recognition, indoor navigation and other fields. For example, in the elderly care assistance scenario of the medical and health scenario, when the elderly issue a command to find elderly care items (such as "find my hearing aid"), the command and the photo of the elderly's indoor environment (such as the bedroom, the nursing home room) can be input into the visual language model to identify whether the target item exists in the photo, thereby providing indoor navigation services for the elderly. For another example, in the banking business processing scenario of the financial technology scenario, when the user issues a financial item search command (such as "find my bank card"), the command and the picture in the bank business hall where the user is located can be input into the visual language model to find the item.
[0003] However, current visual language models struggle to accurately identify items needed by specific user groups, such as the elderly, from images. Specifically, current visual language models are primarily oriented toward general life scenarios and struggle to understand the semantic relationship between item search instructions and specific scene images (such as photos of the elderly's surroundings), making it difficult to accurately identify and find items in these images.
[0004] Therefore, how to improve the accuracy of item search in specific scenarios has become a technical problem that needs to be solved urgently. Summary of the Invention
[0005] The main purpose of the embodiments of the present application is to propose an image recognition method, controller, robot and medium based on visual language, aiming to improve the accuracy of object search in specific scenarios.
[0006] To achieve the above objectives, a first aspect of an embodiment of the present application proposes an image recognition method based on visual language, the method comprising:
[0007] Obtaining item usage record information of a target user; wherein the item usage record information includes the name of the target item, the location of the item, and the user name of the target user;
[0008] Generating a knowledge graph based on the item usage record information to obtain an auxiliary object-finding knowledge graph; wherein the auxiliary object-finding knowledge graph includes at least two element nodes and a relationship between the element nodes, each element node corresponding to one of the target item name, the item location, and the user name;
[0009] Performing graph encoding on the auxiliary object-finding knowledge graph to obtain graph structured semantic features;
[0010] Capturing an image of the scene where the target user is located to obtain a current scene image, and visually encoding the current scene image to obtain current image features;
[0011] Performing attention fusion processing on the graph structured semantic features and the current picture features to obtain a graph-picture joint feature;
[0012] Obtain a natural language search instruction containing target item description information, and perform item search based on the natural language search instruction and the joint features of the atlas screen to obtain an item search status; wherein, the item search status represents whether the target item described by the target item description information exists in the current scene screen.
[0013] In some embodiments, the item usage record information also includes usage time and item function;
[0014] The generating of a knowledge graph based on the item usage record information to obtain an auxiliary object-finding knowledge graph includes:
[0015] Generate an item element node based on the target item name to obtain an item element node;
[0016] Generating a function element node based on the item function to obtain a function element node, and creating a first edge connecting the function element node and the item element node;
[0017] Generating a user element node based on the user name to obtain a user element node, and creating a second edge connecting the user element node and the function element node;
[0018] Generating a location element node based on the item location to obtain a location element node, and creating a third edge connecting the location element node and the user element node;
[0019] A time element node is generated based on the usage time to obtain a time element node, and a fourth edge connecting the time element node and the location element node is created to generate the auxiliary object-finding knowledge graph.
[0020] In some embodiments, there are at least two location element nodes;
[0021] After generating a time element node based on the usage time to obtain the time element node and creating a fourth edge connecting the time element node and the location element node, the method further includes:
[0022] Among at least two of the position element nodes connected to the same item element node, counting the number of the position element nodes having the same item position, to obtain a position occurrence frequency of each item position;
[0023] Determine the item position with the highest position occurrence frequency as the target item position;
[0024] The auxiliary object-finding knowledge graph is updated according to the location of the target item.
[0025] In some embodiments, performing item search based on the natural language instruction for finding the item and the combined features of the atlas screen to obtain item search status includes:
[0026] In response to the object-finding natural language instruction, performing object name recognition on the target object description information to obtain a first object name recognition result;
[0027] If the first item name recognition result indicates that the item name does not exist in the target item description information, feature extraction is performed on the target item description information to obtain a fuzzy description feature of the item;
[0028] Determine a predicted item name based on the fuzzy description features of the item and the structured semantic features of the graph, and perform feature extraction on the predicted item name to obtain a clear description feature of the item;
[0029] The item feature recognition is performed based on the clear description features of the item and the combined features of the atlas screen to obtain the item search situation.
[0030] In some embodiments, the graph structured semantic feature includes at least two graph node features, each of which corresponds to one of the element nodes in the auxiliary object-finding knowledge graph;
[0031] Determining the predicted item name based on the fuzzy description features of the item and the structured semantic features of the graph includes:
[0032] Perform feature correlation analysis based on the fuzzy description features of the item and the features of each graph node to obtain the node description feature correlation;
[0033] Determine the element node corresponding to the graph node feature with the largest node description feature association degree as the target association node;
[0034] If the target associated node is an item element node, determining the target item name corresponding to the target associated node as the predicted item name; wherein the item element node is the element node corresponding to the target item name;
[0035] If the target association node is not the item element node, the target item name corresponding to the item element node connected to the target association node is determined as the predicted item name.
[0036] In some embodiments, the object-finding natural language instruction is a target object-finding text instruction;
[0037] The step of obtaining the natural language instruction for finding an object including description information of the target object includes:
[0038] Obtaining the target user's voice command for finding the object;
[0039] Converting the object-finding voice command into text to obtain an original object-finding text command; wherein the original object-finding text command includes original object description information;
[0040] Performing item name recognition on the original item description information to obtain second item name recognition information;
[0041] If the second item name recognition condition indicates that the item name exists in the original item description information, determining the item name in the original item description information as the original item name;
[0042] According to the association between the original object name and the object name description information in the preset object name knowledge database, the original object-finding text instruction is rewritten to obtain the target object-finding text instruction.
[0043] In some embodiments, the item name description information includes an item alias and a standard item name of the same item;
[0044] The step of rewriting the original object-finding text instruction according to the association between the original object name and the object name description information in the preset object name knowledge database to obtain the target object-finding text instruction includes:
[0045] Performing text semantic similarity analysis based on the original item name and the item alias to obtain item name similarity;
[0046] The standard item name corresponding to the item alias with the greatest item name similarity and a preset object-finding text instruction template are subjected to text fusion to generate the target object-finding text instruction.
[0047] To achieve the above-mentioned purpose, the second aspect of an embodiment of the present application proposes a controller, which includes a memory and a processor, the memory stores a computer program, and the processor implements the method described in the first aspect when executing the computer program.
[0048] To achieve the above-mentioned purpose, a third aspect of an embodiment of the present application proposes a robot, which includes the controller described in the second aspect.
[0049] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method described in the first aspect.
[0050] The image recognition method, controller, robot and medium based on visual language proposed in this application obtain the target user's item usage record information and generate an auxiliary object-finding knowledge graph based on the item usage record information. The auxiliary object-finding knowledge graph can characterize the relationship between the target user (such as the elderly) and the items he uses, such as the relationship between the target item name, the item location and the user name, so as to provide sufficient background knowledge for subsequent item search. The auxiliary object-finding knowledge graph is encoded to obtain graph structured semantic features, and the picture (that is, the current scene picture) collected from the scene where the target user is located is encoded to obtain the current picture features. Then, instead of directly searching for items based on the image of the scene where the target user is located, the graph structured semantic features and the current picture features are first fused through the attention mechanism to form a graph-picture joint feature, so as to form a joint representation containing user habit knowledge and real-time scene information, and then the graph-picture joint feature and the object-finding natural language instruction are combined to search for items. In this way, based on the association information contained in the auxiliary object-finding knowledge graph, we can more accurately understand the description of the object in the natural language instruction of the object-finding instruction (that is, the target object description information), as well as the semantic association between the target object description information and the image of the scene where the target user is located, and strengthen the semantic association between the target object description information and the personalized behavior pattern of the user in the specific scene, thereby reducing the semantic understanding error of the general visual language model in the specific user group scene due to the lack of personalized background knowledge, and improving the accuracy and reliability of object search in the specific scene where the target user is located. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is a flowchart of the image recognition method based on visual language provided by an embodiment of the present application;
[0052] Figure 2 This is a schematic diagram of a specific implementation of the image recognition method based on visual language provided by another embodiment of the present application;
[0053] Figure 3 yes Figure 1 Flowchart of step 102 in FIG.
[0054] Figure 4 is a flowchart of an image recognition method based on visual language provided by another embodiment of the present application;
[0055] Figure 5 This is provided in one embodiment of the present application Figure 1 Flowchart of step 106 in FIG.
[0056] Figure 6 yes Figure 5 Flowchart of step 403 in FIG.
[0057] Figure 7 Another embodiment of the present application provides Figure 1 Flowchart of step 106 in FIG.
[0058] Figure 8 yes Figure 7 Flowchart of step 605 in FIG.
[0059] Figure 9 This is a schematic diagram of the hardware structure of the controller provided in an embodiment of the present application. DETAILED DESCRIPTION
[0060] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0061] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.
[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0063] First, let’s analyze some of the terms used in this application:
[0064] Artificial Intelligence (AI) is a new technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. A branch of computer science, AI attempts to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thinking. It can also refer to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. Basic AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning. This application can acquire and process relevant data based on AI technologies.
[0065] Vision-Language Models (VLMs): These are AI models that combine computer vision and natural language processing techniques. Combining vision and language, VLMs can learn from both images and text to solve cross-modal tasks. VLMs are widely used in areas such as image recognition and indoor navigation.
[0066] A knowledge graph is a semantic network based on a graph structure. It describes knowledge resources and their interconnections. A knowledge graph consists of multiple nodes and edges connecting them. Nodes represent entities (such as people or places), concepts, or attributes, while edges represent relationships between nodes.
[0067] The visual language-based image recognition method, controller, robot, and medium provided in the embodiments of the present application are specifically illustrated through the following embodiments. First, the visual language-based image recognition method in the embodiments of the present application is described.
[0068] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0069] Figure 1 This is an optional flowchart of the image recognition method based on visual language provided in an embodiment of the present application. Figure 1 The method may include but is not limited to steps 101 to 106.
[0070] Step 101: Acquire the target user's item usage record information; wherein the item usage record information includes the target item name, item location, and the target user's user name;
[0071] Step 102: Generate a knowledge graph based on the item usage record information to obtain an auxiliary object finding knowledge graph; wherein the auxiliary object finding knowledge graph includes at least two element nodes and the relationship between the element nodes, and each element node corresponds to one of the target item name, item location, and user name;
[0072] Step 103: Graph encoding is performed on the auxiliary object-finding knowledge graph to obtain graph structured semantic features;
[0073] Step 104: Capture an image of the scene where the target user is located to obtain a current scene image, and perform visual encoding on the current scene image to obtain current image features;
[0074] Step 105: Perform attention fusion processing on the graph structured semantic features and the current picture features to obtain the graph-picture joint features;
[0075] Step 106, obtain the natural language search instruction containing the target item description information, and search for the item based on the natural language search instruction and the joint features of the atlas screen to obtain the item search status; wherein, the item search status represents whether the target item described by the target item description information exists in the current scene screen.
[0076] The beneficial effects of the embodiments of the present application include but are not limited to: by obtaining the target user's item usage record information, and generating an auxiliary item-finding knowledge graph based on the item usage record information, the auxiliary item-finding knowledge graph can characterize the relationship between the target user (such as the elderly) and the items he uses, such as the relationship between the target item name, the item location and the user name, so as to provide sufficient background knowledge for subsequent item searches. The auxiliary item-finding knowledge graph is encoded to obtain graph-structured semantic features, and the picture (that is, the current scene picture) collected from the scene where the target user is located is encoded to obtain the current picture features. Then, instead of directly searching for items based on the image of the scene where the target user is located, the graph-structured semantic features and the current picture features are first fused through the attention mechanism to form a graph-picture joint feature, so as to form a joint representation containing user habit knowledge and real-time scene information, and then the graph-picture joint feature and the item-finding natural language instructions are combined to perform item search. In this way, based on the association information contained in the auxiliary object-finding knowledge graph, we can more accurately understand the description of the object in the natural language instruction of the object-finding instruction (that is, the target object description information), as well as the semantic association between the target object description information and the image of the scene where the target user is located, and strengthen the semantic association between the target object description information and the personalized behavior pattern of the user in the specific scene, thereby reducing the semantic understanding error of the general visual language model in the specific user group scene due to the lack of personalized background knowledge, and improving the accuracy and reliability of object search in the specific scene where the target user is located.
[0077] In step 101 of some embodiments, the item usage record information refers to the record information generated by the target user in the process of using the target item. For example, in the elderly care assistance scenario of the medical and health scenario, the target user may be an elderly person (also known as an elderly user), and the target item may be an elderly care item. In this case, the item usage record information is a record of the elderly using the elderly care item, such as the record information of the elderly using a cane. Specifically, the item usage record information includes the name of the target item, the location of the item, and the user name of the target user. Among them, the location of the item may be the common location of the target item. The user name can be used to distinguish and identify the identity of the user.
[0078] In some embodiments, the target item name is the standard item name of the target item. For example, in the elderly care assistance scenario of the medical and health scenario, the target item name may include any of the following names: hearing aid, crutches, antihypertensive medication, blood pressure monitor, etc. The target item name may also include the names of other medicines, health monitoring equipment, or elderly care assistance devices, but is not limited to these. For another example, in the banking transaction scenario of the financial technology scenario, the target item name may include any of the following names: bank card, ID card, calculator, seal, etc.
[0079] In some embodiments, taking the elderly care assistance scenario of the medical and health scenario as an example, the target user's item usage record information can be obtained from multiple elderly care record source channels such as elderly care institutions, home care environments, health equipment catalogs, and elderly behavior records. For example, the name, location, and other information of the items in the elderly user's environment (such as elderly care institutions, home care environments) can be collected and combined with the elderly user's own name information to generate item usage record information. In another embodiment, the target user's item usage record information can also be obtained by other means, which is not limited to this.
[0080] In step 102 of some embodiments, the auxiliary object-finding knowledge graph refers to a knowledge graph used to assist in the object-finding task. It should be noted that the auxiliary object-finding knowledge graph includes at least two element nodes, and the relationship between the element nodes. Specifically, the relationship between the element nodes can be an edge connecting any two element nodes. Each element node corresponds to an item-related data in the item usage record information, such as any one of the target item name, item location and user name. For example, the data corresponding to multiple element nodes in the auxiliary object-finding knowledge graph may include "crutch" (target item name), "bedroom door" (item location) and "Grandma Zhang" (user name).
[0081] In some embodiments, the data format of the data corresponding to each element node may include text, for example, the target item name, item location, and user name are all text data. In another embodiment, the data format of the data corresponding to the element node may also include an image. For example, the element node may correspond to an image of the shape of the target item, such as a standard image of a "crutch", to support visual alignment. For another example, the element node may also correspond to a link to the image of the shape of the target item, such as a URL (Uniform Resource Locator), to reduce the storage resources occupied by the knowledge graph, thereby enabling the image recognition method to run on embedded devices with limited resources.
[0082] In some embodiments, the auxiliary object-finding knowledge graph can be continuously expanded and updated based on new data, such as new item usage record information, so as to adapt to the habits of different user groups in different specific scenarios and improve the personalization of the image recognition method.
[0083] In step 103 of some embodiments, the graph structured semantic features refer to the features encoded by the auxiliary object-finding knowledge graph. Figure 2As shown, the graph (i.e., the auxiliary object-finding knowledge graph) can be graph-encoded by a graph encoder. Specifically, the graph encoder can include a graph encoder of any one of the graph convolutional network (GCN) model and the graph attention network (GAT) model. In another embodiment, the structured semantic features of the graph can also be obtained by encoding in other ways, such as texting the auxiliary object-finding knowledge graph and inputting it into a text encoder for encoding, but is not limited to this.
[0084] In some embodiments, specifically, the graph structured semantic feature may be a sequence of features including multiple element nodes.
[0085] In step 104 of some embodiments, the current scene image is an image captured of the scene where the target user is located. For example, in a healthcare scenario involving elderly care assistance, the scene where the elderly user (i.e., the target user) is located may include a hospital, a nursing home, a home (e.g., a bedroom), or other elderly care scenes. The current scene image may be a video frame from a video captured in the elderly care scene where the elderly user is located.
[0086] In some embodiments, it should be noted that the current picture feature is a feature obtained by encoding the current scene picture. Figure 2 As shown, the current scene image can be visually encoded by a visual encoder. Specifically, the visual encoder can be a visual language model (VLM), such as the visual encoder (also called image encoder) of the CLIP-ViT model.
[0087] In some embodiments, specifically, the current picture feature corresponding to the current scene picture may be a sequence of block picture features corresponding to multiple image blocks in the current scene picture. For example, the current scene picture may be segmented into image blocks to obtain at least two image blocks, and then image feature extraction may be performed on each image block to obtain the block picture feature of each image block.
[0088] In step 105 of some embodiments, it should be noted that the graph-screen joint feature is a feature vector obtained by fusing the graph structured semantic feature and the current screen feature. Figure 2 As shown in the figure, a domain-aware adapter can be used to perform attention fusion processing on the graph structured semantic features and the current image features to obtain the graph-image joint features. Specifically, the domain-aware adapter can be a model based on an attention mechanism (such as a cross-attention mechanism), such as the Transformer model.
[0089] In some embodiments, a query can be generated based on the structured semantic features of the graph, and a key and value can be generated based on the current screen features. Attention calculation can be performed based on the query, key, and value to obtain a graph-screen joint feature. The graph-screen joint feature is used to indicate the correlation between each element node in the graph and each local area of the image (i.e., the current scene screen), so that the image area related to the graph in the current scene screen can be focused.
[0090] In step 106 of some embodiments, for example, Figure 2 As shown, the visual language model (VLM) can be used to search for objects by combining the features of text (i.e., natural language instructions for finding objects) and the atlas image, and output the object search status.
[0091] In some embodiments, the format of the find-thing natural language instruction can be a text format. The target item description information is the text information used to describe the target item in the find-thing natural language instruction. It should be noted that the target item description information may include any one or more of the following information: the name of the item, function, usage time, location, etc. For example, assuming that the find-thing natural language instruction C1 is "Where is the thing that I often use to amplify sound", then the target item description information D1 is "hearing aid". For another example, assuming that the find-thing natural language instruction C2 is "Find the blood pressure-lowering pills I take every morning", then the target item description information D2 is "blood pressure-lowering pills taken every morning".
[0092] In some embodiments, prompt information can be generated based on the item search situation, such as voice prompts or. For example, in the elderly care assistance scenario of the medical and health scenario, assuming that the target item is a crutch, and the item search situation indicates that there is a crutch in the current scene, then the prompt information "Crutch is placed at the bedroom door" can be generated based on the specific location of the target item in the current scene (such as the bedroom door). Assuming that the item search situation indicates that there is no crutch in the current scene, the prompt information "The target item is not seen in the current picture" can be generated.
[0093] In some embodiments, the image recognition method based on visual language can be applied to robots, such as object-finding and navigation robots. When the object-finding and navigation robot is located in the scene where the target user is located, the camera in the object-finding and navigation robot can be used to capture the current scene image to identify whether there is a target object (such as a crutch) in the current scene image, and then perform object navigation based on the object search situation. For example, if the object search situation indicates that the target object described by the target object description information exists in the current scene image, a navigation voice can be played to assist the target user in finding the object. For another example, if the target object does not exist, the object-finding and navigation robot can play a voice prompt that the object has not been identified.
[0094] In some embodiments, it should be noted that visual language models (VLMs) are primarily used for multimodal tasks, such as indoor navigation, assisted conversation, and image question-answering. However, current visual language models generally rely on large-scale, general sample data from the internet for training. The visual objects and language descriptions involved in this sample data are mostly related to everyday life scenes, resulting in the trained visual language models lacking the ability to adapt to specific groups and specific tasks. For example, in the field of elderly care assistance, especially when robots provide personalized services to the elderly, specific language expressions (such as instructions for finding objects) and specific items used by the elderly in their daily lives, such as commonly used medications, crutches, and hearing aids, are often involved. However, current visual language models rarely, if ever, are optimized for these specific scenario tasks (such as object search and navigation tasks). Due to the lack of such sample data during visual language model training, current visual language models also have difficulty understanding the semantic relationship between specific language expressions and the visual information of objects in specific scene images, resulting in difficulty in accurately recognizing images. For example, robots are prone to recognition failures, misunderstandings, or action errors when performing object recognition or navigation tasks, seriously affecting the usability and reliability of robots in elderly care assistance scenarios.
[0095] Based on this, the embodiments of this application conduct targeted domain adaptation and multimodal alignment optimization for specific scenarios such as elderly care scenarios to address the shortcomings of the image recognition method based on the general VLM model in its insufficient recognition ability in specific scenarios. Specifically, by combining the knowledge graph with the image, the image recognition method can significantly improve the recognition accuracy of uncommon items in specific scenarios (such as hearing aids, specific medicines, etc.), thereby improving the robot's response to search instructions involving such items.
[0096] In some embodiments, the domain-aware adapter (see the detailed description of step 105) can be trained based on the object search results and preset standard object presence conditions, such as the location of the object in the image. Specifically, the model training can be performed using a contrastive loss function, or other methods, without limitation.
[0097] See also Figure 3 ,In some embodiments, the item usage record information also includes the ,usage time and item function;
[0098] Step 102 may include but is not limited to steps 201 to 205:
[0099] Step 201: Generate an item element node based on the target item name to obtain an item element node;
[0100] Step 202: Generate a function element node based on the item function, obtain the function element node, and create a first edge connecting the function element node and the item element node;
[0101] Step 203: Generate a user element node based on the user name to obtain the user element node, and create a second edge connecting the user element node and the function element node;
[0102] Step 204: Generate a location element node based on the item location to obtain the location element node, and create a third edge connecting the location element node and the user element node;
[0103] Step 205: Generate a time element node based on the usage time to obtain the time element node, and create a fourth edge connecting the time element node and the location element node to generate an auxiliary object-finding knowledge graph.
[0104] The advantage of this embodiment is that corresponding element nodes are generated for the item name, item function, user name, item location and usage time respectively, and an item usage association is established between the item element node and the function element node through the first edge, a user usage behavior association is established between the function element node and the user element node through the second edge, a location association of the used item is established between the user element node and the location element node through the third edge, and a time regularity association of the item usage is established between the location element node and the time element node through the fourth edge. In this way, an auxiliary object-finding knowledge graph can be established based on the multi-dimensional information of the item and its association, providing sufficient background knowledge support for subsequent object search, thereby improving the accuracy and reliability of object search in specific scenarios, and better meeting the personalized object-finding needs of users (i.e., target users).
[0105] In step 201 of some embodiments, the item element node is an element node (referred to as node) corresponding to the target item name. For example, if the target item name is a crutch, the item element node may be a "crutch" node.
[0106] In step 202 of some embodiments, the function element node is the element node corresponding to the item function. For example, if the item function is to assist walking, the function element node may be the "assist walking" node. It should be noted that the first edge is the edge connecting the function element node and the item element node.
[0107] In step 203 of some embodiments, the user component node is a component node corresponding to the user name. The second edge is an edge used to connect the user component node and the function component node.
[0108] In step 204 of some embodiments, the location element node is the element node corresponding to the item location. For example, if the item location is the bedroom doorway, the function element node may be the "bedroom doorway" node. It should be noted that the third edge is the edge connecting the location element node and the user element node.
[0109] In step 205 of some embodiments, the time element node is the element node corresponding to the usage time. For example, if the usage time is used in the morning before going out, the time element node may be the "used in the morning before going out" node. It should be noted that the fourth edge is the edge used to connect the time element node and the location element node.
[0110] In some embodiments, it should be noted that at least two element nodes of the auxiliary object-finding knowledge graph include an item element node, a function element node, a user element node, a location element node, and a time element node; and the relationship between the element nodes includes a first edge, a second edge, a third edge, and a fourth edge. For example, the auxiliary object-finding knowledge graph may include: "crutch" (name of the target item) - "assistance in walking" (item function) - "Grandma Zhang" (user name) - "bedroom door" (item location) - "used before going out in the morning" (use time).
[0111] See also Figure 4 ,In some embodiments, there are at least two location element nodes;
[0112] After step 205, the image recognition method based on visual language may further include but is not limited to steps 301 to 303:
[0113] Step 301: Count the number of position element nodes with the same item position among at least two position element nodes connected to the same item element node, and obtain the position occurrence frequency of each item position.
[0114] Step 302: determine the item location with the highest location occurrence frequency as the target item location;
[0115] Step 303: Update the auxiliary object-finding knowledge graph according to the location of the target item.
[0116] The advantage of this embodiment is that by counting the number of location element nodes with the same item location among the multiple location element nodes connected to the same item element node, the frequency of occurrence of each item location is obtained, and the item location with the highest location occurrence frequency is determined as the target item location. This allows the common locations of the same item to be counted, which helps in subsequent item searches. Then, the auxiliary object search knowledge graph is updated based on the target item location, which can dynamically optimize the graph structure. For example, by eliminating low-frequency location nodes, that is, location element nodes corresponding to locations other than the target item location, the interference of low-frequency location information on item searches can be reduced, the probability of misjudgment caused by the redundancy of historical location data can be reduced, and the reliability of item searches can be improved.
[0117] In step 301 of some embodiments, at least two location element nodes connected to the same item element node may correspond to the same or different item locations. For example, assuming the item element node is a crutch node, of the 20 location element nodes connected to the crutch node, 17 of these location element nodes correspond to the first item location, such as the bedroom doorway; and 3 of these location element nodes correspond to the second item location, such as the kitchen doorway. Therefore, the first item location has a frequency of 17, and the second item location has a frequency of 3.
[0118] In step 302 of some embodiments, the target item location is the item location with the highest position occurrence frequency among all item locations. As in the above example, the position occurrence frequency of the first location (17) is greater than the position occurrence frequency of the second location (3), so the target item location is the first location.
[0119] In step 303 of some embodiments, a location element node other than the location element node corresponding to the target item's location may be deleted from the multiple location element nodes connected to the same item element node in the auxiliary object-finding knowledge graph. Similarly to the above example, of the 20 location element nodes connected to the crutch node, the three location element nodes corresponding to the second location and the edges connected to these nodes may be deleted, thereby retaining only the location element nodes corresponding to the target item's location (i.e., the common location), thereby improving the reliability of object search.
[0120] See also Figure 5 In some embodiments, the step 106 of searching for an item based on the natural language command and the combined features of the atlas screen to obtain the item search status may include but is not limited to steps 401 to 404:
[0121] Step 401: In response to the natural language instruction for finding an object, perform object name recognition on the description information of the target object to obtain a first object name recognition result;
[0122] Step 402: If the first item name recognition result indicates that the target item description information does not contain the item name, feature extraction is performed on the target item description information to obtain a fuzzy description feature of the item;
[0123] Step 403: Determine the predicted item name based on the item's fuzzy description features and the graph's structured semantic features, and perform feature extraction on the predicted item name to obtain clear item description features.
[0124] Step 404: perform item feature recognition based on the item's clear description features and the combined features of the atlas image to obtain item search status.
[0125] The advantage of this embodiment is that, for the search scenario where the name of the item is missing from the description information contained in the search instruction, the auxiliary search knowledge graph is used to map the fuzzy description in the search instruction to the item name. Specifically, when it is identified that the target item description information of the search natural language instruction does not contain the item name, the feature in the target item description information, that is, the fuzzy description feature of the item, is extracted, and the predicted item name is determined by combining the feature with the structured semantic feature of the graph. This allows the fuzzy description of the item in the instruction (such as function, time, etc.) to be mapped to the predicted item name based on the item information association path contained in the auxiliary search knowledge graph (such as the "item-function-time" association path). Then, the item search status is obtained based on the clear description feature of the item corresponding to the predicted item name and the joint features of the graph screen. This process uses the semantic association capability of the knowledge graph to perform fuzzy search based on fuzzy instructions that lack item names, effectively solving the problem of incomplete instructions caused by the expression habits of specific user groups (such as the elderly), and improving the accuracy, reliability and compatibility of item search in specific scenarios.
[0126] In step 401 of some embodiments, the first item name recognition condition may indicate whether the item name exists in the target item description information. For example, assuming the natural language instruction to find an item is "find my walking stick," the target item description information is "walking stick," and the first item name recognition condition indicates that the item name exists in the target item description information. For another example, assuming the natural language instruction to find an item is "find something I often use when I go out for a walk," the target item description information is "something I often use when I go out for a walk," and the first item name recognition condition indicates that the item name does not exist in the target item description information.
[0127] In step 402 of some embodiments, the fuzzy description feature of the item refers to a text feature obtained by extracting features from the description information of the target item without the item name.
[0128] In step 403 of some embodiments, the predicted item name is the item name stored in the auxiliary object-finding knowledge graph corresponding to the structured semantic features of the graph, and the predicted item name is similar to the textual semantic features of the item's fuzzy description features. It should be noted that the clear item description features refer to textual features extracted from the predicted item name. Because the clear item description features are features of the predicted item name, subsequent item searches can be performed based on the features of the item name, thereby improving the accuracy of item searches when the search instructions are ambiguous.
[0129] In some embodiments, in step 404, the visual language model can be used to identify the item's features based on the item's clear description features and the combined features of the image. The item search status can also be obtained through other methods, not limited to these.
[0130] See also Figure 6 ,In some embodiments, the graph structured semantic features include at least two ,graph node features, each graph node feature corresponds to an element node in the ,assisted object-finding knowledge graph;
[0131] Step 403 may include but is not limited to steps 501 to 504:
[0132] Step 501: Perform feature correlation analysis based on the item fuzzy description features and each graph node feature to obtain the node description feature correlation;
[0133] Step 502: Determine the element node corresponding to the graph node feature with the maximum node description feature association as the target association node;
[0134] Step 503: If the target associated node is an item element node, the target item name corresponding to the target associated node is determined as the predicted item name; wherein the item element node is an element node corresponding to the target item name;
[0135] Step 504: If the target associated node is not an item element node, the target item name corresponding to the item element node connected to the target associated node is determined as the predicted item name.
[0136] The advantage of this embodiment is that by calculating the correlation between the fuzzy description feature of the item and the feature of each graph node, the element node corresponding to the graph node feature with the highest correlation is screened out, that is, the target association node, so that the name of the item described by the fuzzy description feature of the item can be found based on the correlation between the nodes of the auxiliary object-finding knowledge graph. Specifically, if the target association node corresponds to the name of the item, that is, the target association node is an item element node, then the item name corresponding to the target association node is directly determined as the predicted item name. If the target association node is not an item element node, the item element node to which the target association node is connected is found, and then the predicted item name is obtained. In this way, the node connection relationship in the knowledge graph can be used to deduce the item name, thereby realizing fuzzy search, enhancing the anti-interference ability of the user's non-standardized expression, and improving the accuracy and reliability of item search.
[0137] In step 501 of some embodiments, the node description feature correlation is the feature correlation between the item's fuzzy description feature and each graph node feature. In some embodiments, the node description feature correlation may be a text feature correlation. For example, the node description feature correlation may be obtained through analysis using a natural language model, an attention model, or the like.
[0138] In step 502 of some embodiments, the target association node is the element node corresponding to the graph node feature with the maximum node description feature association degree.
[0139] In some embodiments, in step 503, if the target association node corresponds to an item name, that is, if the target association node is an item element node, the item name corresponding to the target association node is directly determined as the predicted item name. The meaning and function of the item element node can be found in the detailed description of step 201 above and will not be repeated here.
[0140] In some embodiments, in step 504, if the target association node is not an item element node, the item element node to which the target association node is connected is found to obtain the predicted item name. This allows the node connection relationship in the knowledge graph to be used to deduce the item name, thereby achieving fuzzy search.
[0141] See also Figure 7 ,In some embodiments, the object-finding natural language instruction is a target object-finding text instruction;
[0142] The acquisition of the natural language instruction for finding the object containing the description information of the target object in step 106 may include but is not limited to steps 601 to 605:
[0143] Step 601, obtaining the target user's voice command to find the item;
[0144] Step 602: Convert the object-finding voice command to text to obtain an original object-finding text command; wherein the original object-finding text command includes original object description information;
[0145] Step 603: perform item name recognition on the original item description information to obtain a second item name recognition result;
[0146] Step 604: If the second item name recognition condition indicates that the item name exists in the original item description information, the item name in the original item description information is determined as the original item name;
[0147] Step 605 , rewrite the original object-finding text instruction according to the association between the original object name and the object name description information in the preset object name knowledge database to obtain the target object-finding text instruction.
[0148] The advantage of this embodiment is that, considering that the language expression of the target user is often ambiguous and colloquial, the name of the item in his language expression may not be the accurate name of the item. To address the above problem, the original instruction is rewritten to standardize it. Specifically, by converting the target user's voice instruction to find the item into text, and when there is an item name (that is, the original item name) in the item description of the original text instruction to find the item, the original text instruction to find the item is rewritten according to the association between the original item name and the item name description information in the preset item name knowledge database, so as to rewrite the ambiguous instruction into an instruction containing a standard item name. For example, the user's colloquial expression can be rewritten into a standard item name with the same or similar semantics, so that the instruction text input in the subsequent item search stage is consistent with the item name stored in the knowledge graph, thereby improving the ability to process non-standardized instructions, and then improving the accuracy of item search in specific scenarios.
[0149] In step 601 of some embodiments, the voice command for finding an object is a voice command of the target user. For example, the voice command for finding an object of the target user can be obtained through the voice collection module of the robot. Taking the elderly care assistance scenario in the medical and health scenario as an example, the elderly user can say "find my ear machine" to the robot, and the robot can collect the voice and obtain the voice command for finding an object. It should be noted that the language expression of the target user (such as the elderly user) is often vague and colloquial. The name of the item in its language expression may not be the exact name of the item, but an alias of the item, or a colloquial expression caused by personal habit, such as calling a "hearing aid" an "ear machine".
[0150] In step 602 of some embodiments, the original find-the-item text instruction is a text converted from a find-the-item voice instruction. Specifically, the find-the-item voice instruction can be converted to text using a pre-trained speech recognition model. It should be noted that the original item description information refers to the item description information in the original find-the-item text instruction. For example, if the original find-the-item text instruction is "Find my commonly used earphone," the original item description information is "commonly used earphone."
[0151] In step 603 of some embodiments, the second item name identification condition may indicate whether the item name exists in the original item description information. It should be noted that when the second item name identification condition indicates that the item name exists in the original item description information, the item name in the original item description information may be an item alias. Furthermore, the item name may also be other types of colloquial, non-standardized item names, such as item names in a dialect.
[0152] In step 604 of some embodiments, the original item name is the item name in the original item description information. In the same example as above, assuming that the original item description information is "common earphones", the original item name is "earphones".
[0153] In step 605 of some embodiments, the item name knowledge database is a database used to store multiple names of the same item (including aliases and standard names). In some embodiments, the item name knowledge database may include multiple item name descriptions. It should be noted that each item name description represents a mapping relationship between an alias and a standard name for the same item.
[0154] It should be noted that the target object search text instruction is a text instruction obtained by rewriting the original object search text instruction. The target object search text instruction contains a standardized item name, which is the standard item name below. In the same example as above, assuming that the original item name is "ear machine", the standard item name in the target object search text instruction can be "hearing aid". In some embodiments, it should be noted that the language expressions of elderly users are often vague and colloquial, such as "take the medicine I take every day", "find the ear machine I usually wear", etc. Based on this, the embodiment of the present application can convert the language expressions of elderly users into standard instructions that can be recognized by the robot, replace the colloquial and vague language expressions, and let the rewritten instructions retain the original meaning, thereby improving the naturalness and fault tolerance of the interaction between the robot and the elderly users.
[0155] See also Figure 8 In some embodiments, the item name description information includes an item alias and a standard item name of the same item;
[0156] Step 605 may include but is not limited to steps 701 to 702:
[0157] Step 701: Perform text semantic similarity analysis based on the original item name and the item alias to obtain item name similarity;
[0158] Step 702 : Perform text fusion on the standard item name corresponding to the item alias with the greatest item name similarity and the preset object-finding text instruction template to generate a target object-finding text instruction.
[0159] The advantage of this embodiment is that by analyzing the textual semantic similarity between the original item name and the item alias to obtain item name similarity, the standard item name corresponding to the item alias with the greatest item name similarity is screened out, thereby finding the standard item name of the item based on the original item name in the instruction. The standard item name is then embedded into the search text instruction template to generate the target search text instruction, thereby rewriting the colloquial instruction containing the item alias into an instruction containing the standard item name. This reduces the impact of entity reference ambiguity caused by differences in expression among user groups and thereby improves the accuracy of item search in specific scenarios.
[0160] In step 701 of some embodiments, the item name similarity refers to the textual semantic similarity between the original item name and the item alias.
[0161] In step 702 of some embodiments, it should be noted that the item name similarity can be used to analyze which item the original item name is an alias for. For example, assuming that the original item name is "ear machine", and item name description information 1 includes: "hearing aid" (standard item name) - "ear machine" (item alias); item name description information 2 includes: "nitroglycerin tablets" (standard item name) - "medicine for treating angina pectoris" (item alias). Then, among the multiple item name description information, the item alias "ear machine" in item name description information 1 has the greatest similarity with the above-mentioned original item name, and the standard item name "hearing aid" corresponding to the item alias "ear machine" is the standard item name obtained after screening. This standard item name and the original item name are different references to the same item entity.
[0162] The present application also provides a controller comprising a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned visual language-based image recognition method. In some embodiments, the controller may comprise a control chip in a robot. In another embodiment, the controller may also comprise another intelligent terminal.
[0163] See also Figure 9 , Figure 9 The hardware structure of a controller of another embodiment is shown. The controller includes:
[0164] The processor 901 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0165] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called by the processor 901 to execute the visual language-based image recognition method of the embodiments of this application.
[0166] Input / output interface 903, used to implement information input and output;
[0167] Communication interface 904, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0168] Bus 905 , which transmits information between various components of the device (e.g., processor 901 , memory 902 , input / output interface 903 , and communication interface 904 );
[0169] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .
[0170] An embodiment of the present application also provides a robot, which includes the above-mentioned controller.
[0171] It should be noted that the image recognition method based on visual language can be applied to robots. In one embodiment, the robots include, but are not limited to, object-finding and navigation robots, elderly care assistance robots, and other types. The robots may also include robots with other functions or types, without limitation.
[0172] The specific implementation of the robot is basically the same as the specific embodiment of the above-mentioned image recognition method based on visual language, and will not be repeated here.
[0173] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, which implements the above-mentioned visual language-based image recognition method when executed by a processor.
[0174] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0175] It should be noted that the non-Company's software tools or components that appear in the embodiments of this application are merely examples and do not represent actual use.
[0176] The embodiments described in this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0177] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0178] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0179] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0180] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0181] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0182] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0183] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0184] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0185] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0186] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A method for image recognition based on visual language, characterized in that: The method comprises: Obtaining item usage record information of a target user; wherein the item usage record information includes the name of the target item, the location of the item, and the user name of the target user; Generating a knowledge graph based on the item usage record information to obtain an auxiliary object-finding knowledge graph; wherein the auxiliary object-finding knowledge graph includes at least two element nodes and a relationship between the element nodes, each element node corresponding to one of the target item name, the item location, and the user name; Performing graph encoding on the auxiliary object-finding knowledge graph to obtain graph structured semantic features; Capturing an image of the scene where the target user is located to obtain a current scene image, and visually encoding the current scene image to obtain current image features; Performing attention fusion processing on the graph structured semantic features and the current picture features to obtain a graph-picture joint feature; Obtain a natural language search instruction containing target item description information, and perform item search based on the natural language search instruction and the joint features of the atlas screen to obtain an item search status; wherein, the item search status represents whether the target item described by the target item description information exists in the current scene screen.
2. The method according to claim 1, characterized in that The item usage record information also includes usage time and item function; The generating of a knowledge graph based on the item usage record information to obtain an auxiliary object-finding knowledge graph includes: Generate an item element node based on the target item name to obtain an item element node; Generating a function element node based on the item function to obtain a function element node, and creating a first edge connecting the function element node and the item element node; Generating a user element node based on the user name to obtain a user element node, and creating a second edge connecting the user element node and the function element node; Generating a location element node based on the item location to obtain a location element node, and creating a third edge connecting the location element node and the user element node; A time element node is generated based on the usage time to obtain a time element node, and a fourth edge connecting the time element node and the location element node is created to generate the auxiliary object-finding knowledge graph.
3. The method according to claim 2, characterized in that There are at least two position element nodes; After generating a time element node based on the usage time to obtain the time element node and creating a fourth edge connecting the time element node and the location element node, the method further includes: In at least two of the position element nodes connected to the same item element node, counting the number of the position element nodes having the same item position, to obtain the position occurrence frequency of each item position; Determine the item position with the highest position occurrence frequency as the target item position; The auxiliary object-finding knowledge graph is updated according to the location of the target item.
4. The method according to any one of claims 1 to 3, characterized in that The item search is performed based on the natural language instruction for finding the item and the combined features of the atlas screen to obtain the item search status, including: In response to the object-finding natural language instruction, performing object name recognition on the target object description information to obtain a first object name recognition result; If the first item name recognition result indicates that the item name does not exist in the target item description information, feature extraction is performed on the target item description information to obtain a fuzzy description feature of the item; Determine a predicted item name based on the fuzzy description features of the item and the structured semantic features of the graph, and perform feature extraction on the predicted item name to obtain a clear description feature of the item; The item feature recognition is performed based on the clear description features of the item and the combined features of the atlas screen to obtain the item search situation.
5. The method according to claim 4, characterized in that The graph structured semantic feature includes at least two graph node features, each of which corresponds to one of the element nodes in the auxiliary object-finding knowledge graph; Determining the predicted item name based on the fuzzy description features of the item and the structured semantic features of the graph includes: Perform feature correlation analysis based on the fuzzy description features of the item and the features of each graph node to obtain the node description feature correlation; Determine the element node corresponding to the graph node feature with the largest node description feature association degree as the target association node; If the target associated node is an item element node, determining the target item name corresponding to the target associated node as the predicted item name; wherein the item element node is the element node corresponding to the target item name; If the target associated node is not the item element node, the target item name corresponding to the item element node connected to the target associated node is determined as the predicted item name.
6. The method according to any one of claims 1 to 3, characterized in that The natural language instruction for finding an object is a text instruction for finding an object; The step of obtaining the natural language instruction for finding an object including description information of the target object includes: Obtaining the target user's voice command for finding the object; Converting the object-finding voice command into text to obtain an original object-finding text command; wherein the original object-finding text command includes original object description information; Performing item name recognition on the original item description information to obtain second item name recognition information; If the second item name recognition condition indicates that the item name exists in the original item description information, determining the item name in the original item description information as the original item name; According to the association between the original object name and the object name description information in the preset object name knowledge database, the original object-finding text instruction is rewritten to obtain the target object-finding text instruction.
7. The method according to claim 6, characterized in that The item name description information includes an item alias and a standard item name of the same item; The step of rewriting the original object-finding text instruction according to the association between the original object name and the object name description information in the preset object name knowledge database to obtain the target object-finding text instruction includes: Performing text semantic similarity analysis based on the original item name and the item alias to obtain item name similarity; The standard item name corresponding to the item alias with the greatest item name similarity and a preset object-finding text instruction template are subjected to text fusion to generate the target object-finding text instruction.
8. A controller, characterized in that: The controller includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
9. A robot, characterized in that: The robot comprises the controller according to claim 8.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Robot question and answer method for article search
CN113516055A
Indoor navigation method, indoor navigation device, indoor navigation equipment and storage medium
CN113984052A
Large language model training method and device, reasoning method and device, equipment and storage medium
CN118673325A
Pet-searching geographic information matching method based on machine learning and natural language processing
CN120180157A
Real time salient object detection in images and videos
US20230410465A1
Cited By
Graph structure-based intelligent decision-making method, apparatus and device for body, and medium
CN121456805A
Graph structure-based embodied intelligent decision-making method, device, equipment and medium
CN121456805B