Robot target navigation method, device, equipment and medium

By combining the family knowledge graph and the visual language model, accurate target navigation instructions are generated, which solves the problem of processing ambiguous instructions for the elderly and improves the accuracy of robot navigation and service efficiency.

CN120686826APending Publication Date: 2025-09-23PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510827858.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In existing technologies, the visual language model (VLM) system has difficulty processing ambiguous instructions from the elderly, resulting in low efficiency in robot navigation and task execution, and inability to effectively serve the elderly. In addition, traditional models fail to flexibly respond to scene changes and user habits.

Method used

By constructing a family knowledge graph and combining it with the preset language model and visual language model, target navigation instructions are generated, the information expansion of fuzzy instructions and the association of environmental images are achieved, and accurate target navigation instructions are generated.

Benefits of technology

It improves the accuracy of robot target recognition and the reliability of spatial positioning, and enhances the service efficiency and intelligence level of indoor assistive robots for the elderly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120686826A_ABST
    Figure CN120686826A_ABST
Patent Text Reader

Abstract

The invention discloses a robot target navigation method, device and equipment and a medium. The method comprises the following steps: constructing a basic language instruction through a preset language model according to a fuzzy instruction input by a user and a preset family knowledge graph; performing information expansion through the preset language model according to the basic language instruction and preset reference information, and generating a target language instruction; generating a target navigation instruction from the target language instruction and a collected environment image through a preset visual language model; and controlling the robot to move according to the target navigation instruction. The method can be applied to robot navigation scenes in the business fields of medical treatment, health, old-age care and the like. By generating the target navigation instruction to control the robot to move, the problem that an indoor auxiliary robot with a visual language model cannot efficiently provide services for old people can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robot control and can be applied to fields such as medical care, health care, and elderly care. In particular, it relates to a robot target navigation method, device, equipment, and medium. Background Art

[0002] With the aging population, integrated financial and elderly care service models are emerging, aiming to provide comprehensive care for seniors. Intelligent elderly care assistive devices are key to improving service quality, and indoor assistive robots equipped with visual language models (VLMs) are attracting significant attention. However, existing technologies present numerous challenges. Some elderly individuals experience cognitive and language impairments, leading to ambiguous commands. For example, phrases like "Help me find that," which VLM systems, which rely on precise language and visual matching, struggle to process, often leading to navigation and task execution errors. Most navigation systems view commands in isolation, failing to integrate information such as room layout and past interactions, resulting in low efficiency for repetitive robot operations. Traditional VLM models have shortcomings in semantic understanding, making it difficult to identify semantic associations between objects and insufficiently reasoning about implicit commands. Furthermore, these systems use static prompt templates, failing to account for elderly habits, home layout, and changing scenarios, resulting in inflexible object recognition strategies. These factors hinder the efficient and effective use of indoor assistive robots with existing technology. Summary of the Invention

[0003] Embodiments of the present invention provide a robot target navigation method, apparatus, device, and medium, aiming to solve the problem in the prior art that indoor assistive robots equipped with visual language models cannot efficiently provide services for the elderly.

[0004] In a first aspect, an embodiment of the present invention provides a robot target navigation method, which includes: constructing basic language instructions through a preset language model based on fuzzy instructions input by a user and a preset family knowledge graph; expanding information through the preset language model based on the basic language instructions and preset reference information to generate target language instructions; generating target navigation instructions from the target language instructions and the collected environmental images through a preset visual language model; and controlling the robot to move according to the target navigation instructions.

[0005] In the second aspect, an embodiment of the present invention also provides a robot target navigation device, which includes: a construction unit, which is used to construct basic language instructions through a preset language model based on fuzzy instructions input by a user and a preset family knowledge graph; an expansion unit, which is used to expand information through the preset language model based on the basic language instructions and preset reference information to generate target language instructions; a generation unit, which is used to generate target navigation instructions through a preset visual language model based on the target language instructions and the collected environmental images; and a control unit, which is used to control the robot to move according to the target navigation instructions.

[0006] In a third aspect, an embodiment of the present invention further provides a computer device, which includes a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above method when executing the computer program.

[0007] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the above method can be implemented.

[0008] Embodiments of the present invention provide a robot target navigation method, apparatus, device, and medium. The method includes: constructing basic language instructions based on fuzzy instructions input by a user and a preset family knowledge graph using a preset language model; expanding the basic language instructions and preset reference information using the preset language model to generate target language instructions; combining the target language instructions with a captured environmental image using a preset visual language model to generate target navigation instructions; and controlling the robot to move according to the target navigation instructions. In this embodiment of the present invention, the fuzzy instructions input by the user are constructed into basic instructions that the robot can understand using a family knowledge graph and a language model to avoid navigation errors caused by incomplete user input. The basic instructions are expanded using a language model and reference information to integrate dynamic information, enabling the robot to cope with complex scenarios. The target language instructions are then associated and integrated with real-time environmental images using a visual language model to generate target navigation instructions, effectively compensating for the spatial uncertainty of pure language instructions and enabling the robot to cope with dynamic changes in the real environment. By converting the fuzzy language input by users (elderly users) into machine instructions for visual language models, the accuracy of target recognition and the reliability of spatial positioning can be improved, so that indoor assistive robots can efficiently provide accurate services to users such as the elderly. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0010] Figure 1 A schematic diagram of a flow chart of a robot target navigation method provided by an embodiment of the present invention;

[0011] Figure 2 A schematic diagram of a sub-process of a robot target navigation method provided by an embodiment of the present invention;

[0012] Figure 3A schematic diagram of a sub-process of a robot target navigation method provided by an embodiment of the present invention;

[0013] Figure 4 A schematic diagram of a sub-process of a robot target navigation method provided by an embodiment of the present invention;

[0014] Figure 5 A schematic diagram of a sub-process of a robot target navigation method provided by an embodiment of the present invention;

[0015] Figure 6 A schematic diagram of a sub-process of a robot target navigation method provided by an embodiment of the present invention;

[0016] Figure 7 A schematic diagram of a sub-process of a robot target navigation method provided by an embodiment of the present invention;

[0017] Figure 8 A schematic block diagram of a robot target navigation device provided by an embodiment of the present invention;

[0018] Figure 9 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0020] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0021] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0022] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0023] See also Figure 1, Figure 1 A flow chart of the robot target navigation method provided in an embodiment of the present invention. The robot target navigation method in this embodiment can be applied to robots, especially elderly care robots. Its application not only greatly improves the level of intelligence in elderly care services, but also brings unprecedented convenience and safety to the daily lives of the elderly. In specific application scenarios, the elderly care robot can perceive the surrounding environment in real time and accurately identify obstacles and voice commands of the elderly by integrating advanced sensor arrays, high-definition cameras and intelligent algorithms. Combined with the target navigation method of this embodiment, the robot can autonomously convert the fuzzy language input by the elderly user into machine instructions for the visual language model to improve the accuracy of target recognition and the reliability of spatial positioning, so that the indoor auxiliary robot can efficiently provide accurate services to users such as the elderly.

[0024] Figure 1 FIG. 1 is a flow chart of a robot target navigation method provided by an embodiment of the present invention. As shown in the figure, the method includes the following steps S110-S140.

[0025] S110. Construct basic language instructions through a preset language model according to the fuzzy instructions input by the user and the preset family knowledge graph.

[0026] In this embodiment, fuzzy instructions refer to instructions expressed by the user in natural language that are not precise or specific enough. Such instructions may contain ambiguity, omit key information, or use non-technical terminology. For example, instructions such as "I want something hot to drink" or "Dim the lights." The user is the user served by the robot, in this embodiment, the elderly user served by the elderly care robot. The preset family knowledge graph is a structured knowledge base that stores entities, relationships, and attributes related to the home environment, user habits, device functions, and so on. The language model is a statistical or deep learning-based model that can understand and generate natural language, such as BERT, GPT, or T5. Based on the fuzzy instruction input by the user and the preset family knowledge graph, the preset language model is used to construct basic language instructions. Specifically, upon receiving the fuzzy instruction, the elderly care robot performs a preliminary analysis of it using the preset language model to extract keywords and potential intent. For example, if the fuzzy instruction is "Bring me what I drink every morning," the user's intent is determined to be to retrieve an item. The family knowledge graph is queried to determine that the candidate item is milk, and the basic language instruction "Get milk" is generated. By constructing basic language instructions based on fuzzy instructions, the robot can efficiently convert the user's fuzzy instructions into executable operations.

[0027] In one embodiment, if Figure 2 As shown, the step S110 also includes steps S1101-S1102 before the step S110.

[0028] S1101, obtaining object information, placement information, and usage information through user's personalized input and the environment image;

[0029] S1102. Associating the object information, the placement information, and the usage information according to the semantic association and spatial relationship of the objects to generate the preset family knowledge graph.

[0030] In this embodiment, the user's personalized input refers to personalized information related to the home environment provided by the user through voice, text, or other interactive methods. For example, it includes input information such as the user's description of household items (e.g., "This vase is a wedding anniversary gift"), usage habits (e.g., "I usually turn on the table lamp in the living room at 7 p.m."), and preferences (e.g., "I like warm lighting in my bedroom"). The environmental image is visual information of the home environment captured by the robot through a camera or other image acquisition device, including room layout, furniture placement, and item locations. Object information, placement information, and usage information are obtained through the user's personalized input and the environmental image. Object information includes commonly used items such as medicine boxes, remote controls, and water cups. Placement information refers to the placement of items, such as bedside tables and living room tables. Usage information refers to usage scenarios and times, such as taking morning medicine, watching TV, and drinking water. Information such as family members and room layout can also be determined based on information such as the user's personalized input. The object information, placement information, and usage information are associated based on the semantic associations and spatial relationships of the objects to generate the preset family knowledge graph. Specifically, semantic associations are connections between objects based on semantic aspects such as their function, usage, and category. For example, in a home environment, there is a semantic association between a TV and a remote control because the remote control is used to control the TV. Spatial relationships are the relative positions of objects in physical space. For example, in a living room, the sofa may be located in front of the TV, and the coffee table may be located between the sofa and the TV. The preset family knowledge graph is composed of nodes and edges, with object information, placement information, and usage information as nodes and edges representing semantic associations or spatial co-occurrence relationships between items. For example, the preset family knowledge graph may include: (morning medicine, located on the bedside table), (remote control, used for watching TV), etc. By constructing a family knowledge graph, the elderly care robot is provided with rich contextual information, enabling it to more accurately understand user commands, provide personalized services, and achieve more intelligent navigation and interaction in the home environment.

[0031] In one embodiment, if Figure 3 As shown, step S110 includes steps S111-S112.

[0032] S111, converting the fuzzy instruction into text content, performing semantic analysis based on the text content and the preset family knowledge graph, and determining user intention and candidate items;

[0033] S112: Construct the basic language instruction according to the user intention and the candidate through a preset prompt framework in the preset language model.

[0034] In this embodiment, the user intention is the goal that the user hopes to achieve through the instruction, and the candidate is an entity related to the user intention. The preset language model is the language model used in the visual language model. A dynamic prompt construction module is set in the preset language model, which includes a template layer, and the template layer provides a basic preset prompt framework. The fuzzy instruction is converted into text content, and semantic analysis is performed based on the text content and the preset family knowledge graph to determine the user intention and the candidate. Specifically, the fuzzy language instruction input by the elderly and other users is converted into text content through natural language processing and other technologies, and the text is preprocessed to remove noise (such as interjections, repeated words) and unify the format (such as case normalization). The text content is semantically analyzed and intent recognized through a classification model (such as BERT, TextCNN) and the preset family knowledge graph to determine the user's intention and the candidate. , and the user intention and candidate objects are input into the template layer of the dynamic prompt construction module in the preset language model, and the basic language instructions are constructed through the corresponding prompt framework. Specifically, the input of the template layer is: user intention type (such as "pick up", "search", "view", etc.), candidate objects (such as "milk", "remote control", "teacup"), and the preset prompt framework includes: "Please locate {item name} in {place}", "Find {item name} in {room}", "Help me get {item name}, it may be in {area}", "Find {item name} that is commonly used at {time}, it may be in {location}", etc. For example, the fuzzy instruction is "Bring what I drank that morning", and the user intention is identified as "pick up", and the candidates are "milk" and "cup", then the output basic instruction is: "Please locate the milk and cup in the kitchen". By efficiently converting the fuzzy instructions input by elderly users into executable basic language instructions, the elderly care robot can achieve more intelligent and personalized home services.

[0035] S120 : Expand information using the preset language model according to the basic language instruction and preset reference information to generate a target language instruction.

[0036] In this embodiment, the basic language instruction is a preliminary structured instruction generated based on the user's fuzzy instruction, the home knowledge graph, and semantic analysis. It clarifies the user's core intent and the target of the operation, but may lack complete parameters or contextual details. Therefore, information expansion is required to generate the final target language instruction. Specifically, the preset reference information is pre-stored structured or unstructured data related to the home environment, device functions, user habits, etc., which is used to supplement the missing information in the basic instruction. For example, contextual information such as historical interaction records, the current time (evening), and the elderly person's preference (usually watching TV in front of the sofa) is used. The preset language model includes a dynamic prompt construction module, which includes a rule layer. Through this rule layer, the information is expanded to generate the target language instruction. Specifically, the template is detailed and modified based on the current time, spatial context, user preferences, home layout, and other information. For example, the basic instruction "Adjust the brightness of the 'living room main light' to 50%" is expanded into the target language instruction "Adjust the brightness to 50% with a transition time of 2 seconds, enable night mode, and require user confirmation." By combining basic language instructions with preset reference information, smarter and more humane target language instructions are generated, significantly improving the interactive experience and execution efficiency of the elderly care robot.

[0037] In one embodiment, if Figure 4 As shown, step S120 includes steps S121-S122.

[0038] S121: Analyze the basic language instruction in combination with the preset reference information through a preset rule engine of the preset language model to generate supplementary information, wherein the preset reference information includes family structure, current time, and user history information;

[0039] S122: Expand the basic language instruction according to the supplementable information to generate the target language instruction.

[0040] In this embodiment, the preset language model includes a dynamic prompt construction module, and analysis and information expansion are performed through the rule layer in this module. Specifically, the main function of the rule layer is to complete the details and semantic modification of basic instructions based on dynamic context (such as time and space) and user personalized data (such as preferences and historical behavior), thereby generating more accurate and executable target instructions. Its core logic includes: context awareness: combining dynamic information such as current time, spatial structure, and user habits to understand the implicit requirements of the instruction; knowledge fusion: matching the most likely candidates and operation scenarios through the family knowledge graph and user history; rule-driven: generating supplementary parameters based on preset rules (such as time trigger rules and user preference rules). The basic language instruction is combined with the preset reference information and analyzed by the preset rule engine of the preset language model to generate supplementary information. Specifically, the basic language instruction is combined with parameter information such as family structure, current time, and user history information and analyzed by the rule engine to generate supplementary information. For example, the user inputs "Bring me what I drank that morning." If the user's intent for a basic command is to retrieve an item but the candidate item and location are not explicitly specified, the rule engine performs inference based on pre-defined reference information (time, space, history, and user preferences). Specifically, the input "morning" triggers the time rule: "morning" is associated with the breakfast scene, prioritizing beverage items. Semantic analysis of "drinks" matches beverage entities in the family knowledge graph (milk, soy milk, coffee). Based on user history (milk is often consumed in the morning) and preferences (elderly people tend to drink milk), "milk" is prioritized. Spatial semantics are used to query the location: milk's common location: According to historical records, milk is often placed on the "kitchen table." The location parameter is generated: "on the kitchen table." The item is then refined. For example, the specific form of milk: According to the family knowledge graph, milk may exist in the form of a "milk bottle." The item parameter is generated: "milk bottle." The rule engine converts the inference result into supplementary information, such as: "location: on the kitchen table," and the candidate item is a milk bottle. This supplementary information is incorporated into the basic command to generate a complete target language instruction, such as "Please locate the milk bottle on the kitchen table." By using a preset rule engine with a preset language model to complete the context and modify the semantics of basic instructions, the robot can convert ambiguous natural language instructions into precise and executable target instructions.

[0041] In one embodiment, if Figure 5 As shown, after step S120, steps S1201-S1202 are also included.

[0042] S1201, recording all operation data of interaction with the user, and determining whether the operation data meets the preset learning conditions;

[0043] S1202: If the requirement is met, the preset prompt framework and the preset rule engine are updated through a preset learning model.

[0044] In this embodiment, the preset language model includes a dynamic prompt construction module, which includes a learning layer, records all operation data of the interaction between the elderly care robot and the user, such as instructions, feedback, execution results, etc., and determines whether the operation data meets the preset learning conditions, wherein the preset learning conditions can be set according to the number of interactions or the accumulated time, and there is no limitation on this. If the operation data meets the preset learning conditions, the preset prompt framework and the preset rule engine are fine-tuned and updated through conversational instructions using a preset learning model, such as a BERT model. Specifically, the input of the preset learning model is the user's original instruction, context, and family knowledge graph, and the output is a structured prompt word expression (used to strengthen the construction of the preset prompt framework). For example, if the user's spoken language is input, "Where did I put the water before my last meal?", the template and rule layers may have difficulty processing the logic of "before dinner" and "that water". The learning model uses contextual reasoning to identify the potential temporal sequence and item correspondence between "last time", "dinner", and "water", and infers that the user refers to "water to drink before dinner". Combined with the graph, it identifies the "kitchen kettle" and generates an instruction: "Please find the stainless steel kettle on the kitchen countertop". The logic of the generated instruction is applied to the preset prompt framework of the template layer and the preset rule engine of the rule layer, thereby updating them. By updating the preset prompt framework and the preset rule engine, the mapping between language and graph reasoning can be learned from real interaction data, thereby improving the ability of the elderly care robot to handle unknown expressions or situations that are difficult for rules to cover, and significantly improving user experience and task completion rate.

[0045] S130: Generate target navigation instructions by combining the target language instructions and the collected environment image through a preset visual language model.

[0046] In this embodiment, the environmental image is acquired in real time by an image acquisition device, such as a camera, onboard the elderly care robot. This image captures visual information about the robot's surroundings, including room layout, furniture placement, and obstacle locations, and presents it in the form of an image. The preset visual language model is a trained visual language model (VLM). The preset visual language model combines the target language instructions with the captured environmental image to generate target navigation instructions. Specifically, the preset visual language model first analyzes and interprets the captured environmental image. It utilizes computer vision technology to identify various objects in the image, such as furniture, doors, windows, and walls, and determine their positions and spatial relationships. Simultaneously, the model performs semantic parsing on the target language instructions to understand the user's intent and goals. The model correlates the results of the image understanding with the results of the language instruction parsing. Based on the position and spatial relationship of the target object in the image, combined with the target location and objects in the user's instruction, the model determines whether there is a corresponding object at the target location to determine the robot's desired action path. This final action path serves as the target navigation instruction. By automatically generating accurate target navigation instructions, the elderly care robot can provide more intelligent services to meet user needs in various scenarios.

[0047] In one embodiment, if Figure 6 As shown, step 130 includes steps S131-S133.

[0048] S131, converting the target language instruction into an operating unit in a preset format;

[0049] S132, extracting image features from the environment image through a preset visual encoder;

[0050] S133: Fusing the operating unit with the image features through the preset visual language model to generate the target navigation instruction.

[0051] In this embodiment, the operation unit in the preset format is a standardized format that the model can process, such as a token, and the target language instruction is converted into an operation unit in the preset format. Specifically, the target language instruction is mapped to a token sequence that the model can process through a word embedding technology (such as Word2Vec, GloVe or BERT pre-training model). The environment image is extracted with image features through a preset visual encoder. Specifically, the collected environment image is preprocessed, including resizing, normalization, denoising, etc., to ensure that the image meets the input requirements of the model. The preprocessed image is input into a preset visual encoder (such as ResNet, VGG, etc.) to extract high-level semantic features of the image. These features are usually represented in the form of feature vectors or feature maps, which contain information such as objects, scenes, spatial relationships, etc. in the image. Among them, the image feature output by the visual encoder is a high-dimensional vector that can summarize the main content and structure of the image. The preset visual language model is used to fuse the operation unit with the image features to generate the target navigation instructions. Specifically, the operation unit (tokenized language instructions) and image features are input into the preset visual language model for multimodal feature fusion. The model combines language and visual information through an attention mechanism or other fusion techniques to understand the user's intention and the layout of the environment. Based on the fused features, it understands the specific meaning of the user's instructions and the spatial relationships in the environmental image. For example, the model needs to determine the specific location of "next to the sofa" in the image and plan a path to reach that location. Based on the fused features and contextual understanding, the model generates target navigation instructions. These instructions typically include information such as the robot's movement direction, distance, and speed to guide the robot to complete the task. By fusing language instructions with collected images through multimodal feature fusion, the language instructions and environmental information are comprehensively considered to generate more accurate navigation instructions, improving the accuracy and reliability of the elderly care robot in performing tasks.

[0052] In one embodiment, if Figure 7 As shown, step 130 includes steps S134-S136.

[0053] S134: If the preset visual language model generates several candidate navigation instructions, then voice query the user;

[0054] S135: If no user feedback information is received within the preset feedback time, calculating the executable degree of each candidate navigation instruction;

[0055] S136: Determine the candidate navigation instruction with the highest executable degree as the target navigation instruction.

[0056] In this embodiment, after receiving the target language instructions and the environmental image, the preset visual language model processes and analyzes them, potentially generating multiple navigation instructions depending on the actual environment. These generated navigation instructions are considered candidate navigation instructions and are then queried by the elderly user via voice. If the system does not receive user feedback within a set timeframe, alternative strategies are employed to determine the final target navigation instruction. The preset visual language model calculates the executability of each candidate navigation instruction. This executability calculation takes into account multiple factors, such as path length, the number and complexity of obstacles, and safety, though the specific considerations are not limited. The executability of each candidate navigation instruction is compared, and the instruction with the highest executability is selected as the target navigation instruction, which is then transmitted to the actuator of the elderly care robot. By querying the user or calculating the executability when generating multiple navigation instructions to determine the target navigation instruction, the user's sense of control and engagement with the elderly care robot is enhanced, while also improving the robot's robustness and reliability.

[0057] S140: Control the robot to move according to the target navigation instruction.

[0058] In this embodiment, upon receiving a target navigation instruction, the robot first parses the instruction. Because target navigation instructions are encoded in a specific format, the robot must convert them into an internal representation that it can understand and execute. Based on the parsed instructions or the planned path, the robot controls its internal motors to move. By controlling the robot's movement according to the target navigation instruction, the robot can accurately reach the user's designated target location, providing precise service and improving the quality and efficiency of elderly care services.

[0059] Figure 8 FIG is a schematic block diagram of a robot target navigation device 200 provided by an embodiment of the present invention. Figure 8 As shown, corresponding to the above robot target navigation method, the present invention also provides a robot target navigation device. The robot target navigation device includes a unit for executing the above robot target navigation method, and the device can be configured in a terminal such as a desktop computer, tablet computer, laptop computer, etc. Specifically, please refer to Figure 8 The robot target navigation device includes a construction unit 210, an expansion unit 220, a generation unit 230 and a control unit 240.

[0060] The construction unit 210 is used to construct basic language instructions through a preset language model according to the fuzzy instructions input by the user and the preset family knowledge graph.

[0061] In one embodiment, the construction unit 210 includes an acquisition unit and an association unit.

[0062] an acquisition unit, configured to acquire object information, placement information, and usage information through a user's personalized input and the environment image;

[0063] An association unit is used to associate the object information, the placement information and the usage information according to the semantic association and spatial relationship of the objects to generate the preset family knowledge graph.

[0064] In one embodiment, the construction unit 210 includes a semantic analysis unit and a construction sub-unit.

[0065] a semantic analysis unit, configured to convert the fuzzy instruction into text content, perform semantic analysis based on the text content and the preset family knowledge graph, and determine user intent and candidate items;

[0066] A construction subunit is used to construct the basic language instruction according to the user intention and the candidate through a preset prompt framework in the preset language model.

[0067] The expansion unit 220 is configured to expand information based on the basic language instruction and the preset reference information using the preset language model to generate a target language instruction.

[0068] In one embodiment, the expansion unit 220 includes an engine analysis unit and an expansion sub-unit.

[0069] an engine analysis unit, configured to analyze the basic language instruction in combination with the preset reference information through a preset rule engine of the preset language model to generate supplementary information, wherein the preset reference information includes family structure, current time, and user history information;

[0070] The expansion subunit is configured to expand the basic language instruction according to the supplementable information to generate the target language instruction.

[0071] In one embodiment, the expansion unit 220 includes a determination unit and an update unit.

[0072] A judgment unit, configured to record all operation data interacting with the user and judge whether the operation data meets a preset learning condition;

[0073] An updating unit is configured to update the preset prompt framework and the preset rule engine through a preset learning model if the requirement is met.

[0074] The generating unit 230 is configured to generate a target navigation instruction by combining the target language instruction and the collected environment image through a preset visual language model.

[0075] In one embodiment, the generating unit 230 includes a converting unit, an extracting unit, and a fusing unit.

[0076] a conversion unit, configured to convert the target language instruction into an operation unit in a preset format;

[0077] An extraction unit, configured to extract image features from the environment image through a preset visual encoder;

[0078] A fusion unit is used to perform feature fusion on the operation unit and the image feature through the preset visual language model to generate the target navigation instruction.

[0079] In one embodiment, the generating unit 230 includes a converting unit, an extracting unit, and a fusing unit.

[0080] an inquiry unit, configured to make a voice inquiry to the user if the preset visual language model generates a plurality of candidate navigation instructions;

[0081] a calculation unit, configured to calculate the executable degree of each candidate navigation instruction if no user feedback information is received within a preset feedback time;

[0082] The determining unit is configured to determine the candidate navigation instruction with the highest executable degree as the target navigation instruction.

[0083] The control unit 240 is used to control the robot to move according to the target navigation instruction.

[0084] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned robot target navigation device 200 and each unit can refer to the corresponding description in the aforementioned method embodiment. For the convenience and brevity of description, it will not be repeated here.

[0085] The above-mentioned robot target navigation device can be implemented in the form of a computer program. The computer program can be used in Figure 9 Runs on the computer equipment shown.

[0086] See also Figure 9 , Figure 9 This is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 500 can be a terminal or a server. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, personal digital assistant, wearable device, or other electronic device with communication capabilities. The server can be a standalone server or a server cluster consisting of multiple servers.

[0087] See Figure 9The computer device 500 includes a processor 502 , a memory, and a network interface 505 connected via a system bus 501 , wherein the memory may include a non-volatile storage medium 503 and an internal memory 504 .

[0088] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, which, when executed, can enable the processor 502 to execute a robot target navigation method.

[0089] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.

[0090] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a robot target navigation method.

[0091] The network interface 505 is used to communicate with other devices through the network. Figure 9 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0092] The processor 502 is configured to run a computer program 5032 stored in the memory to implement the steps of the above method.

[0093] It should be understood that in the embodiment of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0094] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is computer-readable. The program instructions are executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.

[0095] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor performs the steps of the above method.

[0096] The storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.

[0097] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0098] In the several embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the various units is merely a logical functional division, and actual implementation may employ other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented.

[0099] The steps in the methods of the embodiments of the present invention may be adjusted in order, combined, or deleted as needed. The units in the devices of the embodiments of the present invention may be combined, divided, or deleted as needed. Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.

[0100] If this integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, terminal, or network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present invention.

[0101] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A robot target navigation method, characterized in that: include: Construct basic language instructions through a preset language model based on the fuzzy instructions input by the user and the preset family knowledge graph; According to the basic language instructions and the preset reference information, information expansion is performed using the preset language model to generate target language instructions; The target language instruction and the collected environment image are used to generate a target navigation instruction through a preset visual language model; The robot is controlled to move according to the target navigation instruction.

2. The method according to claim 1, characterized in that Before the step of constructing the basic language instruction, the method includes: Obtaining object information, placement information, and usage information through the user's personalized input and the environmental image; The object information, the placement information, and the usage information are associated according to the semantic association and spatial relationship of the objects to generate the preset family knowledge graph.

3. The method according to claim 1, characterized in that The step of constructing a basic language instruction through a preset language model based on the fuzzy instruction input by the user and the preset family knowledge graph includes: Convert the fuzzy instruction into text content, perform semantic analysis based on the text content and the preset family knowledge graph, and determine the user intention and candidate items; The basic language instruction is constructed according to the user intention and the candidate through a preset prompt framework in the preset language model.

4. The method according to claim 3, characterized in that The step of performing information expansion using the preset language model according to the basic language instructions and the preset reference information to generate the target language instructions includes: Analyzing the basic language instruction in combination with the preset reference information through a preset rule engine of the preset language model to generate supplementary information, wherein the preset reference information includes family structure, current time, and user history information; The basic language instruction is expanded according to the supplementable information to generate the target language instruction.

5. The method according to claim 4, characterized in that After the step of generating the target language instruction, the method further includes: Record all operational data of interactions with the user and determine whether the operational data meets the preset learning conditions; If the requirement is met, the preset prompt framework and the preset rule engine are updated through the preset learning model.

6. The method according to claim 1, characterized in that The step of generating a target navigation instruction by combining the target language instruction and the collected environment image with a preset visual language model includes: Converting the target language instruction into an operation unit in a preset format; Extracting image features from the environment image through a preset visual encoder; The preset visual language model is used to perform feature fusion on the operating unit and the image feature to generate the target navigation instruction.

7. The method according to claim 6, characterized in that The step of performing feature fusion of the operating unit and the image feature through the preset visual language model to generate the target navigation instruction further includes: If the preset visual language model generates a number of candidate navigation instructions, then voice query the user; If no user feedback information is received within the preset feedback time, the executable degree of each candidate navigation instruction is calculated; The candidate navigation instruction with the highest executable degree is determined as the target navigation instruction.

8. A robot target navigation device, characterized in that: include: A construction unit, configured to construct a basic language instruction using a preset language model according to the fuzzy instruction input by the user and the preset family knowledge graph; an expansion unit, configured to expand information using the preset language model according to the basic language instruction and preset reference information to generate a target language instruction; A generating unit, configured to generate a target navigation instruction by combining the target language instruction and the collected environment image through a preset visual language model; A control unit is used to control the robot to move according to the target navigation instruction.

9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the method according to any one of claims 1 to 7 can be implemented.

Citation Information

Cited By

  • Robot control method based on voice analysis

    CN121148376A

  • A robot control method based on voice analysis

    CN121148376B

  • Robot vision language navigation method and device based on multi-modal knowledge graph

    CN122149435A