Robot sensing system and method based on artificial intelligence
By combining speech sensors and vision sensors, using large language models and machine learning models to build semantic raster maps and control mapping relationships, and generating navigation control instructions, it solves the problem that robots are difficult to achieve coordinated control of global path planning and local obstacle avoidance in dynamic changing environments, and improves the robot's flexibility and navigation efficiency.
Patent Information
- Application Number
- CN202510062236.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing robot path planning algorithm performs well in static environments, but it is difficult to maintain the effectiveness of global goals and the real-time local obstacle avoidance in dynamically changing indoor environments, making it difficult for robots to achieve flexible coordinated control in complex environments.
By combining speech sensors and vision sensors, semantic analysis and machine learning models are used to use large language models to perform semantic analysis, semantic raster maps are constructed, the control mapping relationship between the robot's moving position and the navigation path is determined, and navigation control instructions are generated through boundary constraints, so as to realize the coordinated control of the robot's global path planning and local obstacle avoidance in a dynamically changing environment.
It improves the flexibility of the robot in dynamically changing indoor environments, ensures that the robot can successfully complete global path planning and local obstacle avoidance tasks in complex environments, and avoid navigation failures caused by environmental changes.
Smart Images

Figure CN119987361A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of robot perception technology, and more specifically, to an artificial intelligence robot perception system and method. Background Art
[0002] Robot perception is the core of achieving navigation and autonomous decision-making. Common sensors include lidar, visual perception, ultrasonic sensors and infrared sensors. Lidar generates accurate three-dimensional point cloud maps by scanning the surrounding environment and is widely used in environmental mapping and obstacle avoidance. Visual perception uses cameras and computer vision technology to identify objects in the environment and assist in path planning. Ultrasonic sensors are suitable for close-range obstacle avoidance with high accuracy but limited range. Infrared sensors can navigate in low-light environments through temperature difference detection. By fusing data from multiple sensors, robots can achieve precise positioning, dynamic obstacle avoidance and efficient path planning, thereby completing tasks autonomously and safely in complex environments.
[0003] Existing AI-based robot control mainly relies on technologies such as deep learning, reinforcement learning and imitation learning. Deep learning is used for perception tasks such as visual recognition and speech understanding. Reinforcement learning optimizes control strategies by interacting with the environment. Imitation learning enables robots to perform tasks by learning through observation. Robot control methods usually need to combine image recognition, path planning and local obstacle avoidance technologies to enhance the robot's autonomy. However, in indoor environments, dynamic obstacles (such as people, moving objects or other robots) are constantly changing. Robots not only need to avoid dynamic obstacles around them in real time, but also need to balance between global path planning and local obstacle avoidance. Existing path planning algorithms perform well in static environments, but it is difficult for robots to maintain the effectiveness of global goals and the real-time nature of local obstacle avoidance in rapidly changing environments. Therefore, how to achieve coordinated control between global path planning and local obstacle avoidance in dynamically changing indoor environments, thereby improving the flexibility of robots, has become a difficult problem faced by the industry. Summary of the invention
[0004] The present application provides an artificial intelligence robot perception system and method, which can achieve coordinated control between global path planning and local obstacle avoidance of the robot in a dynamically changing indoor environment, thereby improving the flexibility of the robot.
[0005] In a first aspect, the present application provides a robot navigation method based on artificial intelligence, which is used for robot navigation by a robot perception system based on artificial intelligence, wherein the robot comprises a voice sensor and a visual sensor, and the method comprises the following steps: Receive the user's voice command through the voice sensor, perform semantic analysis on the voice command according to the preset large language model, and obtain the robot's task instruction; All objects in the room are visually measured by visual sensors, and then a semantic grid map of the room is obtained; Based on the preset machine learning model, semantic analysis is performed on all visual images collected by the visual sensor during the visual measurement process to obtain the semantic association features of each indoor object. The control mapping relationship between the robot's moving position and the navigation path during the visual measurement process is determined according to the robot's moving position during the visual measurement process and the semantic association features of each indoor object. Determining the boundary constraint amount for controlling the robot to avoid each indoor object through the contour features of each grid unit in the semantic grid map and the control mapping relationship; A navigation control instruction of the robot is generated according to the task instruction and the boundary constraint amount of controlling the robot to avoid each indoor object.
[0006] In some embodiments, semantic parsing of the voice command according to a preset large language model to obtain the robot's task instruction specifically includes: Converting the voice command into voice text; Extracting the user's semantic intention from the voice text based on a preset large language model; The robot's behavioral action sequence is determined according to the semantic intention, and then the robot's task instructions are generated.
[0007] In some embodiments, visually measuring all objects in a room by a visual sensor to obtain a semantic grid map of the room specifically includes: The robot's visual sensor collects visual images and depth maps of each indoor object; Perform semantic segmentation on each visual image to obtain the regional edges and semantic labels of each target object in each visual image; Determine the spatial position distribution of each indoor object based on the depth map of each indoor object and the regional edges and semantic labels of each target object in each visual image; The spatial position distribution of all indoor objects is spatially mapped to obtain the indoor semantic grid map.
[0008] In some embodiments, semantic analysis is performed on all visual images collected by the visual sensor during the visual measurement process based on a preset machine learning model to obtain the semantic association features of each indoor object, specifically including: Obtain the regional edges and semantic labels of each target object in each visual image; Based on the preset machine learning model, combined with the regional edges and semantic labels of each target object in each visual image, all indoor objects in the indoor environment are analyzed for association, and the spatial adjacency status of each indoor object is obtained; The semantic association features of each indoor object are extracted from all spatial adjacency states.
[0009] In some embodiments, determining the control mapping relationship between the moving position of the robot during the visual measurement process and the navigation path according to the moving position of the robot during the visual measurement process and the semantic association features of each indoor object specifically includes: Determine the robot's spatial perception of each indoor object based on the robot's moving position during the visual measurement process; The control mapping relationship between the robot's moving position and navigation path during visual measurement is determined by the robot's spatial perception of each indoor object and the semantic association characteristics of each indoor object.
[0010] In some embodiments, determining the boundary constraint amount for controlling the robot to avoid each indoor object by using the contour features of each grid unit in the semantic grid map and the control mapping relationship specifically includes: Determining a constraint distance for the robot to move around each indoor object by using contour features of each grid cell in the semantic grid map; The boundary constraint amount for controlling the robot to avoid each indoor object is determined according to the constraint distance that the robot moves beside each indoor object and the control mapping relationship.
[0011] In some embodiments, generating the navigation control instruction of the robot according to the task instruction and the boundary constraint amount of controlling the robot to avoid each indoor object specifically includes: Get the current position and motion status of the robot; Determine an obstacle avoidance path for the robot to reach a final target position according to the current position, motion state and the task instruction of the robot; The navigation control instructions of the robot are determined by the obstacle avoidance path of the robot to the final target position and the boundary constraint amount of controlling the robot to avoid each indoor object.
[0012] In a second aspect, the present application provides an artificial intelligence robot perception system, the artificial intelligence robot perception system includes a navigation control unit, and the navigation control unit includes: The command receiving module is used to receive the user's voice command through the voice sensor, perform semantic analysis on the voice command according to the preset large language model, and obtain the robot's task command; A processing module is used to perform visual measurement of all objects in the room through a visual sensor, thereby obtaining a semantic grid map of the room; The processing module is further used to perform semantic analysis on all visual images collected by the visual sensor during the visual measurement process based on a preset machine learning model to obtain semantic association features of each indoor object, and determine a control mapping relationship between the robot's moving position during the visual measurement process and the navigation path according to the robot's moving position during the visual measurement process and the semantic association features of each indoor object; The processing module is further used to determine the boundary constraint amount for controlling the robot to avoid each indoor object through the contour features of each grid unit in the semantic grid map and the control mapping relationship; The execution module is used to generate a navigation control instruction of the robot according to the task instruction and the boundary constraint amount of controlling the robot to avoid each indoor object.
[0013] In a third aspect, the present application provides a computer device, comprising a memory and a processor, wherein the memory stores codes, and the processor is configured to obtain the codes and execute the above-mentioned artificial intelligence-based robot navigation method.
[0014] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned artificial intelligence-based robot navigation method is implemented.
[0015] The technical solution provided by the embodiments disclosed in this application has the following beneficial effects: In the artificial intelligence robot perception system and method provided by the present application, first, a user's voice command is received through a voice sensor, and the voice command is semantically parsed according to a preset large language model to obtain a task instruction of the robot; visual measurement of all objects in the room is performed through a visual sensor to obtain a semantic grid map of the room; semantic analysis is performed on all visual images collected by the visual sensor during the visual measurement process based on a preset machine learning model to obtain semantic association features of each indoor object, and a control mapping relationship between the robot's moving position and the navigation path during the visual measurement process is determined according to the robot's moving position during the visual measurement process and the semantic association features of each indoor object; the boundary constraint amount for controlling the robot to avoid each indoor object is determined according to the contour features of each grid unit in the semantic grid map and the control mapping relationship; and the navigation control instruction of the robot is generated according to the task instruction and the boundary constraint amount for controlling the robot to avoid each indoor object.
[0016] It can be seen that in the present application, the navigation control instructions of the robot can be generated according to the task instructions and the boundary constraint amount of each indoor object controlled by the robot; wherein, firstly, the robot receives the user's voice instructions through the voice sensor, uses the large language model to perform semantic analysis on these instructions, and generates clear task instructions. This step provides the robot with an efficient user interaction method, so that it can flexibly adjust the task objectives according to the user's needs; then, the robot uses the visual sensor to perform a comprehensive visual measurement of the indoor environment and generate a semantic raster map. The semantic raster map provides semantic information for subsequent path planning and dynamic obstacle avoidance. In a dynamically changing indoor environment, objects may continue to move or change. Traditional path planning methods are often limited by changes in obstacles and are difficult to respond flexibly. By constructing a map with semantic information, the robot can more accurately perceive and understand the surrounding objects, including static objects and dynamic objects, and provide more comprehensive environmental information for global path planning; secondly, through the machine learning model The model performs semantic analysis on the images collected by the visual sensor. The robot can not only identify and locate indoor objects, but also obtain the semantic association features of each object. Based on these semantic association features, the control mapping relationship between the robot's moving position and the navigation path during the visual measurement process is obtained. The robot is guided to move accurately in a complex environment through the control mapping relationship to avoid collision with objects, so that the robot can not only successfully complete the global path planning, but also flexibly respond to the needs of local obstacle avoidance, and avoid navigation failure due to sudden changes of objects; then, the boundary constraint amount of the robot to avoid indoor objects is calculated, and the boundary constraint amount is used to enable the robot to correct the robot's motion state in time during the movement to avoid the robot colliding with the boundary of indoor objects; finally, the robot generates a navigation control instruction by combining the task instruction with the boundary constraint amount of the control robot to avoid each indoor object; in summary, the solution of the present application can realize the coordinated control between the global path planning and local obstacle avoidance of the robot in a dynamically changing indoor environment, thereby improving the flexibility of the robot. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is an exemplary flow chart of an artificial intelligence-based robot navigation method according to some embodiments of the present application; Figure 2 is a schematic diagram of a process for determining a semantic grid map according to some embodiments of the present application; Figure 3 is a schematic diagram of a process for determining semantic association features according to some embodiments of the present application; Figure 4 is a schematic diagram of the structure of a navigation control unit according to some embodiments of the present application; Figure 5It is a structural diagram of a computer device for implementing an artificial intelligence-based robot navigation method according to some embodiments of the present application. DETAILED DESCRIPTION
[0018] In order to better understand the technical solution of the present application, the technical solution of the present application will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0019] refer to Figure 1 , which is an exemplary flow chart of a robot navigation method based on artificial intelligence according to some embodiments of the present application. The robot navigation method 100 based on artificial intelligence mainly includes the following steps: In step 101, a user's voice command is received through a voice sensor, and the voice command is semantically parsed according to a preset large language model to obtain a task instruction of the robot.
[0020] In some embodiments, semantic parsing of the voice command according to a preset large language model to obtain the robot's task command can be achieved by the following steps: Converting the voice command into voice text; Extracting the user's semantic intention from the voice text based on a preset large language model; The robot's behavioral action sequence is determined according to the semantic intention, and then the robot's task instructions are generated.
[0021] It should be noted that the large language model described in this application is a generative language model (such as GPT-4). The generative language model is based on the Transformer architecture. It learns the structure and semantics of the language through pre-training of large-scale text data, and uses autoregressive generation to generate natural language text word by word according to the input context.
[0022] In the specific implementation, first, the user's voice command is received through the voice sensor on the robot, and the user's voice command is converted into voice text through voice recognition technology (such as long short-term memory network model); secondly, the voice text is input into a preset large language model, and the output result of the large language model is used as the user's semantic intention, wherein the semantic intention represents a set of actions that the user requires the robot to perform; then, the user's semantic intention is decomposed into tasks using the preset large language model, and the sequence of subtasks obtained after the task decomposition is used as the robot's behavior action sequence, and then the robot's motion controller converts the behavior action sequence into a control signal of the robot, and the obtained control signal is used as the robot's task instruction. Other methods can also be used in other embodiments, which are not limited here.
[0023] In step 102, all objects in the room are visually measured by a visual sensor to obtain a semantic grid map of the room.
[0024] In some embodiments, reference Figure 2 As shown in FIG. 1 , this figure is a schematic diagram of a process for determining a semantic grid map in some embodiments of the present application. In this embodiment, all objects in the room are visually measured by a visual sensor, and then a semantic grid map of the room is obtained, which can be implemented by the following steps: The robot's visual sensor collects visual images and depth maps of each indoor object; Perform semantic segmentation on each visual image to obtain the regional edges and semantic labels of each target object in each visual image; Determine the spatial position distribution of each indoor object based on the depth map of each indoor object and the regional edges and semantic labels of each target object in each visual image; The spatial position distribution of all indoor objects is spatially mapped to obtain the indoor semantic grid map.
[0025] It should be noted that the depth map described in the present application represents an image of the distance from the pixel point of each target object in the image to the visual sensor. It should also be noted that in the present application, one target object corresponds to one indoor object, that is, the target object is judged by a semantic label.
[0026] In addition, it should be noted that the spatial position distribution described in the present application represents the position coordinates of indoor objects distributed in three-dimensional space.
[0027] In the specific implementation, first, use the pre-trained semantic segmentation model (such as DeepLabV3, etc.) to segment each visual image, identify the regional edges of each target object in each visual image through the output of the semantic segmentation model, and assign semantic labels (for example: "table") to each target object in each visual image; secondly, for each indoor object, use the intrinsic parameter matrix of the robot's visual sensor to convert all pixel coordinates in the visual image of the indoor object with the depth value of each pixel point in the depth map of the indoor object, and then use the three-dimensional reconstruction algorithm (such as stereo matching or structured light method) to map the regional edges of each target object in the visual image corresponding to the indoor object and the depth value of each pixel point in the depth map of the indoor object to the three-dimensional space, so as to obtain the spatial position distribution of the indoor object, and then proceed. The spatial position distribution of all indoor objects is obtained; then, the spatial position distribution of each indoor object is taken as point cloud data, and all point cloud data are aligned using a point cloud stitching algorithm (such as an iterative nearest point algorithm). The spatial position of each indoor object can be unified into the same three-dimensional space through the alignment, and the entire three-dimensional space is further gridded using a voxelization method (such as a voxel grid method), and the three-dimensional space is divided into multiple voxel units, each voxel unit contains an indoor object, and then the indoor objects in each voxel unit are labeled with the semantic label of the indoor object corresponding to the target object, thereby generating a spatial map of the distribution of different indoor objects in the indoor environment, and the obtained spatial map is used as a semantic grid map. Other methods can also be used in other embodiments, which are not limited here.
[0028] It should be noted that the semantic grid map described in the present application represents a map that divides the three-dimensional space of the indoor environment into multiple grid units, each grid unit corresponds to an indoor object, wherein the grid unit also contains geometric information and semantic categories of the indoor objects.
[0029] In step 103, semantic analysis is performed on all visual images collected by the visual sensor during the visual measurement process based on a preset machine learning model to obtain semantic association features of each indoor object, and the control mapping relationship between the robot's moving position and the navigation path during the visual measurement process is determined according to the robot's moving position during the visual measurement process and the semantic association features of each indoor object.
[0030] In some embodiments, reference Figure 3 As shown in FIG. 1 , this figure is a schematic diagram of a process for determining semantic association features in some embodiments of the present application. In this embodiment, semantic analysis is performed on all visual images collected by the visual sensor during the visual measurement process based on a preset machine learning model, and the semantic association features of each indoor object can be obtained by the following steps: First, in step 1031, the regional edges and semantic labels of each target object in each visual image are obtained; Secondly, in step 1032, based on the preset machine learning model and combining the regional edges and semantic labels of each target object in each visual image, all indoor objects in the indoor environment are subjected to association analysis to obtain the spatial adjacency status of each indoor object; Then, in step 1033, the semantic association features of each indoor object are extracted from all spatial adjacency states.
[0031] It should be noted that the machine learning model used in the present application is a support vector machine, which obtains the area, overlap and distance between each target object in multiple visual images, and marks the adjacency status of each target object in each visual image. The adjacency status indicates whether the target object in the visual image is adjacent to other target objects in the visual image, and further trains the support vector machine, so that the machine learning model obtained after the training is used as the preset machine learning model.
[0032] In the specific implementation, first, a visual image is selected as the selected visual image, and OpenCV is used to calculate the area, overlap and distance between each target object in the selected visual image, and the area, overlap and distance between each target object in the selected visual image are used as the input of the preset machine learning model, so that the output result of the preset machine learning model is used as the adjacency state of each target object in the selected visual image, and then the adjacency state of each target object in the remaining visual images is obtained, and the values of the adjacency states of target objects with the same semantic label in all visual images are added, and the added value is used as the spatial adjacency state of indoor objects corresponding to the target objects with the same semantic label, and then the spatial adjacency state of all indoor objects is obtained. The spatial adjacency state represents the total number of other indoor objects adjacent to the indoor object; then, a clustering algorithm (such as DBSCAN) is used to cluster the spatial adjacency states of all indoor objects, and the clusters obtained by clustering are used as adjacency state clusters, wherein each adjacency state cluster contains multiple spatial adjacency states, and each spatial adjacency state corresponds to an indoor object. Further, for each adjacency state cluster, a spatial adjacency state is selected from the adjacency state cluster as the selected spatial adjacency state, and the distance between the selected spatial adjacency state and the cluster center of the adjacency state cluster is used as the semantic association feature of the indoor object corresponding to the selected spatial adjacency state, and the semantic association features of the indoor objects corresponding to the remaining spatial adjacency states in the adjacency state cluster are further determined. In other embodiments, other methods can also be used to achieve this, which will not be repeated here.
[0033] It should be noted that the semantic association features described in the present application represent features that reflect the semantic interdependence of the spatial layout of indoor objects. The semantic association features help to improve the robot's perception of the indoor environment.
[0034] In some embodiments, determining the control mapping relationship between the moving position of the robot during the visual measurement process and the navigation path according to the moving position of the robot during the visual measurement process and the semantic association features of each indoor object can be implemented by the following steps: Determine the robot's spatial perception of each indoor object based on the robot's moving position during the visual measurement process; The control mapping relationship between the robot's moving position and navigation path during visual measurement is determined by the robot's spatial perception of each indoor object and the semantic association characteristics of each indoor object.
[0035] It should be noted that the spatial perception described in this application represents the robot's perception strength of indoor objects in three-dimensional space.
[0036] In specific implementation, first, at the moving position in the visual measurement process, the robot obtains the distance, angle, and viewing angle of each indoor object, wherein the angle represents the offset angle of the robot toward the object, and the viewing angle represents the angle of the object within the robot's field of view. For each indoor object, the angle of the indoor object is divided by the viewing angle, and the value obtained by the division is multiplied by the distance from the robot to the indoor object, and the multiplied value is used as the robot's spatial perception of the indoor object, thereby obtaining the robot's spatial perception of all indoor objects; secondly, an indoor object is selected as the selected indoor object, and a negative exponential function with the natural logarithm e as the base is calculated for the semantic association feature of the selected indoor object, and then the value obtained after calculating the negative exponential function with the natural logarithm e as the base is multiplied by the robot's spatial perception of the selected indoor object, and the multiplied value is used as the navigation path control amount between the selected indoor object and the robot, thereby obtaining the navigation path control amount between the remaining indoor objects and the robot, and further the navigation path control amounts between all indoor objects and the robot are summed, and the summed value is used as the control mapping relationship between the robot's moving position and the navigation path during the visual measurement process. Other methods may also be used in other embodiments, which are not limited here.
[0037] It should be noted that the control mapping relationship described in the present application represents the conversion parameters for controlling the robot to fit the navigation path after the robot perceives an indoor object.
[0038] In step 104, the boundary constraint amount for controlling the robot to avoid each indoor object is determined based on the contour features of each grid unit in the semantic grid map and the control mapping relationship.
[0039] In some embodiments, determining the boundary constraint amount for controlling the robot to avoid each indoor object by using the contour features of each grid unit in the semantic grid map and the control mapping relationship can be implemented by the following steps: Determining a constraint distance for the robot to move around each indoor object by using contour features of each grid cell in the semantic grid map; The boundary constraint amount for controlling the robot to avoid each indoor object is determined according to the constraint distance that the robot moves beside each indoor object and the control mapping relationship.
[0040] It should be noted that the contour features of the grid cells in the semantic grid map in the present application represent the boundary shape features of the grid cells corresponding to indoor objects mapped in the semantic grid map. The convex hull algorithm can be used to calculate the outer boundary of the grid cell, and then the safety distance is added to the outer boundary to obtain the contour features of the grid cell, where the safety distance can be set to the physical size of the robot (such as width, height, length).
[0041] In the specific implementation, first, a grid cell is selected from the semantic grid map as the selected grid cell, the distance between the contour feature of the selected grid cell and the contour features of other grid cells is calculated, the minimum distance is selected from all distances, and the minimum distance is used as the constraint distance for the robot to move around the indoor object corresponding to the selected grid cell, and the constraint distance for the robot to move around the indoor objects corresponding to the remaining grid cells in the semantic grid map is continued to be determined; secondly, for each indoor object, the constraint distance for the robot to move around the indoor object is multiplied by the control mapping relationship, and the multiplied value is used as the boundary constraint amount for controlling the robot to avoid indoor objects, and then the boundary constraint amount for controlling the robot to avoid all indoor objects is obtained. Other methods can also be used in other embodiments, which are not limited here.
[0042] It should be noted that the constraint distance described in the present application represents the safe distance that limits the robot from passing through indoor objects during the navigation control process; in addition, it should also be noted that the boundary constraint amount described in the present application represents the minimum adjustment amount that needs to be applied to the robot's motion state in order to avoid colliding with the boundaries of indoor objects during the robot's movement, wherein the motion state includes speed, direction and acceleration.
[0043] In step 105, a navigation control instruction of the robot is generated according to the task instruction and the boundary constraint amount for controlling the robot to avoid each indoor object.
[0044] In some embodiments, generating the navigation control instruction of the robot according to the task instruction and the boundary constraint amount of controlling the robot to avoid each indoor object can be implemented by the following steps: Get the current position and motion status of the robot; Determine an obstacle avoidance path for the robot to reach a final target position according to the current position, motion state and the task instruction of the robot; The navigation control instructions of the robot are determined by the obstacle avoidance path of the robot to the final target position and the boundary constraint amount of controlling the robot to avoid each indoor object.
[0045] It should be noted that the obstacle avoidance path described in the present application refers to a safe path for the robot to follow a predetermined target position while avoiding surrounding obstacles during movement.
[0046] In the specific implementation, first, the current position and motion state are obtained through the robot's positioning system, and the robot's instruction parser parses the task instruction to obtain different intermediate target positions and final target positions. Then, based on the robot's current position, motion state, different intermediate target positions and final target positions, the robot uses a local obstacle avoidance algorithm to generate an obstacle avoidance path for the robot to the final target position; then, for each indoor object on the obstacle avoidance path, an indoor object on the obstacle avoidance path is selected as the selected indoor object, the boundary constraint amount that controls the robot to avoid the selected indoor object is added to the robot's motion state at the current position, and the values obtained by adding are used as the robot's obstacle avoidance path in the selected indoor object. The motion state next to the object, that is: the motion state (speed, direction and acceleration) next to the selected indoor object = the motion state (speed, direction and acceleration) of the current position + the boundary constraint of controlling the robot to avoid the selected indoor object. The motion state of the robot next to the selected indoor object is further converted into instructions through the robot's motion control device, and the obtained instructions are used as local control instructions of the robot next to the selected indoor object, and then the local control instructions of the robot next to all indoor objects on the obstacle avoidance path are obtained. Finally, all local control instructions are sorted according to the order in which the robot needs to pass through the obstacle avoidance path, and the obtained local control instruction sequence is used as the navigation control instruction of the robot.
[0047] It should be noted that the navigation control instructions described in the present application represent instructions for controlling the motion state of the robot when passing through indoor objects on an obstacle avoidance path.
[0048] In addition, in another aspect of the present application, in some embodiments, the present application provides an artificial intelligence robot perception system, the artificial intelligence robot perception system includes a navigation control unit, reference Figure 4 , which is a schematic diagram of the structure of a navigation control unit according to some embodiments of the present application, the navigation control unit 400 includes: an instruction receiving module 401, a processing module 402 and an execution module 403, which are described as follows: The instruction receiving module 401 in the present application is mainly used to receive the user's voice instructions through the voice sensor, perform semantic analysis on the voice instructions according to the preset large language model, and obtain the robot's task instructions; Processing module 402, in this application, processing module 402 is used to perform visual measurement of all objects in the room through a visual sensor, thereby obtaining a semantic grid map of the room; It should be noted that the processing module 402 described in the present application is also used to perform semantic analysis on all visual images collected by the visual sensor during the visual measurement process based on a preset machine learning model to obtain the semantic association features of each indoor object, and determine the control mapping relationship between the robot's moving position and the navigation path during the visual measurement process according to the robot's moving position during the visual measurement process and the semantic association features of each indoor object; In addition, the processing module 402 in the present application is also used to determine the boundary constraint amount for controlling the robot to avoid each indoor object through the contour features of each grid unit in the semantic grid map and the control mapping relationship; Execution module 403, in the present application, the execution module 403 is mainly used to generate navigation control instructions for the robot according to the task instructions and the boundary constraint amount for controlling the robot to avoid each indoor object.
[0049] In addition, the present application also provides a computer device, which includes a memory and a processor, the memory stores code, and the processor is configured to obtain the code and execute the above-mentioned artificial intelligence-based robot navigation method.
[0050] In some embodiments, reference Figure 5 , which is a schematic diagram of the structure of a computer device for implementing an artificial intelligence-based robot navigation method according to some embodiments of the present application. The artificial intelligence-based robot navigation method in the above embodiment can be Figure 5 The computer device 500 shown in the figure is implemented, and the computer device 500 includes at least one processor 501, a communication bus 502, a memory 503 and at least one communication interface 504.
[0051] The processor 501 may be a general-purpose central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more for controlling the execution of the artificial intelligence-based robot navigation method in the present application.
[0052] The communication bus 502 may be used to transmit information between the above-mentioned components.
[0053] The memory 503 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CDROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 503 may exist independently and be connected to the processor 501 through the communication bus 502. The memory 503 may also be integrated with the processor 501.
[0054] The memory 503 is used to store the program code for executing the solution of the present application, and the execution is controlled by the processor 501. The processor 501 is used to execute the program code stored in the memory 503. The program code may include one or more software modules. The method described in the above method embodiment can be implemented by the processor 501 and one or more software modules in the program code in the memory 503.
[0055] The communication interface 504 uses any transceiver or other device for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.
[0056] In a specific implementation, as an embodiment, a computer device may include multiple processors, each of which may be a single-core (singleCPU) processor or a multi-core (multiCPU) processor. The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0057] The above-mentioned computer device may be a general-purpose computer device or a special-purpose computer device. In a specific implementation, the computer device may be a desktop computer, a portable computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device or an embedded device. The embodiment of the present application does not limit the type of computer device.
[0058] In addition, the present application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned artificial intelligence-based robot navigation method.
[0059] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0060] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A robot navigation method based on artificial intelligence, wherein a robot perception system based on artificial intelligence performs robot navigation, wherein the robot comprises a voice sensor and a visual sensor, and wherein: The method comprises the following steps: Receive the user's voice command through the voice sensor, perform semantic analysis on the voice command according to the preset large language model, and obtain the robot's task instruction; All objects in the room are visually measured by visual sensors, and then a semantic grid map of the room is obtained; Based on the preset machine learning model, semantic analysis is performed on all visual images collected by the visual sensor during the visual measurement process to obtain the semantic association features of each indoor object. The control mapping relationship between the robot's moving position and the navigation path during the visual measurement process is determined according to the robot's moving position during the visual measurement process and the semantic association features of each indoor object. Determining the boundary constraint amount for controlling the robot to avoid each indoor object through the contour features of each grid unit in the semantic grid map and the control mapping relationship; A navigation control instruction of the robot is generated according to the task instruction and the boundary constraint amount of controlling the robot to avoid each indoor object.
2. The method according to claim 1, characterized in that The voice command is semantically parsed according to the preset large language model to obtain the robot's task instructions, which specifically include: Converting the voice command into voice text; Extracting the user's semantic intention from the voice text based on a preset large language model; The robot's behavioral action sequence is determined according to the semantic intention, and then the robot's task instructions are generated.
3. The method according to claim 1, characterized in that The visual sensor is used to visually measure all objects in the room, and then the semantic grid map of the room is obtained, which specifically includes: The robot's visual sensor collects visual images and depth maps of each indoor object; Perform semantic segmentation on each visual image to obtain the regional edges and semantic labels of each target object in each visual image; Determine the spatial position distribution of each indoor object based on the depth map of each indoor object and the regional edges and semantic labels of each target object in each visual image; The spatial position distribution of all indoor objects is spatially mapped to obtain the indoor semantic grid map.
4. The method according to claim 1, characterized in that Based on the preset machine learning model, semantic analysis is performed on all visual images collected by the visual sensor during the visual measurement process to obtain the semantic association features of each indoor object, including: Obtain the regional edges and semantic labels of each target object in each visual image; Based on the preset machine learning model, combined with the regional edges and semantic labels of each target object in each visual image, all indoor objects in the indoor environment are analyzed for association, and the spatial adjacency status of each indoor object is obtained; The semantic association features of each indoor object are extracted from all spatial adjacency states.
5. The method according to claim 1, characterized in that The control mapping relationship between the robot's moving position and the navigation path during the visual measurement process is determined according to the robot's moving position during the visual measurement process and the semantic association features of each indoor object. Specifically, it includes: Determine the robot's spatial perception of each indoor object based on the robot's moving position during the visual measurement process; The control mapping relationship between the robot's moving position and navigation path during visual measurement is determined by the robot's spatial perception of each indoor object and the semantic association characteristics of each indoor object.
6. The method according to claim 1, characterized in that Determining the boundary constraint amount of controlling the robot to avoid each indoor object by using the contour feature of each grid unit in the semantic grid map and the control mapping relationship specifically includes: Determining a constraint distance for the robot to move around each indoor object by using contour features of each grid cell in the semantic grid map; The boundary constraint amount for controlling the robot to avoid each indoor object is determined according to the constraint distance that the robot moves beside each indoor object and the control mapping relationship.
7. The method according to claim 1, characterized in that Generating the navigation control instruction of the robot according to the task instruction and the boundary constraint amount of controlling the robot to avoid each indoor object specifically includes: Get the current position and motion status of the robot; Determine an obstacle avoidance path for the robot to reach a final target position according to the current position, motion state and the task instruction of the robot; The navigation control instructions of the robot are determined by the obstacle avoidance path of the robot to the final target position and the boundary constraint amount of controlling the robot to avoid each indoor object.
8. An artificial intelligence robot perception system, the artificial intelligence robot perception system includes a navigation control unit, characterized in that: The navigation control unit comprises: The command receiving module is used to receive the user's voice command through the voice sensor, perform semantic analysis on the voice command according to the preset large language model, and obtain the robot's task command; A processing module is used to perform visual measurement of all objects in the room through a visual sensor, thereby obtaining a semantic grid map of the room; The processing module is further used to perform semantic analysis on all visual images collected by the visual sensor during the visual measurement process based on a preset machine learning model to obtain semantic association features of each indoor object, and determine a control mapping relationship between the robot's moving position during the visual measurement process and the navigation path according to the robot's moving position during the visual measurement process and the semantic association features of each indoor object; The processing module is further used to determine the boundary constraint amount for controlling the robot to avoid each indoor object through the contour features of each grid unit in the semantic grid map and the control mapping relationship; The execution module is used to generate a navigation control instruction of the robot according to the task instruction and the boundary constraint amount of controlling the robot to avoid each indoor object.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the artificial intelligence-based robot navigation method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the artificial intelligence-based robot navigation method according to any one of claims 1 to 7 is implemented.