Semantic mapping method and intelligent terminal

The image frames are obtained through the intelligent terminal, the keyframes are determined and their semantic information and scale information are obtained, and the semantic mapping is constructed, which solves the problems of high equipment requirements and high computing power requirements in the existing technology, and realizes low-cost and low-demand 3D mapping and semantic understanding of courtyard scenes.

CN119992418APending Publication Date: 2025-05-13元鼎智能创新(国际)有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510075518.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When the prior art is conducting 3D mapping of courtyard environments, the equipment requirements are high and the demand for processing dense point cloud computing power is high, making it difficult to achieve low-cost and low-demand scenario 3D mapping.

Method used

N image frames are obtained through the application of the smart terminal, keyframes are determined, semantic information and scale information of the keyframes are obtained, and semantic graphs are constructed based on this to generate semantic point clouds and semantic images.

Benefits of technology

It effectively reduces the equipment requirements for mapping construction, and realizes low-cost and low-demand 3D mapping and semantic understanding of courtyard scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992418A_ABST
    Figure CN119992418A_ABST
Patent Text Reader

Abstract

The invention provides a semantic mapping method and a robot, the semantic mapping method is applied to an intelligent terminal, the intelligent terminal comprises an image acquisition unit, the method comprises the following steps: starting an application program, the application program being pre-stored in the intelligent terminal; acquiring N image frames through an image acquisition unit; performing visual synchronous positioning on the N image frames and obtaining pose information; determining a key frame in the N image frames; obtaining depth information and semantic information of an image in the key frame; calculating scale information of the image in the key frame based on the pose information and the depth information; and performing semantic mapping according to the scale information and the semantic information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the technical field of semantic mapping, and more particularly to a semantic mapping method and a smart terminal. Background Art

[0002] When the robot executes service instructions for lawns, corridors, swimming pools and other courtyard environments that need to be cleaned, it needs to improve its 3D understanding of the courtyard environment, that is, to perform 3D positioning and mapping of the courtyard scene. The 3D understanding of the courtyard environment is mostly carried out using multi-line laser radar, but it has high equipment requirements and high computing power requirements for processing dense point clouds. Summary of the invention

[0003] The technical problem to be solved by the present application is to provide a method for semantic mapping in view of the deficiencies of the above-mentioned prior art, by utilizing the application of the intelligent terminal to obtain N image frames, determining key frames for the N image frames, obtaining the semantic information and scale information of the key frames and performing semantic mapping based on the information, thereby effectively reducing the requirements for the equipment for mapping.

[0004] According to another aspect of the present disclosure, a robot applying the above method is provided.

[0005] In one aspect of the present application, a method for semantic mapping is provided, which is applied to an intelligent terminal, wherein the intelligent terminal includes an image acquisition unit, and the method includes:

[0006] Starting an application program, the application program being pre-stored in the smart terminal;

[0007] Acquire N image frames through an image acquisition unit;

[0008] Performing visual synchronous positioning on the N image frames and obtaining posture information;

[0009] Determining a key frame among the N image frames;

[0010] Acquiring depth information and semantic information of the image in the key frame;

[0011] Calculating scale information of the image in the key frame based on the pose information and the depth information; and

[0012] A semantic map is constructed according to the scale information and the semantic information.

[0013] Further, the performing semantic mapping according to the scale information and the semantic information includes: generating a semantic point cloud for a plurality of pixel points of the image in the key frame according to the scale information and the semantic information.

[0014] Furthermore, the intelligent terminal also includes a display unit, and the method also includes: visualizing the semantic point cloud and generating a semantic image, and displaying the semantic image on the display unit, wherein the semantic image includes one or more display objects and semantic labels corresponding to the one or more display objects.

[0015] Furthermore, the method further includes: receiving a predetermined operation of the user on the semantic image.

[0016] Furthermore, the predetermined operation includes:

[0017] Selecting a display object with an error or a region where the display object with an error is located from the one or more display objects; and

[0018] The semantic label corresponding to the display object with the error is changed.

[0019] Further, the performing visual synchronous positioning on the N image frames and obtaining the position and posture information includes:

[0020] Measuring the acceleration and angular velocity of the smart terminal by an inertial measurement unit;

[0021] Obtaining global positioning information of the smart terminal through a global positioning system; and

[0022] The posture information is determined based on the N image frames, the acceleration, the angular velocity, and the global positioning information.

[0023] Further, the performing visual synchronous positioning on the N image frames and obtaining the position and posture information includes:

[0024] Measuring the acceleration and angular velocity of the smart terminal by an inertial measurement unit;

[0025] And determining the posture information based on the N image frames, the acceleration and the angular velocity.

[0026] Furthermore, acquiring the depth information of the image in the key frame includes: predicting depth values ​​of multiple pixel points of the image in the key frame by using a predetermined deep learning model.

[0027] Furthermore, the semantic information includes one or more of roads, swimming pools, lawns, pedestrians, furniture, street lights, buildings, and obstacles.

[0028] The present application also discloses an intelligent terminal, including a memory and a processor, wherein the memory stores computer program instructions, and the processor executes the method described in any embodiment of the present application when processing the program instructions.

[0029] The embodiments described in this application have the following beneficial effects:

[0030] The robot and semantic mapping method provided in the present application can be applied to smart terminals, and image frames are obtained through the application of the smart terminal, key frames in the image frames are determined, semantic information and scale information of the images of the key frames are obtained, and semantic mapping is performed based on the scale information and semantic information. 3D mapping of the scene is achieved through the application terminal at low cost and low requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solution of the embodiment of the present disclosure, the following briefly introduces the drawings required for the description of the embodiment. The drawings described below are only exemplary embodiments of the present disclosure.

[0032] Figure 1 A flow chart of a method for semantic mapping according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0033] The technical solutions in this application will be described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work belong to the protection scope of this application. It should be noted that the embodiments in this application and the features in the embodiments can be combined with each other without conflict.

[0034] The present application provides a method for semantic mapping, which is applied to smart terminals. The smart terminal may be a terminal such as a smart phone, a tablet computer, a smart wearable device, an augmented reality device, etc. The smart terminal may also be a robot equipped with various cameras and sensors (e.g., a sweeping robot, a weeding robot, etc.), which can automatically perform related tasks (e.g., cleaning tasks). Taking a sweeping robot as an example, it senses the environment through a sensor unit, plans a cleaning path using a control unit, and is equipped with cleaning tools, etc. to remove pollutants such as dust and garbage. The present application does not limit the specific presentation of the robot, as long as the principle of the present application can be realized. The smart terminal can collect three-dimensional spatial data of the environment through the cameras, sensors, etc. carried by it, and can use a specific application to perform data processing and semantic mapping on the collected three-dimensional spatial data. The smart terminal of the present application may also include components such as a processing unit, a storage unit, a display unit, a sensor, a communication unit, a battery, and a housing. In the present application, if there is no additional description, a smart phone will be used as an example of a smart terminal for explanation.

[0035] Semantic map is a map representation that integrates semantic information. Compared with traditional maps based only on geometric shapes and positional relationships, it contains a richer and deeper understanding of various objects and places in the environment. Semantic map is constructed by adding semantic labels, semantic relationships between objects, and related attribute descriptions on the basis of traditional maps (such as two-dimensional coordinate maps, three-dimensional space maps, etc.). It presents the elements in the environment in a structured way, such as using nodes to represent different objects (such as buildings, roads, lawns, swimming pools, etc.) or places (such as schools, hospitals, supermarkets, etc.), and using edges to describe the relationship between them (such as adjacent, contained, belonging, etc.), and at the same time giving each node and edge corresponding semantic attributes, such as the category, function, color, size, etc. of the object. For example, in a semantic map of a courtyard, the swimming pool node will be marked with the category attribute of "swimming pool", and may also have other related attributes such as "swimming pool area", "swimming pool depth", "swimming pool orientation", etc. In other words, the semantic map contains both environmental spatial information and environmental semantic information. By giving the map, the smart terminal can know that there is an object in the environment and what the object is, just like a human. In this application, unless otherwise specified, "courtyard" will be used as the basic scene for semantic mapping.

[0036] The semantic mapping method 100 of the present application is described in detail below with reference to the accompanying drawings. Figure 1 A flowchart of a method 100 for semantic mapping according to an embodiment of the present application is shown. The method 100 includes steps S101 to S107, and the method is applied to an intelligent terminal, and the intelligent terminal includes an image acquisition unit. Steps S101 to S107 are described below.

[0037] In step S101, an application is started, and the application is pre-stored in the smart terminal.

[0038] This application uses the smart terminal to perform semantic mapping, which requires the use of applications pre-stored or set in the smart terminal. The application can be a computer program composed of computer code, and running the application enables the smart terminal to complete one or several specific tasks. Taking a smart phone as an example, the application (i.e., mobile application, also known as mobile app) is stored in the smart phone in advance, and the user can start the mobile application when semantic mapping is needed. The user can start the application by manual triggering. The application can also be automatically started when predetermined conditions are met. For example, if the smart phone senses that its position is in the courtyard, the smart phone automatically starts the application.

[0039] The application program using the semantic mapping method 100 of the present application can be installed by the user, or can be pre-stored in the smart terminal when the smart terminal leaves the factory. The present application does not specifically limit the source and installation method of the application program, as long as the technical principle of the present application can be implemented.

[0040] The image acquisition unit of the smart terminal may include components such as a lens, an image sensor, an image signal processor (ISP), and a control circuit. The lens focuses the external scene by refracting light, and the image sensor performs photoelectric conversion to convert the light signal into an electrical signal. Then the image sensor reads the electrical signal of the pixel unit in a progressive scanning or interlaced scanning manner. The electrical signal is converted into a digital signal by an analog-to-digital conversion circuit, and the digital signal is subjected to noise reduction processing, color correction and enhancement, and resolution adjustment and scaling by an image signal processor. For example, pixels that are significantly deviated from the brightness or color average of surrounding pixels are removed, and the overall color balance of the image is adjusted using a white balance algorithm to restore the true color, increase the saturation and contrast of the image, etc. to make the image more vivid, vivid, and clear, and adjust the resolution of the image. The image data processed by the image signal processor can be output in a specific format (such as JPEG, PNG, etc.). The above description of the components and functions of the image acquisition unit is only exemplary, and those skilled in the art can selectively set the image acquisition unit according to the technical principles of the present application.

[0041] Next, the process proceeds to step S102. In step S102, N image frames are acquired by an image acquisition unit.

[0042] Specifically, after the application is automatically executed, the image acquisition unit may acquire images of the surrounding environment at predetermined time intervals, thereby acquiring N image frames. For example, N image frames may be acquired by a camera, and the image frames may include RGB images and depth images.

[0043] Next, the process proceeds to step S103. In step S103, visual synchronization positioning is performed on the N image frames to obtain position and posture information.

[0044] Visual Simultaneous Localization and Mapping (VSLAM) is an algorithm that combines computer vision and robotics technology, and is used to simultaneously perform position positioning and environmental map construction. For example, by analyzing image information, the intelligent terminal can determine its own position (i.e., positioning) in the environment while building a map of the surrounding environment (i.e., mapping). Visual simultaneous positioning is performed on the N acquired image frames to obtain posture information, and the position and posture of the image acquisition unit at each frame of the image can be inferred from these image frames, i.e., posture. Position is usually expressed in three-dimensional space coordinates, such as x, y, and z coordinates in the world coordinate system, and posture can be expressed in terms of direction and rotation angle, such as Euler angles, quaternions, rotation matrices, etc., to describe the rotation around three coordinate axes. For example, important feature points such as edges and corners can be extracted from the image frames, matched, and the posture of the image acquisition unit, that is, the smart terminal, is calculated using the matched feature points. The environmental map is updated based on the calculated posture, and then it can be detected whether the image acquisition unit has returned to its previous position to reduce the cumulative error and improve the accuracy of the map, and finally complete the visual synchronous positioning of the N image frames and obtain the posture information.

[0045] In one embodiment, the visual synchronous positioning of the N image frames and obtaining the posture information includes: measuring the acceleration and angular velocity of the smart terminal through an inertial measurement unit; obtaining the global positioning information of the smart terminal through a global positioning system; and determining the posture information based on the N image frames, the acceleration, the angular velocity, and the global positioning information.

[0046] Specifically, the smart terminal also includes an inertial measurement unit and a global positioning system. The inertial measurement unit is usually composed of an accelerometer and a gyroscope. The accelerometer is used to measure the acceleration change of the smart terminal in three coordinate axes, usually set as the x, y, and z axes. The global positioning system determines the location information of the smart terminal on the earth by receiving signals transmitted by multiple satellites, that is, the three-dimensional coordinates of the smart terminal in longitude, latitude, and altitude. For example, in outdoor areas, the global positioning system module in the smart terminal can receive satellite signals from different orbits, and after computer calculation, it can obtain the longitude and latitude of its own location, and at the same time obtain altitude information. This information is in a global unified geographic coordinate system, which is convenient for integration with other geographic data. Therefore, the global positioning information of the smart terminal, that is, the absolute position coordinates, is obtained through the global positioning system.

[0047] The acceleration, angular velocity, and global positioning information of the acquired smart terminal are fused with N image frames, that is, the scene features are extracted using N image frames to realize visual positioning, calculate the local posture, estimate the position and direction changes based on the acceleration and angular velocity, and obtain the geographical location information based on the global positioning information. In the process of visual synchronous positioning, when the image frame has unclear features, rapid motion causes image blur, etc., which affects the visual positioning effect, the high-frequency and real-time relative motion data of the inertial measurement unit can play a complementary role. For example, when the smart terminal is rotating rapidly, the image may be temporarily blurred and the posture cannot be accurately determined by vision, but the inertial measurement unit can roughly infer the posture changes within this time period based on the angular velocity information to ensure the consistency of the posture information. Moreover, visual synchronous positioning may have cumulative errors during long-term operation, resulting in large deviations in posture estimation, and the global positioning coordinates provided by the global positioning system as an absolute position reference can regularly correct the posture calculated based on the image frame and the inertial measurement unit, pull it back to the correct global coordinate system, and avoid excessive accumulation of errors. At the same time, when the smart terminal starts or enters a new area, the global positioning system positioning information can provide an initial position reference for posture determination. Finally, the posture information is determined based on the N image frames, the acceleration, the angular velocity, and the global positioning information.

[0048] It should be understood that in the process of fusing acceleration, angular velocity, global positioning information with N image frames, the relative posture change output by the inertial measurement unit, the absolute position coordinates output by the global positioning system, the local posture calculated by the visual synchronous positioning, etc. can be fused according to certain rules. For example, through algorithms such as Kalman filtering, the posture information from different sources is weighted averaged with their respective measurement accuracy and reliability as weights to obtain a comprehensive posture estimation result; the acceleration and angular velocity data measured by the inertial measurement unit can also be directly integrated as constraints into the optimization process of the visual synchronous positioning algorithm based on N image frames, and the positioning information of the global positioning system is also combined to jointly constrain the calculation of the posture, so that the cooperation between different data sources is closer, and the advantages of each data can be more fully utilized to improve the accuracy and reliability of the posture information.

[0049] In one embodiment, the visual synchronous positioning of the N image frames and obtaining the posture information includes: measuring the acceleration and angular velocity of the smart terminal through an inertial measurement unit; and determining the posture information based on the N image frames, the acceleration and the angular velocity.

[0050] Specifically, the acceleration and angular velocity of the smart terminal, that is, the relative posture change, are measured by an inertial measurement unit; visual positioning is achieved through N image frames, and the local posture is calculated; the inertial measurement unit and the image are fused, and the relative position relationship can be obtained by recording images at different positions before and after, combined with the moving distance calculated by the inertial measurement unit, that is, a map of the relative position relationship is obtained using the relative posture change and the local posture, thereby determining the posture information based on the N image frames, the acceleration and the angular velocity.

[0051] Next, the process proceeds to step S104 and step S105 in sequence. In step S104, a key frame is determined in the N image frames. In step S105, depth information and semantic information of the image in the key frame are obtained.

[0052] The semantic information may include one or more of roads, swimming pools, lawns, pedestrians, furniture, street lights, buildings, obstacles, etc.

[0053] By visually locating N image frames synchronously and obtaining position information, the surrounding environment map is constructed. At this time, it is necessary to fuse with semantic information, that is, to fuse the surrounding environment map with one or more semantic information of roads, swimming pools, lawns, pedestrians, furniture, street lights, buildings, obstacles, etc. in the environment to achieve semantic mapping. At this time, it is necessary to determine the key frames in the N image frames, that is, to select some image frames from the N image frames that are important for subsequent operations such as positioning and mapping. The selection process can be considered based on the angle of motion change, the angle of scene richness, and the angle of positioning accuracy.

[0054] In one embodiment, when the motion state of the smart terminal changes significantly (for example, the rotation angle exceeds a certain threshold, the translation distance reaches a certain value, etc.), the corresponding image frame is selected as the key frame. For example, when the smart terminal is moving, the image frame collected after each rotation of 90 degrees or translation of a certain distance (for example, 1 meter) can be used as a key frame, because the images at these moments can reflect the information of different perspectives and different areas in the environment, which is critical for building a complete map.

[0055] In another embodiment, an image frame containing rich scene information, a large number of feature points and a relatively uniform distribution can be used as a key frame, or an image frame that helps improve positioning accuracy and can better match features and estimate poses with other frames can also be used as a key frame. For example, if the quality of feature points in some image frames is high, that is, the feature points have good repeatability and are insensitive to changes in illumination and viewing angle, and the pose can be calculated more stably through these feature points, then the image frame can be selected as a key frame.

[0056] After determining the key frames in the N image frames, the depth information and semantic information of the images in the key frames can be obtained. The depth information refers to the distance from the object corresponding to each pixel in the image to the imaging plane, that is, the plane where the image acquisition unit of the intelligent terminal is located. Obtaining the depth information can convert the two-dimensional image into a three-dimensional representation that is closer to reality, which is very important for understanding the spatial structure of the environment, the positional relationship of objects, and the subsequent three-dimensional map construction. For example, knowing the depth values ​​corresponding to each part of the lawn, aisle, furniture, street lamps and other objects in the image, it is possible to accurately restore their actual layout in space and determine the front and back position relationship between objects. Semantic information is a description of the categories, attributes and relationships between objects, scenes, etc. in the image. For example, if the object in the image is identified as a "table", its attributes may include "wood material", "rectangular shape", and there is a "matching use" relationship with the surrounding "chairs". Obtaining semantic information allows the intelligent terminal to better understand the environment, which helps to make more intelligent decisions, task planning, and interaction with users. For example, a robot performing cleaning tasks in the courtyard can know that "the aisle is between the lawn and the swimming pool" by recognizing semantic information. Therefore, the robot can accurately clean the aisle according to the requirements of the cleaning task.

[0057] In one embodiment, acquiring the depth information of the image in the key frame includes: predicting the depth values ​​of a plurality of pixels of the image in the key frame by using a predetermined deep learning model.

[0058] The deep learning model is a machine learning model based on an artificial neural network. It can automatically learn complex patterns and features from a large amount of data. Based on its nonlinear fitting ability, it can automatically mine the potential relationship between image appearance features and depth information from a large amount of data, thereby realizing the prediction of the depth value of the key frame image pixel. For example, the deep learning model can be trained by data set preparation, loss function selection, and training optimization algorithm. When the trained deep learning model is applied to the depth prediction of the key frame image, the key frame image is input into the model. After the various layers of the model are operated, a depth map with the same size as the input image or after appropriate scaling and other processing is finally output. The value of each pixel in the depth map represents the depth prediction value of the pixel in the real scene corresponding to the image. For example, for a key frame image of a courtyard scene, after the model is input, the depth value corresponding to the pixel area where the table is located in the output depth map is shallow, while the depth value corresponding to the courtyard wall in the distance is deep, which intuitively reflects the spatial position relationship of objects in the courtyard. In this way, the depth information of the image in the key frame can be obtained by predicting the depth values ​​of multiple pixels of the image in the key frame through the deep learning model.

[0059] It should be understood that the architecture of the deep learning model can be set according to actual conditions. For example, the deep learning model can be a convolutional neural network architecture, an encoder-decoder architecture, or a Transformer-based architecture. This application does not limit this, as long as it can implement the technical principles of this application.

[0060] Next, the process proceeds to step S106. Step S106: Calculate scale information of the image in the key frame based on the pose information and the depth information.

[0061] Scale information reflects the actual size of objects or scenes in the image in the real world and the relative size relationship between them. Accurate scale information is important for building maps that conform to the actual spatial layout. For example, knowing the actual size of the lawn in the image and its size ratio with the surrounding buildings can help smart terminals more accurately restore the spatial structure of the entire scene, rather than just knowing the relative position and category between the lawn and the surrounding buildings.

[0062] Through step S103 to step S105, the pose information of N image frames and the depth information of the key frame are obtained, that is, the pose information and depth information of the key frame are obtained. According to the pose information and depth information, the size and proportional relationship of the object in the key frame image in the real world, that is, the scale information, can be calculated. The specific calculation process can use the triangulation principle or the scale assessment method combining monocular vision with pose and depth. The triangulation principle is: assuming that there are key frame images taken from different poses, and the pose information corresponding to each key frame is known, including position and pose, that is, the translation and rotation of the pose, and the depth information of the corresponding pixel point in the image, the actual size of the object and the scale relationship between them can be inferred through the geometric relationship of triangulation. For example, in a binocular vision system, two cameras (corresponding to different postures) shoot the same scene. For a feature point on an object in the scene, the relative posture of the two cameras is known, that is, obtained through the posture information, and the depth value of the feature point in the two camera images is obtained through the depth information. Using the triangulation formula, that is, based on geometric relationships such as the principle of similar triangles, the actual distance from the feature point to the camera plane can be calculated, and then combined with information such as the imaging size of the object in the image, the actual scale of the object, such as length, width, etc., can be inferred. The scale estimation method of monocular vision combined with posture and depth is: for monocular images, although there is a lack of direct stereoscopic vision information, the scale can also be estimated by combining posture changes and depth information. For example, when the smart terminal moves a distance and records the depth changes of the same object in the key frame images before and after the movement, the actual scale of the object can be inferred based on geometric principles such as similar triangles and the changes in the imaging position of the object in the image.

[0063] Next, the process proceeds to step S107. In step S107, semantic mapping is performed according to the scale information and the semantic information.

[0064] Semantic mapping is performed based on the scale information and the semantic information, that is, the scale information and the semantic information are integrated to obtain a map that includes not only the location information of objects in the environment, but also the semantic information of the objects, such as categories, functions, and relationships, as well as the actual size and spatial layout of the scale information. That is, semantic mapping is performed based on the scale information and the semantic information.

[0065] In one embodiment, the performing semantic mapping according to the scale information and the semantic information includes: generating a semantic point cloud for a plurality of pixel points of the image in the key frame according to the scale information and the semantic information.

[0066] Semantic point cloud is a data representation that integrates semantic information. On the basis of traditional 3D point cloud, it gives corresponding semantic attributes to each point, so that these points not only represent the position in space, but also reflect the category, function, etc. of the object. For example, in a semantic point cloud containing a courtyard scene, some points in the point cloud not only have their specific coordinate positions in the courtyard, but can also be labeled as specific semantic categories such as "lawn", "swimming pool", "chair", "tree", etc., which clearly show the distribution and spatial relationship of objects in the courtyard environment.

[0067] In one embodiment, a semantic point cloud is generated for a plurality of pixel points of the image in the key frame according to the scale information and the semantic information. First, a map may be initialized and a unified coordinate system may be established. The pose information, depth information and semantic information of the key frame may be converted to a unified coordinate system to facilitate integration and map construction. According to the depth information of the key frame image and the pixel point coordinates that have been converted to the unified coordinate system, the two-dimensional coordinates of each pixel point are converted to three-dimensional space coordinates using geometric relationships to perform initial point cloud construction. Then, the object category in the key frame image is identified based on the semantic information, and its actual size is determined according to the scale information. These objects are instantiated as specific entities in the map, and the converted three-dimensional space coordinate points are annotated with corresponding semantic attributes including category, size, function, etc. Next, the spatial relationship between the objects in the key frame image is analyzed, such as determining the positional relationship of the objects such as front and back, left and right, and up and down, and the semantic relationship (such as chairs are placed around the table, and there is an aisle between the lawn and the swimming pool) through the depth information, and the spatial structure of the semantic point cloud is further optimized to make it more consistent with the size and spatial layout of the actual objects.

[0068] As new keyframes are added, the semantic map can be continuously updated. On the one hand, the newly appeared objects can be identified, their scales calculated, and their attributes annotated, and integrated into the existing map structure; on the other hand, when the location, attributes, and other information of the annotated objects change, the map can be modified and optimized accordingly.

[0069] In one embodiment, the intelligent terminal also includes a display unit, and the semantic mapping method also includes: visualizing the semantic point cloud and generating a semantic image, and displaying the semantic image on the display unit, wherein the semantic image includes one or more display objects and semantic labels corresponding to the one or more display objects.

[0070] Although the semantic point cloud contains rich spatial and semantic information, it exists in the form of many three-dimensional coordinate points and corresponding semantic attributes, which is not convenient for users to intuitively understand and interact. To address this problem, the semantic point cloud is visualized and a semantic image is generated, that is, for each object represented by the semantic point cloud, the object surface is constructed based on the point cloud through a certain algorithm, and then after the geometric shape of the object surface is constructed, the object surface is drawn using graphics rendering technology, so that the drawn object has more realism and three-dimensionality in the image. According to the point cloud position corresponding to each object in the semantic point cloud and the layout of the image, the display position of each object semantic label can be reasonably determined, and the style of the semantic label can be designed, including font, font size, color, background, etc., so that it is highlighted in the image and easy to read, and the visualization of the semantic point cloud is completed. After completing the above visualization processing steps, the generated image containing the object geometry and semantic labels can be stored in a suitable image format. Common image formats include JPEG, PNG, BMP, etc. When selecting the format, factors such as image quality, compression ratio, and compatibility with the display unit should be considered. Finally, the generated semantic image can be adapted and adjusted according to the resolution, size, display ratio and other parameters of the display unit (such as the screen of the smart terminal, the display, etc.) to ensure that the image can be displayed completely and clearly on the screen. For example, for a high-resolution large-screen display unit, the image size can be appropriately enlarged while ensuring that the image clarity is not affected; for screens with different aspect ratios, the image can be adapted to the screen by scaling or cropping to avoid image deformation; and then the image data is transmitted to the display unit for display through the corresponding display driver and software interface, so that the user can intuitively see the image content containing rich semantic information.

[0071] In this way, the semantic point cloud is visualized and a semantic image is generated, and the semantic image is displayed on the display unit; by converting it into a semantic image through visualization processing, objects in the environment and their semantic relationships can be presented in a more intuitive and easy-to-understand way, making it convenient for users to view and perform subsequent interactive operations with smart terminals.

[0072] In one embodiment, the method for semantic mapping further includes: receiving a predetermined operation of the user on the semantic image, wherein the predetermined operation includes: selecting a display object with an error or a region where the display object with an error is located from the one or more display objects; and changing a semantic label corresponding to the display object with an error.

[0073] After the semantic map is completed, that is, after the spatial information of the courtyard environment and the environmental semantic information are fused, errors may occur and need to be corrected, that is, receiving the user's predetermined operation on the semantic image. The predetermined operation may include a selection operation, a zoom operation, a review operation, an editing operation, and a query operation. The selection operation may, for example, be selecting an erroneous display object from one or more display objects or selecting the area where the erroneous display object is located. The editing operation may, for example, be changing the semantic label corresponding to the erroneous display object.

[0074] The present application also provides an intelligent terminal, including a memory and a processor, wherein the memory stores computer program instructions, and the processor executes the method described in any embodiment when processing the program instructions.

[0075] The processor may be a central processing unit (CPU) or other form of processing unit having data processing capability and / or instruction execution capability, and may control other components in the robot to perform desired functions.

[0076] The memory may include one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor may execute the program instructions to implement the semantic mapping method of each embodiment of the present application described above and / or other desired functions.

[0077] In one embodiment, the robot may further include an input unit and an output unit, and these components are interconnected via a bus system and / or other forms of connection mechanisms. The output device may output various information to the outside. The output unit may include, for example, a display, a speaker, and a communication network and a remote output device connected thereto.

[0078] It should be noted that the sequence of the above embodiments of the present application is only for description and does not represent the advantages or disadvantages of the embodiments.

[0079] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.

[0080] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0081] The above is only an exemplary embodiment of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various changes or substitutions within the technical scope recorded in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A method for semantic mapping, applied to an intelligent terminal, wherein the intelligent terminal includes an image acquisition unit, and the method includes: Starting an application program, the application program being pre-stored in the smart terminal; Acquire N image frames through an image acquisition unit; Performing visual synchronous positioning on the N image frames and obtaining posture information; Determining a key frame among the N image frames; Acquiring depth information and semantic information of the image in the key frame; Calculating scale information of the image in the key frame based on the pose information and the depth information; as well as A semantic map is constructed according to the scale information and the semantic information.

2. The method according to claim 1, wherein: The performing semantic mapping according to the scale information and the semantic information includes: generating a semantic point cloud for a plurality of pixel points of the image in the key frame according to the scale information and the semantic information.

3. The method according to claim 2, wherein: The smart terminal also includes a display unit, and the method also includes: visualizing the semantic point cloud and generating a semantic image, and displaying the semantic image on the display unit, wherein the semantic image includes one or more display objects and semantic labels corresponding to the one or more display objects.

4. The method according to claim 3, further comprising: A predetermined operation of a user on the semantic image is received.

5. The method according to claim 4, wherein: The predetermined operation includes: Selecting a display object with an error or a region where the display object with an error is located from the one or more display objects; and The semantic label corresponding to the display object with the error is changed.

6. The method according to claim 1, wherein: The performing visual synchronous positioning on the N image frames and obtaining the position information comprises: Measuring the acceleration and angular velocity of the smart terminal by an inertial measurement unit; Obtaining global positioning information of the smart terminal through a global positioning system; and The posture information is determined based on the N image frames, the acceleration, the angular velocity, and the global positioning information.

7. The method according to claim 1, wherein: The performing visual synchronous positioning on the N image frames and obtaining the position information comprises: Measuring the acceleration and angular velocity of the smart terminal by an inertial measurement unit; And determining the posture information based on the N image frames, the acceleration and the angular velocity.

8. The method according to claim 1, wherein: The acquiring the depth information of the image in the key frame includes: predicting the depth values ​​of a plurality of pixels of the image in the key frame by using a predetermined deep learning model.

9. The method according to claim 1, wherein: The semantic information includes one or more of roads, swimming pools, lawns, pedestrians, furniture, street lights, buildings, and obstacles.

10. An intelligent terminal, comprising a memory and a processor, wherein the memory stores computer program instructions, and the processor executes the method according to any one of claims 1 to 9 when processing the program instructions.