Image data processing method and device, electronic equipment and readable storage medium
By combining semantic analysis of environmental images and depth images, a 3D Gaussian point cloud is generated and a target map is constructed, which solves the problem of map error in traditional SLAM under light changes, and achieves more accurate and semantic map construction, improving the robot's environment perception and task execution capabilities.
Patent Information
- Application Number
- CN202510573137.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-15
AI Technical Summary
Traditional SLAM methods are prone to map errors under light changes, resulting in problems such as unsmoothing on the ground.
The first semantic information is obtained by analyzing the environmental image, and the target 3D Gaussian point cloud is generated in combination with the depth image, and the target map is constructed based on the point cloud. The 3DGS-SLAM technology is used to adapt to illumination changes, and the Gaussian point density is optimized in combination with semantic segmentation technology to distinguish illumination artifacts from real objects.
It realizes more accurate map construction under light changes. The map is rich in semantic information, adapts to complex lighting environments, and improves the robot's environmental perception and task execution capabilities.
Smart Images

Figure CN120495557A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an image data processing method, device, electronic device, and readable storage medium. Background Art
[0002] Traditional SLAM (Simultaneous Localization and Mapping) methods are prone to map errors (for example, uneven surfaces on the ground) when exposed to changes in lighting. Therefore, how to alleviate these problems has become a technical problem that has become an urgent need for those skilled in the art. Summary of the Invention
[0003] The embodiments of the present application provide an image data processing method, device, electronic device and readable storage medium, which can realize more accurate and semantically rich map construction.
[0004] The embodiments of the present application can be implemented as follows:
[0005] In a first aspect, an embodiment of the present application provides an image data processing method, the method comprising:
[0006] Analyzing the environment image to obtain first semantic information corresponding to the environment image, wherein the first semantic information is used to indicate the semantic type of the target object and the position of the target object in the environment image;
[0007] Obtaining a target 3D Gaussian point cloud based on the first semantic information, the environment image, and a depth image corresponding to the environment image, wherein the target 3D Gaussian point cloud includes second semantic information obtained based on the first semantic information, and the second semantic information is used to indicate a semantic type of the target object and a position in the target 3D Gaussian point cloud;
[0008] A target map is obtained based on the target 3D Gaussian point cloud, wherein the target map includes third semantic information corresponding to the second semantic information, and the third semantic information is used to indicate the semantic type of the target object and the position in the target map.
[0009] In a second aspect, an embodiment of the present application provides an image data processing device, the device comprising:
[0010] An analysis module, configured to analyze the environment image to obtain first semantic information corresponding to the environment image, wherein the first semantic information is used to indicate a semantic type of a target object and a position of the target object in the environment image;
[0011] a point cloud acquisition module, configured to obtain a target 3D Gaussian point cloud based on the first semantic information, the environment image, and a depth image corresponding to the environment image, wherein the target 3D Gaussian point cloud includes second semantic information obtained based on the first semantic information, and the second semantic information is used to indicate the semantic type of the target object and its position in the target 3D Gaussian point cloud;
[0012] A map acquisition module is used to obtain a target map based on the target 3D Gaussian point cloud, wherein the target map includes third semantic information corresponding to the second semantic information, and the third semantic information is used to indicate the semantic type of the target object and the position in the target map.
[0013] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor can execute the machine-executable instructions to implement the image data processing method described in the aforementioned embodiment.
[0014] In a fourth aspect, an embodiment of the present application provides a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image data processing method as described in the aforementioned embodiment.
[0015] The image data processing method, device, electronic device and readable storage medium provided in the embodiments of the present application first obtain corresponding first semantic information by analyzing the environmental image, and then obtain the target 3D Gaussian point cloud based on the first semantic information, the environmental image and the depth image corresponding to the environmental image, and finally obtain the target map based on the target 3D Gaussian point cloud. The target 3D Gaussian point cloud includes second semantic information obtained based on the first semantic information, and the target map includes third semantic information corresponding to the second semantic information. The first semantic information is used to indicate the semantic type of the target object and its position in the environmental image, the second semantic information is used to indicate the semantic type of the target object and its position in the target 3D Gaussian point cloud, and the third semantic information is used to indicate the semantic type of the target object and its position in the target map. In this way, the situation where map errors are prone to occur due to changes in lighting can be alleviated. At the same time, the semantic information is combined with the semantics to build the map so that the obtained map is rich in semantic information, which facilitates the subsequent acquisition of more accurate information from the map. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0017] Figure 1 A block diagram of an electronic device provided in an embodiment of the present application;
[0018] Figure 2 This is a flowchart of an image data processing method according to an embodiment of the present application;
[0019] Figure 3 for Figure 2 Schematic diagram of the flow of sub-steps included in step S120;
[0020] Figure 4 for Figure 3 A schematic flow chart of the sub-steps included in sub-step S123;
[0021] Figure 5 for Figure 2 Schematic diagram of the flow of sub-steps included in step S130;
[0022] Figure 6 The second flowchart of the image data processing method provided in the embodiment of the present application;
[0023] Figure 7 for Figure 6 Schematic diagram of the flow of sub-steps included in step S140;
[0024] Figure 8 A data processing diagram provided in an embodiment of the present application;
[0025] Figure 9 This is a block diagram of an image data processing device according to an embodiment of the present application;
[0026] Figure 10 This is a second block diagram of the image data processing device provided in an embodiment of the present application.
[0027] Icon: 100 - electronic device; 110 - memory; 120 - processor; 130 - communication unit; 200 - image data processing device; 210 - analysis module; 220 - point cloud acquisition module; 230 - map acquisition module; 240 - processing module. DETAILED DESCRIPTION
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Generally, the components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.
[0029] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for protection, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present application.
[0030] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
[0031] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features therein may be combined with each other.
[0032] Please refer to Figure 1 , Figure 1 This is a block diagram of an electronic device 100 provided in an embodiment of the present application. The electronic device 100 may be, but is not limited to, a computer, a server, or the like. The electronic device 100 may include a memory 110, a processor 120, and a communication unit 130. The memory 110, the processor 120, and the communication unit 130 are electrically connected to each other, directly or indirectly, to enable data transmission or interaction. For example, these components may be electrically connected to each other via one or more communication buses or signal lines.
[0033] The memory 110 is used to store programs or data. The memory 110 may be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.
[0034] The processor 120 is used to read / write data or programs stored in the memory 110 and execute corresponding functions. For example, the memory 110 stores an image data processing device 200, which includes at least one software function module stored in the memory 110 in the form of software or firmware. The processor 120 executes software programs and modules stored in the memory 110, such as the image data processing device 200 in the embodiment of the present application, to perform various functional applications and data processing, thereby implementing the image data processing method in the embodiment of the present application.
[0035] The communication unit 130 is used to establish a communication connection between the electronic device 100 and other communication terminals through a network, and to send and receive data through the network.
[0036] It should be understood that Figure 1 The structure shown is only a schematic diagram of the structure of the electronic device 100. The electronic device 100 may also include Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown. Figure 1 Each component shown in the figure can be implemented by hardware, software or a combination thereof.
[0037] Please refer to Figure 2 , Figure 2 This is one of the flow diagrams of the image data processing method provided in an embodiment of the present application. The method can be applied to the above-mentioned electronic device. The specific flow of the image data processing method is described in detail below. In this embodiment, the method may include steps S110 to S130.
[0038] Step S110 : Analyze the environment image to obtain first semantic information corresponding to the environment image.
[0039] In this embodiment, the environmental image is an image required for mapping, and can be a grayscale image or a color image, etc., which can be determined in combination with the actual application scenario. The environmental image can be analyzed in any manner to obtain the first semantic information corresponding to the environmental image. The first semantic information can indicate the semantic type (e.g., table, keyboard, display) of the specific object (i.e., target object, specific object) included in the environmental image and the position of the object in the environmental image, that is, the label and position used to indicate the object.
[0040] Step S120 : obtaining a target 3D Gaussian point cloud according to the first semantic information, the environment image, and a depth image corresponding to the environment image.
[0041] In this embodiment, a point cloud is generated based on the environmental image and the depth image corresponding to the environmental image, and the semantic information corresponding to the generated point cloud is determined based on the first semantic information, thereby obtaining a target 3D Gaussian point cloud. The target 3D Gaussian point cloud includes second semantic information obtained based on the first semantic information. The second semantic information is used to indicate the specific object included in the target 3D Gaussian point cloud and the position of the object in the target 3D Gaussian point cloud, that is, the second semantic information is used to indicate the semantic type of the target object and its position in the target 3D Gaussian point cloud. The depth image is used to describe the depth corresponding to the corresponding pixel point in the environmental image.
[0042] Among them, the specific method of obtaining the target 3D Gaussian point cloud can be determined in combination with actual needs. Optionally, the target 3D Gaussian point cloud can be obtained according to the first semantic information, the environmental image and the depth image corresponding to the environmental image through 3DGS (3D Gaussian Splatting) technology. 3DGS technology is a 3D scene representation and rendering technology, the core of which is to use Gaussian functions to approximate and represent objects and surfaces in 3D scenes. Unlike traditional pixel- or point-based rendering methods, 3DGS describes each object in the scene through Gaussian Spheres, which contain attributes such as position, covariance, color coefficients, and opacity.
[0043] Step S130: obtaining a target map based on the target 3D Gaussian point cloud.
[0044] In this embodiment, the required target map can be determined based on the target 3D Gaussian point cloud in combination with map requirements. For example, assuming that only a local point cloud map corresponding to the current environment image and depth image is required, the above-mentioned target 3D Gaussian point cloud can be directly used as the target map; if other forms of maps are required, the target 3D Gaussian point cloud can be further processed to obtain the target map. The target map includes third semantic information corresponding to the second semantic information, and the third semantic information is used to indicate the semantic type of the target object and its position in the target map.
[0045] In this embodiment, a target map is obtained through a target 3D Gaussian point cloud. This method can adapt to changes in lighting appearance to build a more accurate map; and can achieve semantic segmentation of three-dimensional objects, so that the target map can provide more accurate information.
[0046] In this embodiment, the environment image can be processed using semantic segmentation technology to obtain the first semantic information. Semantic segmentation can be performed using a machine learning method or a deep learning-based method. As a possible implementation method, the environment image can be analyzed using YOLO semantic segmentation technology to obtain the first semantic information. In this implementation method, the first semantic information can be obtained based on the color features, texture features, and shape features of the environment image using YOLO semantic segmentation technology. Other methods can also be used to obtain the first semantic information of the environment image. The image features used in other methods can be determined in combination with the specific semantic information acquisition method used.
[0047] The environment image and the depth image may be obtained simultaneously by a camera, or the depth image corresponding to the environment image may be obtained by other means after obtaining the environment image. Figure 3 As shown, in the process of generating the target 3D Gaussian point cloud, Gaussian point resources are allocated based on the first semantic information, thereby obtaining the target 3D Gaussian point cloud. Figure 3 , Figure 3 for Figure 2 Schematic diagram of the flow of sub-steps included in step S120. In this embodiment, step S120 may include sub-steps S121 to S123.
[0048] Sub-step S121 : obtaining an initial 3D Gaussian point cloud according to the environment image and the depth image.
[0049] Sub-step S122 , in a case where the spatial position of the Gaussian point is constrained according to the depth image, the parameters of the Gaussian point are iteratively updated by an interleaved optimization algorithm based on the initial 3D Gaussian point cloud to obtain the target parameters of the Gaussian point.
[0050] Sub-step S123 : determining the target density of Gaussian points according to the first semantic information, and obtaining the second semantic information.
[0051] In this embodiment, an initial 3D Gaussian point cloud can be first generated based on the environmental image and the depth image, that is, the initial parameters of the Gaussian points can be determined based on the environmental image and the depth image. The initial parameters may include position, etc. (excluding density). Afterwards, based on the initial 3D Gaussian point cloud, the parameters of the Gaussian points are iteratively updated through an interlaced optimization algorithm to obtain the target parameters of the Gaussian points. During the optimization process, the precise distance information provided by the depth image is fully utilized to constrain the spatial position of the Gaussian points and improve the accuracy of the scene representation. In addition, the distribution optimization of the Gaussian points is guided by the semantic information of the environmental image to determine the density of the Gaussian points; and the corresponding second semantic information is determined in combination with the first semantic information, that is, the Gaussian points corresponding to the object corresponding to the first semantic information are marked. Among them, optimizing the parameters of the Gaussian points and determining the density of the Gaussian points can be two steps performed in parallel. The method of determining the target density of the Gaussian points based on the semantic information can be determined in combination with actual needs.
[0052] As a possible implementation, Figure 4 The target density of Gaussian points is obtained by the method shown. Figure 4 , Figure 4 for Figure 3 Flowchart of sub-steps included in sub-step S123. In this embodiment, sub-step S123 may include sub-steps S1231 to S1233.
[0053] Sub-step S1231: determining a second object from the first objects corresponding to the first semantic information.
[0054] Sub-step S1232: increasing the initial density of Gaussian points corresponding to the second object to obtain a target density of Gaussian points corresponding to the second object.
[0055] Sub-step S1233 , reducing the initial density of other Gaussian points to obtain the target density of other Gaussian points.
[0056] In this embodiment, some objects may be pre-specified for use in determining the density of corresponding Gaussian points, and this information may be combined to determine the second object from the first object (i.e., the target object) corresponding to the first semantic information. Alternatively, all first objects corresponding to the first semantic information may be directly used as the second object.
[0057] As a possible implementation method, some initial objects may be prepared in advance as objects expected to be recognized from the environment image, and the initial objects may be directly used as objects for determining the density of corresponding Gaussian points.
[0058] The initial density of each pre-specified Gaussian point can be obtained. The initial density of all Gaussian points can be the same or different, and can be determined in combination with actual needs. After determining the second object, the initial density of the Gaussian point corresponding to the second object can be increased, and the increased initial density can be used as the Gaussian point corresponding to the second object. In addition, the Gaussian points other than the Gaussian point corresponding to the second object are used as other Gaussian points. For each other Gaussian point, the initial density of the other Gaussian point is reduced, and the reduced initial density is used as the initial density of the other Gaussian point.
[0059] The above point cloud generation method optimizes the distribution of Gaussian points based on the semantic information of the environment image. The density of Gaussian points is appropriately increased in areas of the scene with rich details and semantic importance (such as object edges and key structures), and these point clouds are labeled (subsequent reconstruction can then reveal the semantics of different objects). The density is reduced in relatively smooth areas with simple semantics (walls, floors, and other Gaussian points). This semantically-based allocation of Gaussian points distinguishes between lighting artifacts and true dynamic objects, suppressing noise artifacts and maintaining map quality.
[0060] Optionally, you can Figure 5 The target map is obtained in the manner shown. Figure 5 , Figure 5 for Figure 2 Schematic diagram of the flow of sub-steps included in step S130. In this embodiment, step S130 may include sub-steps S131 and S132.
[0061] In sub-step S131 , the current global 3D Gaussian point cloud is obtained by stitching the historical 3D Gaussian point cloud and the target 3D Gaussian point cloud.
[0062] Sub-step S132 , obtaining a grid map as the target map through surface reconstruction based on the current global 3D Gaussian point cloud.
[0063] In this embodiment, the latest stored global 3D Gaussian point cloud can be used as the historical 3D Gaussian point cloud at this time. If the current environment image is the first image obtained, the historical 3D Gaussian point cloud can be considered to be empty. Based on the historical 3D Gaussian point cloud and the target 3D Gaussian point cloud, the current global 3D Gaussian point cloud can be obtained by splicing. Afterwards, based on the current global 3D Gaussian point cloud obtained, it is processed by an efficient surface reconstruction algorithm to convert the current global 3D Gaussian point cloud into a Mesh map (i.e., a grid map) to obtain the target map, thereby constructing a three-dimensional map, and the target map is a grid map at this time. It can be understood that when building a map based on the image collected next time, the current global 3D Gaussian point cloud is used as the historical 3D Gaussian point cloud.
[0064] In the case where the target 3D Gaussian point cloud is obtained by using the 3DGS technology, the target map can be obtained according to the target 3D Gaussian point cloud based on the 3DGS-SLAM technology.
[0065] As the image acquisition device continues to move and new data is collected, the above steps are repeated based on Bayesian update theory to update and optimize the constructed map in real time, ensuring that the map always reflects the latest status of the environment in which the image acquisition device is located, and realizing closed-loop optimization of map construction.
[0066] The inventors of this application have discovered that current robot training typically involves manually constructing a map model, which is then used within a robot training platform to train the robot. This manual construction of the map model results in a high workload and can result in distorted map models.
[0067] To alleviate the above situation, in this embodiment, Figure 6 As shown, the method may further include step S140.
[0068] Step S140: Perform robot training according to the preset target task and the target map.
[0069] In this embodiment, the target task can be specifically set based on actual needs, such as a target-oriented navigation task or a grasping task. The target map can be used as a map for training the robot on the target task. The training can be performed on an intelligent decision-making model used by the robot to perform the task. For example, reinforcement learning training can be performed on the robot based on the target map. This eliminates the need for manual map model construction and mitigates poor training results caused by distorted maps.
[0070] Optionally, you can Figure 7 The method shown is based on the target map and robot training. Please refer to Figure 7 , Figure 7 for Figure 6 Schematic diagram of the flow of sub-steps included in step S140. In this embodiment, step S140 may include sub-steps S141 and S142.
[0071] Sub-step S141 , performing feature extraction on the target map to obtain map features.
[0072] Sub-step S142, obtaining strategy information obtained by the strategy network in the robot based on the map features and the target map, obtaining reward and punishment information corresponding to the strategy information, and updating the strategy network according to the reward and punishment information.
[0073] In this embodiment, the target map can be a map obtained based on a depth image and an environment image collected once, and the map can be a mesh map (i.e., a grid map). Alternatively, the target map can be the target map obtained as a mesh map by combining SLAM technology as described above (for example, a target map obtained by 3DGS-SLAM technology). In this way, by updating and optimizing the constructed map in real time, it is ensured that the map always reflects the latest state of the robot's environment, achieving closed-loop optimization of map construction, facilitating the provision of accurate and up-to-date environmental information to the robot, and significantly improving its environmental perception capabilities.
[0074] The target map can be analyzed to extract map features of the target map, and then the features and the target map can be combined for robot training. The map features include at least one of terrain geometric features, object spatial distribution features, and object semantic features. The terrain geometric features are used to describe the terrain conditions; the object spatial distribution features can be used to describe the distribution of objects corresponding to the third semantic information; the object semantic features can be used to describe the objects present in the target map, that is, to indicate which objects are included. The indicated objects are the objects determined by semantic analysis of the environment image during mapping, that is, the summary set of objects indicated by the third semantic information.
[0075] During training, the robot's policy network can obtain policy information based on the map features and the target map, as well as reward and penalty information corresponding to the policy information. The policy network is then updated based on the reward and penalty information, and the above process is repeated to complete training. The reward and penalty information can be determined based on the target map. For example, if the target map indicates a table, the reward and penalty information corresponding to this policy information is determined based on the requirement that the robot cannot move to the center of the table after executing the policy information. Reinforcement learning algorithms can be used to optimize the robot's action selection, enabling it to make the most appropriate decisions based on different scene states, thereby achieving functions such as autonomous navigation and precise grasping in complex environments, further improving the robot's adaptability and task execution capabilities in real-world scenarios. Reinforcement learning algorithms such as Deep Q Network (DQN) or Proximal Policy Optimization (PPO) can be used, and the specific algorithm can be determined based on actual needs. Optionally, the target map can be placed in ISAAC for reinforcement learning by the humanoid robot. In this way, map- and feature-based reinforcement learning can effectively enhance the robot's task execution capabilities, enabling it to perform better in tasks such as navigation and grasping, and improving the robot's autonomy and adaptability in complex environments.
[0076] In this embodiment, the mapping process is optimized to achieve more accurate and semantically rich map construction; and the constructed map is provided to a humanoid robot for reinforcement learning, and the robot interacts with the features contained in the map to achieve more complex robot planning learning.
[0077] Traditional SLAM is prone to feature matching failure and map errors when encountering lighting changes. It also has difficulty handling lighting noise artifacts and poor rendering. This embodiment, however, uses 3DGS-SLAM to construct a target map, representing the scene with a 3D Gaussian function. Using a staggered optimization algorithm, it accurately captures geometric and appearance details, adapting to lighting and appearance changes to build a more accurate map. Furthermore, it can semantically allocate Gaussian point resources, distinguish lighting artifacts from true dynamic objects, and suppress noise artifacts to protect map quality. It excels in complex lighting and geometric rendering, overcoming traditional shortcomings and making maps more realistic under varying lighting conditions, enabling robots to operate precisely in environments with varying lighting conditions.
[0078] The following combination Figure 8 The above data processing method is described with an example.
[0079] The main process is Figure 8As shown, an RGB-D image is input and 3DGS reconstruction is performed to obtain a mesh model of the scene and its feature information. The results are then fed into ISAAC for reinforcement learning training of a humanoid robot, enabling functions such as object-oriented navigation and grasping. The following example illustrates the specific process.
[0080] The variational inference framework can be used to preprocess the input RGB-D image (for example, resizing the image to the same size), and then input the preprocessed RGB-D image into the 3DGS model for 3DGS scene reconstruction. In the 3DGS model, the parameters of the Gaussian points are iteratively updated through an interleaved optimization algorithm to obtain the target parameters of the Gaussian points. During the optimization process, the precise distance information provided by the depth image is fully utilized to constrain the spatial position of the Gaussian points, improving the accuracy of the scene representation. YOLO semantic segmentation technology is also introduced to guide the distribution optimization of Gaussian points based on the semantic information of the environmental image. The density of Gaussian points is appropriately increased in areas with rich scene details and semantic importance, and their point clouds are labeled. The density is reduced in relatively smooth and semantically simple areas. In this way, the target density of Gaussian points and secondary semantic information can be obtained. Based on the target parameters of the Gaussian points, the target density of the Gaussian points, and the secondary semantic information, the target 3D Gaussian point cloud can be obtained.
[0081] Based on the historical 3D Gaussian point cloud and the target 3D Gaussian point cloud obtained this time, the current global 3D Gaussian point cloud is obtained by splicing. Afterwards, the current global 3D Gaussian point cloud is converted into a mesh graph using an efficient surface reconstruction algorithm to construct a three-dimensional map (i.e., the target map described above). This method can utilize 3DGS technology to achieve high-precision map reconstruction and combine it with YOLO technology to achieve semantic segmentation of three-dimensional objects. As the device where the image acquisition device is located (for example, a robot) continues to move and collect new data, based on the Bayesian update theory, the above steps are repeated continuously to update and optimize the constructed map in real time, ensuring that the map always reflects the latest status of the environment where the image acquisition device is located, thereby achieving closed-loop optimization of map construction. Feature analysis is also performed on the reconstructed mesh map, integrating existing semantic features, spatial distribution features, ground features, etc. for subsequent reinforcement learning.
[0082] The constructed mesh map, serving as the target map, and its feature information can be input into the ISAAC platform, where a humanoid robot undergoes reinforcement learning training. This feature information enables the robot to better plan its movement paths, anticipate potential risks, and efficiently execute tasks. In goal-oriented navigation tasks, the robot utilizes the spatial distribution features in the map to plan the optimal path from its current location to the target. It also uses semantic features to identify obstacles and key objects along the way, avoiding collisions and ensuring progress toward the target. For grasping tasks, the robot uses semantic features to determine the target object's category and location, and combines spatial distribution features with ground features to adjust its posture and grasping angle to accurately grasp the object. Through continuous exploration and trial in a simulated environment, the robot gradually optimizes its behavioral strategy based on a reward mechanism, improving its success rate and efficiency in completing tasks. During training, reinforcement learning algorithms such as deep Q-networks or proximal policy optimization algorithms can be used to optimize the robot's action selection, enabling it to make the most appropriate decisions based on different scene conditions. This enables autonomous navigation and precise grasping in complex environments, further enhancing the robot's adaptability and task execution capabilities in real-world scenarios.
[0083] The above-mentioned image data processing method can realize 3DGS scene reconstruction, provide the humanoid robot with accurate and up-to-date environmental information, and significantly improve its environmental perception ability; reinforcement learning based on Mesh map and features effectively enhances the robot's task execution capability, making it perform better in tasks such as navigation and grasping, and improving the robot's autonomy and adaptability in complex environments.
[0084] In order to execute the corresponding steps in the above embodiments and various possible methods, an implementation method of an image data processing device 200 is given below. Optionally, the image data processing device 200 can adopt the above Figure 1 The device structure of the electronic device 100 is shown in FIG. Figure 9 , Figure 9 This is a block diagram of an image data processing device 200 provided in an embodiment of the present application. It should be noted that the basic principles and technical effects of the image data processing device 200 provided in this embodiment are the same as those of the above-mentioned embodiments. For the sake of brevity, any details not mentioned in this embodiment are referred to the corresponding contents of the above-mentioned embodiments. In this embodiment, the image data processing device 200 may include: an analysis module 210, a point cloud acquisition module 220, and a map acquisition module 230.
[0085] The analysis module 210 is used to analyze the environment image to obtain first semantic information corresponding to the environment image, wherein the first semantic information is used to indicate the semantic type of the target object and the position of the target object in the environment image.
[0086] The point cloud acquisition module 220 is configured to obtain a target 3D Gaussian point cloud based on the first semantic information, the environment image, and a depth image corresponding to the environment image. The target 3D Gaussian point cloud includes second semantic information derived based on the first semantic information, the second semantic information being used to indicate the semantic type of the target object and its position in the target 3D Gaussian point cloud.
[0087] The map acquisition module 230 is configured to obtain a target map based on the target 3D Gaussian point cloud, wherein the target map includes third semantic information corresponding to the second semantic information, and the third semantic information is configured to indicate the semantic type of the target object and its location in the target map.
[0088] Please refer to Figure 10 , Figure 10 This is a second block diagram of the image data processing device 200 provided in an embodiment of the present application. In this embodiment, the image data processing device 200 may further include a processing module 240. The processing module 240 is configured to perform robot training according to a preset target task and the target map.
[0089] Optionally, the above modules can be stored in the form of software or firmware. Figure 1 The memory 110 shown in FIG. 110 may be solidified in the operating system (OS) of the electronic device 100 and may be Figure 1 Meanwhile, the data, program codes, etc. required to execute the above modules may be stored in the memory 110.
[0090] An embodiment of the present application further provides a readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the image data processing method is implemented.
[0091] In summary, the embodiments of the present application provide an image data processing method, device, electronic device and readable storage medium. First, the corresponding first semantic information is obtained by analyzing the environmental image, and then a target 3D Gaussian point cloud is obtained based on the first semantic information, the environmental image and the depth image corresponding to the environmental image. Finally, a target map is obtained based on the target 3D Gaussian point cloud. The target 3D Gaussian point cloud includes second semantic information obtained based on the first semantic information, and the target map includes third semantic information corresponding to the second semantic information. The first semantic information is used to indicate the semantic type of the target object and its position in the environmental image, the second semantic information is used to indicate the semantic type of the target object and its position in the target 3D Gaussian point cloud, and the third semantic information is used to indicate the semantic type of the target object and its position in the target map. In this way, the situation where map errors are prone to occur due to changes in lighting can be alleviated. At the same time, the semantic information is combined with the semantics to build the map so that the obtained map is rich in semantic information, which facilitates the subsequent acquisition of more accurate information from the map.
[0092] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.
[0093] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0094] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0095] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A method for processing image data, characterized in that: The method comprises: Analyzing the environment image to obtain first semantic information corresponding to the environment image, wherein the first semantic information is used to indicate the semantic type of the target object and the position of the target object in the environment image; Obtaining a target 3D Gaussian point cloud based on the first semantic information, the environment image, and a depth image corresponding to the environment image, wherein the target 3D Gaussian point cloud includes second semantic information obtained based on the first semantic information, and the second semantic information is used to indicate a semantic type of the target object and a position in the target 3D Gaussian point cloud; A target map is obtained based on the target 3D Gaussian point cloud, wherein the target map includes third semantic information corresponding to the second semantic information, and the third semantic information is used to indicate the semantic type of the target object and the position in the target map.
2. The method according to claim 1, characterized in that The obtaining of a target 3D Gaussian point cloud according to the first semantic information, the environment image, and a depth image corresponding to the environment image includes: Obtaining an initial 3D Gaussian point cloud based on the environment image and the depth image; Under the condition that the spatial position of the Gaussian point is constrained according to the depth image, the parameters of the Gaussian point are iteratively updated by an interleaved optimization algorithm based on the initial 3D Gaussian point cloud to obtain the target parameters of the Gaussian point; According to the first semantic information, the target density of Gaussian points is determined, and the second semantic information is obtained.
3. The method according to claim 2, characterized in that The determining the target density of Gaussian points according to the first semantic information includes: Determining a second object from the first object corresponding to the first semantic information; increasing an initial density of Gaussian points corresponding to the second object to obtain a target density of Gaussian points corresponding to the second object; Reduce the initial density of other Gaussian points to obtain the target density of other Gaussian points.
4. The method according to claim 1, wherein The method further comprises: The robot is trained according to the preset target tasks and the target map.
5. The method according to any one of claims 1 to 4, characterized in that The target map is a grid map, and the robot training is performed according to the preset target task and the target map, including: Extracting features from the target map to obtain map features, wherein the map features include at least one of terrain geometric features, object spatial distribution features, and object semantic features; Obtaining strategy information obtained by a strategy network in the robot based on the map features and the target map, obtaining reward and punishment information corresponding to the strategy information, and updating the strategy network according to the reward and punishment information.
6. The method according to claim 5, characterized in that Obtaining a target map according to the target 3D Gaussian point cloud includes: According to the historical 3D Gaussian point cloud and the target 3D Gaussian point cloud, the current global 3D Gaussian point cloud is obtained by stitching; According to the current global 3D Gaussian point cloud, a grid map serving as the target map is obtained through surface reconstruction.
7. An image data processing device, characterized in that: The device comprises: An analysis module, configured to analyze the environment image to obtain first semantic information corresponding to the environment image, wherein the first semantic information is used to indicate a semantic type of a target object and a position of the target object in the environment image; a point cloud acquisition module, configured to obtain a target 3D Gaussian point cloud based on the first semantic information, the environment image, and a depth image corresponding to the environment image, wherein the target 3D Gaussian point cloud includes second semantic information obtained based on the first semantic information, and the second semantic information is used to indicate the semantic type of the target object and its position in the target 3D Gaussian point cloud; A map acquisition module is used to obtain a target map based on the target 3D Gaussian point cloud, wherein the target map includes third semantic information corresponding to the second semantic information, and the third semantic information is used to indicate the semantic type of the target object and the position in the target map.
8. The device according to claim 7, characterized in that The point cloud acquisition module is specifically used to: Obtaining an initial 3D Gaussian point cloud based on the environment image and the depth image; Under the condition that the spatial position of the Gaussian point is constrained according to the depth image, the parameters of the Gaussian point are iteratively updated by an interleaved optimization algorithm based on the initial 3D Gaussian point cloud to obtain the target parameters of the Gaussian point; According to the first semantic information, the target density of Gaussian points is determined, and the second semantic information is obtained.
9. An electronic device, characterized in that: The image data processing method comprises a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor can execute the machine executable instructions to implement the image data processing method according to any one of claims 1 to 6.
10. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image data processing method according to any one of claims 1 to 6 is implemented.