Vector map construction method based on autoregression depth estimation and related device
By using the autoregressive depth estimation model in the autonomous driving system, the problem of not being able to directly acquire depth images is solved, the ability to accurately construct vector maps in autonomous driving vehicles is realized, and the accuracy of path planning and environmental perception is improved.
Patent Information
- Application Number
- CN202510518181.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-24
AI Technical Summary
In the field of autonomous driving, it is impossible to directly acquire depth images, such as inconvenient installation of light field cameras for obtaining depth images, resulting in the inability to build accurate vector maps.
Using the method based on autoregressive depth estimation, the corresponding depth information is accurately estimated based on the scene images taken by the camera, thereby constructing a vector map.
It realizes accurate estimation of depth information and construction of accurate vector maps without direct depth image acquisition, improving the path planning and environmental perception capabilities of autonomous driving vehicles.
Smart Images

Figure CN120047640A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a vector map construction method based on autoregressive depth estimation and related devices. Background Art
[0002] In the field of autonomous driving, vector maps can provide accurate road information, lane lines, traffic signs and other details, providing basic support for path planning, decision-making and control of autonomous vehicles. At the same time, vehicles need to accurately perceive the surrounding environment through vector maps in various weather conditions to ensure safe driving.
[0003] At present, it is generally necessary to construct a vector map based on the image captured by the camera and the depth image corresponding to the image. However, in actual scenarios, due to various factors, it is impossible to directly collect the depth image, for example, it is not convenient to install a light field camera for obtaining the depth image. Therefore, how to obtain the depth image corresponding to the image captured by the camera in order to construct a vector map has become a technical problem that technicians in this field need to solve urgently. Summary of the invention
[0004] The embodiments of the present application provide a vector map construction method and related devices based on autoregressive depth estimation, which can accurately estimate the corresponding depth information based on the scene image taken by the camera, and then construct a vector map by combining the scene image and the depth information.
[0005] The embodiments of the present application can be implemented as follows: In a first aspect, an embodiment of the present application provides a method for constructing a vector map based on autoregressive depth estimation, the method comprising: Obtain a scene image obtained by image acquisition through a camera; According to the scene image and a pre-trained autoregressive depth estimation model, a depth image corresponding to the scene image is obtained, wherein the loss used by the depth estimation model in the image pair training process is calculated based on depth prediction maps of different scales and corresponding second target depth maps, the image pair includes a corresponding sample scene image and a true depth image, the second target depth map is a map obtained by enhancing residual features for the first target depth map, the first target depth map is a fixed-size map obtained by processing the true depth image, and the residual features are calculated based on the first target depth map and the obtained depth prediction map; A vector map is constructed based on the scene image and the depth image.
[0006] In a second aspect, an embodiment of the present application provides a vector map construction device based on autoregressive depth estimation, the device comprising: An image acquisition module is used to acquire a scene image acquired by a camera; A depth estimation module, configured to obtain a depth image corresponding to the scene image according to the scene image and a pre-trained autoregressive depth estimation model, wherein the loss used by the depth estimation model in the image pair training process is calculated based on depth prediction maps of different scales and corresponding second target depth maps, the image pair includes a corresponding sample scene image and a true depth image, the second target depth map is a map obtained by enhancing residual features for the first target depth map, the first target depth map is a map of a fixed size obtained by processing the true depth image, and the residual features are calculated based on the first target depth map and the obtained depth prediction map; The construction module is used to construct a vector map according to the scene image and the depth image.
[0007] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor can execute the machine executable instructions to implement the vector map construction method based on autoregressive depth estimation described in the aforementioned embodiment.
[0008] In a fourth aspect, an embodiment of the present application provides a readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method for constructing a vector map based on autoregressive depth estimation as described in the aforementioned embodiment is implemented.
[0009] The method and related device for constructing a vector map based on autoregressive depth estimation provided in the embodiment of the present application first obtain a scene image obtained by image acquisition by a camera; then, based on the scene image and a pre-trained autoregressive depth estimation model, obtain a depth image corresponding to the scene image, the loss used by the depth estimation model in the image pair training process is calculated based on depth prediction maps of different scales and the corresponding second target depth map, the image pair includes the corresponding sample scene image and the real depth image, the second target depth map is a map obtained by enhancing the residual features of the first target depth map, the first target depth map is a fixed-size map obtained by processing the real depth image, the residual features are calculated based on the first target depth map and the obtained depth prediction map; finally, construct a vector map based on the scene image and the depth image. In this way, the corresponding depth information can be accurately estimated based on the scene image captured by the camera, and then a vector map can be constructed by combining the scene image and the depth information. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0011] Figure 1 A block diagram of an electronic device provided in an embodiment of the present application; Figure 2 One of the flow charts of the method for constructing a vector map based on autoregressive depth estimation provided in an embodiment of the present application; Figure 3 The second flowchart of the method for constructing a vector map based on autoregressive depth estimation provided in an embodiment of the present application; Figure 4 for Figure 3 A schematic flow chart of the sub-steps included in step S100; Figure 5 A schematic diagram of a vector map construction process provided in an embodiment of the present application; Figure 6 for Figure 2 A schematic flow chart of the sub-steps included in step S400; Figure 7 for Figure 6 One of the flow charts of the sub-steps included in sub-step S420; Figure 8 for Figure 6 The second schematic flow chart of the sub-steps included in the neutron step S420; Fig. 9 for Figure 8 A schematic flow chart of the sub-steps included in sub-step S425; Fig.10 for Figure 8 A schematic flow chart of the sub-steps included in sub-step S426; Fig.11 One of the block diagrams of a vector map construction device based on autoregressive depth estimation provided in an embodiment of the present application; Fig.12 The second block diagram of the vector map construction device based on autoregressive depth estimation provided in an embodiment of the present application.
[0012] Icon: 100 - electronic device; 110 - memory; 120 - processor; 130 - communication unit; 200 - vector map construction device based on autoregressive depth estimation; 210 - training module; 220 - image acquisition module; 230 - depth estimation module; 240 - construction module. DETAILED DESCRIPTION
[0013] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations.
[0014] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for which protection is sought, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.
[0015] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0016] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0017] Please refer to Figure 1 , Figure 1 A block diagram of an electronic device 100 provided in an embodiment of the present application. The electronic device 100 may be, but is not limited to, a vehicle-mounted terminal device, a cloud device, a computer, a server, etc. The electronic device 100 may include a memory 110, a processor 120, and a communication unit 130. The memory 110, the processor 120, and the communication unit 130 are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these components can be electrically connected to each other via one or more communication buses or signal lines.
[0018] The memory 110 is used to store programs or data. The memory 110 may be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.
[0019] The processor 120 is used to read / write data or programs stored in the memory 110 and execute corresponding functions. For example, the memory 110 stores a vector map construction device 200 based on autoregressive depth estimation, and the vector map construction device 200 based on autoregressive depth estimation includes at least one software function module that can be stored in the memory 110 in the form of software or firmware. The processor 120 executes various functional applications and data processing by running software programs and modules stored in the memory 110, such as the vector map construction device 200 based on autoregressive depth estimation in the embodiment of the present application, that is, realizing the vector map construction method based on autoregressive depth estimation in the embodiment of the present application.
[0020] The communication unit 130 is used to establish a communication connection between the electronic device 100 and other communication terminals through a network, and to send and receive data through the network.
[0021] It should be understood that Figure 1 The structure shown is only a schematic diagram of the structure of the electronic device 100. The electronic device 100 may also include Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown. Figure 1 Each component shown in the figure can be implemented by hardware, software or a combination thereof.
[0022] Please refer to Figure 2 , Figure 2 One of the flow diagrams of the method for constructing a vector map based on autoregressive depth estimation provided in the embodiment of the present application. The method can be applied to the above-mentioned electronic device. The specific process of the method for constructing a vector map based on autoregressive depth estimation is described in detail below. In this embodiment, the method may include steps S200 to S400.
[0023] Step S200, obtaining a scene image acquired by image acquisition by a camera.
[0024] Step S300: obtaining a depth image corresponding to the scene image according to the scene image and a pre-trained autoregressive depth estimation model.
[0025] Step S400: constructing a vector map based on the scene image and the depth image.
[0026] In this embodiment, a scene image obtained through a camera for constructing a map may be obtained first. The scene image may be an RGB image, that is, a color image. It is understandable that the scene image may be an image obtained in real time so as to construct an online vector map; or a plurality of images may be obtained in advance and saved in an image library, and when a map needs to be constructed, images are obtained from the image library as scenes to construct a vector map. The number of scene images obtained may be one or more, which may be determined in combination with actual needs. When there are multiple scene images, the multiple scene images may be images of the same scene obtained at different shooting angles.
[0027] For example, if six cameras are installed on a vehicle, the images obtained by each of the six cameras can be used as the obtained scene images, and then an online vector map can be constructed based on the six scene images, so that the vehicle can perceive the environment based on the constructed online vector map.
[0028] The electronic device may store a pre-trained autoregressive depth estimation model. The depth estimation model may be pre-trained by the electronic device, or may be trained by other devices and sent to the electronic device. The depth estimation model is trained based on multiple image pairs, and parameter adjustment is performed based on loss during the training process. Each image pair includes a corresponding sample scene image and a real depth image. In the process of training based on an image pair, the loss of this time can be calculated in the following way: the loss is calculated based on the depth prediction map of different scales and the second target depth map corresponding to the depth prediction map of each scale, and the depth prediction map of one scale and the corresponding second target depth map have the same scale. Among them, the second target depth map is a map obtained by enhancing the residual features of the first target depth map, and the first target depth map is a fixed-size map obtained by processing the real depth image, and the residual features are calculated based on the first target depth map and the obtained depth prediction map. In the training process of the depth estimation model, in order to make the prediction result at the current scale closer to the true value, let the model focus on learning the gap between the prediction result and the true value to better compensate for the prediction error. In this way, it is convenient to improve the prediction accuracy of the obtained depth estimation model.
[0029] After obtaining the scene image, the depth estimation model may be used to obtain a depth image corresponding to the scene image based on the obtained scene image. If multiple scene images are obtained, multiple depth images are obtained, and the number of the two is the same. Furthermore, the scales of the scene image and the corresponding depth image are the same.
[0030] After obtaining the scene image and the depth image, a vector map can be constructed by setting the map construction method. The vector map can well describe the geometric shape and semantic information of the map elements. Among them, the specific method of constructing the map based on the scene image and the depth image can be set according to the actual needs, and is not specifically limited here. In this way, when the depth image is not directly obtained through a camera, etc., the above-mentioned depth estimation model can be used to accurately predict the corresponding depth image, and then the vector map can be constructed by combining the scene image and the depth information.
[0031] Please refer to Figure 3 , Figure 3 The second flowchart of the method for constructing a vector map based on autoregressive depth estimation provided in the embodiment of the present application is as follows. In this embodiment, before step S300, the method may further include step S100.
[0032] Step S100: pre-training to obtain the depth estimation model.
[0033] The initial depth estimation model parameters may be set first, and a plurality of image pairs may be acquired, and then training may be performed based on the plurality of image pairs, and the initial depth estimation model parameters may be adjusted until the training is completed, thereby obtaining a trained depth estimation model.
[0034] As a possible implementation, Figure 4 The depth estimation model is obtained in the manner shown. Figure 4 , Figure 4 for Figure 3 Schematic diagram of the flow of sub-steps included in step S100. In this embodiment, step S100 may include sub-steps S110 to S180.
[0035] Sub-step S110: for each image pair among the obtained multiple image pairs, obtain a first image sequence corresponding to the sample scene image in the image pair.
[0036] Sub-step S120 , for each sample mapping image, perform depth prediction using the initial depth estimation model to obtain a depth prediction map of the corresponding scale.
[0037] Sub-step S130: performing image encoding on the real depth image to obtain the first target depth map.
[0038] Sub-step S140, for each scale, according to each depth prediction map smaller than the scale corresponding to the sample scene image, through upsampling processing and weighted summation processing, a depth map to be used with the same scale as the first target depth map is obtained.
[0039] Sub-step S150, obtaining residual features according to the first target depth map and the depth map to be used.
[0040] Sub-step S160: obtaining a second target depth map at the scale according to the first target depth map and the residual feature.
[0041] Sub-step S170, calculating the loss corresponding to the image pair according to the depth prediction map of each scale and the corresponding second target depth map.
[0042] Sub-step S180, optimizing the initial depth estimation model according to the calculated loss, and continuing training until the depth estimation model is obtained.
[0043] In this embodiment, for an image pair, the sample scene image in the image pair is used as the image for predicting the depth map. The sample scene image in the image pair can be processed to obtain an image sequence, and the image sequence is used as the first image sequence corresponding to the sample scene image. The first image sequence includes a plurality of sample mapping images with different scales, wherein the scale of one sample mapping image can be the same as the scale of the sample scene image. The sample mapping images in the first image sequence are multi-scale discretized representations, representing image information at different scales (resolutions). Among them, the above-mentioned sample scene images can be processed by an encoder to obtain the first image sequence. The encoder used can be determined in combination with actual needs, for example, Vector Quantized Variational Auto Encoder (VQ-VAE), and the encoder used is not specifically limited here.
[0044] The first image sequence may be input into an initial depth estimation model to obtain a plurality of depth prediction maps, wherein the scales corresponding to the plurality of depth prediction maps are the same as the scale corresponding to the first image sequence.
[0045] Optionally, taking a scale as the current scale, the depth prediction map at the current scale can be predicted in the following manner: according to each depth prediction map smaller than the current scale obtained based on the first image sequence, and each sample mapping image less than or equal to the current scale in the first image sequence, the depth prediction map of the current scale is predicted, that is: ,in, Represents the depth prediction map of the current scale, represents an autoregressive deep prediction model, represents depth prediction images smaller than the current scale obtained based on the first image sequence, Indicates each sample mapping image in the first image sequence that is smaller than or equal to the current scale. In this way, accurate prediction is facilitated.
[0046] Alternatively, the depth prediction map at the current scale may be predicted in the following manner: according to the depth prediction map corresponding to the maximum scale smaller than the current scale obtained based on the first image sequence, and the sample mapping image equal to the current scale in the first image sequence, the depth prediction map of the current scale is predicted, that is: This makes it easier to complete predictions quickly.
[0047] The real depth image in the image pair may also be image encoded to obtain a first target depth map of a fixed scale. The scale of the first target depth map may be greater than the maximum scale of the scales corresponding to the first image sequence. The scale of the first target depth map may be determined in combination with actual needs. The encoder used when obtaining the first image sequence may be used for processing, but the encoding is not discretized and is directly encoded as a continuous feature.
[0048] For each scale, a second target depth map of the scale can be constructed according to the first target depth map and the predicted depth prediction maps below the scale. The process can be expressed by the following formula:
[0049] in, Represents the depth prediction graph of all scales before the current scale (i.e., scales smaller than the current scale) of the initial depth estimation model Through upsampling operation After being mapped to the resolution of the scale k of the depth of the first target image (this operation includes upsampling processing and weighting processing), a depth map to be used is obtained by calculating the cumulative sum thereof; represents a first target depth map; Represents the true depth feature With cumulative forecast The residual of (i.e., residual feature); It is represented as the target feature after enhancing the residual feature. represents the weighting coefficient; Indicates the current scale (resolution); Represents a scaling operation, where the target feature Adjust to current scale ; Represents the quantization operation, discretizing continuous features into labeled maps; Represents the second target depth map at the current scale n.
[0050] In this embodiment, in order to make the prediction results at the current scale closer to the true value, the initial depth estimation model is focused on learning the gap between the prediction results and the true value to better compensate for the prediction error. Therefore, the weighted residual features are summed with the true value to increase the training weight of the gap part, so that the gap part can be trained more emphatically at the current scale.
[0051] The above autoregressive-based depth estimation method can generate a high-precision depth map by modeling the dependencies between pixels. Depth estimation can provide three-dimensional geometric information of the scene, helping the model used to generate maps to better understand the shape and structure of the road environment. By fusing depth features with image features, richer feature representations can be generated, thereby improving the detection and classification performance of map elements.
[0052] Next, the loss of the depth prediction map and the corresponding second target depth map at each scale is calculated to obtain the loss corresponding to the image pair, and then the parameters of the initial depth estimation model can be optimized based on the loss corresponding to the image pair. After that, the above process is repeated until the training is completed, thereby obtaining a trained autoregressive depth estimation model. Among them, the loss corresponding to an image pair can be obtained by calculating the cross entropy loss, and the process can be expressed by the following formula: , wherein the number of sample mapping images in the first image sequence is N.
[0053] To facilitate understanding of the above training process, a simple example is given below based on an image pair to illustrate the above training process.
[0054] First, the image to be predicted (i.e., a sample scene image) is taken as input, and a series of label mappings are obtained through the vector quantization variational autoencoder (VQ-VAE) ,in It is a multi-scale discretization representation, representing image information at different scales (resolutions), and the image details contained increase in sequence. Specifically, is the label map at the lowest scale (resolution), containing the coarsest information of the sample scene image. Increase, tag mapping The resolution gradually increases ( It is used to represent any label map), and then contains more detailed information. Among them, in a series of label maps, the scale of one of the label maps is the same as the scale of the sample scene image, so that it is easy to predict the depth prediction map with the same scale as the sample scene image.
[0055] Mapping image tags As the starting sequence, the starting sequence is used as the input of the initial depth estimation model as a condition for depth estimation, and the predicted low-scale depth marker map (ie, depth prediction map) is also used as input to autoregressively predict the depth marker map at the current n scale .
[0056] Then, the real depth image corresponding to the aforementioned sample scene image is used as input and passed through the vector quantization variational autoencoder (VQ-VAE) which is exactly the same as the image encoding, but this time the encoding is not discretized and is directly encoded into a continuous feature of a fixed size. Next, based on the continuous features The second target depth map at the current scale n is constructed by combining the predicted low-scale depth marker map .
[0057] Finally, the depth label map of each scale corresponding to the image pair and the corresponding second target depth map are calculated, and the cross entropy loss corresponding to the image pair is calculated. The formula can be recorded as: .
[0058] This model building method uses dynamic targets (i.e., the model's own predictions) during training, rather than relying on predefined static targets (the quantized labeled maps provided by VQ-VAE, i.e., the continuous features mentioned above). ). This allows the model to self-correct during the training phase and has a multimodal solution, thus improving the accuracy of depth estimation.
[0059] As a possible implementation, Figure 5 As shown, the scene images obtained are multiple (i.e., multi-view images of the same scene). In this case, the depth image corresponding to each scene image can be obtained using the depth estimation model. Each scene image can be processed first to obtain a second image sequence corresponding to the scene image, and then the second image sequence is input into the depth estimation model to obtain a depth image with the same scale as the scene image.
[0060] As a possible implementation, Figure 6 Build a vector map as shown. Figure 6 , Figure 6 for Figure 2 Schematic diagram of the flow of sub-steps included in step S400. In this embodiment, step S400 may include sub-steps S410 to S430.
[0061] Sub-step S410, obtaining a global bird's-eye view feature according to the scene image and the depth image.
[0062] In this embodiment, the scene image can be combined with the depth image to project it into the BEV (Bird's Eye View) coordinate system, and then processed to obtain the global bird's eye view feature. On the basis of the above method, other methods can be used to process the scene image, and the processing result can be merged with the processing result of the above method to obtain the global bird's-eye view feature. The following combination Figure 5 , how to obtain the global bird's-eye view features through the latter Give an example.
[0063] Assume that at time t, a scene image is obtained from each of the multiple cameras, thus obtaining multiple scene images. The depth information (i.e., depth image) corresponding to the multiple scene images can be obtained using the depth estimation model. , assuming there are 6 cameras on the car, the depth information obtained is , Indicates different cameras on the car. The intrinsic parameters of the camera (i.e., the optical center) can be obtained ,focal length , external parameter rotation matrix and translation vector .
[0064] The coordinate transformation is then performed to convert each pixel in the image coordinate system of each camera (i.e., the pixel in the scene image) Convert to a 3D point in the camera coordinate system , the specific conversion formula is as follows: , , .
[0065] Then the 3D points are converted to the world coordinate system, and the obtained world coordinates are projected into the BEV coordinate system. Finally, the BEV features of the depth projection are obtained through feature extraction and feature fusion. BEV Characteristics It can be regarded as a feature map.
[0066] The surround multi-view images (i.e., the 6 scene images obtained by the 6 cameras mentioned above) are extracted through the backbone network Resnet-101 to obtain the features of different camera views. ,in is the feature of the i-th view; without using depth information, each view feature is projected into the BEV space, and the BEV feature of the image at the current moment is obtained through the spatial cross attention mechanism and feedforward neural network Among them, the view features can be projected into the BEV space through BEVformer. Image BEV features It can be regarded as a feature map.
[0067] Fusion of deep-projected BEV features via attention mechanism and image BEV features , and perform weighted fusion on the feature maps according to the attention weights. After that, the final global bird’s-eye view feature map is obtained through the maximum pooling and upsampling layers. .
[0068] Sub-step S420, obtaining a target decoding result through a map decoder according to the target bird's-eye view features obtained from the global bird's-eye view features.
[0069] In this embodiment, based on the global bird's-eye view feature Determine the target bird's-eye view features, and then use the pre-trained map decoder to perform multi-granularity queries based on the target bird's-eye view features to obtain the target decoding results. Among them, the target decoding results include the corresponding target instance-level query results and target point-level query results. The target instance-level query results are used to indicate the category of the predicted map element, and the target point-level query results are used to indicate the final predicted point of interest location, which is the location in the target bird's-eye view features. The multi-granularity query method can simultaneously capture the local geometric details and overall category information of the map elements, which is convenient for improving the accuracy of the generated vector map.
[0070] Optionally, as a possible implementation method, the above global bird's-eye view feature can be directly As a target bird's eye view feature , which makes it easy to quickly determine the target's bird's-eye view features.
[0071] Since the lengths of map elements vary, relying solely on BEV features at a single scale cannot meet the requirements for detecting elements of different lengths. Therefore, as another possible implementation method, Figure 7 The target bird's-eye view features are obtained in the manner shown. Figure 7 , Figure 7 for Figure 6 One of the flow diagrams of the sub-steps included in sub-step S420. In this embodiment, sub-step S420 may include sub-steps S4211 and S4212.
[0072] Sub-step S4211, down-sampling the global bird's-eye view feature to obtain a down-sampled global bird's-eye view feature.
[0073] Sub-step S4212: Use a flattened tensor to concatenate the global bird's-eye view feature and the downsampled global bird's-eye view feature to obtain the target bird's-eye view feature.
[0074] In this embodiment, a downsampling module can be used to convert the global bird's-eye view features The spatial resolution is reduced by half, and the global bird's-eye view feature is obtained after downsampling . Then use the flattened tensor concatenation in the encoder and Get multi-scale BEV features as target bird's-eye view features Such dual scales can capture both global features and local details.
[0075] In this embodiment, the map decoder includes a plurality of decoding layers connected in sequence. The target bird's-eye view features and the initial instance-level query results can be input into the map decoder, and the multiple decoding layers are used to iteratively query to obtain the target decoding result. Among them, the initial instance-level query result includes the instance determined at the beginning, which can be specifically a type, etc. For two adjacent decoding layers, the output result of the previous decoding layer (i.e., the decoding result of the previous decoding layer) is optimized by the next decoding layer, and the output result of the last decoding layer is the target decoding result, and each decoding layer includes a multi-granularity attention layer.
[0076] As a possible implementation, Figure 5 As shown, the map decoder has L layers in total, and each decoding layer includes a self-attention layer, a multi-granularity attention layer, and a feedforward neural network connected in sequence. That is, each decoding layer can be composed of a self-attention mechanism, a multi-granularity attention mechanism, and a feedforward neural network. The self-attention layer, based on the self-attention mechanism, allows different instance-level queries to interact and capture the relationship between different instances, which helps the map construction model (which may include a module for obtaining the aforementioned global bird's-eye view features, a map decoder, a module for generating a map based on the target decoding results, etc.) to better understand the overall structure of the map elements. The multi-granularity attention layer is used to perform multi-granularity queries through a multi-granularity attention mechanism. Feedforward neural network: further nonlinear transformation is performed on the current instance-level query results and point-level query results (i.e., the output results of the multi-granularity attention layer in the decoding layer) to obtain a more complex representation; at the same time, the intrinsic connection between the instance-level query and the point-level query is further integrated and optimized.
[0077] Instance-level queries can effectively capture the overall category information of road elements, but lack accurate geometric position representation; point-level queries can provide accurate geometric position information, but multiple queries need to be aggregated to represent an instance. In summary, in order to capture comprehensive and accurate instance features, a multi-granularity attention mechanism is used in this embodiment. The multi-granularity attention mechanism consists of two components: a multi-granularity aggregator and a point instance interaction. That is, the multi-granularity attention layer includes a multi-granularity aggregator and a point instance interaction component.
[0078] In the multi-granularity aggregator, instance-level query results and target bird’s-eye view features Specifically, multiple reference points are introduced for each instance-level query to obtain the bird’s-eye view of the target features. Aggregate distant features in the map, so that the map building model can see the features from a bird's eye view of the target Aggregate a wider range of information.
[0079] As a possible implementation, Figure 8 The output result of a decoding layer is obtained in the following way. Figure 8 , Figure 8 for Figure 6 The second flow chart of the sub-steps included in sub-step S420. In this embodiment, sub-step S420 may include sub-steps S422 to S426.
[0080] Sub-step S422, obtaining a first reference point set of the current decoding layer, and obtaining a position code according to the first reference point set and the target bird's-eye view feature.
[0081] In this embodiment, the multi-granularity aggregator in the first decoding layer (i.e., the first decoding layer) obtains the initial instance-level query results As input. The instances in the initial instance-level query result can be pre-specified manually or obtained by other means. The multi-granularity aggregator in other decoding layers takes the point-level query result output by the previous decoding layer (hereinafter referred to as the first point-level query result) and the reference point set used for point-level query in the previous decoding layer (hereinafter referred to as the second reference point set) As input. Among them, Indicates the total number of instance-level queries, is the total number of points belonging to an instance (the points of an instance are subsequently referred to as interest points). The reference points of the first decoding layer are given by Prediction, the reference points of subsequent layers are updated by the reference points of the previous layer.
[0082] That is, when the current decoding layer is the first decoding layer, the first reference point set is determined according to the initial instance-level query result. When the current decoding layer is not the first decoding layer, the first reference point set is determined according to the second reference point set used in the point-level query of the previous decoding layer and the first point-level query result output by the previous decoding layer.
[0083] The determination method of the first reference point set can be expressed by the following formula:
[0084] in, Indicates the current layer, represents the sigmoid activation function, represents the inverse sigmoid activation function, Indicates The reference point set of the layer, Represents a multilayer perceptron.
[0085] Since an instance is represented as a sequence of points, position encoding is added to the instance-level query. In the case of location, use Generate positional encoding :
[0086] in, is a projection layer used to generate position embedding features from the reference points. The above position encoding is also combined with the target bird’s-eye view features get.
[0087] Sub-step S423, calculating the position offset corresponding to the first reference point set according to the position code and the first instance-level query result output by the previous decoding layer.
[0088] Sub-step S424: calculating, according to the first reference point set and the position offset, the sampling points corresponding to the first reference points in the first reference point set.
[0089] In this embodiment, each reference point is assigned sampling points, which are used to aggregate features to enhance the features of the reference points. and the first instance-level query result output by the previous decoding layer (i.e., the instance-level query result input to this decoding layer) , calculate the position offset of each first reference point in the first reference point set . No. The position offset of each first reference point in the multi-granular attention layer in the layer decoding layer The calculation process can be shown as the following formula:
[0090] in, Expand accordingly to match By using the position offset and the first reference point, that is, by and , get the sampling position (i.e., obtaining the sampling point of the first reference point).
[0091] Sub-step S425 , obtaining a second instance-level query result and a second point-level query result according to the sampling points of each first reference point and the target bird's-eye view feature.
[0092] Afterwards, feature extraction and analysis may be performed based on the sampling points of each first reference point and the target bird's-eye view features, thereby obtaining a second instance-level query result and a second point-level query result. Fig. 9 The second instance-level query result and the second point-level query result are obtained in the manner shown. Fig. 9 , Fig. 9 for Figure 8 Schematic diagram of the flow of sub-steps included in sub-step S425. In this embodiment, sub-step S425 may include sub-steps S4251 to S4254.
[0093] Sub-step S4251, obtaining a first weight corresponding to each sampling point according to the position code and the first instance-level query result output by the previous decoding layer.
[0094] Sub-step S4252, according to the first weight corresponding to each sampling point, calculate the second weight of each instance corresponding to the sampling point and the third weight of each interest point corresponding to the sampling point.
[0095] Sub-step S4253, obtaining the second instance-level query result according to the target bird's-eye view features, the sampling points corresponding to the instances, and the second weight corresponding to the instances.
[0096] Sub-step S4254, obtaining the third point-level query result according to the bird's-eye view features of the target, the sampling points corresponding to the points of interest, and the third weights corresponding to the points of interest.
[0097] In this embodiment, after obtaining the position code After that, you can also encode according to the position And the first instance-level query result output by the previous decoding layer , analyze and obtain the weight corresponding to each sampling point as the first weight. This process can be expressed by the following formula: .
[0098] Then, the second instance-level query result can be generated by the weighted sum of the sampled features And the third point level query results , that is, the instance-level query results and point-level query results generated for the first time in this layer are generated. This process can be expressed by the following formula:
[0099] in, For example The index of the point, Assigned to the reference point The index between sampling points, express exist on Normalized weight (i.e. the second weight corresponding to the instance), express exist on Normalized weight (i.e. the third weight corresponding to the point of interest), is a bilinear sampling operator.
[0100] Sub-step S426, obtaining a third instance-level query result and a third point-level query result through point-instance interaction according to the second instance-level query result and the second point-level query result.
[0101] In this embodiment, the second instance-level query result and the second point-level query result can be input into the point-instance interaction component to process and obtain a third instance-level query result and a third point-level query result. The purpose of point-instance interaction is to enhance the interaction of location information and category information between two queries of different granularities. Point-instance interaction includes two different attention operators: P2P (point-to-point) attention and P2I (point-to-instance) attention. The third instance-level query result and the third point-level query result are obtained through the processing of the above two operators. The output result of this decoding layer is determined based on the third instance-level query result and the third point-level query result. For example, in Figure 5 Under the structure shown, a decoding layer also includes a feedforward neural network after the multi-granularity attention layer in the decoding layer, and the multi-granularity attention layer outputs a third instance-level query result and a third point-level query result to the feedforward neural network, and the output of the feedforward neural network based on the third instance-level query result and the third point-level query result is the output result of the decoding layer.
[0102] Optionally, you can Fig.10 The third instance-level query result and the third point-level query result are obtained in the manner shown. Fig.10 , Fig.10 for Figure 8 Schematic diagram of the flow of sub-steps included in sub-step S426. In this embodiment, sub-step S426 may include sub-steps S4261 to S4265.
[0103] Sub-step S4261, obtaining the first position information feature corresponding to the instance according to the obtained sampling points, the second weight corresponding to each instance and the target bird's-eye view feature.
[0104] Sub-step S4262, obtaining the second position information feature corresponding to the point of interest according to the obtained sampling points, the third weight corresponding to each point of interest and the target bird's-eye view feature.
[0105] Sub-step S4263, according to the second point-level query result and the first location information feature, an initial third point-level query result is calculated through the P2P attention mechanism. Wherein, when the current decoding layer is not the first decoding layer, the initial third point-level query result is obtained based on the second point-level query result, the first location information feature, the first point-level query result and the location information feature corresponding to the first point-level query result.
[0106] Sub-step S4264, obtaining the third point-level query result through the P2I attention mechanism according to the initial third point-level query result, the second location information feature, the second instance query result and the first location information feature.
[0107] Sub-step S4265, according to the third point-level query result, the third instance-level query result is obtained by aggregating the interest points corresponding to the same instance in the third point-level query result.
[0108] In this embodiment, the The sampling position (i.e. sampling point) obtained by multi-granular aggregation in the layer decoding layer and attention weights (i.e., the second weight corresponding to the instance), (i.e., the third weight corresponding to the point of interest) is tiled and cascaded to encode the location information features in the P2P attention mechanism and the P2I attention mechanism:
[0109] in, is a multi-layer perceptron for instance-level queries, It is a multi-layer perceptron for point-level queries. , is the corresponding generated position embedding. Represents the first position information feature corresponding to the instance, The second location information feature corresponding to the point of interest is represented by the above location information feature, which is further combined with the target bird's-eye view feature get.
[0110] In the P2P attention mechanism, since the coordinates of map elements are based on the previous layer ( layer), so these point-level queries are more important in predicting the current layer ( The P2P attention module is designed to take the coordinates from the current The point-level query of the layer and the previous layer is used as the input of the P2P attention module. The processing based on the P2P attention mechanism is as follows:
[0111] Since the multi-granular attention layer in the first decoding layer has no previous decoding layer, a self-attention operation is performed in the first multi-granular attention layer. In subsequent multi-granular attention layers, the previous point-level query results With the currently generated point-level query results Hybrid performs cross-attention operations. Represents the initial third-level query result.
[0112] That is, when the current decoding layer is the first decoding layer, the initial third point-level query result is obtained through a self-attention operation based on the second point-level query result and the first position information feature. When the current decoding layer is not the first decoding layer, the initial third point-level query result is obtained through a cross-attention operation based on the second point-level query result, the first position information feature, the first point-level query result and the first point-level query result.
[0113] In the P2I attention mechanism, following the P2P attention mechanism, the P2I attention operation realizes information interaction between different granularities. Point-level queries use cross-attention to exchange geometric information with instance-level queries: , represents the third point-level query result.
[0114] Finally, the third point-level query results belonging to the same instance-level query are aggregated to update the corresponding instance-level query, thereby obtaining the third instance-level query result of this decoding layer. The process can be shown as follows:
[0115] in, express The index of .
[0116] Point-granularity queries are used to predict point locations using MLP as a regression head, while instance-granularity queries are used to predict the category of a map element using another MLP. By leveraging the multi-granularity aggregator and point-instance interaction components, multi-granularity queries can be generated and updated. At the same time, the geometry and category of individual map elements can be effectively perceived.
[0117] Each decoding layer performs the above processing, and when the output of the last decoding layer is obtained, the target decoding result is obtained.
[0118] Sub-step S430, obtaining the vector map according to the target decoding result.
[0119] After obtaining the target decoding result, a vector map may be generated based on the target decoding result in any manner.
[0120] The above method provided in the embodiment of the present application is a multi-granularity vector map construction method based on autoregressive depth estimation. By constructing an autoregressive depth model, this method can effectively improve the accuracy of depth estimation information; and, by utilizing multi-granularity query, it can simultaneously capture global features and local information, which is conducive to more comprehensive identification of map information; and by fusing depth information, the geometric position information of the global bird's-eye view features is improved, thereby constructing a more accurate vector map.
[0121] In order to execute the corresponding steps in the above embodiments and various possible methods, an implementation method of a vector map construction device 200 based on autoregressive depth estimation is given below. Optionally, the vector map construction device 200 based on autoregressive depth estimation can adopt the above Figure 1 The device structure of the electronic device 100 is shown. Fig.11 , Fig.11 This is one of the block diagrams of the vector map construction device 200 based on autoregressive depth estimation provided in the embodiment of the present application. It should be noted that the basic principle and technical effect of the vector map construction device 200 based on autoregressive depth estimation provided in this embodiment are the same as those in the above-mentioned embodiment. For the sake of brief description, for parts not mentioned in this embodiment, please refer to the corresponding contents in the above-mentioned embodiment. In this embodiment, the vector map construction device 200 based on autoregressive depth estimation may include: an image acquisition module 220, a depth estimation module 230 and a construction module 240.
[0122] The image acquisition module 220 is used to obtain a scene image acquired by a camera.
[0123] The depth estimation module 230 is used to obtain a depth image corresponding to the scene image according to the scene image and a pre-trained autoregressive depth estimation model. The loss used by the depth estimation model in the image pair training process is calculated based on depth prediction images of different scales and corresponding second target depth images, the image pair includes a corresponding sample scene image and a true depth image, the second target depth image is a map obtained by enhancing the residual features of the first target depth map, the first target depth map is a fixed-size map obtained by processing the true depth image, and the residual features are calculated based on the first target depth map and the obtained depth prediction map.
[0124] The construction module 240 is used to construct a vector map according to the scene image and the depth image.
[0125] Please refer to Fig.12 , Fig.12The second block diagram of the vector map construction device 200 based on autoregressive depth estimation provided in the embodiment of the present application. In this embodiment, the vector map construction device 200 based on autoregressive depth estimation may further include a training module 210. The training module 210 is used to pre-train the depth estimation model.
[0126] Optionally, the above modules can be stored in the form of software or firmware. Figure 1 The memory 110 shown in the figure may be fixed in the operating system (OS) of the electronic device 100 and may be Figure 1 Meanwhile, the data and program codes required for executing the above modules may be stored in the memory 110.
[0127] An embodiment of the present application also provides a readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the vector map construction method based on autoregressive depth estimation is implemented.
[0128] In summary, the embodiment of the present application provides a method and related device for constructing a vector map based on autoregressive depth estimation. First, a scene image is obtained by acquiring an image through a camera; then, based on the scene image and a pre-trained autoregressive depth estimation model, a depth image corresponding to the scene image is obtained. The loss used in the depth estimation model in the image pair training process is calculated based on depth prediction maps of different scales and corresponding second target depth maps. The image pair includes a corresponding sample scene image and a true depth image. The second target depth map is a map obtained by enhancing the residual features of the first target depth map. The first target depth map is a fixed-size map obtained by processing the true depth image. The residual features are calculated based on the first target depth map and the obtained depth prediction map; finally, a vector map is constructed based on the scene image and the depth image. In this way, the corresponding depth information can be accurately estimated based on the scene image taken by the camera, and then a vector map can be constructed in combination with the scene image and the depth information.
[0129] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0130] In addition, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.
[0131] If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program codes.
[0132] The above description is only an optional embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A vector map construction method based on autoregressive depth estimation, characterized in that: The method comprises: Obtain a scene image obtained by image acquisition through a camera; According to the scene image and a pre-trained autoregressive depth estimation model, a depth image corresponding to the scene image is obtained, wherein the loss used by the depth estimation model in the image pair training process is calculated based on depth prediction maps of different scales and corresponding second target depth maps, the image pair includes a corresponding sample scene image and a true depth image, the second target depth map is a map obtained by enhancing residual features for the first target depth map, the first target depth map is a fixed-size map obtained by processing the true depth image, and the residual features are calculated based on the first target depth map and the obtained depth prediction map; A vector map is constructed based on the scene image and the depth image.
2. The method according to claim 1, characterized in that The method further comprises: Pre-training to obtain the depth estimation model; The pre-training to obtain the depth estimation model includes: For each image pair among the obtained multiple image pairs, a first image sequence corresponding to the sample scene image in the image pair is obtained, wherein the first image sequence includes a plurality of sample mapping images with different scales; For each sample mapping image, the initial depth estimation model is used to perform depth prediction to obtain a depth prediction map of the corresponding scale; Performing image encoding on the real depth image to obtain the first target depth map; For each scale, according to each depth prediction map smaller than the scale corresponding to the sample scene image, a depth map to be used having the same scale as the first target depth map is obtained through upsampling processing and weighted summation processing; Obtaining residual features according to the first target depth map and the depth map to be used; Obtaining a second target depth map at the scale according to the first target depth map and the residual feature; According to the depth prediction map of each scale and the corresponding second target depth map, the loss corresponding to the image pair is calculated; The initial depth estimation model is optimized according to the calculated loss, and training is continued until the depth estimation model is obtained.
3. The method according to claim 1 or 2, characterized in that: The constructing of a vector map according to the scene image and the depth image comprises: Obtaining a global bird's-eye view feature according to the scene image and the depth image; According to the target bird's-eye view features obtained from the global bird's-eye view features, a target decoding result is obtained through a map decoder, wherein the target decoding result includes a corresponding target instance-level query result and a target point-level query result; The vector map is obtained according to the target decoding result.
4. The method according to claim 3, characterized in that The map decoder includes a plurality of decoding layers connected in sequence, and obtaining a target decoding result through the map decoder according to the target bird's-eye view feature obtained from the global bird's-eye view feature includes: According to the target bird's-eye view features and the initial instance-level query results, the multiple decoding layers are used to iteratively query to obtain the target decoding result, wherein, for two adjacent decoding layers, the output result of the previous decoding layer is optimized by the next decoding layer, and the output result of the last decoding layer is the target decoding result, and each decoding layer includes a multi-granularity attention layer.
5. The method according to claim 4, characterized in that The querying using the multiple decoding layers iteratively according to the target bird's-eye view features and the initial instance-level query result to obtain the target decoding result includes: Obtaining a first reference point set of the current decoding layer, and obtaining a position code according to the first reference point set and the target bird's-eye view features, wherein when the current decoding layer is the first decoding layer, the first reference point set is determined according to the initial instance-level query result; when the current decoding layer is not the first decoding layer, the first reference point set is determined according to the second reference point set used in the point-level query of the previous decoding layer and the first point-level query result output by the previous decoding layer; Calculate the position offset corresponding to the first reference point set according to the position code and the first instance-level query result output by the previous decoding layer; Calculate, according to the first reference point set and the position offset, the sampling point corresponding to each first reference point in the first reference point set; Obtaining a second instance-level query result and a second point-level query result according to the sampling points of each first reference point and the bird's-eye view features of the target; According to the second instance-level query result and the second point-level query result, a third instance-level query result and a third point-level query result are obtained through point-instance interaction, wherein the output result of this decoding layer is determined based on the third instance-level query result and the third point-level query result.
6. The method according to claim 5, characterized in that The obtaining of the second instance-level query result and the second point-level query result according to the sampling points of each first reference point and the target bird's-eye view feature includes: Obtaining a first weight corresponding to each sampling point according to the position code and a first instance-level query result output by a previous decoding layer; According to the first weight corresponding to each sampling point, the second weight of each instance corresponding to the sampling point and the third weight of each point of interest corresponding to the sampling point are calculated; Obtaining the second instance-level query result according to the bird's-eye view features of the target, the sampling points corresponding to the instances, and the second weight corresponding to the instances; The third point-level query result is obtained according to the bird's-eye view features of the target, the sampling points corresponding to the points of interest, and the third weight corresponding to the points of interest.
7. The method according to claim 6, characterized in that The obtaining of a third instance-level query result and a third point-level query result through point-instance interaction according to the second instance-level query result and the second point-level query result includes: Obtaining a first position information feature corresponding to the instance according to the obtained sampling points, the second weight corresponding to each instance, and the target bird's-eye view feature; Obtaining second position information features corresponding to the points of interest according to the obtained sampling points, the third weights corresponding to the points of interest, and the bird's-eye view features of the target; According to the second point-level query result and the first location information feature, an initial third point-level query result is calculated through a P2P attention mechanism, wherein when the current decoding layer is not the first decoding layer, the initial third point-level query result is obtained based on the second point-level query result, the first location information feature, the first point-level query result, and the location information feature corresponding to the first point-level query result; Obtaining the third point-level query result through a P2I attention mechanism according to the initial third point-level query result, the second location information feature, the second instance query result, and the first location information feature; According to the third point-level query result, the third instance-level query result is obtained by aggregating interest points corresponding to the same instance in the third point-level query result.
8. The method according to claim 3, characterized in that The target bird's-eye view feature is obtained by: Downsampling the global bird's-eye view feature to obtain a downsampled global bird's-eye view feature; The global bird's-eye view feature and the downsampled global bird's-eye view feature are connected in series using a flattened tensor to obtain the target bird's-eye view feature.
9. A vector map construction device based on autoregressive depth estimation, characterized in that: The device comprises: An image acquisition module is used to acquire a scene image acquired by a camera; A depth estimation module, configured to obtain a depth image corresponding to the scene image according to the scene image and a pre-trained autoregressive depth estimation model, wherein the loss used by the depth estimation model in the image pair training process is calculated based on depth prediction maps of different scales and corresponding second target depth maps, the image pair includes a corresponding sample scene image and a true depth image, the second target depth map is a map obtained by enhancing residual features for the first target depth map, the first target depth map is a map of a fixed size obtained by processing the true depth image, and the residual features are calculated based on the first target depth map and the obtained depth prediction map; The construction module is used to construct a vector map according to the scene image and the depth image.
10. An electronic device, characterized in that: It includes a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor can execute the machine executable instructions to implement the vector map construction method based on autoregressive depth estimation as described in any one of claims 1-8.
Citation Information
Patent Citations
Aerial image geographic positioning method based on spatial scale attention mechanism and vector map
CN113239952A
Depth map acquisition method and device, and storage medium
CN113724311A
Image-laser radar data fusion method based on mixed attention mechanism
CN114398937A
Monocular vision inertial navigation positioning method based on self-supervised deep learning
CN114526728A
Self-supervised depth estimation framework for indoor environment
CN116745813A
Cited By
Peer-to-peer network federal collaborative decision-making method and device
CN120499189A
A peer-to-peer network federation collaborative decision-making method and device
CN120499189B