Data processing method, point cloud data processing method, computing device and storage medium

By combining the data processing model of the data processing module and the result adjustment module, the problem of inaccurate point cloud data processing is solved, and the accuracy and applicability of the target object recognition results are achieved.

CN120849883APending Publication Date: 2025-10-28ALIBABA (CHINA) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410515948.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-25
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing neural network models have the problem of inaccurate processing results when processing point cloud data and cannot meet the needs of actual application scenarios.

Method used

A data processing model combining a data processing module and a result adjustment module is adopted. The data processing module is used to perform initial object recognition. The result adjustment module is used to determine the result adjustment parameters according to the initial object recognition results, and the initial object recognition results are adjusted to obtain the target object recognition results.

Benefits of technology

The accuracy of point cloud data processing is achieved, the needs of actual application scenarios are met, and the applicability of the model and the accuracy of recognition results are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849883A_ABST
    Figure CN120849883A_ABST
Patent Text Reader

Abstract

Embodiments of the present specification provide a data processing method, a point cloud data processing method, a computing device and a storage medium, the point cloud data processing method comprising: determining to-be-processed point cloud data, and inputting the to-be-processed point cloud data into a data processing model, the data processing model comprises a data processing module and a result adjusting module; using the data processing module to perform identification processing on the to-be-processed point cloud data to obtain an initial object identification result of the to-be-processed point cloud data; determining a result adjustment parameter by using the result adjustment module according to the to-be-processed point cloud data and the initial object recognition result, and performing result adjustment processing on the initial object recognition result according to the result adjustment parameter to obtain a target object recognition result of the to-be-processed point cloud data; therefore, the target object recognition result corresponding to the to-be-processed point cloud data is accurately obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of artificial intelligence technology, and in particular to a point cloud data processing method. One or more embodiments of this specification also relate to a data processing method, another point cloud data processing method, a machine learning model training method, a data processing model training method, a traffic data processing method, a point cloud data processing device, a data processing device, another point cloud data processing device, a machine learning model training device, a data processing model training device, a traffic data processing device, a computing device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] With the continuous development of artificial intelligence technology, when it is necessary to process point cloud data, neural network models can be used to process the point cloud data and obtain the corresponding processing results.

[0003] However, existing neural network models suffer from inaccurate processing results when processing point cloud data, failing to meet the demands of real-world applications. Therefore, accurately obtaining the processing results for point cloud data is a pressing issue that needs to be addressed. Summary of the Invention

[0004] In view of this, embodiments of this specification provide a point cloud data processing method. One or more embodiments of this specification simultaneously relate to a data processing method, another point cloud data processing method, a machine learning model training method, a data processing model training method, a traffic data processing method, a point cloud data processing device, a data processing device, another point cloud data processing device, a machine learning model training device, a data processing model training device, a traffic data processing device, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies in the prior art that prevent accurate acquisition of processing results corresponding to point cloud data.

[0005] According to a first aspect of the embodiments of this specification, a point cloud data processing method is provided, comprising:

[0006] The point cloud data to be processed is determined, and the point cloud data to be processed is input into the data processing model, wherein the data processing model includes a data processing module and a result adjustment module:

[0007] The data processing module is used to perform recognition processing on the point cloud data to be processed, and the initial object recognition result of the point cloud data to be processed is obtained.

[0008] Using the result adjustment module, result adjustment parameters are determined based on the point cloud data to be processed and the initial object recognition result. The initial object recognition result is then adjusted according to the result adjustment parameters to obtain the target object recognition result of the point cloud data to be processed.

[0009] According to a second aspect of the embodiments of this specification, a point cloud data processing apparatus is provided, comprising:

[0010] The data determination module is configured to determine the point cloud data to be processed and input the point cloud data to be processed into the data processing model, wherein the data processing model includes a data processing module and a result adjustment module;

[0011] The result acquisition module is configured to use the data processing module to perform recognition processing on the point cloud data to be processed, and obtain the initial object recognition result of the point cloud data to be processed.

[0012] The result adjustment module is configured to determine result adjustment parameters based on the point cloud data to be processed and the initial object recognition result, and to perform result adjustment processing on the initial object recognition result based on the result adjustment parameters to obtain the target object recognition result of the point cloud data to be processed.

[0013] According to a third aspect of the embodiments of this specification, a data processing method is provided, comprising:

[0014] The data to be processed is determined and input into the data processing model, wherein the data processing model includes a data processing module and a result adjustment module;

[0015] The data processing module is used to perform recognition processing on the data to be processed to obtain the initial recognition result of the data to be processed:

[0016] Using the result adjustment module, result adjustment parameters are determined based on the data to be processed and the initial recognition result, and the initial recognition result is adjusted according to the result adjustment parameters to obtain the target recognition result of the data to be processed.

[0017] According to a fourth aspect of the embodiments of this specification, a data processing apparatus is provided, comprising:

[0018] The data determination module is configured to determine the data to be processed and input the data to be processed into the data processing model, wherein the data processing model includes a data processing module and a result adjustment module;

[0019] The result acquisition module is configured to use the data processing module to identify and process the data to be processed, and obtain the initial identification result of the data to be processed.

[0020] The result adjustment module is configured to determine result adjustment parameters based on the data to be processed and the initial recognition result, and to perform result adjustment processing on the initial recognition result based on the result adjustment parameters to obtain the target recognition result of the data to be processed.

[0021] According to a fifth aspect of the embodiments of this specification, another point cloud data processing method is provided, including:

[0022] Identify the point cloud data to be processed;

[0023] The point cloud data to be processed is input into a data processing model to obtain the target object recognition result of the point cloud data to be processed. The data processing model includes a data processing module and a result adjustment module. The data processing module is used to perform recognition processing on the point cloud data to be processed to obtain an initial object recognition result of the point cloud data to be processed. The result adjustment module is used to determine a result adjustment parameter based on the point cloud data to be processed and the initial object recognition result, and to perform result adjustment processing on the initial object recognition result based on the result adjustment parameter to obtain the target object recognition result of the point cloud data to be processed.

[0024] According to a sixth aspect of the embodiments of this specification, another point cloud data processing apparatus is provided, comprising:

[0025] The data determination module is configured to determine the point cloud data to be processed.

[0026] The result acquisition module is configured to input the point cloud data to be processed into a data processing model to obtain the target object recognition result of the point cloud data to be processed. The data processing model includes a data processing module and a result adjustment module. The data processing module is used to perform recognition processing on the point cloud data to be processed to obtain an initial object recognition result of the point cloud data to be processed. The result adjustment module is used to determine result adjustment parameters based on the point cloud data to be processed and the initial object recognition result, and perform result adjustment processing on the initial object recognition result based on the result adjustment parameters to obtain the target object recognition result of the point cloud data to be processed.

[0027] According to a seventh aspect of the embodiments of this specification, a machine learning model training method is provided, comprising:

[0028] The machine learning model to be trained and the training data associated with the target task are determined, wherein the machine learning model to be trained includes a data processing module and a result adjustment module.

[0029] The data processing module is used to perform recognition processing on the training data to obtain the initial recognition result of the training data;

[0030] Using the result adjustment module, result adjustment parameters are determined based on the training data and the initial recognition result, and the initial recognition result is adjusted according to the result adjustment parameters to obtain the target recognition result of the training data;

[0031] Based on the target recognition result, the model parameters of the machine learning model are adjusted to obtain a trained target model, wherein the target model is used to perform the target task.

[0032] According to an eighth aspect of the embodiments of this specification, a machine learning model training apparatus is provided, comprising:

[0033] The training data determination module is configured to determine the machine learning model to be trained and the training data associated with the target task, wherein the machine learning model to be trained includes a data processing module and a result adjustment module.

[0034] The result acquisition module is configured to use the data processing module to perform recognition processing on the training data to obtain the initial recognition result of the training data:

[0035] The result adjustment module is configured to determine result adjustment parameters based on the training data and the initial recognition result, and to perform result adjustment processing on the initial recognition result based on the result adjustment parameters to obtain the target recognition result of the training data.

[0036] The model training module is configured to adjust the model parameters of the machine learning model based on the target recognition result to obtain a trained target model, wherein the target model is used to perform the target task.

[0037] According to a ninth aspect of the embodiments of this specification, a data processing model training method is provided, comprising:

[0038] The data processing model to be trained and the training point cloud data associated with the point cloud data processing task are determined, wherein the object recognition model to be trained includes a data processing module and a result adjustment module.

[0039] The data processing module is used to perform object recognition processing on the training point cloud data to obtain the initial object recognition result of the training point cloud data.

[0040] Using the data processing module, based on the training point cloud data and the initial object recognition result, a result adjustment parameter is determined, and the initial object recognition result is adjusted according to the result adjustment parameter to obtain the target object recognition result of the training point cloud data;

[0041] Based on the target object recognition result, the model parameters of the data processing model are adjusted to obtain a trained data processing model, wherein the data processing model is used to perform the point cloud data processing task.

[0042] According to a tenth aspect of the embodiments of this specification, a data processing model training apparatus is provided, comprising:

[0043] The training data determination module is configured to determine the data processing model to be trained and the training point cloud data associated with the point cloud data processing task, wherein the object recognition model to be trained includes a data processing module and a result adjustment module.

[0044] The result acquisition module is configured to use the data processing module to perform object recognition processing on the training point cloud data to obtain the initial object recognition result of the training point cloud data.

[0045] The result adjustment module is configured to use the data processing module to determine result adjustment parameters based on the training point cloud data and the initial object recognition result, and to perform result adjustment processing on the initial object recognition result based on the result adjustment parameters to obtain the target object recognition result of the training point cloud data.

[0046] The model training module is configured to adjust the model parameters of the data processing model based on the target object recognition result to obtain a trained data processing model, wherein the data processing model is used to perform the point cloud data processing task.

[0047] According to the eleventh aspect of the embodiments of this specification, a traffic data processing method is provided, including:

[0048] The point cloud data to be processed corresponding to the target traffic scene is determined, and the point cloud data to be processed is input into the data processing model, wherein the data processing model includes a data processing module and a result adjustment module;

[0049] The data processing module is used to identify and process the point cloud data to be processed, and the initial traffic object identification result of the point cloud data to be processed is obtained.

[0050] Using the result adjustment module, result adjustment parameters are determined based on the point cloud data to be processed and the initial traffic object recognition result. The initial traffic object recognition result is then adjusted according to the result adjustment parameters to obtain the target traffic object recognition result of the point cloud data to be processed.

[0051] According to a twelfth aspect of the embodiments of this specification, a traffic data processing apparatus is provided, comprising:

[0052] The data determination module is configured to determine the point cloud data to be processed corresponding to the target traffic scene and input the point cloud data to be processed into the data processing model, wherein the data processing model includes a data processing module and a result adjustment module;

[0053] The result acquisition module is configured to use the data processing module to identify and process the point cloud data to be processed, and obtain the initial traffic object identification result of the point cloud data to be processed.

[0054] The result adjustment module is configured to determine result adjustment parameters based on the point cloud data to be processed and the initial traffic object recognition result, and to perform result adjustment processing on the initial traffic object recognition result based on the result adjustment parameters to obtain the target traffic object recognition result of the point cloud data to be processed.

[0055] According to a thirteenth aspect of the embodiments of this specification, a computing device is provided, comprising:

[0056] Memory and processor;

[0057] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, they implement the steps of the above-described point cloud data processing method, data processing method, another point cloud data processing method, machine learning model training method, data processing model training method, or traffic data processing method.

[0058] According to a fourteenth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions, which, when executed by a processor, implement the steps of the above-described point cloud data processing method, a data processing method, another point cloud data processing method, a machine learning model training method, a data processing model training method, or a traffic data processing method.

[0059] According to a fifteenth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described point cloud data processing method, a data processing method, another point cloud data processing method, a machine learning model training method, a data processing model training method, or a traffic data processing method.

[0060] In one or more embodiments provided in this specification, the point cloud data processing method first inputs the point cloud data to be processed into a data processing model, and uses the data processing module in the data processing model to perform recognition processing on the point cloud data to be processed to obtain the initial object recognition result of the point cloud data to be processed. Then, in order to avoid the problem of inaccurate initial object recognition result, the result adjustment module in the data processing model is used to determine the result adjustment parameters based on the point cloud data to be processed and the initial object recognition result, and the initial object recognition result is adjusted according to the result adjustment parameters. This achieves accurate acquisition of the target object recognition result corresponding to the point cloud data to be processed, and avoids the problem that the processing result of the point cloud data is inaccurate, which leads to the inability to meet the needs of point cloud data processing in actual application scenarios. Attached Figure Description

[0061] Figure 1 This is an application diagram of a point cloud data processing method provided in one embodiment of this specification;

[0062] Figure 2 This is a flowchart illustrating a point cloud data processing method provided in one embodiment of this specification;

[0063] Figure 3 This is a flowchart illustrating the processing procedure of a point cloud data processing method provided in one embodiment of this specification.

[0064] Figure 4 This is a flowchart illustrating a data processing method provided in one embodiment of this specification;

[0065] Figure 5 This is a flowchart of another point cloud data processing method provided in one embodiment of this specification;

[0066] Figure 6 This is a flowchart illustrating a machine learning model training method provided in one embodiment of this specification;

[0067] Figure 7 This is a flowchart illustrating a data processing model training method provided in one embodiment of this specification;

[0068] Figure 8 This is a flowchart illustrating a traffic data processing method provided in one embodiment of this specification;

[0069] Figure 9 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation

[0070] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0071] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0072] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0073] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0074] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of model parameters. A large model can also be called a foundation model. It is pre-trained using large-scale unlabeled corpora to produce a pre-trained model with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.

[0075] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and Computer Vision. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as NLP tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.

[0076] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0077] 3D: three-dimensional.

[0078] Transformer: A network model used for sequence modeling.

[0079] MLP: Multi-Layer Perceptron.

[0080] RPN: The full English name is Region Proposal Network. It is used for two-stage object detectors and is responsible for outputting coarse-grained object detection boxes. RPN is a fully convolutional network that can simultaneously predict the location, class, and score of the object's detection box. For a point cloud within a preset range of input, it can output a series of rectangular candidate regions and determine the corresponding object class and score for each rectangular candidate region. This score can be used to evaluate the probability of each proposal being the target object class.

[0081] Proposal: The output of the RPN network, usually referring to a coarse bounding box.

[0082] Foreground features refer to the points in point cloud data that represent target objects, objects of interest, or important components of a scene, along with their related geometric features and physical attributes. For example, foreground features can include objects such as cars, pedestrians, and non-motorized vehicles, along with their corresponding shape and contour information, and surface reflectivity information.

[0083] Background features refer to points in point cloud data other than foreground features, as well as their related geometric features and physical attributes. For example, background features could be vegetation or greenery along a road.

[0084] Bird's Eye View (BEV) feature map: This refers to a bird's-eye view feature, which is a top-down view of an object or scene viewed from above. It is formed by projecting 3D point cloud data onto a 2D plane. BEV features are usually represented as a pixel matrix, where each pixel value represents the point cloud density or other feature information within that region.

[0085] Semantic features: These can be semantic features on the BEV feature map, referring to features associated with specific object categories, scene elements, or behaviors. For example, semantic features may include: object category labels, road layout and infrastructure, dynamic states, space occupancy, traffic rules and behaviors, and / or environmental attributes. Object category labels refer to the different semantic labels that different types of objects (such as vehicles, pedestrians, bicycles, traffic signs, streetlights, etc.) may be assigned in the BEV feature map. Road layout and infrastructure refer to the location and shape of road elements such as road boundaries, lane lines, pedestrian crossings, curbs, and medians. Dynamic states refer to the dynamic attributes of moving objects, such as speed, acceleration, and direction of travel. Space occupancy refers to the volume of space occupied by each object in the environment. Traffic rules and behaviors refer to the more complex semantic information that the BEV feature map may contain in high-level autonomous driving scenarios, such as the status of traffic lights, the indication of traffic signs, and the behavioral intentions of other road users (such as turn signal signals). Environmental attributes can refer to other semantic features of the scene, such as road surface type (e.g., asphalt, gravel, grass), weather conditions (e.g., rain, snow, fog), lighting conditions, etc. These attributes may affect the vehicle's driving performance or the effectiveness of sensors.

[0086] Geometric features: Raw point cloud data has not been processed by neural networks and contains coordinate information generated by reflections from the object's surface. It can directly reflect the object's shape, outline, position, and other information. Therefore, the features of the raw point cloud are geometric features.

[0087] With the continuous development of artificial intelligence technology, when it is necessary to process point cloud data, neural network models can be used to process the point cloud data and obtain the corresponding processing results.

[0088] For example, LiDAR can actively scan its surrounding environment to obtain point cloud data. For traffic perception systems using LiDAR solutions, a fast and high-performance point cloud target detector (neural network model) is essential. The performance of the target detector directly affects the effectiveness of subsequent target tracking, path planning, and other higher-level systems. Point cloud data acquired by LiDAR typically has the following two characteristics: 1. Sparsity: Point clouds only exist on the surface of objects; there are no point clouds inside objects or in large open areas. Furthermore, the more distant the object, the sparser the point cloud data reflected. 2. Large-area 3D perception: Typically, the coverage range of LiDAR reaches hundreds of meters in length and width, and about 3-5 meters in height. This is much larger than the perception range of pixel-space-based images. Therefore, 3D target detection based on point clouds usually requires compressing the feature space to a smaller dimension, and the blurring of foreground and background features due to environmental changes is also a challenge in 3D target detection.

[0089] Based on the characteristics of point cloud data, a point cloud 3D target detection scheme can first voxelize the point cloud, then compress the voxelized features into a BEV plane feature, and use a 2D neural network model (such as CNN) to extract features on the plane feature, finally obtaining the detection result on the BEV plane; however, this scheme has the drawback of low detection accuracy, and the process of voxelizing the point cloud is a process of information loss, which inevitably leads to a loss of accuracy.

[0090] Based on this, this specification provides a point cloud data processing method. This specification also relates to a data processing method, another point cloud data processing method, a machine learning model training method, a data processing model training method, a point cloud data processing device, a data processing device, another point cloud data processing device, a machine learning model training device, a data processing model training device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.

[0091] See Figure 1 , Figure 1 This diagram illustrates an application illustration of a point cloud data processing method according to an embodiment of this specification, based on... Figure 1It can be seen that the user can send point cloud data to the server 104 through the terminal 102. The server 104 can process the point cloud data using an object processing model. Specifically, firstly, the point cloud data is input into the RPN model for object recognition to obtain an initial object recognition result. Then, to avoid inaccurate initial object recognition results, the point cloud data and the initial object recognition result are input into a transformer model. The transformer model determines a correction value for the initial object recognition result based on the point cloud data, and corrects the initial object recognition result based on the correction value to obtain an accurate target object detection result. Then, the server 104 can send the target object detection result to the terminal 102 for display to the user.

[0092] See Figure 2 , Figure 2 A flowchart of a point cloud data processing method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0093] Step 202: Determine the point cloud data to be processed and input the point cloud data to be processed into the data processing model, wherein the data processing model includes a data processing module and a result adjustment module.

[0094] The point cloud data to be processed can be understood as point cloud data that needs to be processed using a data processing model. This point cloud data can be the point cloud data of a target scene. By performing point cloud data acquisition operations on the target scene, the corresponding point cloud data to be processed can be obtained. In one or more embodiments provided in this specification, the point cloud data to be processed differs depending on the scenario in which the point cloud data processing method is applied. For example, when the point cloud data processing method is applied to a traffic scenario, the target scene can be a traffic scenario (such as an intersection area), and the point cloud data to be processed can be the point cloud data of that traffic scenario. Correspondingly, the target object recognition result of the point cloud data to be processed can be the recognition result of target objects in the traffic scenario (such as cars violating traffic rules, pedestrians violating traffic rules, etc.). The data processing model can be understood as a model capable of processing the point cloud data of a traffic scenario to identify the target object recognition result in the traffic scenario. In one or more embodiments provided in this specification, the point cloud data to be processed is different when the point cloud data processing method is applied to different scenarios. For example, when the point cloud data processing method is applied to an animal protection scenario, the target scenario can be an animal protection area (such as a wetland, forest, or other area), and the point cloud data to be processed can be the point cloud data of that animal protection area. Correspondingly, the target object recognition result of the point cloud data to be processed can be the recognition result of the target object (such as an injured animal, a pedestrian who has illegally entered the area, etc.) in the animal protection area. The data processing model can be understood as a model that can process the point cloud data of the animal protection area to identify the target object recognition result in the animal protection area.

[0095] A data processing model can be understood as a model capable of processing point cloud data. This data processing model can be a model composed of a data processing module and a result adjustment module. For example, the data processing model can be a target object detection model composed of an RPN model and a Transformer model.

[0096] The data processing module can be understood as the module in the data processing model that processes the point cloud data to be processed. Through this module, object recognition processing is performed on the point cloud data to obtain the initial object recognition results. This data processing module can be a sub-model in the data processing model, or one or more network layers within the model. For example, this data processing module can be an RPN model.

[0097] The result adjustment module can be understood as a module in the data processing model that adjusts the initial object recognition results. This module determines the result adjustment parameters based on the point cloud data to be processed and the initial object recognition results, and then adjusts the initial object recognition results according to these parameters to obtain the target object recognition results for the point cloud data to be processed. This result adjustment module can be a sub-model in the data processing model, or one or more network layers within the model. For example, this data processing module can be a Transformer model.In one or more embodiments provided in this specification, the result adjustment module includes a feature processing unit, a feature fusion unit, an encoding unit, and a decoding unit. The feature processing unit can perform feature processing, determining a first data feature set and a second data feature set based on the point cloud data to be processed and the initial object recognition result. The feature fusion unit can perform feature fusion processing, fusing first data features in the first data feature set to obtain a first fused feature corresponding to the first data feature set, and fusing second data features in the second data feature set to obtain a second fused feature corresponding to the second data feature set. The encoding unit can perform feature encoding processing, using an attention mechanism to encode the first fused feature and the second fused feature to obtain a fused feature code. The decoding unit can perform decoding processing, decoding the fused feature code to obtain the result adjustment parameters. In other words, the result adjustment module can use the feature processing unit, feature fusion unit, encoding unit, and decoding unit to process the point cloud data to be processed and the initial object recognition result to obtain the result adjustment parameters. The unit includes a multi-head attention subunit (e.g., a multi-head attention module), a residual connection subunit (e.g., an Add&Norm layer), a feedforward neural network subunit (e.g., a feedforward neural network), and an encoding subunit (e.g., an encoder). The multi-head attention subunit is used to determine the attention matrix. It utilizes an attention mechanism to determine a first attention matrix for the first fused feature based on the similarity matrix, and also utilizes an attention mechanism to determine a second attention matrix for the second fused feature based on the similarity matrix. The residual connection subunit performs residual connection processing on the attention matrix. The feedforward neural network subunit updates the attention matrix. It utilizes a first matrix update strategy to update the first attention matrix based on the first and second update features to obtain a first updated attention matrix, and utilizes a second matrix update strategy to update the second attention matrix based on the first and second update features to obtain a second updated attention matrix. The encoding subunit performs encoding processing on the first and second updated attention matrices to obtain the fused feature encoding. In other words, the encoding unit can use the attention mechanism to encode the first fused feature and the second fused feature, and obtain the fused feature encoding, based on the multi-head attention subunit, residual connection subunit, feedforward neural network subunit and encoding subunit.

[0098] In one or more embodiments provided in this specification, determining the point cloud data to be processed includes:

[0099] The device receives point cloud data to be processed sent by a point cloud data acquisition device, wherein the point cloud data acquisition device is a device configured by the user through a client for performing point cloud data acquisition operations for a target scene, and the point cloud data to be processed is obtained by the point cloud data acquisition device performing point cloud data acquisition operations for the target scene.

[0100] The point cloud data acquisition device can be understood as a device used to acquire the point cloud data to be processed. For example, the point cloud data acquisition device can be a lidar, structured light scanner, etc.

[0101] Taking the application of the point cloud data processing method in a traffic scenario in one or more embodiments provided in this specification as an example, the point cloud data processing method will be described. The point cloud data acquisition device is a LiDAR, the target scenario is a road area, and the point cloud data to be processed is the raw point cloud data collected from the road area. Therefore, in the process of detecting vehicles violating traffic rules in a road area, it is necessary to collect the raw point cloud data corresponding to that road area using the LiDAR corresponding to the road area. This LiDAR can be configured by the user through a client. The user configures the LiDAR through the client so that it can collect the raw point cloud data corresponding to the road area and send the raw point cloud data to a server for point cloud data processing. This server is the server used by the point cloud data processing method in one or more embodiments of this specification.

[0102] In the above embodiments, by receiving point cloud data to be processed sent by the point cloud data acquisition device, the point cloud data to be processed can be processed in a timely manner, thereby solving the problem of needing to process point cloud data in real-world scenarios.

[0103] Step 204: Using the data processing module, perform recognition processing on the point cloud data to be processed to obtain the initial object recognition result of the point cloud data to be processed.

[0104] The initial object recognition result can be understood as the initial object recognition result obtained after object recognition processing of the target objects contained in the point cloud data to be processed. The initial object recognition result can determine the target objects in the target scene. For example, the initial object recognition result can be the detection box, coordinate information, and position information of the target object. In one or more embodiments provided in this specification, the initial object recognition result can be the 3D Proposals output by the RPN model, which are coarse and inaccurate. The target object will differ depending on the application of the point cloud data processing method in the one or more embodiments provided in this specification to different scenarios. For example, when the point cloud data processing method is applied to a traffic scenario, the target scene can be a road area (such as an intersection area), and the target object can be a car violating traffic rules, a pedestrian violating traffic rules, etc. in the road area. When the point cloud data processing method is applied to an animal protection scenario, the target scene can be an animal protection area (such as a wetland, forest, etc.), and the target object can be an injured animal in the animal protection area, a pedestrian illegally entering the area, etc.

[0105] It should be noted that, since the initial object recognition result may be inaccurate, in order to ensure the accuracy of the object recognition result of the data processing model, the initial object recognition result needs to be corrected by the result adjustment module to obtain an accurate target object recognition result; that is, the initial object recognition result may be an incorrect, inaccurate, or incomplete object recognition result; correspondingly, the target object recognition result may be a correct, accurate, or complete object recognition result. The object recognition result can be the result used to identify target objects in the target scene, for example, the object recognition result may be a detection bounding box, coordinate information, position information, etc., for the target object; in one or more embodiments provided in this specification, the object recognition result may be 3D Proposals.

[0106] Specifically, during the process of inputting the point cloud data to be processed into the data processing model, the point cloud data to be processed can be input into the data processing module in the data processing model. The data processing module is used to perform object recognition processing on the point cloud data to be processed, thereby obtaining the initial object recognition result of the point cloud data to be processed.

[0107] Following the previous example, after inputting the original point cloud data into the object processing model, the RPN model in the object processing model can be used to process the original point cloud data to obtain coarse 3D proposals. It should be noted that the data processing module (such as the RPN model) in the object processing model can be selected according to the actual application scenario. For example, a more accurate RPN model can be selected, or a faster RPN model can be selected.

[0108] Step 206: Using the result adjustment module, determine the result adjustment parameters based on the point cloud data to be processed and the initial object recognition result, and perform result adjustment processing on the initial object recognition result according to the result adjustment parameters to obtain the target object recognition result of the point cloud data to be processed.

[0109] It should be noted that in one or more embodiments provided in this specification, considering the potential inaccuracy of the initial object recognition result, the initial object recognition result can be adjusted to ensure the accuracy of the object recognition result of the data processing model, thereby obtaining an accurate target object recognition result. For example, in traffic scenarios, since traffic perception has high requirements for the accuracy of detection results, the 3D target object detection scheme for point clouds provided in this specification can design the 3D target detector as a two-stage target detector. However, the problem with this 3D target object detection scheme for point clouds is that the two-stage point cloud 3D target detector design has a complex coupling structure. This 3D target object detection scheme for point clouds will choose to build a two-stage neural network model on a fixed single-stage target detector (i.e., the first-stage neural network model, such as the RPN model) to form a two-stage target detector. In the target detection process, the second-stage neural network model in the two-stage target detector will use the intermediate layer features of the first-stage neural network model (such as the first-stage neural network model). (Feature map determined by the network model); Because the second-stage neural network model needs to share the intermediate layer features of the first-stage neural network model, there is a strong coupling relationship between the first-stage and second-stage neural network models. This results in the two-stage neural network models being mutually bound and unable to be easily separated. In other words, this two-stage network structure cannot arbitrarily replace the first-stage RPN model. Therefore, in practical applications, it is impossible to flexibly adjust the first-stage neural network model according to actual needs, resulting in poor model applicability. For example, it is impossible to flexibly choose a high-accuracy RPN model or a high-efficiency RPN model as the first-stage neural network model. Furthermore, this two-stage network has a complex structural module design, resulting in slow operation and significant time consumption.

[0110] The aforementioned problem of the interdependent neural network models in the two stages being difficult to separate, coupled with the fact that different first-stage neural network models may be selected based on the specific needs of real-world applications, means that the 3D target object detection scheme for point clouds described above cannot meet the requirements of practical applications. For example, in practical applications, a faster RPN model needs to be deployed to improve the efficiency of object recognition, or a more accurate RPN model needs to be deployed to improve the accuracy of object recognition. However, the aforementioned 3D target object detection scheme for point clouds does not allow for the flexible selection of different RPN models for deployment.

[0111] To address the shortcomings of the aforementioned 3D target object detection schemes for point clouds, the point cloud data processing method provided in one or more embodiments of this specification offers a data processing model comprising a data processing module and a result adjustment module. In determining the result adjustment parameters, the result adjustment module can do so based on the point cloud data to be processed and the initial object recognition result, without using the feature data (e.g., feature maps) of the data processing module. This avoids the problem of strong coupling between the data processing module and the result adjustment module, which is difficult to separate due to the need for the result adjustment module to use intermediate layer features of the data processing module. This allows for flexible deployment of different data processing modules according to the needs of actual applications, meeting diverse requirements and improving the applicability of the point cloud data processing method provided in one or more embodiments of this specification. For example, the point cloud data processing method provided in one or more embodiments of this specification designs a network structure (i.e., a data processing model) that can utilize any RPN. This network structure proposes a Transformer network model with different channels after the RPN model to correct the result of the RPN model (i.e., the initial object recognition result). The Transformer network (i.e., the result adjustment module) with different channels comprises three parts: a feature processing unit (embedding layer), an encoding unit, and a decoding unit. The feature processing unit can be the input embedding, which combines the original point cloud and the last layer of BEV features of the RPN model as input features, effectively preserving the geometric features of the original point cloud data and the semantic features obtained after processing by the RPN model. The encoding unit is the encoder, which proposes a new bidirectional attention mechanism to enhance the data of the embedding input features. The decoding unit is the decoder, which compresses the features encoded by the encoder into a global feature representation. Finally, this global feature representation is used to obtain the corrected result of the RPN output. This design can greatly improve the effect of point cloud 3D object detection.

[0112] The result adjustment parameter can be understood as a parameter that adjusts the initial object recognition result. This result adjustment parameter can be a correction value; for example, a correction value for correcting a 3D proposal.

[0113] The target object recognition result can be understood as the object recognition result obtained after adjusting the initial object recognition result by adjusting the result adjustment parameters; the target object in the target scene can be determined by the target object recognition result; for example, the target object recognition result can be the detection box, bounding box, coordinate information, position information, etc. of the target object; in one or more embodiments provided in this specification, the target object recognition result can be 3D Proposals, which are accurate 3D Proposals.

[0114] In one or more embodiments provided in this specification, the step of using the result adjustment module to determine the result adjustment parameters based on the point cloud data to be processed and the initial object recognition result includes steps one to three:

[0115] Step 1: Input the point cloud data to be processed and the initial object recognition result into the result adjustment module in the data processing model. Using the result adjustment module, determine the first data feature set and the second data feature set based on the point cloud data to be processed and the initial object recognition result.

[0116] The first data feature set can be understood as a set containing first data features, which can be point cloud data in the point cloud data to be processed that corresponds to the initial object recognition result, and semantic features corresponding to the initial object recognition result.

[0117] The second data feature set can be understood as a set containing second data features, which can be the target point cloud data in the point cloud data corresponding to the initial object recognition result, and the target point cloud features corresponding to the target point cloud data in the semantic features corresponding to the initial object recognition result.

[0118] Specifically, after obtaining the initial object recognition result, the point cloud data to be processed and the initial object recognition result are input into the result adjustment module in the data processing model. The result adjustment module is used to extract features based on the point cloud data to be processed and the initial object recognition result, thereby determining the first data feature set and the second data feature set.

[0119] In one or more embodiments provided in this specification, the step of using the result adjustment module to determine a first data feature set and a second data feature set based on the point cloud data to be processed and the initial object recognition result includes:

[0120] Using the result adjustment module, the point cloud data to be processed is dimensionally transformed to determine the point cloud feature map corresponding to the point cloud data to be processed.

[0121] Based on the object location information of the initial object recognition result, object point cloud data is determined from the point cloud data to be processed, and object point cloud features are determined from the point cloud feature map based on the object location information.

[0122] Determine the target point cloud location information for the initial object recognition result, and based on the target point cloud location information, determine the target point cloud data from the object point cloud data, and based on the target point cloud location information, determine the target point cloud features from the object point cloud features;

[0123] The object point cloud data and the object point cloud features are determined as first data features, and a first data feature set is determined based on the first data features. The target point cloud data and the target point cloud features are determined as second data features, and a second data feature set is determined based on the second data features.

[0124] In this context, a point cloud feature map can be understood as a two-dimensional feature map corresponding to the point cloud data to be processed. This point cloud feature map contains the semantic features corresponding to the point cloud data to be processed. For example, the point cloud feature map can be a BEV feature map.

[0125] Object location information can be understood as the coordinate location information corresponding to the initial object recognition result. For example, if the initial object recognition result is a 3D Proposal, the object location information can be the coordinate information of the 3D Proposal or the spatial location of the 3D Proposal.

[0126] Object point cloud data can be understood as point cloud data in the point cloud data to be processed that corresponds to the object position information of the initial object recognition result. For example, when the object position information is the coordinate information of a 3D proposal, the object point cloud data can be the point cloud data corresponding to the spatial position or coordinate position of the 3D proposal. In one or more embodiments provided in this specification, the object point cloud data can be the original point cloud geometric features.

[0127] Object point cloud features can be understood as semantic features in the point cloud feature map corresponding to the object location information of the initial object recognition result. For example, using the spatial location of coarse 3D proposals, semantic features corresponding to the coarse 3D proposals can be extracted from the BEV feature map.

[0128] The target point cloud location information can be understood as pre-set location information based on the initial object recognition result. This target point cloud location information is used to obtain target point cloud data from the object point cloud data corresponding to the initial object recognition result. For example, the target point cloud location information can be the center point position, corner point position, etc. of the initial object recognition result. In the case of the initial object recognition result being a 3D proposal, the target point cloud location information can be the corner point and center point of the 3D proposal.

[0129] Target point cloud data can be understood as point cloud data in object point cloud data that corresponds to the location information of the target point cloud. For example, based on the corner points and center points of the 3D Proposal, the 9 key points corresponding to the 3D Proposal are obtained from the original point features.

[0130] The target point cloud features can be understood as the semantic features in the object point cloud features that correspond to the location information of the target point cloud.

[0131] Specifically, the point cloud data processing method provided in this specification, in the process of determining the first data feature set and the second data feature set using the result adjustment module, can be executed by the feature processing unit in the result adjustment module. The specific method is as follows: the feature processing unit in the result adjustment module performs dimensional transformation on the point cloud data to be processed to determine the point cloud feature map corresponding to the point cloud data to be processed; determines the object location information of the initial object recognition result, and obtains the corresponding object point cloud data from the point cloud data to be processed based on the object location information; and then extracts the corresponding object point cloud features from the point cloud feature map based on the object location information.

[0132] After determining the object point cloud data and object point cloud features, the feature processing unit can determine the preset target point cloud location information based on the initial object recognition result, and obtain the corresponding target point cloud data from the object point cloud data based on the target point cloud location information, and extract the corresponding target point cloud features from the object point cloud features based on the target point cloud location information.

[0133] Finally, the feature processing unit determines the object point cloud data and object point cloud features as the first data features, and constructs a first data feature set based on the first data features; and determines the target point cloud data and target point cloud features as the second data features, and constructs a second data feature set based on the second data features.

[0134] Following the previous example, the point cloud data processing method provided in this manual fuses the original point cloud geometric features with BEV semantic features through embedding. Specifically, this step is executed as follows:

[0135] 1. Using the spatial location of coarse 3D proposals, extract (fetch) the original point cloud geometric features corresponding to the spatial location from the original point cloud data.

[0136] The original point cloud geometric features refer to the positional information of the original point cloud corresponding to the spatial position of the 3D Proposals in the original point cloud data; the positional information of the original point cloud can be a three-dimensional array in the three-dimensional world coordinate system.

[0137] 2. Perform dimensionality transformation on the original point cloud data to obtain the BEV feature map, and use the spatial location of the coarse 3D Proposals to extract (interpolate) the semantic features corresponding to the coarse 3D Proposals from the BEV feature map. For example, N=255 points can be selected as the original point features (referring to the original point cloud corresponding to the spatial location of the 3D Proposals). It retains the geometric location (referring to geometric features) and the semantic information of the corresponding location (referring to semantic features) of the original points. Then, based on the corner points and center points of the 3D Proposals, the features of 9 key points (referring to geometric features and semantic features) are obtained.

[0138] In the above embodiments, by utilizing the result adjustment module to perform dimensional transformation on the point cloud data to be processed and determining the point cloud feature map corresponding to the point cloud data to be processed, the result adjustment module can determine the result adjustment parameters based on the point cloud feature map it has determined. This avoids the problem that the result adjustment module needs to use the intermediate layer features of the data processing module, which leads to a strong coupling relationship between the data processing module and the result adjustment module, making it difficult to separate them. Furthermore, by using the location information of the initial object recognition result and the target point cloud location information for the initial object recognition result, the first data feature set and the second data feature set are determined, which facilitates the accurate determination of the result adjustment parameters for the initial object recognition result based on the similarity comparison between the two sets.

[0139] Based on the above embodiments, it can be seen that the point cloud data processing method in one or more embodiments provided in this specification can use a result adjustment module to determine result adjustment parameters according to the point cloud data to be processed, the initial object recognition result, and the point cloud feature map. The point cloud feature map can be determined by performing dimensional transformation on the point cloud data to be processed using the result adjustment module.

[0140] Step 2: Fuse the first data features in the first data feature set to obtain the first fused feature corresponding to the first data feature set, and fuse the second data features in the second data feature set to obtain the second fused feature corresponding to the second data feature set.

[0141] In one or more embodiments provided in this specification, fusing first data features in the first data feature set to obtain a first fused feature corresponding to the first data feature set, and fusing second data features in the second data feature set to obtain a second fused feature corresponding to the second data feature set, includes:

[0142] Determine a first data feature in the first data feature set, and determine a second data feature in the second data feature set;

[0143] The feature fusion unit in the result adjustment module is used to fuse the first data features to obtain the first fused feature corresponding to the first data feature set, and the feature fusion unit is used to fuse the second data features to obtain the second fused feature corresponding to the second data feature set.

[0144] The feature fusion unit can be understood as a unit in the result adjustment module used for feature fusion. By using the feature fusion unit, the first data features can be fused to obtain the first fused feature; or, the second data features can be fused to obtain the second fused feature. In one or more embodiments provided in this specification, the feature fusion unit can be a sub-model in the result adjustment module, or the result adjustment module can be one or more network layers in the result adjustment module; for example, the feature fusion unit can be an MLP network model.

[0145] The first fusion feature can be understood as the fusion feature obtained after performing feature fusion on the first data feature; the second fusion feature can be understood as the fusion feature obtained after performing feature fusion on the second data feature.

[0146] Following the previous example, after obtaining the original point cloud geometric features and semantic features corresponding to the 3D Proposals, and obtaining the features of 9 key points (referring to geometric and semantic features), the MLP network can be used to fuse the original point cloud geometric features and the semantic features corresponding to the 3D Proposals to obtain the corresponding fused features. The fused original point features and key point features are represented as F (i.e., the first fused feature) and Fk (i.e., the second fused feature), respectively.

[0147] In the above embodiments, by utilizing the result adjustment module, the first data features in the first data feature set are quickly fused to obtain the first fused feature, and the second data features in the second data feature set are fused to obtain the second fused feature. This facilitates the accurate determination of the result adjustment parameters for the initial object recognition result based on the similarity comparison between the two fused features.

[0148] In one or more embodiments provided in this specification, the step of fusing the first data features using the feature fusion unit in the result adjustment module to obtain the first fused feature corresponding to the first data feature set, and fusing the second data features using the feature fusion unit to obtain the second fused feature corresponding to the second data feature set, includes:

[0149] Using the feature fusion unit, the center point coordinate data of the initial object recognition result is determined;

[0150] Subtract the center point coordinate data from the object point cloud data in the first data feature to obtain updated object point cloud data. Project the updated object point cloud data and the object point cloud features in the first data feature set onto the first target dimension for feature fusion to obtain the first fused feature corresponding to the first data feature set.

[0151] Subtract the center point coordinate data from the target point cloud data in the second data feature set to obtain updated target point cloud data. Project the updated target point cloud data and the target point cloud features in the second data feature set onto the first target dimension for feature fusion to obtain the second fused feature corresponding to the second data feature set.

[0152] The center point coordinate data of the initial object recognition result can be understood as the coordinate data corresponding to the center point of the initial object recognition result. For example, if the initial object recognition result is 3D Proposals, the center point coordinate data can be the center point position of the 3D Proposals.

[0153] Updating object point cloud data can be understood as obtaining point cloud data after normalizing the object point cloud data in the first data feature. Specifically, the normalization process can be performed by subtracting the center point coordinate data from the object point cloud data in the first data feature.

[0154] Updating the target point cloud data can be understood as obtaining the point cloud data after normalizing the target point cloud data in the second data feature. Specifically, the normalization process can be performed by subtracting the center point coordinate data from the target point cloud data in the second data feature.

[0155] The first target dimension can be understood as the dimension required during feature fusion. This first target dimension can be set according to the actual application scenario; for example, it can be 256-dimensional, 128-dimensional, etc. During feature fusion, the first or second data feature needs to be projected onto a preset dimension (such as the first target dimension) for fusion. For example, when the feature fusion unit is an MLP network, the MLP network can map the features of the point cloud (i.e., map the features of the point cloud to D dimensions) to D dimensions (i.e., the first target dimension) for fusion, thereby obtaining the corresponding fused features.

[0156] Following the previous example, in the process of fusing the geometric features of the original point cloud and the semantic features corresponding to the 3D proposals using an MLP network, the method of feature fusion using an MLP network is as follows:

[0157] First, determine the original point cloud features (referring to the original point cloud corresponding to the spatial location of 3D Proposals), as well as the geometric location of the original points (i.e., object point cloud data) and the semantic information of the corresponding locations (referring to object point cloud features).

[0158] Secondly, based on the corner and center points of the 3D Proposal, the features of 9 key points (referring to the target point cloud data and target point cloud features) are obtained.

[0159] Finally, the geometric and semantic features of each point cloud are concatenated, and the concatenated features are fused to obtain the fused features; the method for obtaining the fused features can be expressed as the following formula (1):

[0160]

[0161] Where S(·) represents the normalization operation between points. Specifically, it subtracts the coordinates of the center point of the corresponding 3D Proposal from the position coordinates of each point in the point cloud P to obtain the normalized coordinates.

[0162] Where T represents the reflectance features of the point cloud and the high-dimensional deep features (i.e., semantic features) extracted from the BEV feature map. It is an MLP network that maps the features of points to D dimensions (e.g., 256 dimensions).

[0163] After obtaining the fused features based on the above formula, the fused original point features and keypoint features can be represented as F and F, respectively. k .

[0164] In the above embodiments, by utilizing the result adjustment module, the first data features in the first data feature set are quickly fused to obtain the first fused feature, and the second data features in the second data feature set are fused to obtain the second fused feature. This facilitates the accurate determination of the result adjustment parameters for the initial object recognition result based on the similarity comparison between the two fused features.

[0165] Furthermore, the method of fusing the first data features using the feature fusion unit in the result adjustment module to obtain the first fused feature corresponding to the first data feature set, and fusing the second data features using the feature fusion unit to obtain the second fused feature corresponding to the second data feature set, can also be used to determine the center point coordinate data of the initial object recognition result; subtract the center point coordinate data from the result data in the first data features to obtain updated result data, and subtract the center point coordinate data from the target result data in the second data features to obtain updated target result data; using the feature fusion unit, projecting the updated result data and the result features in the first data features onto a first target dimension for feature fusion to obtain the first fused feature corresponding to the first data feature set; using the feature fusion unit, projecting the updated target result data and the target result features in the second data features onto a first target dimension for feature fusion to obtain the second fused feature corresponding to the second data feature set; no specific limitations are imposed here.

[0166] Step 3: Encode the first fusion feature and the second fusion feature using an attention mechanism to obtain the fusion feature encoding, and decode the fusion feature encoding to obtain the result adjustment parameters.

[0167] Specifically, in the point cloud data processing method provided in this specification, the first fusion feature and the second fusion feature can be encoded using an attention mechanism through the encoding unit in the result adjustment module to obtain the fusion feature encoding; the encoding unit in the result adjustment module is used to encode the features, for example, the encoding unit can be an encoder.

[0168] In one or more embodiments provided in this specification, the step of encoding the first fused feature and the second fused feature using an attention mechanism to obtain fused feature encoding includes:

[0169] Project the first fusion feature and the second fusion feature onto the second target dimension to obtain the first target dimension feature corresponding to the first fusion feature and the second target dimension feature corresponding to the second fusion feature;

[0170] Determine the similarity matrix between the first feature of the target dimension and the second feature of the target dimension;

[0171] Using an attention mechanism, a first attention matrix for the first fused feature is determined based on the similarity matrix; and using an attention mechanism, a second attention matrix for the second fused feature is determined based on the similarity matrix.

[0172] The fused feature encoding is obtained based on the first attention matrix and the second attention matrix.

[0173] The second target dimension can be set according to the actual application scenario. This manual does not impose specific restrictions on it. For example, the second target dimension can be 256 dimensions, 128 dimensions, etc.

[0174] The first feature of the target dimension can be understood as the feature information represented by the first fused feature when projected onto the second target dimension. Correspondingly, the second feature of the target dimension can be understood as the feature information represented by the second fused feature when projected onto the second target dimension.

[0175] Specifically, the point cloud data processing method provided in this specification firstly projects the first fused feature and the second fused feature onto a pre-set second target dimension, thereby obtaining a first target dimension feature corresponding to the first fused feature in the second target dimension and a second target dimension feature corresponding to the second fused feature in the second target dimension; secondly, it calculates the similarity between the first target dimension feature and the second target dimension feature to obtain a similarity matrix between them; finally, it uses an attention mechanism to calculate a first attention matrix for the first fused feature based on the similarity matrix, and uses the same attention mechanism to calculate a second attention matrix for the second fused feature based on the similarity matrix. Furthermore, it obtains the fused feature encoding based on the first attention matrix and the second attention matrix. This attention mechanism allows for the acquisition of a high-value attention matrix using contextual information, improving the accuracy of subsequent parameter adjustments.

[0176] Following the previous example, the point cloud data processing method provided in this specification can encode fused features using a bidirectional attention encoder from points (i.e., object point cloud data) to keypoints (i.e., target point cloud data). Specifically, the point cloud data processing method proposes a bidirectional attention encoder from points to keypoints. The core of this encoder is a bidirectional attention mechanism used to encode fused features. The specific steps for performing the encoding process are as follows:

[0177] 1. For the inputs F and F k F and F kLinearly projected onto a preset dimension (i.e., the second target dimension), we obtain the projected Q (i.e., the first feature of the target dimension) and Q. k (i.e., the second feature of the target dimension), and then calculate Q and Q. k The similarity matrix between them (i.e., the similarity matrix) can be calculated using the following formula (2):

[0178] R = Q·(Q k ) T Formula (2)

[0179] In the formula, T represents the reflectance feature of the point cloud and the high-dimensional deep features (i.e., semantic features) extracted from the BEV feature map.

[0180] 2. After determining the similarity matrix, the original point F and keypoint F' are calculated based on the similarity matrix using the multi-head attention module. k Each attention matrix; the specific method for determining the attention matrix can be found in the following formula (3).

[0181]

[0182] Where σ(·) is the softmax function calculated along the N-dimensional plane (i.e., D in the formula). It is the softmax function calculated along 9 dimensions (i.e., D in the formula).

[0183] Using formulas 2 and 3 above, we obtain the original point F and the key point F. k Each has its own attention matrix.

[0184] In one or more embodiments provided in this specification, obtaining the fused feature encoding based on the first attention matrix and the second attention matrix includes:

[0185] A first updated feature corresponding to the first fusion feature is determined, and a second updated feature corresponding to the second fusion feature is determined, wherein the first updated feature is obtained by projecting the first fusion feature onto a third target dimension, and the second updated feature is obtained by projecting the second fusion feature onto the third target dimension;

[0186] The first attention matrix is ​​updated using the first matrix update strategy, based on the first update feature and the second update feature, to obtain the first updated attention matrix;

[0187] Using the second matrix update strategy, the second attention matrix is ​​updated based on the first update feature and the second update feature to obtain the second updated attention matrix;

[0188] The first updated attention matrix and the second updated attention matrix are encoded to obtain the fused feature encoding.

[0189] The first updated feature can be understood as the feature obtained by projecting the first fused feature onto the third target dimension. The first updated feature can be the feature information represented by the first fused feature projected onto the third target dimension. The second updated feature can be understood as the feature obtained by projecting the second fused feature onto the third target dimension. The second updated feature can be the feature information represented by the second fused feature projected onto the third target dimension. The third target dimension can be set according to the actual application scenario and is not specifically limited in this specification. For example, the third target dimension can be 256-dimensional, 128-dimensional, 64-dimensional, etc.

[0190] The first matrix update strategy can be understood as a strategy for updating the first attention matrix. For example, the first matrix update strategy can be: multiplying the first attention matrix and the second update feature, and adding the attention matrix obtained by the multiplication process to the first update feature to obtain the first updated attention matrix.

[0191] The second matrix update strategy can be understood as a strategy for updating the second attention matrix. For example, the second matrix update strategy can be: multiply the second attention matrix with the first update feature, and add the attention matrix obtained by multiplication to the second update feature to obtain the second updated attention matrix.

[0192] It should be noted that the first matrix update strategy calculates the attention matrix differently from the second matrix update strategy.

[0193] Specifically, the point cloud data processing method provided in this specification first projects the first fused feature onto the third target dimension to obtain the first updated feature, and projects the second fused feature onto the third target dimension to obtain the second updated feature; secondly, using the operation method provided by the first matrix update strategy, the first attention matrix is ​​updated based on the first updated feature and the second updated feature to obtain the first updated attention matrix; and using the operation method provided by the second matrix update strategy, the second attention matrix is ​​updated based on the first updated feature and the second updated feature to obtain the second updated attention matrix; finally, the first updated attention matrix and the second updated attention matrix are encoded to obtain the fused feature encoding; thus, the first updated attention matrix and the second updated attention matrix are summarized into a global feature expression, and this global feature expression can be used to obtain accurate result correction parameters.

[0194] Continuing with the previous example, in the Multi-Head Attention module, complete the processing of F and F. k After calculating the attention matrix, F and F k The corresponding attention matrices are input into the Add&Norm layer, which combines the original inputs (F and F) k The attention matrix is ​​processed by residual connection, and the attention matrix after residual connection is output.

[0195] After the attention matrix is ​​processed by the Add&Norm layer, and then processed by the feedforward network and residual connections (Add&Norm layers), the network layer outputs (first updated attention matrix, second updated attention matrix) can be updated as follows:

[0196]

[0197]

[0198] in, This refers to a feedforward neural network; V k V is also composed of F k The result is obtained by linearly projecting F onto a preset dimension. This preset dimension can be set according to the actual application scenario; where F... update It can be the first update of the attention matrix, (F k ) update It can be the second updated attention matrix.

[0199] For F update and (F) k ) update The data is then fed into a 3-layer encoder. After passing through the 3-layer encoder, the updated features F of the N points are generated in this step. update and (F) k ) update The features are aggregated into a global feature representation, and then this feature is used to obtain the correction value of the 3D Proposal.

[0200] In one or more embodiments provided in this specification, the initial object recognition result is an initial object detection box; the result adjustment parameter is a correction value for the initial object detection box;

[0201] The step of adjusting the parameters based on the results to perform result adjustment processing on the initial object recognition result, and obtaining the target object recognition result of the point cloud data to be processed, includes:

[0202] The initial object detection box is corrected according to the correction value to obtain the target object detection box corresponding to the object to be identified in the point cloud data to be processed.

[0203] Following the previous example, after obtaining the correction value of the 3D Proposal, the coarse 3D Proposal encoding is corrected using this correction value to obtain the corrected 3D Proposal encoding. Then, the corrected 3D Proposal encoding is decoded through the decoding layer (Decoder & Feed forward network) of the transformer network to obtain the accurate 3D Proposal (i.e., the target object detection box). This avoids the problem that inaccurate point cloud data processing results cannot meet the needs of point cloud data processing in practical application scenarios.

[0204] In one or more embodiments provided in this specification, after adjusting the initial object recognition result according to the result adjustment parameters to obtain the target object recognition result of the point cloud data to be processed, the method further includes:

[0205] The client corresponding to the point cloud data acquisition device is determined, and the target object recognition result is sent to the client so that the client can display the target object recognition result to the user.

[0206] Following the previous example, after obtaining an accurate 3D proposal, the accurate 3D proposal is sent to the client, allowing the client to display the accurate 3D proposal to the user.

[0207] In one or more embodiments provided in this specification, the point cloud data processing method first inputs the point cloud data to be processed into a data processing model. Then, using the data processing module in the data processing model, it performs recognition processing on the point cloud data to obtain an initial object recognition result. Next, to avoid inaccurate initial object recognition results, it uses a result adjustment module in the data processing model to determine result adjustment parameters based on the point cloud data to be processed and the initial object recognition result. The initial object recognition result is then adjusted according to these parameters. This achieves accurate acquisition of the target object recognition result corresponding to the point cloud data to be processed, avoiding the problem that inaccurate point cloud data processing results would fail to meet the needs of point cloud data processing in practical application scenarios.

[0208] The following is in conjunction with the appendix Figure 3Taking the application of the point cloud data processing method provided in this specification in a point cloud 3D target detection scenario as an example, the point cloud data processing method will be further explained. Figure 3 This specification illustrates a flowchart of a point cloud data processing method according to an embodiment, based on... Figure 3 As can be seen, the point cloud data processing method provided in this manual includes two parts: a first stage and a second stage.

[0209] The execution process for the first phase is as follows:

[0210] Step 1: Determine the raw point cloud data.

[0211] The raw point cloud data can be obtained by collecting point cloud data for a target scene. The target scene can be set according to the actual application scenario. For example, the target scene can be a traffic road, office area, etc.

[0212] The raw point cloud data can be collected using a point cloud data acquisition device, which can be configured according to the actual application scenario. For example, the point cloud data acquisition device can be a LiDAR.

[0213] Step 2: Input the raw point cloud data into the RPN model to obtain coarse 3D proposals.

[0214] The 3D Proposals can be a 3D bounding box, and the point cloud data defined by the bounding box can be the target object to be detected. For example, in a traffic scenario, the target object can be a vehicle violating traffic rules; the 3D Proposals can define the point cloud data corresponding to the vehicle violating traffic rules, thereby detecting the vehicle violating traffic rules.

[0215] It should be noted that the 3D proposals output by this RPN model are coarse, with defects in accuracy and precision. Therefore, a second-stage model is needed to correct these 3D proposals in order to obtain more accurate and precise 3D proposals.

[0216] The execution process for the second phase is as follows:

[0217] Step 1: Fuse the geometric features of the original point cloud with the semantic features of BEV through embedding.

[0218] The specific execution method for this step is as follows:

[0219] 1. Using the spatial location of coarse 3D proposals, extract (fetch) the original point cloud geometric features corresponding to the spatial location from the original point cloud data.

[0220] The original point cloud geometric features refer to the positional information of the original point cloud corresponding to the spatial position of the 3D Proposals in the original point cloud data; the positional information of the original point cloud can be a three-dimensional array in the three-dimensional world coordinate system.

[0221] 2. Perform dimensional transformation on the original point cloud data to obtain the BEV feature map, and use the spatial location of the coarse 3D Proposals to extract (interpolate) the semantic features corresponding to the coarse 3D Proposals from the BEV feature map.

[0222] 3. Use an MLP network to fuse the geometric features of the original point cloud and the semantic features corresponding to the 3D proposals.

[0223] The following example illustrates how MLP networks can be used for feature fusion.

[0224] For example, firstly, this scheme can randomly select N=255 points as the original point features (referring to the original point cloud corresponding to the spatial location of the 3D Proposals), which preserves the geometric location of the original points (referring to geometric features) and the semantic information of the corresponding locations (referring to semantic features).

[0225] Secondly, based on the corner and center points of the 3D Proposal, the features of 9 key points (referring to geometric and semantic features) are obtained.

[0226] Then, the geometric and semantic features of each point cloud are concatenated, and the concatenated features are fused to obtain the fused features; the method for obtaining the fused features can be found in the above formula (1).

[0227] After obtaining the fused features based on the above formula (1), the fused original point features and keypoint features can be represented as F and F, respectively. k .

[0228] Step 2: Encode the fused features using a point-to-keypoint bidirectional attention encoder.

[0229] Specifically, this solution proposes a point-to-keypoint bidirectional attention encoder. The core of this encoder is a bidirectional attention mechanism used to encode fused features. The specific steps for performing the encoding process are as follows:

[0230] 1. For the inputs F and Fk (Input), F and F k Linearly project onto a preset dimension (set according to the actual scenario) to obtain the projected Q and Q'. k Then calculate Q and Q k The similarity matrix between them can be calculated using formula (2) above.

[0231] 2. After determining the similarity matrix, the original point F and keypoint F' are calculated based on the similarity matrix using the multi-head attention module. k Each attention matrix; for the specific method of determining the attention matrix, please refer to the above formula (3).

[0232] 3. In the Multi-Head Attention module, complete the processing of F and F. k After calculating the attention matrix, F and F k The corresponding attention matrices are input into the Add&Norm layer, which combines the original inputs (F and F) k The attention matrix is ​​processed by residual connection, and the attention matrix after residual connection is output.

[0233] 4. After the attention matrix is ​​processed by the Add&Norm layer, and then processed by the feedforward network and residual connections (Add&Norm layers), the network layer output can be updated as follows:

[0234]

[0235]

[0236] 5. For F update and (F) k ) update The data is then fed into a 3-layer encoder. After passing through the 3-layer encoder, the updated features F of the N points are generated in this step. update and (F) k ) update The features are aggregated into a global feature representation, and then this feature is used to obtain the correction value of the 3D Proposal.

[0237] Step 3: After obtaining the correction value of the 3D Proposal, use the correction value to correct the coarse 3D Proposal encoding to obtain the corrected 3D Proposal encoding (i.e., encoder encoding features).

[0238] Step 4: The corrected 3D proposal encoding is decoded through the decoding layer (Decoder & Feed forward network) of the transformer network to obtain the final prediction result (i.e., the accurate 3D proposal).

[0239] Based on the above, the point cloud data processing method provided in this specification is a Transformer-based 3D point cloud target detection method. This scheme designs a network structure that can utilize any RPN (Real-Time Point Network). Following the RPN, a Transformer network with different channels is proposed to correct the RPN results. The channel-different Transformer network consists of three parts:

[0240] The first part is the input embedding, which combines the original point cloud and the last layer of BEV features of RPN as input features, effectively preserving the geometric features of the original point cloud data and the semantic features obtained after RPN processing.

[0241] The second part is the encoder, which proposes a novel bidirectional attention mechanism to augment the embedding input features.

[0242] The third part is the decoder, which compresses the features encoded by the encoder into a global feature representation. Finally, this global feature representation is used to obtain the corrected RPN. This design can improve the performance of point cloud 3D object detection.

[0243] This design significantly reduces the computational complexity of the encoder compared to traditional encoders using self-attention mechanisms, thereby greatly improving the network's learning ability, accelerating convergence, and ultimately enhancing performance. Furthermore, the method maintains flexibility, allowing replacement with any RPN for various industrial applications, avoiding serious shortcomings in scalability and performance. It also implements sequence modeling capabilities using Transformers, designing a target detection framework for modeling unordered point clouds. Any RPN network can be used to obtain 3D proposals, combined with an improved Transformer network, thus improving the performance of point cloud 3D target detection networks.

[0244] Based on the above, this specification provides a point cloud data processing method that offers a flexible and high-performance point cloud 3D object detection method using an arbitrary RPN and an improved Transformer network. It should be noted that the effects achieved by this method include, but are not limited to:

[0245] First, the point cloud data processing method provided in this specification fully utilizes the original geometric features of the point cloud, without forcibly voxelizing it into a regular shape. Furthermore, the two-stage approach includes an additional correction process, significantly improving the detection performance for target objects.

[0246] Secondly, the point cloud data processing method provided in this specification utilizes the original point cloud and BEV features without utilizing the intermediate layer features of the RPN, thus making this method highly flexible. It can be composed of any RPN and an improved Transformer network to form a flexible and high-performance point cloud 3D target detection method.

[0247] Finally, this specification presents a point cloud data processing method that uses and improves upon the Transformer. It leverages the Transformer to process unordered point clouds and integrates geometric and semantic features to achieve accurate perception of object location and category. The point-to-keypoint attention mechanism in the Transformer significantly accelerates network convergence efficiency and ultimately substantially improves the detection metric mAP.

[0248] In summary, this specification presents a point cloud data processing method, proposing a Transformer-based 3D point cloud object detection network architecture that effectively combines the geometric and semantic features of point clouds, significantly improving the results of 3D point cloud object detection. Furthermore, through a point-to-keypoint attention mechanism, the learning efficiency of the encoder in the Transformer structure can be greatly accelerated.

[0249] See Figure 4 , Figure 4 A flowchart of a data processing method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0250] Step 402: Determine the data to be processed and input the data to be processed into the data processing model, wherein the data processing model includes a data processing module and a result adjustment module.

[0251] The data to be processed can be understood as data that needs to be processed by the data processing model. For example, the data to be processed can be image data to be processed, text data to be processed, code data to be processed, or point cloud data to be processed in the point cloud data processing method mentioned above.

[0252] A data processing model can be understood as a model for processing the data to be processed. For an explanation of the data processing model and how it processes the data to be processed, please refer to the corresponding or relevant content in the above-mentioned point cloud data processing methods, which will not be repeated here.

[0253] In one or more embodiments provided in this specification, determining the data to be processed includes:

[0254] Receive the data to be processed sent by the client, wherein the data to be processed is sent by the user through the data processing interface of the client;

[0255] The data processing interface can be an application interface, webpage, or other form of human-computer interaction between the client and the user. Through this interface, users can process data as needed.

[0256] Step 404: Use the data processing module to perform recognition processing on the data to be processed to obtain the initial recognition result of the data to be processed.

[0257] The initial recognition result can be the recognition result of the target object in the data to be processed; the target object can be an object contained in the data to be processed, for example, the target object can be an object in the image data to be processed, a specific text data in the text data to be processed, or a target object in the point cloud data to be processed; the initial recognition result may be an inaccurate or incorrect recognition result; correspondingly, the target recognition result can be an accurate and correct recognition result; the recognition result can be understood as location information, detection box, 3D proposal, etc.

[0258] Step 406: Using the result adjustment module, determine the result adjustment parameters based on the data to be processed and the initial recognition result, and perform result adjustment processing on the initial recognition result according to the result adjustment parameters to obtain the target recognition result of the data to be processed.

[0259] The adjustment parameters for this result can be found in the corresponding or relevant content in the above point cloud data processing method, and will not be elaborated here.

[0260] It should be noted that in the data processing method provided in one or more embodiments of this specification, the steps of the data processing model using the data processing module and the result adjustment module to process the data to be processed can be referred to in the above-mentioned point cloud data processing method, where the data processing model uses the data processing module and the result adjustment module to process the point cloud data to be processed, and will not be repeated here.

[0261] In one or more embodiments provided in this specification, determining the result adjustment parameters using the result adjustment module based on the data to be processed and the initial recognition result includes:

[0262] The data to be processed and the initial recognition result are input into the result adjustment module. The result adjustment module is used to determine the first data feature set and the second data feature set based on the data to be processed and the initial recognition result.

[0263] The first data features in the first data feature set are fused to obtain a first fused feature corresponding to the first data feature set; and the second data features in the second data feature set are fused to obtain a second fused feature corresponding to the second data feature set.

[0264] The first fusion feature and the second fusion feature are encoded using an attention mechanism to obtain a fusion feature encoding, and the fusion feature encoding is decoded to obtain the result adjustment parameters.

[0265] In one or more embodiments provided in this specification, determining a first data feature set and a second data feature set using the result adjustment module based on the data to be processed and the initial recognition result includes:

[0266] Using the result adjustment module, the data to be processed is dimensionally transformed to determine the data feature map corresponding to the data to be processed;

[0267] Based on the location information of the initial identification result, determine the result data corresponding to the initial identification result from the data to be processed, and based on the location information, determine the result feature corresponding to the initial identification result from the data feature map;

[0268] Determine the target location information for the initial identification result, and based on the target location information, determine the target result data from the result data, and based on the target location information, determine the target result features from the result features;

[0269] The result data and the result features are determined as the first data feature, and the first data feature set is determined based on the first data feature. The target result data and the target result features are determined as the second data feature, and the second data feature set is determined based on the second data feature.

[0270] In one or more embodiments provided in this specification, fusing first data features in the first data feature set to obtain a first fused feature corresponding to the first data feature set, and fusing second data features in the second data feature set to obtain a second fused feature corresponding to the second data feature set, includes:

[0271] Determine a first data feature in the first data feature set, and determine a second data feature in the second data feature set;

[0272] The feature fusion unit in the result adjustment module is used to fuse the first data features to obtain the first fused feature corresponding to the first data feature set, and the feature fusion unit is used to fuse the second data features to obtain the second fused feature corresponding to the second data feature set.

[0273] In one or more embodiments provided in this specification, the step of fusing the first data features using the feature fusion unit in the result adjustment module to obtain the first fused feature corresponding to the first data feature set, and fusing the second data features using the feature fusion unit to obtain the second fused feature corresponding to the second data feature set, includes:

[0274] The feature fusion unit in the result adjustment module is used to determine the center point coordinate data of the initial recognition result;

[0275] Using the feature fusion unit, the center point coordinate data is subtracted from the result data in the first data feature to obtain updated result data. The updated result data and the result features in the first data feature are then projected onto the first target dimension for feature fusion to obtain the first fused feature corresponding to the first data feature set.

[0276] Using the feature fusion unit, the target result data in the second data feature is subtracted from the center point coordinate data to obtain updated target result data. The updated target result data and the target result features in the second data feature are then projected onto the first target dimension for feature fusion to obtain the second fused feature corresponding to the second data feature set.

[0277] In one or more embodiments provided in this specification, the step of encoding the first fused feature and the second fused feature using an attention mechanism to obtain fused feature encoding includes:

[0278] Project the first fusion feature and the second fusion feature onto the second target dimension to obtain the first target dimension feature corresponding to the first fusion feature and the second target dimension feature corresponding to the second fusion feature;

[0279] Determine the similarity matrix between the first feature of the target dimension and the second feature of the target dimension;

[0280] Using an attention mechanism, a first attention matrix for the first fused feature is determined based on the similarity matrix; and using an attention mechanism, a second attention matrix for the second fused feature is determined based on the similarity matrix.

[0281] The fused feature encoding is obtained based on the first attention matrix and the second attention matrix.

[0282] In one or more embodiments provided in this specification, obtaining the fused feature encoding based on the first attention matrix and the second attention matrix includes:

[0283] A first updated feature corresponding to the first fusion feature is determined, and a second updated feature corresponding to the second fusion feature is determined, wherein the first updated feature is obtained by projecting the first fusion feature onto a third target dimension, and the second updated feature is obtained by projecting the second fusion feature onto the third target dimension;

[0284] The first attention matrix is ​​updated using the first matrix update strategy, based on the first update feature and the second update feature, to obtain the first updated attention matrix;

[0285] Using the second matrix update strategy, the second attention matrix is ​​updated based on the first update feature and the second update feature to obtain the second updated attention matrix;

[0286] The first updated attention matrix and the second updated attention matrix are encoded to obtain the fused feature encoding.

[0287] In one or more embodiments provided in this specification, after adjusting the initial recognition result according to the result adjustment parameters to obtain the target recognition result of the data to be processed, the method further includes:

[0288] The target recognition result is sent to the client so that the client can display the target recognition result to the user based on the data processing interface.

[0289] It should be noted that in the data processing methods provided in one or more embodiments of this specification, the data processing module and the result adjustment module in the data processing model can be deployed on the same or different computer hardware devices; wherein, the computer hardware device can be a computer, a client or a server; or the computer hardware device can be a CPU, GPU memory, video memory, etc. in a computer, client or server.

[0290] Therefore, when the data processing module and the result adjustment module in the data processing model are deployed on different computer hardware devices, the data processing steps for the data to be processed are realized through the cooperation between the different computer hardware devices. This avoids the problem of excessive load on the computer hardware device caused by deploying the data processing module and the result adjustment module in the data processing model on the same computer hardware device.

[0291] In this data processing model, when the data processing module and the result adjustment module are deployed on the same computer hardware device, the data processing steps for the data to be processed are implemented through a single computer hardware device, which improves the efficiency of data processing and avoids the problem of low data processing efficiency caused by the cooperation between different computer hardware devices.

[0292] In one or more embodiments provided in this specification, the data processing method first inputs the data to be processed into a data processing model, uses the data processing module in the data processing model to perform recognition processing on the data to be processed, and obtains an initial recognition result of the data to be processed. Then, in order to avoid the problem of inaccurate initial recognition result, the result adjustment module in the data processing model is used to determine the result adjustment parameters based on the data to be processed and the initial recognition result, and performs result adjustment processing on the initial recognition result based on the result adjustment parameters. This achieves the goal of obtaining an accurate target recognition result corresponding to the data to be processed, and avoids the problem that the inaccurate recognition result of the data to be processed may lead to the inability to meet the needs of processing the data in actual application scenarios.

[0293] The above is an illustrative scheme of a data processing method according to this embodiment. It should be noted that the technical solution of this data processing method belongs to the same concept as the technical solution of the point cloud data processing method described above. For details not described in detail in the technical solution of the data processing method, please refer to the description of the technical solution of the point cloud data processing method described above.

[0294] See Figure 5 , Figure 5 A flowchart of another point cloud data processing method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0295] Step 502: Determine the point cloud data to be processed.

[0296] Step 504: Input the point cloud data to be processed into the data processing model to obtain the target object recognition result of the point cloud data to be processed. The data processing model includes a data processing module and a result adjustment module. The data processing module is used to perform recognition processing on the point cloud data to be processed to obtain the initial object recognition result of the point cloud data to be processed. The result adjustment module is used to determine the result adjustment parameters based on the point cloud data to be processed and the initial object recognition result, and perform result adjustment processing on the initial object recognition result based on the result adjustment parameters to obtain the target object recognition result of the point cloud data to be processed.

[0297] In one or more embodiments provided in this specification, this other point cloud data processing method, during the processing of point cloud data to be processed using a data processing model, firstly inputs the point cloud data to be processed into the data processing model, and uses the data processing module in the data processing model to perform recognition processing on the point cloud data to be processed to obtain the initial object recognition result of the point cloud data to be processed. Then, in order to avoid the problem of inaccurate initial object recognition result, the result adjustment module in the data processing model is used to determine the result adjustment parameter based on the point cloud data to be processed and the initial object recognition result, and performs result adjustment processing on the initial object recognition result according to the result adjustment parameter, thereby achieving accurate acquisition of the target object recognition result corresponding to the point cloud data to be processed, avoiding the problem that the inaccurate point cloud data processing result cannot meet the needs of point cloud data processing in actual application scenarios.

[0298] The above is an illustrative scheme of another point cloud data processing method in this embodiment. It should be noted that the technical solution of this other point cloud data processing method belongs to the same concept as the technical solution of the point cloud data processing method described above. For details not described in detail in the technical solution of the other point cloud data processing method, please refer to the description of the technical solution of the point cloud data processing method described above.

[0299] See Figure 6 , Figure 6 A flowchart of a machine learning model training method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0300] Step 602: Determine the machine learning model to be trained and the training data associated with the target task, wherein the machine learning model to be trained includes a data processing module and a result adjustment module.

[0301] Step 604: Use the data processing module to perform recognition processing on the training data to obtain the initial recognition result of the training data.

[0302] Step 606: Using the result adjustment module, determine the result adjustment parameters based on the training data and the initial recognition result, and perform result adjustment processing on the initial recognition result according to the result adjustment parameters to obtain the target recognition result of the training data.

[0303] Step 608: Based on the target recognition result, adjust the model parameters of the machine learning model to obtain a trained target model, wherein the target model is used to perform the target task.

[0304] In one or more embodiments provided in this specification, determining the result adjustment parameters using the result adjustment module based on the training data and the initial recognition result includes:

[0305] The training data and the initial recognition result are input into the result adjustment module. The result adjustment module is used to determine the first data feature set and the second data feature set based on the training data and the initial recognition result.

[0306] The first data features in the first data feature set are fused to obtain a first fused feature corresponding to the first data feature set; and the second data features in the second data feature set are fused to obtain a second fused feature corresponding to the second data feature set.

[0307] The first fusion feature and the second fusion feature are encoded using an attention mechanism to obtain fusion feature encoding, and the fusion feature encoding is decoded to obtain the result adjustment data.

[0308] In one or more embodiments provided in this specification, determining the first data feature set and the second data feature set using the result adjustment module based on the training data and the initial recognition result includes:

[0309] Using the result adjustment module, the training data is dimensionally transformed to determine the data feature map corresponding to the training data;

[0310] Based on the location information of the initial recognition result, the result data corresponding to the initial recognition result is determined from the training data, and based on the location information, the result features corresponding to the initial recognition result are determined from the data feature map;

[0311] Determine the target location information for the initial identification result, and based on the target location information, determine the target result data from the result data, and based on the target location information, determine the target result features from the result features;

[0312] The result data and the result features are determined as the first data feature, and the first data feature set is determined based on the first data feature. The target result data and the target result features are determined as the second data feature, and the second data feature set is determined based on the second data feature.

[0313] In one or more embodiments provided in this specification, fusing first data features in the first data feature set to obtain a first fused feature corresponding to the first data feature set, and fusing second data features in the second data feature set to obtain a second fused feature corresponding to the second data feature set, includes:

[0314] Determine a first data feature in the first data feature set, and determine a second data feature in the second data feature set;

[0315] The feature fusion unit in the result adjustment module is used to fuse the first data features to obtain the first fused feature corresponding to the first data feature set, and the feature fusion unit is used to fuse the second data features to obtain the second fused feature corresponding to the second data feature set.

[0316] In one or more embodiments provided in this specification, the step of fusing the first data features using the feature fusion unit in the result adjustment module to obtain the first fused feature corresponding to the first data feature set, and fusing the second data features using the feature fusion unit to obtain the second fused feature corresponding to the second data feature set, includes:

[0317] The feature fusion unit in the result adjustment module is used to determine the center point coordinate data of the initial recognition result;

[0318] Using the feature fusion unit, the center point coordinate data is subtracted from the result data in the first data feature to obtain updated result data. The updated result data and the result features in the first data feature are then projected onto the first target dimension for feature fusion to obtain the first fused feature corresponding to the first data feature set.

[0319] Using the feature fusion unit, the target result data in the second data feature is subtracted from the center point coordinate data to obtain updated target result data. The updated target result data and the target result features in the second data feature are then projected onto the first target dimension for feature fusion to obtain the second fused feature corresponding to the second data feature set.

[0320] In one or more embodiments provided in this specification, the step of encoding the first fused feature and the second fused feature using an attention mechanism to obtain fused feature encoding includes:

[0321] Project the first fusion feature and the second fusion feature onto the second target dimension to obtain the first target dimension feature corresponding to the first fusion feature and the second target dimension feature corresponding to the second fusion feature;

[0322] Determine the similarity matrix between the first feature of the target dimension and the second feature of the target dimension;

[0323] Using an attention mechanism, a first attention matrix for the first fused feature is determined based on the similarity matrix; and using an attention mechanism, a second attention matrix for the second fused feature is determined based on the similarity matrix.

[0324] The fused feature encoding is obtained based on the first attention matrix and the second attention matrix.

[0325] In one or more embodiments provided in this specification, obtaining the fused feature encoding based on the first attention matrix and the second attention matrix includes:

[0326] A first updated feature corresponding to the first fusion feature is determined, and a second updated feature corresponding to the second fusion feature is determined, wherein the first updated feature is obtained by projecting the first fusion feature onto a third target dimension, and the second updated feature is obtained by projecting the second fusion feature onto the third target dimension;

[0327] The first attention matrix is ​​updated using the first matrix update strategy, based on the first update feature and the second update feature, to obtain the first updated attention matrix;

[0328] Using the second matrix update strategy, the second attention matrix is ​​updated based on the first update feature and the second update feature to obtain the second updated attention matrix;

[0329] The first updated attention matrix and the second updated attention matrix are encoded to obtain the fused feature encoding.

[0330] In one or more embodiments provided in this specification, the machine learning model training method first inputs training data into the machine learning model to be trained during the training process. The data processing module in the machine learning model then performs recognition processing on the training data to obtain initial recognition results. Next, to avoid inaccurate initial recognition results, the result adjustment module in the machine learning model determines result adjustment parameters based on the training data and the initial recognition results. The initial recognition results are then adjusted according to these parameters to obtain accurate target recognition results. Finally, the accurate target recognition results are used to adjust the model parameters of the machine learning model to obtain a target model capable of outputting accurate target recognition results, thus avoiding the problem of inaccurate recognition results output by the target model.

[0331] The above is an illustrative scheme of the machine learning model training method in this embodiment. It should be noted that the technical solution of this machine learning model training method belongs to the same concept as the technical solutions of the data processing method and the point cloud data processing method described above. For details not described in detail in the technical solution of the machine learning model training method, please refer to the descriptions of the technical solutions of the data processing method and the point cloud data processing method described above.

[0332] See Figure 7 , Figure 7 A flowchart of a data processing model training method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0333] Step 702: Determine the data processing model to be trained and the training point cloud data associated with the point cloud data processing task, wherein the object recognition model to be trained includes a data processing module and a result adjustment module.

[0334] Step 704: Use the data processing module to perform object recognition processing on the training point cloud data to obtain the initial object recognition result of the training point cloud data.

[0335] Step 706: Using the data processing module, determine the result adjustment parameters based on the training point cloud data and the initial object recognition result, and perform result adjustment processing on the initial object recognition result based on the result adjustment parameters to obtain the target object recognition result of the training point cloud data.

[0336] Step 708: Based on the target object recognition result, adjust the model parameters of the data processing model to obtain a trained data processing model, wherein the data processing model is used to execute the point cloud data processing task.

[0337] In one or more embodiments provided in this specification, determining the result adjustment parameters using the data processing module based on the training point cloud data and the initial object recognition result includes:

[0338] The training point cloud data and the initial object recognition result are input into the result adjustment module. The result adjustment module is used to determine the first data feature set and the second data feature set based on the training point cloud data and the initial object recognition result.

[0339] The first data features in the first data feature set are fused to obtain the first fused feature corresponding to the first data feature set; and the second data features in the second data feature set are fused to obtain the second fused feature corresponding to the second data feature set.

[0340] The first fusion feature and the second fusion feature are encoded using an attention mechanism to obtain object fusion feature encoding, and the object fusion feature encoding is decoded to obtain object adjustment data.

[0341] 3. The object recognition model training method according to claim 2, wherein the step of using the result adjustment module to determine the first data feature set and the second data feature set based on the training point cloud data and the initial object recognition result includes:

[0342] Using the result adjustment module, the training point cloud data is dimensionally transformed to determine the point cloud feature map corresponding to the training point cloud data.

[0343] Based on the object location information of the initial object recognition result, object point cloud data is determined from the training point cloud data, and object point cloud features are determined from the point cloud feature map based on the object location information.

[0344] Determine the target point cloud location information for the initial object recognition result, and based on the target point cloud location information, determine the target point cloud data from the object point cloud data, and based on the target point cloud location information, determine the target point cloud features from the point cloud feature map;

[0345] The object point cloud data and the object point cloud features are determined as first data features, and a first data feature set is determined based on the first data features. The target point cloud data and the target point cloud features are determined as second data features, and a second data feature set is determined based on the second data features.

[0346] In one or more embodiments provided in this specification, fusing first data features in the first data feature set to obtain a first fused feature corresponding to the first data feature set, and fusing second data features in the second data feature set to obtain a second fused feature corresponding to the second data feature set, includes:

[0347] Determine a first data feature in the first data feature set, and determine a second data feature in the second data feature set;

[0348] The feature fusion unit in the result adjustment module is used to fuse the first data features to obtain the first fused feature corresponding to the first data feature set, and the feature fusion unit is used to fuse the second data features to obtain the second fused feature corresponding to the second data feature set.

[0349] In one or more embodiments provided in this specification, the step of fusing the first data features using the feature fusion unit in the result adjustment module to obtain the first fused feature corresponding to the first data feature set, and fusing the second data features using the feature fusion unit to obtain the second fused feature corresponding to the second data feature set, includes:

[0350] Using the feature fusion unit, the center point coordinate data of the initial object recognition result is determined;

[0351] Subtract the center point coordinate data from the object point cloud data in the first data feature to obtain updated object point cloud data. Project the updated object point cloud data and the object point cloud features in the first data feature set onto the first target dimension for feature fusion to obtain the first fused feature corresponding to the first data feature set.

[0352] Subtract the center point coordinate data from the target point cloud data in the second data feature set to obtain updated target point cloud data. Project the updated target point cloud data and the target point cloud features in the second data feature set onto the first target dimension for feature fusion to obtain the second fused feature corresponding to the second data feature set.

[0353] In one or more embodiments provided in this specification, the step of encoding the first fused feature and the second fused feature using an attention mechanism to obtain fused feature encoding includes:

[0354] Project the first fusion feature and the second fusion feature onto the second target dimension to obtain the first target dimension feature corresponding to the first fusion feature and the second target dimension feature corresponding to the second fusion feature;

[0355] Determine the similarity matrix between the first feature of the target dimension and the second feature of the target dimension;

[0356] Using an attention mechanism, a first attention matrix for the first fused feature is determined based on the similarity matrix; and using an attention mechanism, a second attention matrix for the second fused feature is determined based on the similarity matrix.

[0357] The fused feature encoding is obtained based on the first attention matrix and the second attention matrix.

[0358] In one or more embodiments provided in this specification, obtaining the fused feature encoding based on the first attention matrix and the second attention matrix includes:

[0359] A first updated feature corresponding to the first fusion feature is determined, and a second updated feature corresponding to the second fusion feature is determined, wherein the first updated feature is obtained by projecting the first fusion feature onto a third target dimension, and the second updated feature is obtained by projecting the second fusion feature onto the third target dimension;

[0360] The first attention matrix is ​​updated using the first matrix update strategy, based on the first update feature and the second update feature, to obtain the first updated attention matrix;

[0361] Using the second matrix update strategy, the second attention matrix is ​​updated based on the first update feature and the second update feature to obtain the second updated attention matrix;

[0362] The first updated attention matrix and the second updated attention matrix are encoded to obtain the fused feature encoding.

[0363] In one or more embodiments provided in this specification, the initial object recognition result is an initial object detection box; the result adjustment parameter is a correction value for the initial object detection box;

[0364] The step of adjusting the parameters based on the results to perform result adjustment processing on the initial object recognition result, and obtaining the target object recognition result of the point cloud data to be processed, includes:

[0365] The initial object detection box is corrected according to the correction value to obtain the target object detection box corresponding to the object to be identified in the point cloud data to be processed.

[0366] In one or more embodiments provided in this specification, the data processing model training method first inputs training point cloud data into the data processing model to be trained, and uses the data processing module in the data processing model to perform object recognition processing on the training point cloud data to obtain initial object recognition results. Secondly, to avoid inaccurate initial object recognition results, the data processing module in the data processing model determines result adjustment parameters based on the training point cloud data and the initial object recognition results, and adjusts the initial object recognition results according to the result adjustment parameters to obtain target object recognition results for the training point cloud data. Finally, the accurate target object recognition results are used to adjust the model parameters of the data processing model to obtain a data processing model capable of outputting accurate target object recognition results, thus avoiding the problem of inaccurate target object recognition results output by the data processing model.

[0367] The above is an illustrative scheme of the data processing model training method in this embodiment. It should be noted that the technical solution of this data processing model training method belongs to the same concept as the above-described data processing method and point cloud data processing method. For details not described in detail in the technical solution of the data processing model training method, please refer to the descriptions of the above-described data processing method and point cloud data processing method.

[0368] Corresponding to the above method embodiments, this specification also provides an embodiment of a point cloud data processing device, which includes:

[0369] The data determination module is configured to determine the point cloud data to be processed and input the point cloud data to be processed into the data processing model, wherein the data processing model includes a data processing module and a result adjustment module.

[0370] The result acquisition module is configured to use the data processing module to perform recognition processing on the point cloud data to be processed, and obtain the initial object recognition result of the point cloud data to be processed.

[0371] The result adjustment module is configured to determine result adjustment parameters based on the point cloud data to be processed and the initial object recognition result, and to perform result adjustment processing on the initial object recognition result based on the result adjustment parameters to obtain the target object recognition result of the point cloud data to be processed.

[0372] Optionally, the result adjustment module is further configured to:

[0373] The point cloud data to be processed and the initial object recognition result are input into the result adjustment module in the data processing model. The result adjustment module is used to determine the first data feature set and the second data feature set based on the point cloud data to be processed and the initial object recognition result.

[0374] The first data features in the first data feature set are fused to obtain a first fused feature corresponding to the first data feature set; and the second data features in the second data feature set are fused to obtain a second fused feature corresponding to the second data feature set.

[0375] The first fusion feature and the second fusion feature are encoded using an attention mechanism to obtain a fusion feature encoding, and the fusion feature encoding is decoded to obtain the result adjustment parameters.

[0376] Optionally, the result adjustment module is further configured to:

[0377] Using the result adjustment module, the point cloud data to be processed is dimensionally transformed to determine the point cloud feature map corresponding to the point cloud data to be processed.

[0378] Based on the object location information of the initial object recognition result, object point cloud data is determined from the point cloud data to be processed, and object point cloud features are determined from the point cloud feature map based on the object location information.

[0379] Determine the target point cloud location information for the initial object recognition result, and based on the target point cloud location information, determine the target point cloud data from the object point cloud data, and based on the target point cloud location information, determine the target point cloud features from the object point cloud features;

[0380] The object point cloud data and the object point cloud features are determined as first data features, and a first data feature set is determined based on the first data features. The target point cloud data and the target point cloud features are determined as second data features, and a second data feature set is determined based on the second data features.

[0381] Optionally, the result adjustment module is further configured to:

[0382] Determine a first data feature in the first data feature set, and determine a second data feature in the second data feature set;

[0383] The feature fusion unit in the result adjustment module is used to fuse the first data features to obtain the first fused feature corresponding to the first data feature set, and the feature fusion unit is used to fuse the second data features to obtain the second fused feature corresponding to the second data feature set.

[0384] Optionally, the result adjustment module is further configured to:

[0385] Using the feature fusion unit, the center point coordinate data of the initial object recognition result is determined;

[0386] Subtract the center point coordinate data from the object point cloud data in the first data feature to obtain updated object point cloud data. Project the updated object point cloud data and the object point cloud features in the first data feature set onto the first target dimension for feature fusion to obtain the first fused feature corresponding to the first data feature set.

[0387] Subtract the center point coordinate data from the target point cloud data in the second data feature set to obtain updated target point cloud data. Project the updated target point cloud data and the target point cloud features in the second data feature set onto the first target dimension for feature fusion to obtain the second fused feature corresponding to the second data feature set.

[0388] Optionally, the result adjustment module is further configured to:

[0389] Project the first fusion feature and the second fusion feature onto the second target dimension to obtain the first target dimension feature corresponding to the first fusion feature and the second target dimension feature corresponding to the second fusion feature;

[0390] Determine the similarity matrix between the first feature of the target dimension and the second feature of the target dimension;

[0391] Using an attention mechanism, a first attention matrix for the first fused feature is determined based on the similarity matrix; and using an attention mechanism, a second attention matrix for the second fused feature is determined based on the similarity matrix.

[0392] The fused feature encoding is obtained based on the first attention matrix and the second attention matrix.

[0393] Optionally, the result adjustment module is further configured to:

[0394] A first updated feature corresponding to the first fusion feature is determined, and a second updated feature corresponding to the second fusion feature is determined, wherein the first updated feature is obtained by projecting the first fusion feature onto a third target dimension, and the second updated feature is obtained by projecting the second fusion feature onto the third target dimension;

[0395] The first attention matrix is ​​updated using the first matrix update strategy, based on the first update feature and the second update feature, to obtain the first updated attention matrix;

[0396] Using the second matrix update strategy, the second attention matrix is ​​updated based on the first update feature and the second update feature to obtain the second updated attention matrix;

[0397] The first updated attention matrix and the second updated attention matrix are encoded to obtain the fused feature encoding.

[0398] Optionally, the initial object recognition result is an initial object detection box; the result adjustment parameter is a correction value for the initial object detection box;

[0399] The result adjustment module is further configured to:

[0400] The initial object detection box is corrected according to the correction value to obtain the target object detection box corresponding to the object to be identified in the point cloud data to be processed.

[0401] Optionally, the data determination module is further configured to:

[0402] The device receives point cloud data to be processed sent by a point cloud data acquisition device, wherein the point cloud data acquisition device is a device configured by the user through a client for performing point cloud data acquisition operations for a target scene, and the point cloud data to be processed is obtained by the point cloud data acquisition device performing point cloud data acquisition operations for the target scene.

[0403] The point cloud data processing device further includes a result sending module, configured as follows:

[0404] The client corresponding to the point cloud data acquisition device is determined, and the target object recognition result is sent to the client so that the client can display the target object recognition result to the user.

[0405] The point cloud data processing apparatus in one or more embodiments provided in this specification first inputs the point cloud data to be processed into a data processing model. The data processing module in the data processing model then performs recognition processing on the point cloud data to obtain an initial object recognition result. Next, to avoid inaccurate initial object recognition results, a result adjustment module in the data processing model determines result adjustment parameters based on the point cloud data to be processed and the initial object recognition result. The initial object recognition result is then adjusted according to these parameters. This achieves accurate acquisition of the target object recognition result corresponding to the point cloud data to be processed, avoiding the problem that inaccurate point cloud data processing results would fail to meet the needs of point cloud data processing in practical application scenarios.

[0406] The above is an illustrative scheme of a point cloud data processing device according to this embodiment. It should be noted that the technical solution of this point cloud data processing device and the technical solution of the point cloud data processing method described above belong to the same concept. For details not described in detail in the technical solution of the point cloud data processing device, please refer to the description of the technical solution of the point cloud data processing method described above.

[0407] Corresponding to the above method embodiments, this specification also provides a data processing apparatus embodiment, which includes:

[0408] The data determination module is configured to determine the data to be processed and input the data to be processed into the data processing model, wherein the data processing model includes a data processing module and a result adjustment module;

[0409] The result acquisition module is configured to use the data processing module to identify and process the data to be processed, and obtain the initial identification result of the data to be processed.

[0410] The result adjustment module is configured to determine result adjustment parameters based on the data to be processed and the initial recognition result, and to perform result adjustment processing on the initial recognition result based on the result adjustment parameters to obtain the target recognition result of the data to be processed.

[0411] Optionally, the result adjustment module is further configured to:

[0412] The data to be processed and the initial recognition result are input into the result adjustment module. The result adjustment module is used to determine the first data feature set and the second data feature set based on the data to be processed and the initial recognition result.

[0413] The first data features in the first data feature set are fused to obtain the first fused feature corresponding to the first data feature set; and the second data features in the second data feature set are fused to obtain the second fused feature corresponding to the second data feature set.

[0414] The first fusion feature and the second fusion feature are encoded using an attention mechanism to obtain a fusion feature encoding, and the fusion feature encoding is decoded to obtain the result adjustment parameters.

[0415] Optionally, the result adjustment module is further configured to:

[0416] Using the result adjustment module, the data to be processed is dimensionally transformed to determine the data feature map corresponding to the data to be processed;

[0417] Based on the location information of the initial recognition result, the result data corresponding to the initial recognition result is determined from the data to be processed, and based on the location information, the result features corresponding to the initial recognition result are determined from the data feature map.

[0418] Determine the target location information for the initial identification result, and based on the target location information, determine the target result data from the result data, and based on the target location information, determine the target result features from the result features;

[0419] The result data and the result features are determined as the first data feature, and the first data feature set is determined based on the first data feature. The target result data and the target result features are determined as the second data feature, and the second data feature set is determined based on the second data feature.

[0420] Optionally, the result adjustment module is further configured to:

[0421] Determine a first data feature in the first data feature set, and determine a second data feature in the second data feature set:

[0422] The feature fusion unit in the result adjustment module is used to fuse the first data features to obtain the first fused feature corresponding to the first data feature set, and the feature fusion unit is used to fuse the second data features to obtain the second fused feature corresponding to the second data feature set.

[0423] Optionally, the result adjustment module is further configured to:

[0424] The feature fusion unit in the result adjustment module is used to determine the center point coordinate data of the initial recognition result;

[0425] Using the feature fusion unit, the center point coordinate data is subtracted from the result data in the first data feature to obtain updated result data. The updated result data and the result features in the first data feature are then projected onto the first target dimension for feature fusion to obtain the first fused feature corresponding to the first data feature set.

[0426] Using the feature fusion unit, the target result data in the second data feature is subtracted from the center point coordinate data to obtain updated target result data. The updated target result data and the target result features in the second data feature are then projected onto the first target dimension for feature fusion to obtain the second fused feature corresponding to the second data feature set.

[0427] Optionally, the result adjustment module is further configured to:

[0428] Project the first fusion feature and the second fusion feature onto the second target dimension to obtain the first target dimension feature corresponding to the first fusion feature and the second target dimension feature corresponding to the second fusion feature;

[0429] Determine the similarity matrix between the first feature of the target dimension and the second feature of the target dimension;

[0430] Using an attention mechanism, a first attention matrix for the first fused feature is determined based on the similarity matrix; and using an attention mechanism, a second attention matrix for the second fused feature is determined based on the similarity matrix.

[0431] The fused feature encoding is obtained based on the first attention matrix and the second attention matrix.

[0432] Optionally, the result adjustment module is further configured to:

[0433] A first updated feature corresponding to the first fusion feature is determined, and a second updated feature corresponding to the second fusion feature is determined, wherein the first updated feature is obtained by projecting the first fusion feature onto a third target dimension, and the second updated feature is obtained by projecting the second fusion feature onto the third target dimension;

[0434] The first attention matrix is ​​updated using the first matrix update strategy, based on the first update feature and the second update feature, to obtain the first updated attention matrix;

[0435] Using the second matrix update strategy, the second attention matrix is ​​updated based on the first update feature and the second update feature to obtain the second updated attention matrix;

[0436] The first updated attention matrix and the second updated attention matrix are encoded to obtain the fused feature encoding.

[0437] Optionally, the data determination module is further configured to:

[0438] Receive the data to be processed sent by the client, wherein the data to be processed is sent by the user through the data processing interface of the client;

[0439] The data processing device further includes a result sending module, configured to:

[0440] The target recognition result is sent to the client so that the client can display the target recognition result to the user based on the data processing interface.

[0441] In one or more embodiments provided in this specification, the data processing apparatus first inputs the data to be processed into a data processing model. Using the data processing module within the model, it performs recognition processing on the data to be processed, obtaining an initial recognition result. Then, to avoid inaccurate initial recognition results, the result adjustment module within the data processing model determines result adjustment parameters based on the data to be processed and the initial recognition result. Based on these parameters, the initial recognition result is adjusted, thereby achieving an accurate target recognition result corresponding to the data to be processed. This avoids the problem that inaccurate recognition results of the data to be processed would prevent the processing of the data from failing to meet the needs of actual application scenarios.

[0442] The above is an illustrative scheme of a data processing apparatus according to this embodiment. It should be noted that the technical solution of this data processing apparatus and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the data processing apparatus, please refer to the description of the technical solution of the data processing method described above.

[0443] Corresponding to the above method embodiments, this specification also provides another embodiment of a point cloud data processing apparatus, which includes:

[0444] The data determination module is configured to determine the point cloud data to be processed.

[0445] The result acquisition module is configured to input the point cloud data to be processed into a data processing model to obtain the target object recognition result of the point cloud data to be processed. The data processing model includes a data processing module and a result adjustment module. The data processing module is used to perform recognition processing on the point cloud data to be processed to obtain an initial object recognition result of the point cloud data to be processed. The result adjustment module is used to determine result adjustment parameters based on the point cloud data to be processed and the initial object recognition result, and perform result adjustment processing on the initial object recognition result based on the result adjustment parameters to obtain the target object recognition result of the point cloud data to be processed.

[0446] Another point cloud data processing device in one or more embodiments provided in this specification, in the process of processing point cloud data to be processed using a data processing model, firstly, inputs the point cloud data to be processed into the data processing model, and uses the data processing module in the data processing model to perform recognition processing on the point cloud data to be processed to obtain the initial object recognition result of the point cloud data to be processed. Then, in order to avoid the problem of inaccurate initial object recognition result, the result adjustment module in the data processing model determines the result adjustment parameter according to the point cloud data to be processed and the initial object recognition result, and performs result adjustment processing on the initial object recognition result according to the result adjustment parameter, thereby realizing the accurate acquisition of the target object recognition result corresponding to the point cloud data to be processed, avoiding the problem that the processing result of the point cloud data is inaccurate, thus failing to meet the needs of point cloud data processing in actual application scenarios.

[0447] The above is an illustrative scheme of another point cloud data processing device according to this embodiment. It should be noted that the technical solution of this other point cloud data processing device and the technical solution of the other point cloud data processing method described above belong to the same concept. For details not described in detail in the technical solution of the other point cloud data processing device, please refer to the description of the technical solution of the other point cloud data processing method described above.

[0448] Corresponding to the above method embodiments, this specification also provides embodiments of a machine learning model training apparatus, which includes:

[0449] The training data determination module is configured to determine the machine learning model to be trained and the training data associated with the target task, wherein the machine learning model to be trained includes a data processing module and a result adjustment module.

[0450] The result acquisition module is configured to use the data processing module to identify and process the training data to obtain the initial identification result of the training data;

[0451] The result adjustment module is configured to determine result adjustment parameters based on the training data and the initial recognition result, and to perform result adjustment processing on the initial recognition result based on the result adjustment parameters to obtain the target recognition result of the training data.

[0452] The model training module is configured to adjust the model parameters of the machine learning model based on the target recognition result to obtain a trained target model, wherein the target model is used to perform the target task.

[0453] In one or more embodiments of the machine learning model training apparatus provided in this specification, during the training process of the machine learning model, training data is first input into the machine learning model to be trained. The data processing module in the machine learning model is used to perform recognition processing on the training data to obtain the initial recognition result of the training data. Secondly, in order to avoid the problem of inaccurate initial recognition result, the result adjustment module in the machine learning model is used to determine the result adjustment parameters based on the training data and the initial recognition result, and the initial recognition result is adjusted according to the result adjustment parameters to obtain an accurate target recognition result. Finally, the accurate target recognition result is used to adjust the model parameters of the machine learning model to obtain a target model that can output accurate target recognition results, thus avoiding the problem of inaccurate recognition results output by the target model.

[0454] The above is an illustrative scheme of a machine learning model training device according to this embodiment. It should be noted that the technical solution of this machine learning model training device and the technical solution of the above-described machine learning model training method belong to the same concept. For details not described in detail in the technical solution of the machine learning model training device, please refer to the description of the technical solution of the above-described machine learning model training method.

[0455] Corresponding to the above method embodiments, this specification also provides an embodiment of a data processing model training apparatus, which includes:

[0456] The training data determination module is configured to determine the data processing model to be trained and the training point cloud data associated with the point cloud data processing task, wherein the object recognition model to be trained includes a data processing module and a result adjustment module.

[0457] The result acquisition module is configured to use the data processing module to perform object recognition processing on the training point cloud data to obtain the initial object recognition result of the training point cloud data:

[0458] The result adjustment module is configured to use the data processing module to determine result adjustment parameters based on the training point cloud data and the initial object recognition result, and to perform result adjustment processing on the initial object recognition result based on the result adjustment parameters to obtain the target object recognition result of the training point cloud data.

[0459] The model training module is configured to adjust the model parameters of the data processing model based on the target object recognition result to obtain a trained data processing model, wherein the data processing model is used to perform the point cloud data processing task.

[0460] In the data processing model training apparatus provided in one or more embodiments of this specification, during the training process of the data processing model, firstly, training point cloud data is input into the data processing model to be trained. The data processing module in the data processing model to be trained performs object recognition processing on the training point cloud data to obtain initial object recognition results. Secondly, to avoid the problem of inaccurate initial object recognition results, the data processing module in the data processing model determines result adjustment parameters based on the training point cloud data and the initial object recognition results, and performs result adjustment processing on the initial object recognition results based on the result adjustment parameters to obtain target object recognition results of the training point cloud data. Finally, the accurate target object recognition results are used to adjust the model parameters of the data processing model to obtain a data processing model that can output accurate target object recognition results, thus avoiding the problem of inaccurate target object recognition results output by the data processing model.

[0461] The above is an illustrative scheme of a data processing model training device according to this embodiment. It should be noted that the technical solution of this data processing model training device and the technical solution of the data processing model training method described above belong to the same concept. For details not described in detail in the technical solution of the data processing model training device, please refer to the description of the technical solution of the data processing model training method described above.

[0462] See Figure 8 , Figure 8 A flowchart of a traffic data processing method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0463] Step 802: Determine the point cloud data to be processed corresponding to the target traffic scene, and input the point cloud data to be processed into the data processing model, wherein the data processing model includes a data processing module and a result adjustment module.

[0464] The target traffic scenario can be understood as the traffic environment that requires point cloud data collection. For example, the target traffic environment can be a traffic area, including but not limited to intersections, highways, roads within communities, overpasses, pedestrian crossings, highway toll stations, etc. No specific restrictions are imposed here.

[0465] Point cloud data to be processed can be understood as point cloud data collected from the target traffic scene. By using point cloud data acquisition equipment to collect point cloud data from the target traffic scene, the corresponding point cloud data to be processed can be obtained.

[0466] The data processing model can be found in the embodiments of the above-mentioned data processing method, point cloud data processing method, or data processing model training method, and will not be elaborated here.

[0467] In one or more embodiments provided in this specification, determining the point cloud data to be processed corresponding to the target traffic scene includes:

[0468] The system receives point cloud data corresponding to a target traffic scene from a point cloud data acquisition device. This point cloud data acquisition device is configured by a user through a client to perform point cloud data acquisition operations on the target traffic scene. The point cloud data to be processed is obtained by the point cloud data acquisition device performing point cloud data acquisition operations on the target traffic scene. In this embodiment, by receiving the point cloud data to be processed from the point cloud data acquisition device, the system can process the point cloud data corresponding to the target traffic scene in a timely manner, thereby solving the problem of identifying target objects in the target traffic scene in real-world scenarios.

[0469] Step 804: Using the data processing module, the point cloud data to be processed is identified and processed to obtain the initial traffic object identification result of the point cloud data to be processed.

[0470] The initial traffic object recognition result can be understood as the initial traffic object recognition result obtained after performing object recognition processing on the traffic objects contained in the point cloud data to be processed. The initial traffic object recognition result can determine the traffic objects in the target traffic scene. For example, the initial traffic object recognition result can be the detection bounding box, coordinate information, and location information of the traffic objects. In one or more embodiments provided in this specification, the initial traffic object recognition result can be 3D Proposals output by the RPN model, which are coarse and inaccurate. In one or more embodiments provided in this specification, the target traffic scene can be an intersection area, and the traffic objects can be objects such as cars violating traffic rules and pedestrians violating traffic rules in the intersection area.

[0471] It should be noted that, since the initial traffic object identification result may be inaccurate, in order to ensure the accuracy of the traffic object identification result of the data processing model, the initial traffic object identification result needs to be corrected by the result adjustment module to obtain the accurate target traffic object identification result; that is to say, the initial traffic object identification result may be an incorrect, inaccurate or incomplete object identification result; correspondingly, the target traffic object identification result may be a correct, accurate or complete object identification result.

[0472] Step 806: Using the result adjustment module, determine the result adjustment parameters based on the point cloud data to be processed and the initial traffic object recognition result, and perform result adjustment processing on the initial traffic object recognition result according to the result adjustment parameters to obtain the target traffic object recognition result of the point cloud data to be processed.

[0473] The target traffic object recognition result can be understood as the traffic object recognition result obtained after adjusting the initial traffic object recognition result through result adjustment parameters; the traffic object in the target traffic scene can be determined through the target traffic object recognition result; for example, the target traffic object recognition result can be the detection box, bounding box, coordinate information, position information, etc. for the traffic object; in one or more embodiments provided in this specification, the target traffic object recognition result can be 3D Proposals, and the 3D Proposals are accurate 3D Proposals.

[0474] In one or more embodiments provided in this specification, the step of using the result adjustment module to determine the result adjustment parameters based on the point cloud data to be processed and the initial traffic object recognition result includes:

[0475] The point cloud data to be processed and the initial traffic object recognition result are input into the result adjustment module in the data processing model. The result adjustment module is used to determine the first data feature set and the second data feature set based on the point cloud data to be processed and the initial traffic object recognition result.

[0476] The first data features in the first data feature set are fused to obtain the first fused feature corresponding to the first data feature set; and the second data features in the second data feature set are fused to obtain the second fused feature corresponding to the second data feature set.

[0477] The first fusion feature and the second fusion feature are encoded using an attention mechanism to obtain a fusion feature encoding, and the fusion feature encoding is decoded to obtain the result adjustment parameters.

[0478] In one or more embodiments provided in this specification, the result adjustment module determines a first data feature set and a second data feature set based on the point cloud data to be processed and the initial traffic object recognition result, including:

[0479] Using the result adjustment module, the point cloud data to be processed is dimensionally transformed to determine the point cloud feature map corresponding to the point cloud data to be processed.

[0480] Based on the object location information of the initial traffic object identification result, object point cloud data is determined from the point cloud data to be processed, and object point cloud features are determined from the point cloud feature map based on the object location information.

[0481] Determine the target point cloud location information for the initial traffic object identification result, and based on the target point cloud location information, determine the target point cloud data from the object point cloud data, and based on the target point cloud location information, determine the target point cloud features from the object point cloud features;

[0482] The object point cloud data and the object point cloud features are determined as first data features, and a first data feature set is determined based on the first data features. The target point cloud data and the target point cloud features are determined as second data features, and a second data feature set is determined based on the second data features.

[0483] In the above embodiments, by utilizing the result adjustment module to perform dimensional transformation on the point cloud data to be processed and determining the point cloud feature map corresponding to the point cloud data to be processed, the result adjustment module can determine the result adjustment parameters based on the point cloud feature map it has determined. This avoids the problem that the result adjustment module needs to use the intermediate layer features of the data processing module, which leads to a strong coupling relationship between the data processing module and the result adjustment module, making it difficult to separate them. Furthermore, by using the location information of the initial traffic object recognition result and the target point cloud location information for the initial traffic object recognition result, the first data feature set and the second data feature set are determined, which facilitates the accurate determination of the result adjustment parameters for the initial traffic object recognition result based on the similarity comparison between the two sets.

[0484] In one or more embodiments provided in this specification, after adjusting the initial traffic object identification result according to the result adjustment parameters to obtain the target traffic object identification result of the point cloud data to be processed, the method further includes:

[0485] The client corresponding to the point cloud data acquisition device is determined, and the target traffic object recognition result is sent to the client so that the client can display the target traffic object recognition result to the user.

[0486] Following the previous example, after obtaining a precise 3D proposal corresponding to the traffic object, the precise 3D proposal is sent to the client, allowing the client to display the precise 3D proposal to the user, thereby meeting the need for accurate identification of target objects in the target traffic scene in real-world scenarios.

[0487] The traffic data processing method provided in one or more embodiments of this specification first inputs the point cloud data to be processed corresponding to the target traffic scene into a data processing model. Then, using the data processing module in the data processing model, the point cloud data is processed to identify and obtain an initial traffic object identification result. Next, to avoid inaccurate initial traffic object identification results, a result adjustment module in the data processing model determines result adjustment parameters based on the point cloud data and the initial traffic object identification result. The initial traffic object identification result is then adjusted according to these parameters. This achieves accurate identification of the target traffic object corresponding to the point cloud data, avoiding the problem of inaccurate point cloud data processing results failing to meet the needs of point cloud data processing in practical application scenarios. Ultimately, this satisfies the need for accurate identification of target objects in the target traffic scene in practical scenarios.

[0488] The above is an illustrative scheme of a traffic data processing method according to this embodiment. It should be noted that the technical solution of this traffic data processing method belongs to the same concept as the above-described data processing method, point cloud data processing method, or data processing model training method. For details not described in detail in the technical solution of this traffic data processing method, please refer to the description of the above-described data processing method, point cloud data processing method, or data processing model training method.

[0489] Corresponding to the above method embodiments, this specification also provides an embodiment of a point cloud data processing device, which includes:

[0490] The data determination module is configured to determine the point cloud data to be processed corresponding to the target traffic scene and input the point cloud data to be processed into the data processing model, wherein the data processing model includes a data processing module and a result adjustment module;

[0491] The result acquisition module is configured to use the data processing module to identify and process the point cloud data to be processed, and obtain the initial traffic object identification result of the point cloud data to be processed.

[0492] The result adjustment module is configured to determine result adjustment parameters based on the point cloud data to be processed and the initial traffic object recognition result, and to perform result adjustment processing on the initial traffic object recognition result based on the result adjustment parameters to obtain the target traffic object recognition result of the point cloud data to be processed.

[0493] Optionally, the result adjustment module can also be configured as follows:

[0494] The point cloud data to be processed and the initial traffic object recognition result are input into the result adjustment module in the data processing model. The result adjustment module is used to determine the first data feature set and the second data feature set based on the point cloud data to be processed and the initial traffic object recognition result.

[0495] The first data features in the first data feature set are fused to obtain a first fused feature corresponding to the first data feature set; and the second data features in the second data feature set are fused to obtain a second fused feature corresponding to the second data feature set.

[0496] The first fusion feature and the second fusion feature are encoded using an attention mechanism to obtain a fusion feature encoding, and the fusion feature encoding is decoded to obtain the result adjustment parameters.

[0497] The traffic data processing device provided in one or more embodiments of this specification first inputs the point cloud data to be processed corresponding to the target traffic scene into a data processing model. Using the data processing module in the data processing model, the point cloud data to be processed is identified to obtain an initial traffic object identification result. Then, to avoid inaccurate initial traffic object identification results, the result adjustment module in the data processing model determines result adjustment parameters based on the point cloud data to be processed and the initial traffic object identification result. The initial traffic object identification result is then adjusted according to these parameters. This achieves accurate acquisition of the target traffic object identification result corresponding to the point cloud data to be processed, avoiding the problem that inaccurate point cloud data processing results would fail to meet the needs of point cloud data processing in practical application scenarios. Thus, it meets the need for accurate identification of target objects in the target traffic scene in practical scenarios.

[0498] The above is an illustrative scheme of a traffic data processing device according to this embodiment. It should be noted that the technical solution of this traffic data processing device and the technical solution of the point cloud data processing method described above belong to the same concept. For details not described in detail in the technical solution of this traffic data processing device, please refer to the description of the technical solution of the point cloud data processing method described above.

[0499] Figure 9 A structural block diagram of a computing device 900 according to one embodiment of this specification is shown. The components of the computing device 900 include, but are not limited to, a memory 910 and a processor 920. The processor 920 is connected to the memory 910 via a bus 930, and a database 950 is used to store data.

[0500] The computing device 900 also includes an access device 940, which enables the computing device 900 to communicate via one or more networks 960. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 940 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.

[0501] In one embodiment of this specification, the above-described components of the computing device 900 and Figure 9 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 9 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0502] The computing device 900 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 900 can also be a mobile or stationary server.

[0503] The processor 920 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above-mentioned point cloud data processing method, data processing method, another point cloud data processing method, machine learning model training method, data processing model training method, or traffic data processing method.

[0504] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the computing device embodiments are basically similar to embodiments of a point cloud data processing method, a data processing method, another point cloud data processing method, a machine learning model training method, a data processing model training method, or a traffic data processing method. Therefore, the description is relatively simple, and relevant parts can be referred to in the descriptions of embodiments of a point cloud data processing method, a data processing method, another point cloud data processing method, a machine learning model training method, a data processing model training method, or a traffic data processing method.

[0505] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the aforementioned point cloud data processing method, data processing method, another point cloud data processing method, machine learning model training method, data processing model training method, or traffic data processing method.

[0506] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computer-readable storage medium embodiments are relatively simple in description because they are substantially similar to embodiments of a point cloud data processing method, a data processing method, another point cloud data processing method, a machine learning model training method, a data processing model training method, or a traffic data processing method. Relevant parts can be referred to in the description of the embodiments of a point cloud data processing method, a data processing method, another point cloud data processing method, a machine learning model training method, a data processing model training method, or a traffic data processing method.

[0507] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described point cloud data processing method, data processing method, another point cloud data processing method, machine learning model training method, data processing model training method, or traffic data processing method.

[0508] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product belongs to the same concept as the technical solutions of the aforementioned point cloud data processing method, data processing method, another point cloud data processing method, machine learning model training method, data processing model training method, or traffic data processing method. Details not described in detail in the technical solution of the computer program product can be found in the descriptions of the aforementioned point cloud data processing method, data processing method, another point cloud data processing method, machine learning model training method, data processing model training method, or traffic data processing method.

[0509] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0510] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0511] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.

[0512] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0513] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A data processing method, comprising: The data to be processed is determined and input into the data processing model, wherein the data processing model includes a data processing module and a result adjustment module; The data processing module is used to perform recognition processing on the data to be processed to obtain the initial recognition result of the data to be processed. Using the result adjustment module, result adjustment parameters are determined based on the data to be processed and the initial recognition result, and the initial recognition result is adjusted according to the result adjustment parameters to obtain the target recognition result of the data to be processed.

2. The data processing method according to claim 1, wherein determining the result adjustment parameters using the result adjustment module based on the data to be processed and the initial recognition result includes: The data to be processed and the initial recognition result are input into the result adjustment module. The result adjustment module is used to determine the first data feature set and the second data feature set based on the data to be processed and the initial recognition result. The first data features in the first data feature set are fused to obtain the first fused feature corresponding to the first data feature set; and the second data features in the second data feature set are fused to obtain the second fused feature corresponding to the second data feature set. The first fusion feature and the second fusion feature are encoded using an attention mechanism to obtain a fusion feature encoding, and the fusion feature encoding is decoded to obtain the result adjustment parameters.

3. The data processing method according to claim 2, wherein the step of using the result adjustment module to determine the first data feature set and the second data feature set based on the data to be processed and the initial recognition result includes: Using the result adjustment module, the data to be processed is dimensionally transformed to determine the data feature map corresponding to the data to be processed; Based on the location information of the initial identification result, determine the result data corresponding to the initial identification result from the data to be processed, and based on the location information, determine the result feature corresponding to the initial identification result from the data feature map; Determine the target location information for the initial identification result, and based on the target location information, determine the target result data from the result data, and based on the target location information, determine the target result features from the result features; The result data and the result features are determined as the first data feature, and the first data feature set is determined based on the first data feature. The target result data and the target result features are determined as the second data feature, and the second data feature set is determined based on the second data feature.

4. The data processing method according to claim 2, wherein fusing the first data features in the first data feature set to obtain a first fused feature corresponding to the first data feature set, and fusing the second data features in the second data feature set to obtain a second fused feature corresponding to the second data feature set, comprises: Determine a first data feature in the first data feature set, and determine a second data feature in the second data feature set; The feature fusion unit in the result adjustment module is used to fuse the first data features to obtain the first fused feature corresponding to the first data feature set, and the feature fusion unit is used to fuse the second data features to obtain the second fused feature corresponding to the second data feature set.

5. The data processing method according to claim 4, wherein the step of fusing the first data features using the feature fusion unit in the result adjustment module to obtain the first fused feature corresponding to the first data feature set, and fusing the second data features using the feature fusion unit to obtain the second fused feature corresponding to the second data feature set, comprises: The feature fusion unit in the result adjustment module is used to determine the center point coordinate data of the initial recognition result; Using the feature fusion unit, the center point coordinate data is subtracted from the result data in the first data feature to obtain updated result data. The updated result data and the result features in the first data feature are then projected onto the first target dimension for feature fusion to obtain the first fused feature corresponding to the first data feature set. Using the feature fusion unit, the target result data in the second data feature is subtracted from the center point coordinate data to obtain updated target result data. The updated target result data and the target result features in the second data feature are then projected onto the first target dimension for feature fusion to obtain the second fused feature corresponding to the second data feature set.

6. The data processing method according to claim 2, wherein the step of encoding the first fusion feature and the second fusion feature using an attention mechanism to obtain fusion feature encoding includes: Project the first fusion feature and the second fusion feature onto the second target dimension to obtain the first target dimension feature corresponding to the first fusion feature and the second target dimension feature corresponding to the second fusion feature; Determine the similarity matrix between the first feature of the target dimension and the second feature of the target dimension; Using an attention mechanism, a first attention matrix for the first fused feature is determined based on the similarity matrix; and using an attention mechanism, a second attention matrix for the second fused feature is determined based on the similarity matrix. The fused feature encoding is obtained based on the first attention matrix and the second attention matrix.

7. The data processing method according to claim 6, wherein obtaining the fused feature encoding based on the first attention matrix and the second attention matrix comprises: A first updated feature corresponding to the first fusion feature is determined, and a second updated feature corresponding to the second fusion feature is determined, wherein the first updated feature is obtained by projecting the first fusion feature onto a third target dimension, and the second updated feature is obtained by projecting the second fusion feature onto the third target dimension; The first attention matrix is ​​updated using the first matrix update strategy, based on the first update feature and the second update feature, to obtain the first updated attention matrix; Using the second matrix update strategy, the second attention matrix is ​​updated based on the first update feature and the second update feature to obtain the second updated attention matrix; The first updated attention matrix and the second updated attention matrix are encoded to obtain the fused feature encoding.

8. The data processing method according to claim 1, wherein determining the data to be processed includes: Receive the data to be processed sent by the client, wherein the data to be processed is sent by the user through the data processing interface of the client; After adjusting the initial recognition result according to the result adjustment parameters to obtain the target recognition result of the data to be processed, the method further includes: The target recognition result is sent to the client so that the client can display the target recognition result to the user based on the data processing interface.

9. A point cloud data processing method, comprising: The point cloud data to be processed is determined and input into the data processing model, wherein the data processing model includes a data processing module and a result adjustment module; The data processing module is used to perform recognition processing on the point cloud data to be processed, and the initial object recognition result of the point cloud data to be processed is obtained. Using the result adjustment module, result adjustment parameters are determined based on the point cloud data to be processed and the initial object recognition result. The initial object recognition result is then adjusted according to the result adjustment parameters to obtain the target object recognition result of the point cloud data to be processed.

10. The point cloud data processing method according to claim 9, wherein determining the result adjustment parameters using the result adjustment module based on the point cloud data to be processed and the initial object recognition result includes: The point cloud data to be processed and the initial object recognition result are input into the result adjustment module in the data processing model. The result adjustment module is used to determine the first data feature set and the second data feature set based on the point cloud data to be processed and the initial object recognition result. The first data features in the first data feature set are fused to obtain the first fused feature corresponding to the first data feature set; and the second data features in the second data feature set are fused to obtain the second fused feature corresponding to the second data feature set. The first fusion feature and the second fusion feature are encoded using an attention mechanism to obtain a fusion feature encoding, and the fusion feature encoding is decoded to obtain the result adjustment parameters.

11. The point cloud data processing method according to claim 10, wherein the step of using the result adjustment module to determine the first data feature set and the second data feature set based on the point cloud data to be processed and the initial object recognition result includes: Using the result adjustment module, the point cloud data to be processed is dimensionally transformed to determine the point cloud feature map corresponding to the point cloud data to be processed. Based on the object location information of the initial object recognition result, object point cloud data is determined from the point cloud data to be processed, and object point cloud features are determined from the point cloud feature map based on the object location information. Determine the target point cloud location information for the initial object recognition result, and based on the target point cloud location information, determine the target point cloud data from the object point cloud data, and based on the target point cloud location information, determine the target point cloud features from the object point cloud features; The object point cloud data and the object point cloud features are determined as first data features, and a first data feature set is determined based on the first data features. The target point cloud data and the target point cloud features are determined as second data features, and a second data feature set is determined based on the second data features.

12. The point cloud data processing method according to claim 10, wherein fusing the first data features in the first data feature set to obtain a first fused feature corresponding to the first data feature set, and fusing the second data features in the second data feature set to obtain a second fused feature corresponding to the second data feature set, comprises: Determine a first data feature in the first data feature set, and determine a second data feature in the second data feature set; The feature fusion unit in the result adjustment module is used to fuse the first data features to obtain the first fused feature corresponding to the first data feature set, and the feature fusion unit is used to fuse the second data features to obtain the second fused feature corresponding to the second data feature set.

13. The point cloud data processing method according to claim 12, wherein the step of fusing the first data features using the feature fusion unit in the result adjustment module to obtain the first fused feature corresponding to the first data feature set, and fusing the second data features using the feature fusion unit to obtain the second fused feature corresponding to the second data feature set, comprises: Using the feature fusion unit, the center point coordinate data of the initial object recognition result is determined; Subtract the center point coordinate data from the object point cloud data in the first data feature to obtain updated object point cloud data. Project the updated object point cloud data and the object point cloud features in the first data feature set onto the first target dimension for feature fusion to obtain the first fused feature corresponding to the first data feature set. Subtract the center point coordinate data from the target point cloud data in the second data feature set to obtain updated target point cloud data. Project the updated target point cloud data and the target point cloud features in the second data feature set onto the first target dimension for feature fusion to obtain the second fused feature corresponding to the second data feature set.

14. The point cloud data processing method according to claim 10, wherein the step of encoding the first fusion feature and the second fusion feature using an attention mechanism to obtain fusion feature encoding includes: Project the first fusion feature and the second fusion feature onto the second target dimension to obtain the first target dimension feature corresponding to the first fusion feature and the second target dimension feature corresponding to the second fusion feature; Determine the similarity matrix between the first feature of the target dimension and the second feature of the target dimension; Using an attention mechanism, a first attention matrix for the first fused feature is determined based on the similarity matrix; and using an attention mechanism, a second attention matrix for the second fused feature is determined based on the similarity matrix. The fused feature encoding is obtained based on the first attention matrix and the second attention matrix.

15. The point cloud data processing method according to claim 14, wherein obtaining the fused feature encoding based on the first attention matrix and the second attention matrix comprises: A first updated feature corresponding to the first fusion feature is determined, and a second updated feature corresponding to the second fusion feature is determined, wherein the first updated feature is obtained by projecting the first fusion feature onto a third target dimension, and the second updated feature is obtained by projecting the second fusion feature onto the third target dimension; The first attention matrix is ​​updated using the first matrix update strategy, based on the first update feature and the second update feature, to obtain the first updated attention matrix; Using the second matrix update strategy, the second attention matrix is ​​updated based on the first update feature and the second update feature to obtain the second updated attention matrix; The first updated attention matrix and the second updated attention matrix are encoded to obtain the fused feature encoding.

16. The point cloud data processing method according to claim 9, wherein the initial object recognition result is an initial object detection box; and the result adjustment parameter is a correction value for the initial object detection box; The step of adjusting the parameters based on the results to perform result adjustment processing on the initial object recognition result, and obtaining the target object recognition result of the point cloud data to be processed, includes: The initial object detection box is corrected according to the correction value to obtain the target object detection box corresponding to the object to be identified in the point cloud data to be processed.

17. The point cloud data processing method according to claim 9, wherein determining the point cloud data to be processed includes: The device receives point cloud data to be processed sent by a point cloud data acquisition device, wherein the point cloud data acquisition device is a device configured by the user through a client for performing point cloud data acquisition operations for a target scene, and the point cloud data to be processed is obtained by the point cloud data acquisition device performing point cloud data acquisition operations for the target scene. After adjusting the parameters based on the results to process the initial object recognition result and obtain the target object recognition result of the point cloud data to be processed, the process further includes: The client corresponding to the point cloud data acquisition device is determined, and the target object recognition result is sent to the client so that the client can display the target object recognition result to the user.

18. A point cloud data processing method, comprising: Identify the point cloud data to be processed; The point cloud data to be processed is input into a data processing model to obtain the target object recognition result of the point cloud data to be processed. The data processing model includes a data processing module and a result adjustment module. The data processing module is used to perform recognition processing on the point cloud data to be processed to obtain an initial object recognition result of the point cloud data to be processed. The result adjustment module is used to determine a result adjustment parameter based on the point cloud data to be processed and the initial object recognition result, and to perform result adjustment processing on the initial object recognition result based on the result adjustment parameter to obtain the target object recognition result of the point cloud data to be processed.

19. A method for training a machine learning model, comprising: The machine learning model to be trained and the training data associated with the target task are determined, wherein the machine learning model to be trained includes a data processing module and a result adjustment module. The data processing module is used to perform recognition processing on the training data to obtain the initial recognition result of the training data; Using the result adjustment module, result adjustment parameters are determined based on the training data and the initial recognition result, and the initial recognition result is adjusted according to the result adjustment parameters to obtain the target recognition result of the training data; Based on the target recognition result, the model parameters of the machine learning model are adjusted to obtain a trained target model, wherein the target model is used to perform the target task.

20. A traffic data processing method, comprising: The point cloud data to be processed corresponding to the target traffic scene is determined, and the point cloud data to be processed is input into the data processing model, wherein the data processing model includes a data processing module and a result adjustment module; The data processing module is used to identify and process the point cloud data to be processed, and the initial traffic object identification result of the point cloud data to be processed is obtained. Using the result adjustment module, result adjustment parameters are determined based on the point cloud data to be processed and the initial traffic object recognition result. The initial traffic object recognition result is then adjusted according to the result adjustment parameters to obtain the target traffic object recognition result of the point cloud data to be processed.

21. The traffic data processing method according to claim 20, wherein determining the result adjustment parameters using the result adjustment module based on the point cloud data to be processed and the initial traffic object identification result includes: The point cloud data to be processed and the initial traffic object recognition result are input into the result adjustment module in the data processing model. The result adjustment module is used to determine the first data feature set and the second data feature set based on the point cloud data to be processed and the initial traffic object recognition result. The first data features in the first data feature set are fused to obtain the first fused feature corresponding to the first data feature set; and the second data features in the second data feature set are fused to obtain the second fused feature corresponding to the second data feature set. The first fusion feature and the second fusion feature are encoded using an attention mechanism to obtain a fusion feature encoding, and the fusion feature encoding is decoded to obtain the result adjustment parameters.

22. A computing device, comprising: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, they implement the steps of the data processing method according to any one of claims 1 to 8, the point cloud data processing method according to any one of claims 9 to 17, the point cloud data processing method according to claim 18, the machine learning model training method according to claim 19, or the traffic data processing method according to any one of claims 20 to 21.

23. A computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the data processing method of any one of claims 1 to 8, the point cloud data processing method of any one of claims 9 to 17, the point cloud data processing method of claim 18, the machine learning model training method of claim 19, or the traffic data processing method of any one of claims 20 to 21.

24. A computer program product comprising a computer program / instructions which, when executed by a processor, implement the steps of the data processing method of any one of claims 1 to 8, the point cloud data processing method of any one of claims 9 to 17, the point cloud data processing method of claim 18, the machine learning model training method of claim 19, or the traffic data processing method of any one of claims 20 to 21.