A Point Cloud Processing Method and System Based on a Mixture of Sparse Experts

By converting the original point cloud into high-dimensional features and using the backbone network for multiple feature extraction, the problem of computational complexity and information loss in point cloud processing of sparse MoE architecture is solved, and more efficient computing and more accurate point cloud prediction are achieved.

CN117150266BActive Publication Date: 2025-06-10SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311084717.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-25
Publication Date
2025-06-10
Estimated Expiration
2043-08-25

AI Technical Summary

Technical Problem

The existing sparse hybrid expert model (MoE) has problems with computational complexity and information loss in point cloud processing, especially in the case of sparse data and high-dimensionality, which leads to poor processing results.

Method used

The original point cloud is mapped into high-dimensional features through a point cloud word participle, and a backbone network, including Transformer Block, MoE Block and MoMSE Block subnetwork, perform multiple feature extractions on high-dimensional features to obtain the target point cloud features. Decode the target point cloud features based on the downstream task header of the downstream task to obtain the prediction information of the corresponding downstream task.

Benefits of technology

It improves computing efficiency, enhances the model's perception and understanding ability of point clouds, avoids the loss of key point cloud local information, and thus improves the prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117150266B_ABST
    Figure CN117150266B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a point cloud processing method and system based on a sparse mixture of experts model. The original point cloud is mapped into high-dimensional features through a point cloud tokenizer; the high-dimensional features are input into a backbone network; the backbone network performs multiple feature extractions on the high-dimensional features to obtain target point cloud features; and the target point cloud features are decoded according to a downstream task head of a downstream task to obtain prediction information corresponding to the downstream task. The solution of the present invention has higher computational efficiency in models with the same capacity, avoids the loss of key local point cloud information, and thus improves the prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of computer vision technology, and in particular, to a point cloud processing method and system based on a sparse mixture of experts model. Background Art

[0002] A point cloud is a data set composed of a large number of three-dimensional point coordinates, which is used to represent the spatial geometric structure of an object or a scene. By processing and analyzing the point cloud, applications of the point cloud in fields such as 3D modeling, scene understanding, non-contact measurement, and medical image processing can be realized. The traditional way to process the point cloud is to use a dense network structure to convert the point cloud data into an image or voxel representation form, and then apply models such as a convolutional neural network (CNN) for processing. However, using the above traditional dense network structure will significantly increase the computational amount when increasing the model capacity, and will reduce the efficiency of the general model and increase the computational cost.

[0003] Therefore, currently in the field of point clouds, some deep learning models specifically for point cloud data have emerged. For example, the sparse mixture of experts model (Mixture of Experts, MoE) can better utilize the local and global features of point cloud data by introducing an expert network and a gating network. However, when the existing sparse MoE architecture is applied to the field of point clouds, the following problems exist:

[0004] (1) In order to reduce the computational complexity and storage requirements, the sparse MoE architecture usually needs to downsample the point cloud to reduce the number of points. However, downsampling may cause the loss of key local information of the point cloud.

[0005] (2) Due to problems such as data sparsity, high dimensionality, and insufficient training samples of point cloud data, it is difficult to effectively train the MoE network in the field of point clouds with a small amount of data, resulting in poor point cloud processing effects. Summary of the Invention

[0006] In view of this, the embodiments of the present invention provide a point cloud processing method and system, an electronic device, and a computer storage medium based on a sparse mixture of experts model to at least partially solve the above problems.

[0007] According to the first aspect of the embodiments of the present invention, a point cloud processing method based on a sparse mixture of experts model is provided, including mapping an original point cloud into high-dimensional features through a point cloud tokenizer; inputting the high-dimensional features into a backbone network; performing multiple feature extractions on the high-dimensional features through the backbone network to obtain target point cloud features; and decoding the target point cloud features according to a downstream task head of a downstream task to obtain prediction information corresponding to the downstream task.

[0008] In one implementation, the original point cloud is mapped to high-dimensional features through a point cloud tokenizer, including sampling N points in the original point cloud through the farthest point sampling algorithm; aggregating neighboring points with each of the N points as the center to generate N groups of point clouds; and mapping the point cloud information of each group of point clouds in the N groups of point clouds to the high-dimensional feature space through a multi-layer perceptron network to obtain the high-dimensional features of the point cloud information of each group of point clouds.

[0009] In another implementation, the backbone network includes two Transformer Block sub-networks, one MoEBlock sub-network, and one MoMSE Block sub-network.

[0010] In another implementation, the high-dimensional features are subjected to multiple feature extractions through the backbone network to obtain the target point cloud features, including the following steps:

[0011] Step a:

[0012] (1) Based on the first multi-head self-attention module and the first feed-forward neural network module of the two Transformer Block sub-networks, the high-dimensional features are updated in two layers to obtain the updated high-dimensional features;

[0013] (2) Through the second multi-head self-attention module in a MoE Block sub-network, the updated high-dimensional features are updated again to obtain the further updated high-dimensional features;

[0014] (3) Through the distribution module in a MoE Block sub-network, the further updated high-dimensional features are evenly grouped and weighted to obtain n groups of further updated high-dimensional features and the weight values of each further updated high-dimensional feature in the n groups of further updated high-dimensional features;

[0015] (4) Through the second feed-forward neural network module in a MoE Block sub-network, the n groups of further updated high-dimensional features are updated again to obtain n groups of output high-dimensional features output by a MoE Block sub-network, where, for each group of further updated high-dimensional features in the n groups of further updated high-dimensional features, the architectures of the second feed-forward neural network modules are the same and the parameters are different.

[0016] Step b:

[0017] (1) Input the n groups of output high-dimensional features obtained in step a into a MoMSE Block sub-network respectively;

[0018] (2) Update the n groups of output high-dimensional features through the third multi-head self-attention module and the third feed-forward neural network module in a MoMSE Block sub-network in sequence to obtain the updated n groups of output high-dimensional features. Among them, for the n groups of output high-dimensional features, the architectures of the third multi-head self-attention module and the third feed-forward neural network module are the same, but the parameters are different;

[0019] (3) Based on the weight values obtained in step a, perform weighted summation on all high-dimensional features in the feature extraction process, and keep each high-dimensional feature at its initial position in the feature extraction process.

[0020] Step c: Repeat steps a and b three times to obtain the target point cloud features.

[0021] According to the second aspect of the embodiments of the present invention, there is provided a point cloud processing system based on a mixture of sparse experts model, including a mapping module for mapping the original point cloud into high-dimensional features through a point cloud tokenizer; an input module for inputting the high-dimensional features into the backbone network; an extraction module for performing multiple feature extractions on the high-dimensional features through the backbone network to obtain target point cloud features; and a decoding module for decoding the target point cloud features according to the downstream task head of the downstream task to obtain prediction information corresponding to the downstream task.

[0022] According to the third aspect of the embodiments of the present invention, there is provided an electronic device including a processor and a memory storing a program. Among them, the program includes instructions that, when executed by the processor, cause the processor to execute the steps of the method in the first aspect.

[0023] According to the fourth aspect of the embodiments of the present invention, there is provided a computer storage medium having a computer program stored thereon, and the program, when executed by a processor, implements the method in the first aspect.

[0024] In summary, in the solution of the embodiments of the present invention, the original point cloud is mapped into high-dimensional features through a point cloud tokenizer, the high-dimensional features are subjected to multiple feature extractions through the backbone network to obtain target point cloud features, and the target point cloud features are decoded according to the downstream task head of the downstream task to obtain prediction information corresponding to the downstream task. The solution of the present invention has higher computational efficiency in models with the same capacity, the trained model has stronger perception and understanding capabilities for point clouds, avoids the loss of key local point cloud information, and thus improves the prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0026] Figure 1 It is a step flowchart of a point cloud processing method based on a sparse mixture of experts model according to an embodiment of the present invention.

[0027] Figure 2 It is an overall framework diagram of a point cloud processing method based on a sparse mixture of experts model according to another embodiment of the present invention.

[0028] Figure 3 It is a structural block diagram of a point cloud processing system based on a sparse mixture of experts model according to another embodiment of the present invention.

[0029] Figure 4 It is a structural schematic diagram of an electronic device according to another embodiment of the present invention. Detailed implementation manners

[0030] To have a clearer understanding of the technical features, objectives, and effects of the embodiments of the present application, the following will describe the specific implementation manners of the embodiments of the present application with reference to the accompanying drawings.

[0031] In this article, "exemplarily" means "serving as an instance, example, or illustration". Any illustration or implementation manner described as "schematic" in this article should not be interpreted as a more preferred or more advantageous technical solution.

[0032] To make the drawings concise, only the parts related to the present application are schematically shown in each drawing, and they do not represent their actual structures as products. Additionally, to make the drawings concise and easy to understand, in some drawings, components with the same structure or function are only schematically shown for one or more of them, or only one or more of them are labeled.

[0033] In addition, the descriptions involving "first", "second", etc. in the present invention are only for descriptive purposes and should not be understood as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features.

[0034] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art shall fall within the scope of protection of the embodiments of the present invention.

[0035] The following further illustrates the specific implementation of the embodiments of the present invention in conjunction with the accompanying drawings of the embodiments of the present invention.

[0036] See Figure 1 which is a step flowchart of a point cloud processing method based on a sparse mixture of experts model according to an embodiment of the present invention, mainly including:

[0037] Step S110, mapping the original point cloud into high-dimensional features through a point cloud tokenizer.

[0038] It should be understood that in point cloud processing, the original point cloud is usually composed of a large number of discrete points, and each point contains position information and other attributes. The role of the point cloud tokenizer is to divide the original point cloud into multiple groups, aggregate adjacent points together to form groups, so as to reduce the complexity of the point cloud data and extract the features of some local regions.

[0039] Step S120, inputting the high-dimensional features into the backbone network.

[0040] Step S130, performing multiple feature extractions on the high-dimensional features through the backbone network to obtain the target point cloud features.

[0041] By performing multiple iterative calculations and optimizations on the high-dimensional features through the backbone network, more representative and discriminative point cloud features can be extracted, providing better inputs for subsequent downstream tasks.

[0042] Step S140, decoding the target point cloud features according to the downstream task head of the downstream task to obtain the prediction information corresponding to the downstream task.

[0043] The downstream task head is a module designed according to specific application requirements, such as tasks like classification, segmentation, object detection, or registration. This step uses the already optimized point cloud features as inputs and maps the target point cloud features into the prediction information required for the downstream task through the corresponding decoding network structure.

[0044] In summary, in the solution of the embodiment of the present invention, through the cooperation of steps such as converting the original point cloud into high-dimensional features, performing multiple feature extractions and optimizations, and decoding and predicting according to specific downstream tasks, efficient processing of point cloud data and more accurate task completion can be achieved, making the solution of the present invention more efficient in computing in models with the same capacity, the trained model has stronger perception and understanding capabilities for point clouds, avoiding the loss of key local information of point clouds, and having a higher prediction accuracy.

[0045] In one implementation, the original point cloud is mapped to high-dimensional features through a point cloud tokenizer, including sampling N points in the original point cloud through the farthest point sampling algorithm; aggregating adjacent points with each of the N points as the center to generate N groups of point clouds; and mapping the point cloud information of each group of point clouds in the N groups of point clouds to the high-dimensional feature space through a multi-layer perceptron network to obtain the high-dimensional features of the point cloud information of each group of point clouds.

[0046] In another implementation, the backbone network includes two Transformer Block sub-networks, one MoEBlock sub-network, and one MoMSE Block sub-network.

[0047] In another implementation, see Figure 2 , the high-dimensional features are repeatedly extracted through the backbone network to obtain the target point cloud features, including the following steps:

[0048] Step a:

[0049] (1) Based on the first multi-head self-attention module and the first feed-forward neural network module of the two Transformer Block sub-networks, the high-dimensional features are updated in two layers to obtain the updated high-dimensional features;

[0050] (2) Through the second multi-head self-attention module in a MoE Block sub-network, the updated high-dimensional features are updated again to obtain the further updated high-dimensional features;

[0051] (3) Through the distribution module in a MoE Block sub-network, the further updated high-dimensional features are evenly grouped and weighted to obtain n groups of further updated high-dimensional features and the weight values of each further updated high-dimensional feature in the n groups of further updated high-dimensional features;

[0052] (4) The n groups of further updated high-dimensional features are updated again through the second feed-forward neural network module in a MoE Block sub-network to obtain n groups of output high-dimensional features output by a MoE Block sub-network, where, for each group of further updated high-dimensional features in the n groups of further updated high-dimensional features, the architectures of the second feed-forward neural network modules are the same and the parameters are different.

[0053] Step b:

[0054] (1) Input the n groups of output high-dimensional features obtained in step a into a MoMSE Block sub-network respectively;

[0055] (2) Update the n groups of output high-dimensional features in sequence through the third multi-head self-attention module and the third feed-forward neural network module in a MoMSE Block sub-network to obtain the updated n groups of output high-dimensional features. Among them, for the n groups of output high-dimensional features, the architectures of the third multi-head self-attention module and the third feed-forward neural network module are the same, but the parameters are different;

[0056] (3) Based on the weight values obtained in step a, perform weighted summation on all high-dimensional features in the feature extraction process, and keep each high-dimensional feature at its initial position in the feature extraction process.

[0057] Step c: Repeat steps a and b three times to obtain the target point cloud features.

[0058] Exemplarily, in combination with Figure 1 and Figure 2 describe the implementation steps of the method of the present invention in detail. For a given original point cloud P and a downstream task T, the implementation steps include:

[0059] 1. First, map the point cloud to high-dimensional features through a point cloud tokenizer (Point Tokenizer):

[0060] Use the farthest point sampling algorithm (Farthest Point Sampling, FPS) to sample N points in the original point cloud P, and aggregate its neighboring points with each of these points as the center to form N groups, that is, generate N groups of point clouds;

[0061] Use a deep learning-based multi-layer perceptron network (Multi-layer perceptron, MLP) to map the point cloud information of each group to the high-dimensional feature space to obtain the high-dimensional features of the point cloud information of each group.

[0062] 2. Then input the high-dimensional features into the backbone network, and perform multiple feature extractions on the high-dimensional features through the backbone network to obtain the target point cloud features:

[0063] It should be noted that the backbone network includes two Transformer Block sub-networks, one MoE Block sub-network and one MoMSE Block sub-network;

[0064] After the high-dimensional features input into the backbone network are updated through two layers of standard Transformer Blocks, they are input into the MoE module. The MoE module includes a MoE Block sub-network and a MoMSE Block sub-network;

[0065] The update of the two layers of standard Transformer Blocks is specifically as follows: Based on the first multi-head self-attention module MHSA 1 and the first feed-forward neural network module FFN 1 of the two Transformer Block sub-networks, the high-dimensional features are updated through two layers to obtain updated high-dimensional features.

[0066] In the MoE module:

[0067] Through the second multi-head self-attention module MHSA 2 in a MoE Block sub-network, the updated high-dimensional features are updated again to obtain further updated high-dimensional features;

[0068] Through the distribution module Router in a MoE Block sub-network, the further updated high-dimensional features are evenly grouped and weighted to obtain n groups of further updated high-dimensional features and the weight values of each further updated high-dimensional feature in the n groups of further updated high-dimensional features;

[0069] Through the second feed-forward neural network module FFN 2 in a MoE Block sub-network, the n groups of further updated high-dimensional features are updated again to obtain n groups of output high-dimensional features output by a MoE Block sub-network. Among them, for each group of further updated high-dimensional features in the n groups of further updated high-dimensional features, the architectures of the second feed-forward neural network modules are the same, but the parameters are different;

[0070] The n groups of output high-dimensional features obtained in the previous step are respectively input into a MoMSE Block sub-network. Through the third multi-head self-attention module MHSA 3 and the third feed-forward neural network module FFN 3 in a MoMSE Block sub-network, the n groups of output high-dimensional features are updated in sequence to obtain n groups of updated output high-dimensional features. Among them, for the n groups of output high-dimensional features, the architectures of the third multi-head self-attention module MHSA 3 and the third feed-forward neural network module FFN 3 are the same, but the parameters are different;

[0071] Based on the weight values obtained from the foregoing steps, perform a weighted sum of all the high-dimensional features in the feature extraction process, and keep each high-dimensional feature at its initial position in the feature extraction process, that is, each high-dimensional feature after processing is placed back at its original position in Figure 2 its original position;

[0072] Repeat the foregoing steps three times to obtain the fully optimized point cloud features, that is: the target point cloud features.

[0073] 3. After obtaining the target point cloud features, the target point cloud features can be decoded into the prediction information corresponding to the downstream task by using the downstream task heads of different downstream tasks.

[0074] In summary, in the solution of the embodiment of the present invention, through the synergistic effect of converting the original point cloud into high-dimensional features, performing multiple feature extractions and optimizations, and decoding and predicting according to specific downstream tasks, efficient processing of point cloud data and more accurate task completion can be achieved, making the solution of the present invention more efficient in calculation in models with the same capacity, enabling the trained model to have stronger perception and understanding capabilities for point clouds, avoiding the loss of key local information of point clouds, and having higher prediction accuracy.

[0075] See Figure 3 which is the structural block diagram of a point cloud processing system based on a sparse mixture of experts model according to another embodiment of the present invention, including:

[0076] A mapping module 310, configured to map the original point cloud into high-dimensional features through a point cloud tokenizer.

[0077] An input module 320, configured to input the high-dimensional features into the backbone network.

[0078] An extraction module 330, configured to perform multiple feature extractions on the high-dimensional features through the backbone network to obtain the target point cloud features.

[0079] A decoding module 340, configured to decode the target point cloud features according to the downstream task head of the downstream task to obtain the prediction information corresponding to the downstream task.

[0080] In summary, in the solution of the embodiment of the present invention, through the synergistic effect of converting the original point cloud into high-dimensional features, performing multiple feature extractions and optimizations, and decoding and predicting according to specific downstream tasks, efficient processing of point cloud data and more accurate task completion can be achieved, making the solution of the present invention more efficient in calculation in models with the same capacity, enabling the trained model to have stronger perception and understanding capabilities for point clouds, avoiding the loss of key local information of point clouds, and having higher prediction accuracy.

[0081] The system of this embodiment is used to implement the corresponding methods in the foregoing multiple method embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here. In addition, the implementation of the functions of each module in the system of this embodiment can refer to the description of the corresponding parts in the foregoing method embodiments, which will not be elaborated here either.

[0082] According to another aspect of the embodiments of the present invention, an electronic device is provided. Refer to Figure 4 , and now the structural block diagram of the electronic device 400 that can be used as the server or client of the present application will be described. It is an example of a hardware device that can be applied to various aspects of the present application. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0083] The electronic device 400 may include: a processor 402, a communications interface 404, a memory 406, and a communication bus 408.

[0084] The processor 402, the communications interface 404, and the memory 406 communicate with each other through the communication bus 408. The communications interface 404 is used to communicate with other electronic devices or servers.

[0085] The processor 402 is used to execute the program 410, and specifically can execute the relevant steps in the foregoing method embodiments.

[0086] Specifically, the program 410 may include program code, and the program code includes computer operation instructions.

[0087] The processor 402 may be a processor CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0088] A memory 406 for storing a program 410. The memory 406 may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.

[0089] Specifically, the program 410 may be used to cause the processor 402 to perform the following operations: map the original point cloud to high-dimensional features through a point cloud tokenizer; input the high-dimensional features into a backbone network; perform multiple feature extractions on the high-dimensional features through the backbone network to obtain target point cloud features; decode the target point cloud features according to a downstream task head of a downstream task to obtain prediction information corresponding to the downstream task.

[0090] In addition, for the specific implementation of each step in the program 410, reference may be made to the corresponding steps and descriptions in the corresponding units in the foregoing method embodiments, which will not be elaborated herein. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the devices and modules described above may refer to the corresponding process descriptions in the foregoing method embodiments, which will not be elaborated herein.

[0091] It should be noted that according to the needs of implementation, each component / step described in the embodiments of the present invention may be split into more components / steps, or two or more components / steps or partial operations of components / steps may be combined into new components / steps to achieve the purpose of the embodiments of the present invention.

[0092] An exemplary embodiment of the present invention further provides a computer storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the methods of the embodiments of the present invention.

[0093] The methods according to the embodiments of the present invention described above may be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and to be stored in a local recording medium through network download, so that the methods described herein can be stored in such software processes on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods described herein are implemented. In addition, when a general-purpose computer accesses the code for implementing the methods shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for implementing the methods shown herein.

[0094] It should be understood that although this specification is described according to various embodiments, not every embodiment contains only an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

[0095] The above embodiments are only used to illustrate the embodiments of the present application, rather than to limit the embodiments of the present application. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present application. The patent protection scope of the embodiments of the present application should be defined by the claims.

Claims

1. A point cloud processing method based on a sparse mixture of experts model, characterized in that, it includes: Mapping the original point cloud into high-dimensional features through a point cloud tokenizer; Inputting the high-dimensional features into a backbone network; Performing multiple feature extractions on the high-dimensional features through the backbone network to obtain target point cloud features, including the following steps: Step a: (1) Based on the first multi-head self-attention module and the first feed-forward neural network module of two Transformer Block sub-networks, performing two-layer updates on the high-dimensional features to obtain updated high-dimensional features; (2) Through the second multi-head self-attention module in a MoE Block sub-network, performing another update on the updated high-dimensional features to obtain further updated high-dimensional features; (3) Through a distribution module in a MoE Block sub-network, uniformly grouping and calculating weights for the further updated high-dimensional features to obtain n groups of further updated high-dimensional features and the weight values of each further updated high-dimensional feature in the n groups of further updated high-dimensional features; (4) Performing another update on the n groups of further updated high-dimensional features through the second feed-forward neural network module in a MoE Block sub-network to obtain n groups of output high-dimensional features output by a MoE Block sub-network, where, for each group of further updated high-dimensional features in the n groups of further updated high-dimensional features, the architectures of the second feed-forward neural network modules are the same but the parameters are different; Step b: (1) Inputting the n groups of output high-dimensional features obtained in step a into a MoMSE Block sub-network respectively; (2) Sequentially updating the n groups of output high-dimensional features through the third multi-head self-attention module and the third feed-forward neural network module in a MoMSE Block sub-network to obtain updated n groups of output high-dimensional features, where, for the n groups of output high-dimensional features, the architectures of the third multi-head self-attention module and the third feed-forward neural network module are the same but the parameters are different; (3) Based on the weight values obtained in step a, performing weighted summation on all the high-dimensional features during the feature extraction process and keeping each high-dimensional feature at its initial position during the feature extraction process; Step c: Repeating steps a and b three times to obtain target point cloud features; Decoding the target point cloud features according to the downstream task head of the downstream task to obtain prediction information corresponding to the downstream task.

2. The method according to claim 1, characterized in that, the mapping of the original point cloud into high-dimensional features through a point cloud tokenizer includes: Sampling N points in the original point cloud through the farthest point sampling algorithm; Aggregating adjacent points with each of the N points as the center to generate N groups of point clouds; Mapping the point cloud information of each group of point clouds in the N groups of point clouds into the high-dimensional feature space through a multi-layer perceptron network to obtain the high-dimensional features of the point cloud information of each group of point clouds.

3. The method according to claim 2, characterized in that, The backbone network includes two TransformerBlock sub-networks, one MoE Block sub-network, and one MoMSE Block sub-network.

4. A point cloud processing system based on a sparse mixture of experts model, characterized in that, it includes: A mapping module for mapping the original point cloud into high-dimensional features through a point cloud tokenizer; An input module for inputting the high-dimensional features into the backbone network; An extraction module for performing multiple feature extractions on the high-dimensional features through the backbone network to obtain target point cloud features, including the following steps: Step a: (1) Based on the first multi-head self-attention module and the first feed-forward neural network module of two Transformer Block sub-networks, update the high-dimensional features twice to obtain updated high-dimensional features; (2) Through the second multi-head self-attention module in a MoE Block sub-network, update the updated high-dimensional features again to obtain further updated high-dimensional features; (3) Through the distribution module in a MoE Block sub-network, evenly group and calculate the weights of the further updated high-dimensional features to obtain n groups of further updated high-dimensional features and the weight values of each further updated high-dimensional feature in the n groups of further updated high-dimensional features; (4) Through the second feed-forward neural network module in a MoE Block sub-network, update the n groups of further updated high-dimensional features again to obtain n groups of output high-dimensional features output by a MoE Block sub-network. Among them, for each group of further updated high-dimensional features in the n groups of further updated high-dimensional features, the architectures of the second feed-forward neural network module are the same, but the parameters are different; Step b: (1) Input the n groups of output high-dimensional features obtained in step a into a MoMSE Block sub-network respectively; (2) Update the n groups of output high-dimensional features in sequence through the third multi-head self-attention module and the third feed-forward neural network module in a MoMSE Block sub-network to obtain updated n groups of output high-dimensional features. Among them, for the n groups of output high-dimensional features, the architectures of the third multi-head self-attention module and the third feed-forward neural network module are the same, but the parameters are different; (3) Based on the weight values obtained in step a, perform weighted summation on all high-dimensional features in the feature extraction process, and keep each high-dimensional feature at its initial position in the feature extraction process; Step c: Repeat steps a and b three times to obtain target point cloud features; A decoding module for decoding the target point cloud features according to the downstream task head of the downstream task to obtain prediction information corresponding to the downstream task.

5. An electronic device, characterized in that, it includes: A processor; A memory for storing programs; Among them, the program includes instructions, and when the instructions are executed by the processor, the processor executes the steps of the method according to any one of claims 1-3.

6. A computer storage medium, characterized in that, A computer program is stored thereon, and when the program is executed by a processor, the method described in any one of claims 1 to 3 is implemented.

Citation Information

Patent Citations

  • Method and system for using 2D pre-training model as 3D downstream task backbone network

    CN115719443A

  • Voxel-based feature learning network

    US10970518B1