Semantic segmentation model training method and its device, computing equipment and medium

Through cross-modal processing and training of semantic segmentation models, the problem of semantic segmentation performance degradation caused by domain gap or domain offset is solved, and better domain adaptation and semantic segmentation performance are achieved.

CN115346085BActive Publication Date: 2025-06-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210977064.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-15
Publication Date
2025-06-06
Estimated Expiration
2042-08-15

AI Technical Summary

Technical Problem

In the case of domain gaps or domain offsets in the prior art, it is difficult to achieve efficient semantic segmentation.

Method used

By obtaining the training sample set of the source domain and the target domain, the features of the first and second modalities are used to generate prediction results, imitation results, and recovery results, and cross-modal processing is performed to train the semantic segmentation model.

Benefits of technology

The domain adaptation effect of the model is improved, and better semantic segmentation performance can be achieved when the domain gap or domain offset is large.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115346085B_ABST
    Figure CN115346085B_ABST
Patent Text Reader

Abstract

A semantic segmentation model training method and its apparatus, computing device and medium are provided. The training method includes: obtaining a training sample set including a source domain sample set and a target domain sample set; for each sample in the training sample set: generating a first modality feature based on the first modality data, and generating a second modality feature based on the second modality data, wherein each modality feature includes features of multiple points associated with the positions and quantities of points included in the second modality data; and generating prediction results, imitation results and restoration results corresponding to the first modality and the second modality respectively based on the first modality feature and the second modality feature; and training the model based on cross-modality processing between the prediction results, imitation results and restoration results corresponding to the first modality and the second modality respectively for each sample. It also includes masking the first modality data or the second modality data, and training the model using the mask data of each modality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and more specifically, to a semantic segmentation model training method and device, a semantic segmentation method and device, a computing device, and a computer-readable storage medium. Background Art

[0002] In many applications, especially in robotics, autonomous driving and virtual reality, three-dimensional (3D) scene understanding is required. Common 3D sensors include lidar, millimeter-wave radar, depth camera, 3D scanner, etc. They can obtain the geometry, shape and scale information of objects and environments from the real world to help AI understand the real environment.

[0003] The scanning data of 3D sensors usually save the information of each point in the form of 3D point clouds, including 3D coordinates, reflectivity, size, etc. How to obtain useful information from 3D point clouds is an important research field in artificial intelligence. 3D semantic segmentation of 3D point clouds can assign semantic labels to each point to achieve the interpretation of three-dimensional scenes in reality. 3D sensors can be multimodal, for example, they can simultaneously obtain 2D image data (first modality data), 3D point cloud data (second modality data) and infrared image data (third modality data), etc.

[0004] 3D semantic segmentation may encounter domain gaps or domain shifts, such as domain gaps or domain shifts between day and night, different countries or sample sets. For example, for autonomous driving scenarios, a model trained on a sample set of 2D image data and 3D point cloud data in one country may not perform well when directly applied to a scenario in another country, because the road conditions in the two countries are somewhat different.

[0005] Therefore, a solution is needed that can still achieve good semantic segmentation in the presence of domain gap or domain shift. Summary of the invention

[0006] According to one aspect of the present application, a training method for a model for semantic segmentation is provided, comprising: obtaining a training sample set including a source domain sample set and a target domain sample set, wherein each sample in the source domain sample set includes first modality data, second modality data and a semantic label associated with the second modality, and each sample in at least a subset of the target domain sample set includes first modality data, second modality data and does not include the semantic label; for each sample in the training sample set: generating a first modality feature based on the first modality data of the sample, and generating a second modality feature based on the second modality data of the sample, wherein the second modality data of the sample includes data of multiple points; based on the first modality feature and the second modality feature, generating a first modality feature based on the first modality data of the sample, and generating a second modality feature based on the second modality data of the sample, wherein the second modality data of the sample includes data of multiple points; based on the first modality feature and the second modality feature, generating a second modality feature based on the first modality feature and the second modality feature; Bimodal features, generating prediction results, imitation results and recovery results corresponding to the first modality and the second modality respectively, wherein the prediction result corresponding to each modality indicates the semantic class prediction probability of each of the multiple points obtained based on the features of the current modality, the imitation result corresponding to each modality indicates the semantic class prediction probability of each of the multiple points obtained by imitating the features of another modality based on the features of the current modality, and the recovery result corresponding to each modality indicates the recovery data of the other modality obtained by using the features of the current modality; and training the model based on cross-modal processing between the prediction results, imitation results and recovery results corresponding to the first modality and the second modality respectively for each sample in the training sample set.

[0007] According to another aspect of the present application, a semantic segmentation method is also provided, including: acquiring first modal data and second modal data for the same scene, the second modal data including data of multiple points; using a semantic segmentation model to determine a first prediction result for the first modal data and a second prediction result for the second modal data, wherein the first prediction result and the second prediction result respectively indicate the semantic category prediction probability of each point in the second modal data, wherein the semantic segmentation model is trained according to the above method; and for each point in the second modal data, determining the maximum probability corresponding to each semantic category based on the first prediction result as the first probability, determining the maximum probability of each semantic category based on the second prediction result as the second probability, and taking the semantic category corresponding to the larger one of the first probability and the second probability as the semantic category to which the point belongs.

[0008] According to another aspect of the present application, a training device for a model of semantic segmentation is also provided, comprising an acquisition module, a result generation module and a training module. The acquisition module is used to acquire a training sample set including a source domain sample set and a target domain sample set, wherein each sample in the source domain sample set includes first modality data, second modality data and a semantic label associated with the second modality, and each sample in at least a subset of the target domain sample set includes first modality data, second modality data and does not include the semantic label; the result generation module is used to generate, for each sample in the training sample set: a first modality feature based on the first modality data of the sample, and a second modality feature based on the second modality data of the sample, wherein the second modality data of the sample includes data of multiple points; based on the first modality feature and the second modality feature, generate Prediction results, simulation results and recovery results corresponding to a first modality and a second modality respectively, wherein the prediction result corresponding to each modality indicates the semantic class prediction probability of each point obtained based on the characteristics of the current modality, the simulation result corresponding to each modality indicates the semantic class prediction probability of each point obtained by imitating the characteristics of another modality based on the characteristics of the current modality, and the recovery result corresponding to each modality indicates the recovery data of another modality obtained by using the characteristics of the current modality; and a training module is used to train the model based on cross-modal processing between the prediction results, simulation results and recovery results corresponding to the first modality and the second modality respectively for each sample in the training sample set.

[0009] According to another aspect of the present application, a device for semantic segmentation is also provided, and the training device includes an acquisition module, a prediction module, and a determination module. The acquisition module is used to acquire first modal data and second modal data for the same scene, and the second modal data includes data of multiple points. The prediction module is used to determine a first prediction result for the first modal data and a second prediction result for the second modal data using a model for semantic segmentation, wherein the first prediction result and the second prediction result respectively indicate the probability that each point in the second modal data belongs to each semantic category, wherein the model is trained according to the method described above. The determination module is used to determine, for each point in the second modal data, the maximum probability corresponding to each semantic category based on the first prediction result as the first probability, determine the maximum probability of each semantic category based on the second prediction result as the second probability, and take the semantic category corresponding to the larger one of the first probability and the second probability as the semantic category to which the point belongs.

[0010] According to another aspect of the present application, a computing device is provided, including: a processor; and a memory on which a computer program is stored. When the computer program is executed by the processor, the processor executes the method as described above.

[0011] According to another aspect of the present application, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the processor executes the steps of the training method and the semantic segmentation method as described above.

[0012] According to another aspect of the present application, a computer program product is also provided, including a computer program, which implements the steps of the training method and semantic segmentation method as described above when the computer program is executed by a processor.

[0013] Through the training method of the model for semantic segmentation in the present application, by using the source domain sample set and the target domain sample set to train the model, the domain adaptation effect of the model can be improved; and through the multiple supervision tasks introduced in the cross-modal processing process, while utilizing the complementarity of the modalities, the model can also learn more correspondences between the first modality and the second modality. Even in the case of a large domain gap or domain offset, the network corresponding to the first modality and the second modality in the model still has good performance, thereby further improving the domain adaptation effect, and achieving better semantic segmentation in actual target domain applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0015] Figure 1 An exemplary application scenario diagram of an embodiment of the present application is shown.

[0016] Figure 2 A schematic diagram of a training process of a model for semantic segmentation based on a cross-modal solution according to an embodiment of the present application is shown.

[0017] Figure 3 It shows that when the domain gap or domain offset is too large Figure 2 The output effect diagram of the model obtained through training process.

[0018] Figure 4A-4B A flowchart of a method for training a model for semantic segmentation according to an embodiment of the present application is shown.

[0019] Figure 5 More details are shown for each sub-step in step S430.

[0020] Figure 6 Shown with Figure 5 Figure 1 shows a schematic diagram of the information flow associated with the training process.

[0021] Figure 7 Shown with Figure 5 Schematic diagram of the loss calculation associated with the training process shown.

[0022] Figures 8A-8B A schematic process of generating a corresponding prediction result based on the first modal feature according to an embodiment of the present application is shown.

[0023] Fig. 9 A flowchart of a semantic segmentation method according to an embodiment of the present application is shown.

[0024] Figures 10A-10E Shown by reference Figures 4A-8B Schematic diagram of the output effect of the model trained with the described training method.

[0025] Fig.11 A structural block diagram of a training device for a model for semantic segmentation according to an embodiment of the present application is shown.

[0026] Fig.12 A structural block diagram of a device for semantic segmentation according to an embodiment of the present application is shown.

[0027] Fig.13 A structural block diagram of a computing device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solution and beneficial effects of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0029] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to imitate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0030] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0031] Machine Learning (ML) is a multi-disciplinary interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithmic complexity theory and other disciplines. It specializes in how computers imitate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning generally include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning. For example, in the embodiments of the present application, a domain adaptation solution in machine learning technology is used, based on cross-modal processing, so that the trained model can perform semantic segmentation in the target domain more accurately.

[0032] Autonomous driving technology usually includes high-precision maps, environmental perception, behavioral decision-making, path planning, motion control and other technologies. Autonomous driving technology has broad application prospects. 3D semantic segmentation is one of the core technologies in autonomous driving application scenarios. It is used to classify each point in the 3D point cloud data obtained based on 3D sensors, etc., so as to segment the road environment point cloud data. It can identify objects such as pedestrians and vehicles to help vehicles understand the road environment. This technology can be applied to unmanned vehicles and can greatly improve the vehicle's understanding of the surrounding environment. In this application, under the framework of the domain adaptation solution, the model is trained by cross-modal processing, which can improve the accuracy of the 3D semantic segmentation results in the actual application process of the target domain, thereby improving the reliability and safety of the autonomous driving process. Of course, the solution of this application is not only applied to autonomous driving scenarios, but also to other scenarios that require 3D semantic segmentation, such as robot control, virtual reality, etc.

[0033] Before describing the embodiments of the present application in detail, the terms that may be used in the present application are defined as follows.

[0034] Domain Adaptation: A type of transfer learning where the data distributions of the source domain and the target domain are different, but the two tasks are the same. For example, the task is the learning goal, which can be understood as the functional relationship between the input data and the output data. The source domain sample set can come from daytime road data, and the target domain sample set can come from nighttime road data. The purpose is to use the knowledge of the source domain sample set to improve the prediction effect of the model in the target domain. Both the source domain sample set and the target domain sample set can be used to train the model.

[0035] Source domain sample set: Each sample includes a label, which can be regarded as the party providing transfer knowledge. In this application, each sample includes a first modality data (e.g., 2D image data) and a second modality data (e.g., 3D point cloud data), and a semantic label corresponding to the second modality, such as the semantic category of each point obtained after 3D semantic segmentation of the 3D point cloud data.

[0036] Target domain sample set: Most of the data samples do not have labels. Each data sample is related to the source domain but different, and can be regarded as the party where transfer learning works. In this application, each data sample (target domain) that does not have a label includes a first modality data (e.g., 2D image data) and a second modality data (e.g., 3D point cloud data).

[0037] Modality: Data forms, such as 2D images and 3D point clouds can be considered as two different modalities.

[0038] As mentioned earlier, semantic segmentation encounters the problem of domain gap or domain shift, and domain adaptation-based solutions can solve this problem.

[0039] Figure 1 An exemplary application scenario diagram of an embodiment of the present application is shown. The application scenario is for an autonomous driving scenario.

[0040] like Figure 1 As shown, the vehicle 10 and the server 20 communicate with each other. The vehicle 10 is provided with sensors for collecting point cloud data and image data of the driving scene. The sensors include, for example, image sensors (for collecting image data of each target in the driving scene) and lidar sensors as 3D sensors (for collecting 3D point cloud data of each target in the driving scene). The lidar sensor can also collect 2D image data of the driving scene at the same time. In addition, the sensors may also include millimeter wave radar sensors and ultrasonic radar sensors, etc. This application does not limit the type of sensor, as long as the required 2D image data and 3D point cloud data of the driving scene can be obtained. The vehicle 10 may also include an on-board controller for receiving the driving scene data sent by each sensor and communicating with the server through a communication module.

[0041] The server 20 can receive the current 2D image data and 3D point cloud data from the vehicle 10 via the communication module, and use a pre-trained domain adaptation model that can achieve 3D semantic segmentation (the training method will be described later) to identify the current 2D image data and 3D point cloud data for the current driving scene and determine the current road conditions. For example, the vehicle controller at the vehicle can obtain the recognition result from the server and generate control information for controlling the operation of the vehicle based on the recognition result, or directly obtain the control information from the server and control the operation of the vehicle 10 based on the control information.

[0042] Alternatively, the server 20 may also send a pre-trained domain adaptation model capable of implementing 3D semantic segmentation to the vehicle controller through the communication module, so that the vehicle controller performs similar processing after acquiring the current 2D image data and 3D point cloud data, thereby controlling the operation of the vehicle 10.

[0043] In some other embodiments, the operation of the above-mentioned server 20 can be performed by the terminal, that is, the current 2D image data and 3D point cloud data are received from the vehicle 10 at the terminal, and a pre-trained domain adaptation model capable of 3D semantic segmentation is used to determine the current road conditions, and interact with the vehicle controller to control the operation of the vehicle 10.

[0044] The server may include one or more processors, memory, and I / O interfaces for interacting with terminal devices. In addition, the server may also be configured with a database. The server may be an independent physical server, or a server cluster or distributed system consisting of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0045] The terminal device may include one or more processors, memory, I / O interface for interacting with the server, and display panel, etc. The terminal device may be a smart phone, tablet computer, laptop computer, desktop computer, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected via wired or wireless communication, which is not limited in this application.

[0046] The methods in various embodiments of the present application may be executed by a server, or by a terminal, or by a collaboration between the two.

[0047] Figure 2 A schematic diagram of a training process of a model for semantic segmentation based on a cross-modal solution according to an embodiment of the present application is shown.

[0048] In the context of the present application, the training sample set of the model includes a source domain sample set and a target domain sample set, and both are used for the training of the model, so that the effect of domain adaptation can be achieved, and based on the cross-modal processing scheme for the results of the two modalities, the effect of domain adaptation can be further improved by utilizing the complementarity of the modalities.

[0049] The source domain sample set can be expressed as (X 2D,S ,X 3D,S ,Y 3D,S ), that is, each sample in the source domain sample set includes a first modality data (e.g., 2D image data), a second modality data (e.g., 3D point cloud data) and a label (semantic label) corresponding to the second modality data. The target domain sample set can be expressed as (X 2D,T ,X 3D,T ), that is, each sample in the target domain sample set includes a first modality data (e.g., 2D image data) and a second modality data (e.g., 3D point cloud data), but does not include a semantic label. In most descriptions of this application, the first modality data is 2D image data and the second modality data is 3D point cloud data as an example, but it is not limited to this. For example, infrared data can also be used as a modality data.

[0050] In addition, in the present application, all data dimensions of the source domain sample set and the target domain sample set are the same. For example, the dimension of the 2D image data included in each sample is (H, W, 3), and the dimension of the 3D point cloud data is (N, 3), where H and W are the height and width of the 2D image data, respectively, N is the number of points in the 3D point cloud data, and 3 is the number of channels (for example, RGB three channels).

[0051] like Figure 2 As shown, a 2D image data (dimension is (H, W, 3)) and a 3D point cloud data (dimension is (N, 3)) are obtained for each sample in the training sample set, and are respectively input into the first feature extraction network and the second feature extraction network to obtain the first modal feature map - 2D feature map (Feature Map, dimension is (H, W, F 2D )) and the second modality feature - 3D feature (dimension is (N,F 3D ), where H and W are the height and width of the feature map extracted from the 2D image data, respectively, and F 2D is the number of feature channels based on the first feature extraction network structure, N is the number of points in the 3D point cloud data, and F 3Dis the number of feature channels based on the second feature extraction network. Then, the 2D feature map is sampled, for example, the 3D point cloud data is projected onto the 2D image using the camera coordinates, and only the 2D feature points projected onto the 3D point cloud data are sampled for the 2D feature map to obtain the first modal feature - 2D feature (N, F 2D ), that is, the number of sampling points (feature points) is the same as the number of points N in the 3D point cloud data. For the convenience of description, unless otherwise specified, 2D features refer to the 2D features after sampling (N, F 2D ).

[0052] Optionally, the first feature extraction network may be a U-Net based on ResNet-34, and the second feature extraction network may be a SparseConvNet. Of course, 2D features (2D feature maps) and 3D features may also be obtained by other feature extraction networks well known in the art, and this application does not limit this.

[0053] Then, the 2D features are passed through the first prediction network to obtain the first prediction result of semantic segmentation (using P 2D ), and the 3D features pass through the second prediction network to obtain the second prediction result of semantic segmentation (denoted by P 3D Represented). For example, the first prediction result may be a semantic category prediction result for N points, so that the dimension of the first prediction result may be (N, C), where C is the number of semantic categories, that is, the first prediction result may indicate the probability of belonging to each semantic category corresponding to the 2D feature of each point in the N points. Similarly, the dimension of the second prediction result is also (N, C), and the second prediction result may indicate the probability of belonging to each semantic category corresponding to the 3D feature of each point in the N points. Optionally, the first prediction network and the second prediction network may each be composed of a linear layer and a normalization layer. As is known in the art, the linear layer is also called a fully connected layer, in which each neuron is connected to all neurons in the previous layer to achieve a linear combination or linear transformation of the previous layer; the normalization layer, for example, uses softmax to output each probability.

[0054] After obtaining the first prediction result (for 2D features) and the second prediction result (for 3D features), training can be performed based on them. In addition, since there is a target domain sample set, it can be considered to use it for model training to improve the domain adaptation effect of the model. Since the first prediction result and the second prediction result are both prediction results for semantic segmentation, even if at least part of the samples in the target domain sample set do not have semantic labels, modal complementarity can be used, that is, feature matching of 2D features and 3D features can be performed for model training.

[0055] For example, the 2D feature can also pass through the first imitation network to obtain the first imitation result of the semantic segmentation imitating the 3D feature (using P2D→3D The first imitation result can indicate the prediction result when the 2D feature is used to imitate the 3D feature, and the dimension is also (N, C). The 3D feature can also pass through the second imitation network to obtain the second imitation result of the semantic segmentation imitating the 2D feature (denoted by P 3D→2D The second imitation result may indicate the prediction result (N, C) when the 2D feature is imitated with the 3D feature. Then, the first imitation result (P 2D→3D ) and the second prediction result (P 3D ) is matched, and the second imitation result (P 3D→2D ) and the first prediction result (P 2D ) to match, that is, the first imitation result (P 2D→3D ) and the second simulation result (P 3D→2D ) are respectively composed of the corresponding second prediction results (P 3D ) and the first prediction result (P 2D ) to supervise.

[0056] More specifically, for the training process, since the effect of domain adaptation is to be improved, both the source domain sample set and the target domain sample set can be used for model training, and they can be performed alternately. For example, the model can be trained alternately in batches, and each batch includes multiple samples. For example, after training with 10 samples in the source domain sample set, 10 samples in the target domain sample set are used for training, and the training is performed alternately in this way.

[0057] When training with samples from the source domain sample set (and samples in the target domain sample set that include semantic labels), since these samples all include semantic labels, full supervised training can be performed, that is, the first prediction result P corresponding to the 2D feature of each sample 2D The second prediction result P corresponding to the 3D feature 3D The corresponding semantic label Y 3D,S To supervise, the second imitation result P of the sample's 2D feature to the 3D feature 3D→2D And the first imitation result P of 3D features to 2D features 2D→3D The first prediction result P 2D and the second prediction result P 3D To supervise, such as Figure 2 Indicated by the dashed arrow.

[0058] When training with samples in the target domain sample set that do not include semantic labels, it is impossible to use the semantic label Y 3D,S To supervise the first prediction result P corresponding to the 2D feature of each sample 2D The second prediction result P corresponding to the 3D feature 3D, so the first prediction network and the second prediction network do not work, but the second imitation result P of the 2D feature of each sample to the 3D feature 3D→2D And the first imitation result P of 3D features to 2D features 2D→3D The first prediction result P 2D and the second prediction result P 3D To supervise, that is, the first imitation network and the second imitation network work normally, such as Figure 2 Indicated by the dashed arrow.

[0059] When using samples from the source domain sample set and the target domain sample set for training, the loss can be calculated by introducing a loss function, and the parameters of each network can be adjusted based on the goal of making the loss value as small as possible, for example, by using a gradient descent method.

[0060] As described above, the training process involves two supervision processes, one is to supervise the prediction results using semantic labels, and the other is to supervise the imitation results using the prediction results. For the former, the loss function can be a cross entropy loss function, as shown in formula (1), and for the latter, the loss function can be a KL divergence, as shown in formula (2).

[0061] CE(Y 3D,S ||P 2D ),CE(Y 3D,S ||P 3D ) (1)

[0062] KL(P 2D ||P 3D→2D ),KL(P 3D ||P 2D→3D ) (2)

[0063] Optionally, the calculation results of each loss function can be combined to adjust the model parameters. For example, each loss function can be added or weighted to obtain a comprehensive loss function, and the parameters of each network can be adjusted based on the goal of making the value of the comprehensive loss function as small as possible.

[0064] Optionally, when using the comprehensive loss function to calculate the loss for samples with semantic labels, both the cross entropy loss function and the KL divergence loss function need to be calculated, while when using the comprehensive loss function to calculate for samples without semantic labels, there is no need to calculate the cross entropy loss function part (i.e., it is 0), and only the loss of the KL divergence loss function part is calculated, so that the model parameters are adjusted according to the loss calculated by the comprehensive loss function. For example, the total loss of each batch of samples can be calculated to adjust the model parameters once.

[0065] It can be seen that in the reference Figure 2In the training process of the model described, the cross-modal processing between modalities is realized, for example, the first imitation result (P 2D→3D ) and the second simulation result for the second modality (P 3D→2D ) are respectively composed of the corresponding second prediction results (P 3D ) and the first prediction result (P 2D ) for supervision, so the complementarity between the modalities can be used for semantic segmentation, so that the predicted results of the two prediction networks of the first modality (for example, 2D image data) and the second modality (3D point cloud data) for semantic segmentation can help each other, thereby improving the accuracy of the semantic segmentation of the model when applied in the actual target domain to a certain extent, and improving the domain adaptation effect.

[0066] However, as mentioned above, the source domain sample set and the target domain sample set are alternately used for model training, and based on the cross-modal processing process, the trained model can have a good domain adaptation effect. However, if the domain gap or domain offset is too large, for example, the source domain sample set is road data from developed areas, and the target domain sample set is road data from underdeveloped areas, there will be a situation where the difference between the sample sets of the two domains is too large, so the performance of the model trained with these sample sets may not be good, that is, the performance of the first prediction network that obtains the first prediction result for the first modality and the second prediction network that obtains the second prediction result for the second modality are not good. Therefore, even based on the cross-modal processing process, it is impossible to improve the accuracy of the semantic segmentation of the model when applied in the actual target domain through the complementarity between the two modalities.

[0067] Can be Figure 3 To illustrate the output effect of the model when the domain gap or domain offset is too large.

[0068] Assume that the data in the source domain come from road data in developed countries, and the data in the target domain come from road data in underdeveloped countries. Figure 3 As shown, the first modality data is 2D image data, and the second modality data is 3D point cloud data. After passing through the 2D feature extraction network and the 3D feature extraction network respectively, they enter the semantic segmentation model (for example, Figure 2 The first prediction network, the first imitation network, the second prediction network, and the second imitation network shown in the figure respectively obtain the semantic segmentation prediction results based on the 2D image data and the semantic segmentation prediction results based on the 3D point cloud data (for example, from Figure 2 The outputs of the first prediction network and the second prediction network shown are hereinafter referred to as 2D prediction results and 3D prediction results, respectively). The two outputs of the semantic segmentation model indicate the predicted (recognized) semantic category of each point at the position of each point of the 3D point cloud data.

[0069] from Figure 3 As can be seen in the figure, the vertical poles (white boxes) of the input 2D image data in the left figure are only a small part recognized in the output 2D prediction results (only the category of the points in this part is recognized as vertical poles, as shown in black in the upper left figure), and the vehicle is almost not recognized in the output semantic segmentation prediction results; the vertical poles (black boxes) of the input 3D point cloud data in the right figure are only a small part recognized in the output 3D prediction results (only the category of the points in this part is recognized as vertical poles, as shown in black in the upper right figure), and the vehicle is not completely recognized in the output 3D prediction results.

[0070] It can be seen that when the domain gap or domain offset is too large, Figure 2 The model trained in the described way cannot perform semantic segmentation well, so it is also necessary to train the model based on Figure 2 The described training method was further improved.

[0071] Therefore, an embodiment of the present application also provides a cross-modal modeling method, which enables the model to be trained more fully by additionally introducing a supervision process related to the recovery network, and thus can have a better domain adaptation effect. In addition, a cross-modal mask modeling method is also provided, which masks the first modality data and the second modality data of at least part of the samples, and utilizes the cross-modal processing between one of the modality data of each sample and the other modality mask data, so that the model learns more first modality-second modality correspondences, realizes modal complementarity, and enables the trained model to be fully fitted, thereby further improving the domain adaptation effect of the model.

[0072] Figure 4A-4B A schematic flow chart of a method for training a model for semantic segmentation according to an embodiment of the present application is shown.

[0073] like Figure 4A As shown, in step S410, a training sample set including a source domain sample set and a target domain sample set is obtained, wherein each sample in the source domain sample set includes first modality data, second modality data and a semantic label associated with the second modality, and each sample in at least a subset of the target domain sample set includes first modality data, second modality data and does not include the semantic label.

[0074] For example, as described above, both the source domain sample set and the target domain sample set are used to train the model, so that the domain adaptation effect can be achieved, so that the model can be better applied in the target domain. The samples in the source domain sample set contain pre-labeled semantic labels in addition to the first modality data and the second modality data, while at least a portion of the samples in the target domain sample set only include the first modality data and the second modality data and do not include pre-labeled semantic labels. The types of the first modality data and the second modality data of each sample in the source domain sample set and the target domain sample set are the same, but the application scenarios (domains) are different. For example, each first modality data can be 2D image data during the day, and each second modality data can be 3D point cloud data. For example, the first modality data of the source domain sample can be 2D image data of a scene during the day, and the second modality data can be 3D point cloud data of the scene during the day. The first modality data of the target domain sample can be 2D image data of a scene at night, and the second modality data can be 3D point cloud data of the scene at night.

[0075] In an embodiment of the present application, there may also be a small number of samples containing semantic labels in the target domain sample set. In this case, operations similar to those used to train a model using samples in the source domain sample set can be performed. For example, the first modal features and the second modal features of the target domain samples can be passed through a prediction network and supervised training can be performed using the corresponding semantic labels contained therein.

[0076] Optionally, each first modality data may be 2D image data, and each second modality data may be 3D point cloud data. However, the present application is not limited thereto, and the first modality data or the second modality data may also be infrared data, etc.

[0077] In step S420, for each sample in the training sample set, a first modal feature is generated based on the first modal data of the sample, wherein the second modal data of the sample includes data of multiple points; and a second modal feature is generated based on the second modal data of the sample; and based on the first modal feature and the second modal feature, a prediction result, a simulation result and a recovery result corresponding to the first modality and the second modality are generated respectively.

[0078] Among them, the prediction result corresponding to each modality indicates the predicted probability that each point obtained based on the characteristics of the current modality belongs to each semantic category, the imitation result corresponding to each modality indicates the predicted probability that each point obtained by imitating the characteristics of another modality based on the characteristics of the current modality belongs to each semantic category, and the recovery result corresponding to each modality indicates the recovery data of another modality obtained using the characteristics of the current modality.

[0079] For example, in the context of the present application, the recovered data is the predicted data of the original data of the other modality of the first modality and the second modality based on the features of one of the first modality and the second modality, and as will be described later, the difference between the predicted data and the original data of the other modality (for example, calculated using a loss function) will be used to train the model so that the model learns more correspondences between the two modalities. For example, 3D point cloud data is predicted based on 2D features, and 2D image data is predicted based on 3D features (3D features are features of multiple points included in the point cloud, so what is predicted accordingly is sampled data of 2D image data). In addition, as will be described later, the prediction process here can be performed through a recovery network in the model, such as a multilayer perceptron, a deep learning model, and the like.

[0080] Optionally, in order to obtain the first modal feature and the second modal feature, a first feature extraction network can be used to perform feature extraction on the first modal data to obtain a first feature map; the second modal data is projected according to the first modality to obtain second modality-first modality projection data, and the second modality-first modality projection data includes the projection data of the multiple points; the first feature map is sampled according to the positions of the multiple points in the second modality-first modality projection data to obtain the first modal feature; and a second feature extraction network is used to perform feature extraction on the second modality data to obtain the second modal feature.

[0081] Optionally, the multiple points included in the second modal data may be multiple feature points, multiple position points, or a collection of multiple points carrying data that can be processed by a processor. The second modal feature includes features of multiple points corresponding to the data of the multiple points in the second modal data, and the first modal feature is obtained by sampling the positions projected by the multiple points in the second modal data, that is, it also includes features of the multiple points.

[0082] For example, when the first modality is a 2D image and the second modality is a 3D point cloud, the first feature extraction network can be a U-Net based on ResNet-34, and the second feature extraction network can be a SparseConvNet. Of course, 2D features and 3D features can also be obtained by other feature extraction networks well known in the art, and this application does not limit this. For example, 3D point cloud data includes data of N points, so the extracted 3D features include the 3D features of the N points (as described in the context, the dimensions are (N, F 3D )), and the 2D features include the 2D features of the N points obtained by sampling (as described in the context, the dimension is (N, F 2D )).

[0083] In addition, the first modality feature can be used to generate a first prediction result, a first simulation result, and a first recovery result, and the second modality feature can be used to generate a second prediction result, a second simulation result, and a second recovery result.

[0084] For example, based on the first modal feature, the first prediction network of the model is used to generate a first prediction result corresponding to the first modality, and based on the second modal feature, the second prediction network of the model is used to generate a second prediction result corresponding to the second modality; based on the first modal feature, the first imitation network of the model is used to generate a first prediction result corresponding to the first modality, and based on the second modal feature, the second imitation network of the model is used to generate a second imitation result corresponding to the second modality; and based on the first modal feature, the first recovery network of the model is used to generate a first recovery result corresponding to the first modality, and based on the second modal feature, the second recovery network of the model is used to generate a second recovery result corresponding to the second modality.

[0085] Optionally, each of the first prediction network and the second prediction network may include a linear layer and a normalization layer; each of the first imitation network and the second imitation network may include a linear layer and a normalization layer; and each of the first restoration network and the second restoration network may include a multilayer perceptron.

[0086] For example, the linear layer is used to perform linear processing on the input first modality feature or the second modality feature, and the processed feature is input into a normalization layer (e.g., softmax), so that the probability of each point belonging to each semantic category can be obtained as a prediction result. A multi-layer perceptron (MLP) is used to perform recovery processing based on the input first modality feature or the second modality feature, and the number of intermediate channels is, for example, 4096. Training the model can be interpreted as adjusting the parameters of the included networks.

[0087] In step S430, the model is trained based on cross-modal processing between prediction results, simulation results and recovery results corresponding to the first modality and the second modality respectively for each sample in the training sample set.

[0088] Optionally, when performing cross-modal processing, partial results corresponding to one of the first and second modalities can be matched with the corresponding reference data of the modality to use the matching results in the model training process, wherein at least a portion of the reference data of the modality is obtained based on the modal data of the other modality of the first and second modalities (for example, the other modal data itself or its sampled data, or the prediction results corresponding to the other modal data). That is, when training a model based on cross-modal processing, the data information of the other modality can be cross-utilized. More example details of the model training process will be discussed later in conjunction with Figure 5-8B Describe. As mentioned above, when training the model, the model parameters can be adjusted according to batches, and the source domain sample set and the target domain sample set are used alternately to train the model to improve the effect of domain adaptation. That is, different subsets of the source domain sample set and different subsets of the target domain sample set are used alternately to train the model. For example, the source domain sample set and the target domain sample set can be divided into multiple source domain subsets and multiple target domain subsets according to the batch size (batchsize), each subset can correspond to a batch, and then after training the model with a source domain subset, the model is trained with a target domain subset, and then the model is trained with another source domain subset, and so on. Optionally, the batch size can be selected according to actual needs, for example, the minimum can be 1, and the maximum can be the maximum number of source domain or target domain sample sets. Optionally, after training the model with more than one source domain subset, the target domain subset can be used to train the model, and the present disclosure does not limit this.

[0089] In addition, in some other implementations, mask processing can be used to further improve the domain adaptation effect of the model.

[0090] Therefore, the training method 400 may also include a mask processing process, such as Figure 4B shown. Figure 4B Steps S410-S430 and Figure 4A For example, the corresponding prediction results, simulation results and recovery results generated based on the features of each modality in step S420 and the training process based on these results in step S430 are similar to Figure 4A The description is similar to that in , so it will not be repeated here.

[0091] In step S410', according to a predetermined standard, mask processing is selectively performed on the first modality data and the second modality data of each sample in the training sample set to obtain first modality mask data and second modality mask data corresponding to at least one sample.

[0092] For example, the masking process may be performed on the first modality data and the second modality data of at least one randomly selected sample in the training sample set, and for each of the at least one sample, only one of the first modality data and the second modality data may be masked, or both modality data may be masked. The masking process described in the present application may refer to removing part of the content information included in at least one of the first modality data and the second modality data.

[0093] For example, for each randomly selected sample to be subjected to mask processing, when mask processing is performed on the first modality data and / or the second modality data included in the sample, the first modality data and / or the second modality data can be divided into a plurality of blocks of a predetermined size, respectively. Based on the same concept, each modality data can also be divided into a predetermined number of blocks of the same size. Then, a removal operation is performed on the plurality of blocks of a predetermined size obtained by dividing each modality data based on a predetermined removal ratio to implement mask processing.

[0094] Optionally, the first modality data or the second modality data of each sample to be masked may correspond to at least one mask data respectively. Thus, each sample in the training sample set may include multiple samples for training the model, namely, the first modality data, the second modality data, and at least one first modality mask data and / or at least one second modality mask data.

[0095] In addition, because the semantic integrity must be ensured, each sample input into the model for training cannot all be masked data, that is, at least one sample is the first modal data that has not been masked, and it is also not possible to mask only samples of one modality, for example, it is not possible to mask only the first modal data or only the second modal data, that is, both modalities must have samples that have not been masked and samples that have been masked. That is to say, when used to train the model, for each sample, a first modal data is paired with a second modal data or a corresponding second modal mask data as one input, or a second modal data is paired with a first modal data or a corresponding first modal mask data as another input, thereby realizing cross-modal mask modeling. While ensuring the integrity of the semantics, the training sample set is expanded, and by pairing masked samples of one modality with samples of another modality that have not been masked, the model can learn more first modality-second modality correspondences.

[0096] Therefore, the predetermined standard can be set as follows: by presetting the probability of the first modality data and the second modality data of each sample being masked, for example, the probability of the first modality data of each sample being masked is preset to a first predetermined probability, and the probability of the second modality data of each sample being masked is preset to a second predetermined probability. Alternatively, the predetermined standard can also be set as follows: presetting the number of first modality data and second modality data to be masked, for example, the ratio of the number of samples whose first modality data are masked to the number of samples in the training sample set is preset to a first predetermined ratio, and the ratio of the number of samples whose second modality data are masked to the number of samples in the training sample set is preset to a second predetermined ratio.

[0097] In some embodiments, a hyperparameter set may be introduced to define the predetermined size of the block when dividing, the predetermined removal ratio when removing the block, and the first predetermined probability (or ratio) and the second predetermined probability (or ratio). For example, the hyperparameter set may include the following four hyperparameters (p, mr, m 2D ,m 3D ), wherein p is a predetermined size of each block into which each first modality data or second modality data is divided or a predetermined number of divided blocks, mr is a removal ratio, and m 2D represents the probability that the first modality data of each sample is masked (or the ratio of the samples whose first modality data are masked to the total number (i.e., the number of training sample sets)), m 3D The probability that the second modality data of each sample is subjected to masking is preset to a second predetermined probability (or the ratio of the number of samples whose second modality data is subjected to masking to the total number (i.e., the number of training sample sets)), such that (1-m 2D -m 3D ) represents the probability that both the first modality data and the second modality data of each sample are not masked (or the ratio of the total number of samples whose first modality data and the second modality data are both masked).

[0098] The specific values ​​of the hyperparameter set can be adjusted according to actual conditions, and this application does not impose any restrictions on this. As a specific example, in the experiment of domain adaptation for daytime / nighttime scenes, four relatively suitable hyperparameter values ​​can be obtained, which are (16, 0.15, 0.2, 0.2).

[0099] In addition, Figure 4B In the case where mask processing is added, one of the first modal feature and the second modal feature generated in step S420 may be based on mask data. For example, in this case, when generating the first modal feature and the second modal feature for each sample in the training sample set in step S420, it may also include generating the first modal feature based on the first modal data or the first modal mask data of the sample, generating the second modal feature based on the second modal data or the second modal mask data of the sample, and the sample on which at least one of the first modal feature and the second modal feature of the sample is based is not mask data.

[0100] Optionally, when extracting the first modal feature and the second modal feature, when adding mask processing, the first feature extraction network can be used to perform feature extraction on the first modal data or the first modal mask data to obtain a first feature map; the second modal data or the second modal mask data is projected according to the first modality to obtain second modality-first modality projection data, and the second modality-first modality projection data includes the projection data of the multiple points; the first feature map is sampled according to the positions of the multiple points in the second modality-first modality projection data to obtain the first modal feature; and the second feature extraction network is used to perform feature extraction on the second modality data or the second modality mask data to obtain the second modal feature.

[0101] pass Figure 4A-4B The training method of the model for semantic segmentation shown in the figure can achieve the domain adaptation effect of the model by training the model with the source domain sample set and the target domain sample set; and through the multiple supervision tasks introduced in the cross-modal processing process, the model can also learn more correspondences between the first modality and the second modality while utilizing the complementarity of the modalities. Even in the case of large domain gaps or domain offsets, the network corresponding to the first modality and the second modality in the model still has good performance, so the domain adaptation effect can be further improved, and better semantic segmentation can be achieved in actual target domain applications.

[0102] In addition, when introducing masked data samples, the training sample set is expanded while ensuring the semantic integrity, and by matching masked samples of one modality with unmasked samples of another modality, that is, by realizing cross-modal mask modeling, the trained model can learn more first-modality-second-modality correspondences, further improve the domain adaptation effect, and achieve better semantic segmentation in actual target domain applications.

[0103] The following combination Figure 5-8B The process of training models based on cross-modal processing is described in detail. Figure 5 More details are shown for each sub-step in step S430. Figure 6 Shown with Figure 5 The information flow diagram associated with the training process is shown in Figure 1. Figure 7 Shown with Figure 5 Schematic diagram of the loss calculation associated with the training process shown.

[0104] In step S430-1, for each sample of the training sample set: when the sample does not include the semantic label, a first loss between the first prediction result and the second imitation result, a second loss between the second prediction result and the first imitation result, a third loss between the first recovery result and the second modal data of the sample, and a fourth loss between the second recovery result and the sampling data of the first modal data of the sample are calculated; and when the sample includes the semantic label, a fifth loss between the first prediction result and the semantic label of the sample and a sixth loss between the second prediction result and the semantic label of the sample are further calculated.

[0105] Optionally, when calculating the fourth loss, it involves determining the sampling data of the first modality data. This is because the second modality data includes data of multiple points, and the number and position of the points are determined. Cross-modality processing is required, so the information of the points in the first modality data to be used also needs to correspond to the number and position of the points in the second modality data. Therefore, the sampling data of the first modality data can be obtained by projection. For example, the second modality data of the sample can be projected according to the first modality to obtain second modality-first modality projection data, and the second modality-first modality projection data includes the projection data of the multiple points; then, the first modality data of the sample is sampled according to the position of the multiple points in the second modality-first modality projection data to obtain the sampling data of the first modality data.

[0106] Optionally, the first loss between the first prediction result and the second imitation result and the second loss between the second prediction result and the first imitation result can be calculated using the KL divergence loss function, the third loss between the first recovery result and the second modality data and the fourth loss between the second recovery result and the sampled data of the first modality data can be calculated using the L2 loss function; and the fifth loss between the first prediction result and the semantic label of the sample and the sixth loss between the second prediction result and the semantic label of the sample can be calculated using the cross entropy loss function. The forms of these loss functions are well known in the art, so they are not repeated here.

[0107] Figure 6 An example process of training a model is shown. Figure 6 As shown, still taking the first modality data as 2D image data and the second modality data as 3D point cloud data as an example, the 2D image data or a 2D image data after mask processing enters the first modality branch, and is Figure 2 The first feature extraction network and sampling method shown in the figure obtain 2D features, and the 3D point cloud data or a 3D point cloud data after mask processing enters the second modal branch and passes through Figure 2The second feature extraction network shown in FIG. 3D features are obtained. The 2D features are processed by the first processing branch (eg, Figure 6 After the 2D prediction result, 3D imitation result and 3D restoration result are obtained, the 3D feature is processed by the second processing branch (for example, Figure 6 The 3D prediction result, 2D simulation result and 2D restoration result are obtained after the branch shown by the gray solid line.

[0108] 2D features (dimensions are (N, F 2D )) After passing through the first prediction network, the first imitation network and the first recovery network (first modal branch), the 2D prediction result P is obtained. 2D (dimension is (N, C)), 3D simulation result P 2D→3D (dimension is (N, C)) and 3D recovery result M 2D→3D (dimension is (N, 3)). 3D features (dimension is (N, F 3D )) After passing through the second prediction network, the second imitation network and the second recovery network (the second modal branch), the 3D prediction result P is obtained. 3D (dimension is (N, C)), 2D simulation result P 3D→2D (dimension is (N, C)) and 2D recovery result M 3D→2D (The dimension is (N, 3)).

[0109] Then, based on the 2D prediction result P 2D (First prediction result) and 2D simulation result P 3D→2D (second imitation result) between the imitation loss (first loss, such as Figure 6 The black dashed line in the figure shows the 3D (Second prediction result) and 3D simulation result P 2D→3D (First imitation result) between the imitation loss (second loss, such as Figure 6 The gray dashed line shows the 3D point cloud data X 3D (without mask processing, dimension is (N, 3)) and 3D restoration result M 2D→3D (first restoration result) (dimension is (N, 3)) between the restoration loss (third loss), and using the 2D image data X 2D The sample data Xs 2D (without mask processing, dimension is (N, 3)) and 2D recovery result M 3D→2D (second restored result) (dimension is (N, 3)) restoration loss (fourth loss), and if the sample includes a semantic label, the 2D prediction result P is further used 2D (First prediction result) and the corresponding semantic label Y 3D,SThe prediction loss (fifth loss) is calculated using the 3D prediction result P 3D (Second prediction result) and the corresponding semantic label Y 3D,S The prediction loss between (the sixth loss). Finally, the parameters of the model (including the above-mentioned networks) are adjusted based on these losses.

[0110] It should be noted that the sample data Xs of the 2D image data is used when calculating the restoration loss. 2D , since the 2D restoration result M 3D→2D It is obtained based on 3D features. The 3D features are associated with N points of 3D point cloud data, so the 2D recovery result M 3D→2D It is also associated with the N points, so it is also necessary to first calculate the 3D point cloud data X 3D The point pair 2D image data X 2D Sampling is performed to obtain the sampling data Xs of the 2D image data 2D That is, X 3D is the 3D point cloud data itself, Xs 2D It is a set of points obtained by mapping 3D point cloud data to 2D image data and sampling it.

[0111] In addition, even if the data on which the first modal feature and the second modal feature are based are masked data, the sampling data used to calculate the recovery loss between the recovery results of each modality is also unmasked data (i.e., original modal data) to ensure the information integrity of the benchmark data, thereby better promoting the learning of the first modality-second modality correspondence relationship of the model.

[0112] In summary, the features of each modality (obtained through the feature extraction network based on modal data or mask data) will enter three networks, namely the prediction network, the imitation network, and the recovery network. The prediction network is used to generate the final prediction result (semantic segmentation result), the imitation network performs cross-modal feature matching, and the recovery network performs cross-modal recovery. Therefore, the trained model can learn more first-modality-second-modality correspondences, so that the complementarity of the two modalities can be better applied to the training model, making the model more suitable and better for semantic segmentation in practical applications.

[0113] In addition, as described above, the loss function may be a combination of loss functions of multiple networks, such as addition and weighting, so as to adjust the model parameters (parameters of each network) based on the goal of minimizing the loss value of the loss function.

[0114] Figure 7 A schematic diagram showing the various losses that need to be calculated during training. Figure 7The example in which the first modality data is 2D image data, the second modality data is 3D point cloud data, and the samples in the source domain sample set have semantic labels and the target domain sample set does not have semantic labels is still used for explanation.

[0115] like Figure 7 As shown, the model is trained on the source domain and the target domain. In the source domain (each sample includes a semantic label), the loss can include the 2D prediction result P 2D and the corresponding semantic label Y 3D,S The prediction loss (e.g., CE based on the cross entropy loss function (Y 3D,S ||P 2D ), 3D prediction result P 3D and the corresponding semantic label Y 3D,S The prediction loss between (e.g., CE(Y 3D,S ||P 3D )), 2D prediction result P 2D With 2D simulation results P 3D→2D The imitation loss between (for example, KL based on KL divergence loss function (P 2D ||P 3D→2D ), 3D prediction result P 3D With 3D simulation results P 2D→3D The imitation loss between (for example, KL(P 3D ||P 2D→3D ), the sampling data of 2D image data (without mask processing) and the 2D restoration result M obtained based on 3D features 3D→2D The recovery loss L 3D =L 2 (Xs 2D ||M 3D→2D ), and using 3D point cloud data X 3D (without mask processing, dimension is (N, 3)) and the 3D restoration result M obtained based on 2D features 2D→3D (Dimension is (N, 3)) The recovery loss L between 3d =L 2 (X 3D ||M 2D→3D ).

[0116] Continue back Figure 5 In step S430-2, the model is trained based on each loss calculated for each sample in the training sample set.

[0117] For example, the model can be trained alternately on the source domain and the target domain, for example, the model can be trained alternately on a subset of the source domain and a subset of the target domain, and after the loss value is calculated for a batch of samples using a combination loss function of each loss function, the model parameters are adjusted once.

[0118] Optionally, in other embodiments, in order to enable the model to better learn the correspondence between the first modality and the second modality, so that the complementarity of the two modalities can be better applied to the training model, in the process of generating corresponding prediction results based on one of the first modal features and the second modal features, the information of the other modal features can be added, which can help better predict the prediction results of the one modality.

[0119] Therefore, when generating the first prediction result or the second prediction result based on the first modal feature or the second modal feature of each sample, a dynamic cross-modal filter network (also called a dynamic convolutional layer) can be used to replace the simple linear layer structure used in the above prediction network, that is, the first prediction network (the second prediction network) includes a dynamic convolutional layer and a normalization layer. Since the following operations for generating prediction results for the first modal feature and the second modal feature are the same, the following is combined with Figures 8A-8B The following description is made by taking the case where a corresponding prediction result is generated based on the first modal feature as an example. The case where a corresponding prediction result is generated based on the second modal feature can also be operated similarly.

[0120] First, as shown in operation 810, a weight matrix for performing a dynamic convolution operation on the first modality feature is generated based on the second modality feature.

[0121] Optionally, the weight matrix can be obtained after the second modal feature passes through a linear layer, and the parameters of the linear layer can be adjusted during the training of the model. In other words, the adjustment of the parameters of the dynamic cross-modal filter network can be considered as the adjustment of the linear layer.

[0122] The dimension of the weight matrix is ​​associated with the dimension of the first modality feature and the number of semantic categories. For example, the dimension of the first modality feature is (N, F 2D ), and the number of semantic categories is C, then the dimension of the weight matrix is ​​(N, F 2D , C), where N is the number of sampling points for the first modality data (or the first modality mask data), F 2D is the number of channels of the first modal feature obtained by feature extraction. Optionally, the number of sampling points corresponding to the first modal feature can be determined based on the second modal data. For example, as described above, the number of sampling points corresponding to the 2D feature obtained by feature extraction and sampling of the 2D image data is the number of points N in the 3D point cloud data.

[0123] For example, Figure 8B As shown, based on the 3D feature H 3D Generate the weight matrix W for dynamic convolution operation on 3D features 2DAlternatively, when generating the corresponding prediction result based on the second modal feature, based on the 2D feature H 2D Generate the weight matrix W for dynamic convolution operation on 3D features 3D .

[0124] Then, as shown in operation S820, based on the weight matrix, a dynamic convolution operation is performed on the features of each point in the first modal features to obtain a semantic feature for each point, and based on the semantic feature of each point, the normalization layer is used to obtain a prediction result for each point.

[0125] For example, the semantic features of each point i It can be obtained by the following convolution operation (for the sake of convenience, 2D is regarded as a superscript, which actually has the same meaning as the subscript in the previous text):

[0126]

[0127] Among them, W i 2D is the weight sub-matrix in the weight matrix used to perform dynamic convolution on the point i, is the feature of point i in the first modal feature.

[0128] Then, a normalization layer (e.g., softmax function) is used to normalize the semantic features of each point i. Processing is performed to obtain the prediction result P for each point i i 2D (the probability of belonging to each semantic category).

[0129] Since the weight matrix (convolution kernel) corresponding to the feature of each sampling point is different, the convolution operation on the features of different sampling points using different weight sub-matrices can be considered as a dynamic convolution operation. In addition, since it is well known in the art that the role of convolution operation can be regarded as filtering, the process of dynamically convolving the weight matrix generated by the second modal feature with the first modal feature can be considered as a dynamic cross-modal filtering process.

[0130] Finally, as shown in operation S830, a first prediction result is obtained based on the prediction result of each point.

[0131] For example, the prediction result P obtained for each point i in the first modal feature is i 2D Put together, we can form the final semantic segmentation result P predicted by 2D features. 2D .

[0132] For example, Figure 8B As shown, the weight matrix W 2D and 2D feature H 2DPerform dynamic convolution operation (calculate the features of each sampling point separately and combine them into the final result) to obtain the semantic segmentation result P 2D .

[0133] In this way, since the information of the other modality is introduced when determining the prediction results of each modality (based on the weight matrix) and the "convolution kernel-feature" relationship is established according to the position of the sampling point, this scheme can achieve better first modality-second modality feature matching, so that when each modality is predicted, the feature information of the other modality can be combined, thereby obtaining a more accurate semantic segmentation result.

[0134] It can be seen that through the above-mentioned method of training the semantic segmentation model of the present application, by alternately training the model with the source domain sample set and the target domain sample set, the domain adaptation effect of the model can be improved, and through the cross-modal processing process (for example, by imitating the prediction results of another modality through the imitation network and using the recovery network to restore the original data of the other modality, and calculating the corresponding loss, and adjusting the parameters of the model based on these losses), the model can learn more about the correspondence between the first modality and the second modality, so that the complementarity between the modalities can be better utilized, and the domain adaptation effect of the model can be further improved. In addition, by masking some samples and performing cross-modal processing based on a similar process, the model can learn more about the correspondence between the first modality and the second modality, and can also use the dynamic cross-modal filter as a prediction network to add information about the features of another modality in the prediction process, dynamically perform cross-modal feature matching, and thus better utilize the complementarity between the modalities, so that the trained model can achieve better semantic segmentation in practical applications even when the domain gap or domain offset is large.

[0135] According to another aspect of the present application, a semantic segmentation method is also provided.

[0136] Fig. 9 A flowchart of a semantic segmentation method according to an embodiment of the present application is shown.

[0137] like Fig. 9 As shown, in step S910, first modality data and second modality data for the same scene are acquired, wherein the second modality data includes data of multiple points.

[0138] Optionally, the first modality data may be 2D image data, and the second modality data may be 3D point cloud data. The targeted scene may be a current driving scene.

[0139] In step S920, a semantic segmentation model is used to determine a first prediction result for the first modality data and a second prediction result for the second modality data, wherein the first prediction result and the second prediction result respectively indicate the probability that each point in the second modality data belongs to each semantic category.

[0140] Optionally, the semantic segmentation model can be based on the previous reference Figures 4A-8B Therefore, more details of the semantic segmentation model can be referred to the content described above, so it will not be repeated here.

[0141] For example, after acquiring the first modality data and the second modality data, the first modality feature (for example, 2D feature) (also including the sampling process of point features) and the second modality feature (for example, 3D feature) can be obtained respectively through feature extraction networks (such as the first feature extraction network and the second feature extraction network mentioned above), and then the first prediction result and the second prediction result are obtained based on the first modality feature and the second modality feature.

[0142] In step S930, for each point in the second modal data, the maximum probability corresponding to each semantic category is determined as the first probability based on the first prediction result, and the maximum probability of each semantic category is determined as the second probability based on the second prediction result, and the semantic category corresponding to the larger one of the first probability and the second probability is taken as the semantic category to which the point belongs.

[0143] For example, for point i, the first prediction result indicates that the probabilities that point i belongs to the five semantic categories are (0.05, 0.15, 0.10, 0.70, 0.10) respectively, then the first prediction result indicates that the probability that the candidate category of point i is the fourth category is 0.7, and the second prediction result indicates that the probabilities that point i belongs to the five semantic categories are (0.15, 0.10, 0.35, 0.30, 0.10) respectively, then the second prediction indicates that the probability that the candidate category of point i is the third category is 0.35. Since 0.7 is greater than 0.35, the semantic category of the point is finally determined to be the fourth category.

[0144] In this way, since the semantic segmentation model is based on the previous reference Figures 4A-8B The training method described in the previous section is used to train the image, so its application in the target domain can achieve better semantic segmentation results.

[0145] The following combination Figures 10A-10E A schematic diagram is provided to illustrate the effect of the semantic segmentation model obtained by training using the training method of the present application.

[0146] First, assume that the data in the source domain comes from road data in developed countries, and the data in the target domain comes from road data in underdeveloped countries. Fig. 10AAs shown in the figure, the 2D image data and 3D point cloud data obtained when the target domain sample or the actual target domain application is input to the model are the same as the previous Figure 3 In the left figure, the vertical poles in the 2D prediction results corresponding to the input 2D image data are mostly identified (shown in black), and the vehicle is also mostly identified in the 3D prediction results corresponding to the 3D point cloud data (shown in gray). In the right figure, the vertical poles and vehicles in the input 3D point cloud data are also almost completely identified in the two output prediction results.

[0147] Then, assume that the data in the source domain comes from road data in developed countries, and the data in the target domain comes from road data in underdeveloped countries. Fig. 10B As shown, the upper left picture is the 2D image data of the parking lot, the upper right picture is the corresponding 3D point cloud data, the lower left picture is the 2D prediction result based on the 2D image data, and the lower left picture is the 3D prediction result based on the 3D point cloud data. Objects of different semantic categories identified will be represented by different colors, for example, green represents the natural environment, red represents the parking lot, blue represents the vehicle, and yellow represents the building. It can be seen that the prediction results corresponding to the 2D image data and the 3D point cloud data are mostly correct, and the two prediction results are relatively close, which also proves that the modal complementarity is well utilized.

[0148] Finally, if Figures 10C-10E As shown, three different domain adaptation situations are shown. Figure 2-3 Related models and Figures 4A-8B A comparison chart of the effects of the improved models in .

[0149] Fig. 10C The first row shows the 2D image data and its corresponding semantic labels, with different colors showing different semantic categories. The semantic labels are marked for most points.

[0150] Figure 2-3 The 2D prediction results corresponding to the relevant models only predict part of the natural environment and mistakenly predict the natural environment as objects of other categories (as shown in the box), and the 3D prediction results do not predict the natural environment (as shown in the box). Figures 4A-8B The model trained with the method correctly predicted the natural environment.

[0151] Fig. 10D The first row shows the 2D image data and its corresponding semantic labels, with different colors showing different semantic categories. The semantic labels are marked for vehicles, pedestrians and road posts.

[0152] Figure 2-3 The 3D prediction results corresponding to the relevant models do not predict pedestrians. Figures 4A-8B The model trained with the method correctly predicted pedestrians.

[0153] Fig.10E This is for the case of domain adaptation from daytime road data to nighttime road data. The first row shows 2D image data and its corresponding semantic labels, with vehicles shown in black.

[0154] Figure 2-3 The 2D prediction results and 3D results of the relevant models did not predict the oncoming vehicle, and incorrectly predicted the vertical poles (on the left and between the vehicles). Figures 4A-8B The model trained by this method correctly predicted the vehicle and did not misidentify other objects that did not need to be identified.

[0155] It can be seen that compared with Figure 2-3 Related models, when the domain gap or domain offset is too large, are based on Figures 4A-8B The model trained by this method can improve the accuracy of the model's semantic segmentation.

[0156] According to another aspect of the present application, a training device for a model for semantic segmentation is also provided.

[0157] Fig.11 A structural block diagram of a training device for a model for semantic segmentation according to an embodiment of the present application is shown.

[0158] like Fig.11 As shown, the training device 1100 includes an acquisition module 1110 , a result generation module 1120 and a training module 1130 .

[0159] The acquisition module 1110 is used to perform an operation of acquiring a training sample set, that is, to acquire a training sample set including a source domain sample set and a target domain sample set.

[0160] Each sample in the source domain sample set includes first modality data, second modality data, and a semantic label associated with the second modality, and each sample in at least a subset of the target domain sample set includes first modality data, second modality data, and does not include the semantic label;

[0161] The result generation module 1120 is used to generate results corresponding to the first modality and the second modality for each sample of the training sample set. For example, it is used to generate a first modality feature based on the first modality data of the sample, and to generate a second modality feature based on the second modality data of the sample, wherein the second modality data of the sample includes data of multiple points; based on the first modality feature and the second modality feature, a prediction result, an imitation result, and a recovery result corresponding to the first modality and the second modality are generated, respectively, wherein the prediction result corresponding to each modality indicates the prediction probability that each of the multiple points obtained based on the features of the current modality belongs to each semantic category, the imitation result corresponding to each modality indicates the prediction probability that each point obtained by imitating the features of another modality based on the features of the current modality belongs to each semantic category, and the recovery result corresponding to each modality indicates the recovery data of another modality obtained using the features of the current modality.

[0162] The training module 1130 is used to train the model based on cross-modal processing between prediction results, simulation results and recovery results corresponding to the first modality and the second modality respectively for each sample in the training sample set.

[0163] Optionally, the training device 1100 may also include a mask processing module 1140, which may be used to selectively perform mask processing on the first modality data and the second modality data of each sample in the training sample set according to a predetermined standard, and obtain the first modality mask data and the second modality mask data corresponding to at least one sample. In this way, when generating the first modality feature and the second modality feature, the result generation module 1120 may specifically generate the first modality feature based on the first modality data or the first modality mask data of the sample, and generate the second modality feature based on the second modality data or the second modality mask data of the sample, and the sample on which at least one of the first modality feature and the second modality feature of the sample is based is not mask data.

[0164] More details of each module have been referenced above. Figure 4A-9 It has been described in detail, so it will not be repeated here.

[0165] It can be seen that through the training device of the above-mentioned semantic segmentation model of the present application, by using the source domain sample set and the target domain sample set to train the model, the domain adaptation effect of the model can be improved, and through the cross-modal processing process, the model can better learn the correspondence between the first modality and the second modality, so that the complementarity between the modalities can be better utilized, and the domain adaptation effect of the model can be further improved. In addition, by masking some samples and performing cross-modal processing based on a similar process, the model can better learn the correspondence between the first modality and the second modality, and can also use the dynamic cross-modal filter as a prediction network to add information of the features of another modality in the prediction process, and dynamically perform cross-modal feature matching to better utilize the complementarity between the modalities, so that the trained model can achieve better semantic segmentation in practical applications even when the domain gap or domain offset is large.

[0166] According to another aspect of the present application, a training device for a model for semantic segmentation is also provided.

[0167] Fig.12 A structural block diagram of a device for semantic segmentation according to an embodiment of the present application is shown.

[0168] like Fig.12 As shown, the training device 1200 includes an acquisition module 1210 , a prediction module 1220 and a determination module 1230 .

[0169] The acquisition module 1210 is used to acquire first modality data and second modality data for the same scene, where the second modality data includes data of multiple points.

[0170] The prediction module 1220 is used to determine a first prediction result for the first modal data and a second prediction result for the second modal data using a model for semantic segmentation, wherein the first prediction result and the second prediction result respectively indicate the probability that each point in the second modal data belongs to each semantic category, wherein the model is based on a reference Figures 4A-8B Trained by the method described.

[0171] Determination module 1230 is used to determine, for each point in the second modal data, the maximum probability corresponding to each semantic category based on the first prediction result as the first probability, determine the maximum probability of each semantic category based on the second prediction result as the second probability, and take the semantic category corresponding to the larger one of the first probability and the second probability as the semantic category to which the point belongs.

[0172] More details of each module have been described in detail above, so they will not be repeated here.

[0173] Through such a device, since the semantic segmentation model is obtained by the previous reference Figures 4A-8B The training method described in the previous section is used to train the image, so its application in the target domain can achieve better semantic segmentation results.

[0174] In addition, although Figure 11-12 The above modules are shown in an exemplary manner, but it should be understood that the training device 1100 and the device 1200 can be divided into more or fewer modules according to different functions, or each module can be divided into further sub-modules. In some example embodiments, the module or its sub-module can be implemented by electronic hardware (e.g., a general-purpose processor, DSP, ASIC, FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, etc.), computer software (e.g., can be stored in random access memory (RAM), flash memory, read-only memory (ROM), erasable programmable ROM (EPROM), etc.), or a combination of the two.

[0175] Fig.13 A schematic block diagram of a computing device 1300 according to an embodiment of the present application is shown.

[0176] like Fig.13 As shown, the computing device 1300 includes one or more processors, one or more memories, a network interface, an input device, and a display screen connected via a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the terminal stores an operating system and may also store a computer executable program, which, when executed by the processor, enables the processor to implement the various operations described in the various steps of the training method and the semantic segmentation method described above. The internal memory may also store a computer executable program, which, when executed by the processor, enables the processor to perform the various operations described in the various steps of the training method and the semantic segmentation method.

[0177] The processor can be an integrated circuit chip with signal processing capabilities. The above processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The disclosed methods, steps and logic block diagrams in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., which can be an X84 architecture or an ARM architecture.

[0178] The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable categories of memory.

[0179] The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the terminal housing, or an external keyboard, touchpad or mouse.

[0180] The electronic device may be a terminal or a server. The terminal may include but is not limited to: a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart TV, etc.; a variety of clients (applications, APPs) may run in the terminal, such as a multimedia player client, a social client, a browser client, an information stream client, an education client, etc. The server may be a reference Figure 2 The server described can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.

[0181] According to another aspect of the present application, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the processor executes the steps of the training method and the semantic segmentation method as described above.

[0182] According to another aspect of the present application, a computer program product is also provided, including a computer program, which implements the steps of the training method and semantic segmentation method as described above when the computer program is executed by a processor.

[0183] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the methods and devices according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a part of a code, and the module, program segment, or a part of the code contains at least one executable instruction for realizing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0184] The exemplary embodiments of the present application described in detail above are merely illustrative and not restrictive. It should be understood by those skilled in the art that various modifications and combinations may be made to these embodiments or their features without departing from the principles and spirit of the present application, and such modifications should fall within the scope of the present application.

Claims

1. A training method for a model for semantic segmentation, include: Acquire a training sample set including a source domain sample set and a target domain sample set, wherein each sample in the source domain sample set includes first modality data, second modality data, and a semantic label associated with the second modality, and each sample in at least a subset of the target domain sample set includes first modality data, second modality data, and does not include the semantic label; For each sample in the training sample set: generating a first modal feature based on the first modal data of the sample, and generating a second modal feature based on the second modal data of the sample, wherein the second modal data of the sample includes data of a plurality of points; as well as Based on the first modality features and the second modality features, generating prediction results, imitation results and recovery results corresponding to the first modality and the second modality respectively, wherein the prediction result corresponding to each modality indicates the semantic category prediction probability of each of the multiple points obtained based on the features of the current modality, the imitation result corresponding to each modality indicates the semantic category prediction probability of each of the multiple points obtained by imitating the features of another modality based on the features of the current modality, and the recovery result corresponding to each modality indicates the recovery data of another modality obtained by predicting the data of another modality by using the features of the current modality; as well as The model is trained based on cross-modal processing between prediction results, simulation results and recovery results corresponding to the first modality and the second modality respectively for each sample in the training sample set.

2. The training method according to claim 1, further comprising: include: According to a predetermined standard, selectively masking the first modality data and the second modality data of each sample in the training sample set to obtain the first modality mask data and the second modality mask data corresponding to at least one sample; The step of generating a first modal feature based on the first modal data of the sample and generating a second modal feature based on the second modal data of the sample comprises: The first modal feature is generated based on the first modal data or the first modal mask data of the sample, and the second modal feature is generated based on the second modal data or the second modal mask data of the sample, and the sample on which at least one of the first modal feature and the second modal feature of the sample is based is not mask data.

3. The training method according to claim 2, in, The predetermined criterion includes: the probability that the first modality data of each sample in the training sample set is subjected to masking is a first predetermined probability, and the probability that the second modality data of each sample in the training sample set is subjected to masking is a second predetermined probability.

4. The training method according to claim 2, in, The method selectively performs mask processing on the first modality data and the second modality data of each sample in the training sample set according to a predetermined standard, including: Dividing at least one of the first modality data and the second modality data of each sample selected for mask processing into a plurality of blocks having a predetermined size, respectively; and Based on a predetermined removal ratio, a removal operation is performed on the plurality of blocks of a predetermined size obtained by dividing each modality data to implement mask processing.

5. The training method according to any one of claims 1 to 4, in, Based on the first modality feature and the second modality feature, generating prediction results, simulation results, and recovery results corresponding to the first modality and the second modality respectively, including: Based on the first modality feature, a first prediction result corresponding to the first modality is generated using a first prediction network of the model, and based on the second modality feature, a second prediction result corresponding to the second modality is generated using a second prediction network of the model; Based on the first modality feature, a first imitation network of the model is used to generate a first imitation result corresponding to the first modality, and based on the second modality feature, a second imitation network of the model is used to generate a second imitation result corresponding to the second modality; and Based on the first modal feature, a first restoration result corresponding to the first modality is generated using a first restoration network of the model, and based on the second modal feature, a second restoration result corresponding to the second modality is generated using a second restoration network of the model.

6. The training method according to claim 5, in, The model is trained based on cross-modal processing between prediction results, simulation results, and recovery results corresponding to the first modality and the second modality for each sample in the training sample set, including: For each sample of the training sample set: when the sample does not include the semantic label, a first loss between the first prediction result and the second imitation result, a second loss between the second prediction result and the first imitation result, a third loss between the first recovery result and the second modality data of the sample, and a fourth loss between the second recovery result and the sampling data of the first modality data of the sample are calculated; and when the sample includes the semantic label, a fifth loss between the first prediction result and the semantic label of the sample and a sixth loss between the second prediction result and the semantic label of the sample are further calculated; and The model is trained based on each loss calculated for each sample of the training sample set.

7. The training method according to claim 6, in, The sampling data of the first modality data of each sample in the training sample set is obtained in the following manner: Projecting the second modality data of the sample according to the first modality to obtain second modality-first modality projection data, wherein the second modality-first modality projection data includes projection data of the plurality of points; The first modality data of the sample is sampled according to the positions of the multiple points in the first modality projection data to obtain sampled data of the first modality data.

8. The training method according to claim 6, in, Calculate a first loss between the first prediction result and the second simulation result and a second loss between the second prediction result and the first simulation result using a KL divergence loss function; Calculating a third loss between the first restored result and the second modal data and a fourth loss between the second restored result and the sampled data of the first modal data by using an L2 loss function; A fifth loss between the first prediction result and the semantic label of the sample and a sixth loss between the second prediction result and the semantic label of the sample are calculated using a cross entropy loss function.

9. The training method according to claim 5, in, Each of the first prediction network and the second prediction network includes a dynamic convolution layer and a normalization layer; Each of the first and second mimetic networks includes a linear layer and a normalization layer; and Each of the first restoration network and the second restoration network includes a multilayer perceptron.

10. The training method according to claim 9, in, Based on the first modal feature, generating a first prediction result corresponding to the first modality using a first prediction network of the model includes: Generating a first weight matrix for performing a dynamic convolution operation on the first modal feature based on the second modal feature; Performing a dynamic convolution operation on the features of each point in the first modal features based on the first weight matrix to obtain a semantic feature of each point, and obtaining a semantic category prediction result of each point using the normalization layer based on the semantic feature of each point; and The first prediction result is obtained based on the semantic category prediction result of each point.

11. The training method according to claim 9, in, Based on the second modality feature, generating a second prediction result corresponding to the second modality using a second prediction network of the model includes: Generating a second weight matrix for performing a dynamic convolution operation on the second modal feature based on the first modal feature; Based on the second weight matrix, a dynamic convolution operation is performed on the feature of each point in the second modal feature to obtain a semantic feature of each point, and based on the semantic feature of each point, a semantic category prediction result of each point is obtained by using the normalization layer; and The second prediction result is obtained based on the semantic category prediction result of each point.

12. The training method according to claim 1 or 2, in, The second modality data or the second modality mask data of each sample of the training sample set includes data of the plurality of points, and Generating a first modal feature based on the first modal data of the sample, and generating a second modal feature based on the second modal data of the sample, comprising: Using a first feature extraction network to perform feature extraction on the first modal data or the first modal mask data to obtain a first feature map; Projecting the second modality data or the second modality mask data according to the first modality to obtain second modality-first modality projection data, wherein the second modality-first modality projection data includes projection data of the plurality of points; Sampling the first feature map according to the positions of the plurality of points in the second modality-first modality projection data to obtain the first modality feature; and Using a second feature extraction network to extract features from the second modality data or the second modality mask data to obtain the second modality features, Wherein, the first modal feature and the second modal feature include features of the plurality of points.

13. The training method according to claim 1, in, The source domain sample set is divided into different source domain subsets, and the target domain sample is divided into different target domain subsets. Different source domain subsets of the source domain sample set and different target domain subsets of the target domain sample set are alternately used to train the model.

14. The training method according to claim 1, in, The first modality data is 2D image data, and the second modality data is 3D point cloud data.

15. A semantic segmentation method, include: Acquire first modality data and second modality data for the same scene, where the second modality data includes data of multiple points; Determine a first prediction result for the first modality data and a second prediction result for the second modality data using a semantic segmentation model, wherein the first prediction result indicates a semantic category prediction probability for each point in the first modality data, and the second prediction result indicates a semantic category prediction probability for each point in the second modality data, wherein the semantic segmentation model is trained according to the method according to any one of claims 1 to 14; as well as For each point in the second modal data, the maximum probability corresponding to each semantic category is determined as the first probability based on the first prediction result, and the maximum probability of each semantic category is determined as the second probability based on the second prediction result, and the semantic category corresponding to the larger one of the first probability and the second probability is taken as the semantic category to which the point belongs.

16. A training device for a model for semantic segmentation, include: an acquisition module, configured to acquire a training sample set including a source domain sample set and a target domain sample set, wherein each sample in the source domain sample set includes first modality data, second modality data, and a semantic label associated with the second modality, and each sample in at least a subset of the target domain sample set includes first modality data, second modality data, and does not include the semantic label; A result generation module is used for each sample in the training sample set: generating a first modal feature based on first modal data of the sample, and generating a second modal feature based on second modal data of the sample, wherein the second modal data includes data of a plurality of points; Based on the first modality feature and the second modality feature, generating a prediction result, an imitation result and a recovery result corresponding to the first modality and the second modality respectively, wherein the prediction result corresponding to each modality indicates a semantic category prediction probability of each of the multiple points obtained based on the feature of the current modality, the imitation result corresponding to each modality indicates a semantic category prediction probability of each of the multiple points obtained by imitating the feature of another modality based on the feature of the current modality, and the recovery result corresponding to each modality indicates the recovery data of another modality obtained by predicting the data of another modality by using the feature of the current modality; A training module is used to train the model based on cross-modal processing between prediction results, simulation results and recovery results corresponding to the first modality and the second modality respectively for each sample in the training sample set.

17. A computing device, include: processor; as well as A memory having a computer program stored thereon, wherein when the computer program is executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 15.

18. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 15.