3D Scene Segmentation Domain Migration Method and Device Based on Multi-Source Heterogeneous Data Fusion

By combining LiDAR point cloud data and camera image data, cross-modal comparison learning and supervised learning methods are used to solve the problem of multimodal data complementarity in the existing technology, and the performance of the three-dimensional scene segmentation model on the target domain is significantly improved.

CN116246070BActive Publication Date: 2025-06-24PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310092486.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-29
Publication Date
2025-06-24
Estimated Expiration
2043-01-29

AI Technical Summary

Technical Problem

The existing three-dimensional scene segmentation domain migration method ignores the use of multimodal data and the complementarity of data between modes, resulting in limited domain migration effect.

Method used

By combining data from different sensors and modalities (such as LiDAR point cloud data and camera image data), using complementary information between multimodal data, and using cross-modal contrast learning and supervised learning methods, the migration of the three-dimensional scene segmentation model from the source domain to the target domain is realized.

Benefits of technology

By leveraging the complementarity of multimodal data, the learning and representation of features are enhanced, and the performance of the model on the target domain is significantly improved, so that the good segmentation effect is achieved without the need for expensive labeled data on the target domain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116246070B_ABST
    Figure CN116246070B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a three-dimensional scene segmentation domain transfer method and apparatus based on multi-source heterogeneous data fusion. The method includes: obtaining multi-modal data of a source domain and a target domain and segmentation labels of source domain three-dimensional point cloud data, and extracting features of the multi-modal data; based on the mapping relationship between two-dimensional image data and three-dimensional point cloud data, dividing source domain pixel features and target domain pixel data into pixel features with a mapping relationship and pixel features without a mapping relationship respectively; training a semantic segmentation model through comprehensive cross-modal contrast learning and supervised learning to obtain a trained semantic segmentation model; inputting the three-dimensional point cloud data to be segmented in the target domain into the trained semantic segmentation model to obtain a three-dimensional scene segmentation domain transfer result. The present disclosure realizes the transfer of the semantic segmentation model from the source domain to the target domain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the domain transfer technology for 3D scene segmentation, and particularly to a 3D scene segmentation domain transfer method and device based on multi-source heterogeneous data fusion. Background Art

[0002] 3D scenes are usually represented in the form of LiDAR point clouds scanned by lidar. The goal of the 3D scene segmentation task is to predict the class label of each point in the given LiDAR point cloud for better understanding of the 3D scene. The 3D scene segmentation task is very necessary in many applications such as robotics, autonomous driving, and virtual reality. With the development of deep learning technology, a well-performing 3D scene segmentation model can be obtained by training on a large amount of labeled data.

[0003] However, since the data during training and testing may come from different domains (referred to as the source domain and the target domain respectively), such as different environments, different countries, or different datasets, and there are often significant distribution differences between the data in different domains, the performance of the model trained on the source domain will drop significantly when tested on the target domain. And it is difficult to directly train on the target domain due to the high cost of annotation and the lack of labeled data. To address this problem, researchers have proposed domain transfer methods, hoping to use the labeled source domain data and unlabeled target domain data for training, so as to achieve the transfer of the model from the source domain to the target domain and make the trained model also achieve good performance on the target domain.

[0004] Existing 3D scene segmentation domain transfer methods usually adopt a metric learning-based or adversarial learning-based approach, hoping to learn invariant features between different domains to narrow the difference between the source domain and the target domain. However, these methods only focus on the data of a single LiDAR point cloud modality and ignore the use of multi-modal data and the complementarity between modal data, resulting in limited domain transfer effects. Summary of the Invention

[0005] To overcome the deficiencies of the prior art, the present invention proposes a 3D scene segmentation domain transfer method and device based on multi-source heterogeneous data fusion. This method combines data from different sensors and different modalities (LiDAR point cloud data and camera image data in the present invention), and utilizes the complementary information between multi-modal data to achieve the transfer of the 3D scene segmentation model from the source domain to the target domain and improve the domain transfer effect.

[0006] The technical content of the present invention includes:

[0007] A 3D scene segmentation domain transfer method based on multi-source heterogeneous data fusion, the method comprising:

[0008] Obtain the multi-modal data of the source domain and the target domain, as well as the segmentation labels of the three-dimensional point cloud data in the source domain, and extract the features of the multi-modal data; the multi-modal data includes: three-dimensional point cloud data and two-dimensional image data, and the features of the multi-modal data include: source domain point features, source domain pixel features, target domain point features, and target domain pixel data;

[0009] Based on the mapping relationship between the two-dimensional image data and the three-dimensional point cloud data, divide the source domain pixel features and the target domain pixel data into pixel features with a mapping relationship and pixel features without a mapping relationship respectively;

[0010] Train the semantic segmentation model through comprehensive cross-modal contrast learning and supervised learning to obtain the trained semantic segmentation model; among them, the cross-modal contrast learning refers to increasing the feature representation similarity of point features and pixel features with a mapping relationship between different modalities in the source domain and the target domain, and reducing the feature representation similarity of point features and pixel features without a mapping relationship between different modalities in this feature space; the supervised learning refers to using the segmentation labels of the three-dimensional point cloud data in the source domain for supervised training;

[0011] Input the three-dimensional point cloud data to be segmented in the target domain into the trained semantic segmentation model to obtain the three-dimensional scene segmentation domain transfer result.

[0012] Further, the extraction of the features of the three-dimensional point cloud data in the source domain includes:

[0013] Train the SparseConvNet network based on the first training dataset to obtain a three-dimensional feature extraction model; where the first training dataset consists of several three-dimensional point cloud sample data;

[0014] Input the three-dimensional point cloud data in the source domain into the three-dimensional feature extraction model to obtain the source domain point features.

[0015] Further, the extraction of the features of the two-dimensional image data in the source domain includes:

[0016] Train the ResNet network based on the second training dataset to obtain a two-dimensional feature extraction model; where the second training dataset consists of several image sample data;

[0017] Input the two-dimensional image data in the source domain into the two-dimensional feature extraction model to obtain the source domain pixel features.

[0018] Further, the division of the source domain two-dimensional image features into pixel features with a mapping relationship and pixel features without a mapping relationship based on the mapping relationship between the two-dimensional image data and the three-dimensional point cloud data includes:

[0019] Using the calibration data of the camera, map the three-dimensional point cloud data of the source domain modality onto the two-dimensional image data of the source domain modality;

[0020] Based on the mapping results of the three-dimensional point cloud data and the two-dimensional image data of the source domain modality, divide the source domain two-dimensional image features into pixel features with mapping relationships and pixel features without mapping relationships.

[0021] Furthermore, in the source domain, the loss for cross-modal contrastive learning training of the semantic segmentation model where N represents the number of source domain point features, represents the source domain pixel features with mapping relationships, represents the source domain point features, represents the source domain pixel features and the distance between the source domain point features τ represents the parameter coefficient.

[0022] Furthermore, the loss for supervised learning training of the semantic segmentation model where, represents the true class label of the two-dimensional pixels in the source domain, represents the true class label of the three-dimensional points in the source domain, represents the predicted class label of the two-dimensional pixels in the source domain, represents the predicted class label of the three-dimensional points in the source domain.

[0023] Furthermore, training the semantic segmentation model by comprehensively combining cross-modal contrastive learning and supervised learning to obtain the trained semantic segmentation model, including:

[0024] Calculate the total training loss where, represents the supervised learning loss, represents the cross-modal contrastive learning loss in the source domain, represents the cross-modal contrastive learning loss in the target domain, and α represents the weight coefficient;

[0025] Use the Adam optimizer to optimize the parameters of the semantic segmentation model until the total training loss L during the training process total converges, thereby obtaining the trained semantic segmentation model.

[0026] A three-dimensional scene segmentation domain transfer device based on multi-source heterogeneous data fusion, the device includes:

[0027] A feature extraction module, which is used to obtain multi-modal data of the source domain and the target domain and the segmentation labels of the source domain 3D point cloud data, and extract the features of the multi-modal data; the multi-modal data includes: 3D point cloud data and 2D image data, and the features of the multi-modal data include: source domain point features, source domain pixel features, target domain point features, and target domain pixel data;

[0028] A feature alignment module, which divides the source domain pixel features and the target domain pixel data into pixel features with a mapping relationship and pixel features without a mapping relationship respectively based on the mapping relationship between the 2D image data and the 3D point cloud data;

[0029] A model training module, which is used to train a semantic segmentation model through comprehensive cross-modal contrast learning and supervised learning to obtain a trained semantic segmentation model; wherein, the cross-modal contrast learning means that in the source domain and the target domain, the feature representation similarity between point features and pixel features with a mapping relationship in different modalities is increased, and the feature representation similarity between point features and pixel features without a mapping relationship in different modalities in this feature space is reduced; the supervised learning means using the segmentation labels of the source domain 3D point cloud data for supervised training;

[0030] A semantic segmentation module, which is used to input the 3D point cloud data to be segmented in the target domain into the trained semantic segmentation model to obtain a 3D scene segmentation domain migration result.

[0031] An electronic device, the electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the above-mentioned 3D scene segmentation domain migration method based on multi-source heterogeneous data fusion is implemented.

[0032] A computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above-mentioned 3D scene segmentation domain migration method based on multi-source heterogeneous data fusion is implemented.

[0033] Compared with the prior art, the present invention utilizes cross-modal contrast learning and the complementarity of multi-modal data to strengthen the learning and representation of features, realizes the migration of the model from the source domain to the target domain, so that good segmentation effects can be obtained without using costly labeled data for training in the target domain, which is of great help for the understanding of real scenes. Brief Description of the Drawings

[0034] Figure 1 It is a flowchart of the 3D scene segmentation domain migration method based on multi-source heterogeneous data fusion provided by the present invention.

[0035] Figure 2This is the mapping relationship between multi-source heterogeneous data in the present invention.

[0036] Figure 3 This is the visualization of the scene segmentation result of the present invention in the target domain. Detailed implementation manners

[0037] The following further describes the present invention by way of embodiments in conjunction with the accompanying drawings, but does not limit the scope of the present invention in any way.

[0038] The three-dimensional scene segmentation domain migration method of the present invention uses two-dimensional image data and three-dimensional point cloud data as inputs simultaneously in the source domain and the target domain. After obtaining the corresponding two-dimensional pixel features and three-dimensional point features, cross-modal information fusion is promoted in the feature space in a contrastive learning manner. By using complementary two-dimensional image information, the domain difference between three-dimensional LiDAR point clouds is reduced, thereby effectively improving the domain migration effect of the model.

[0039] Specifically, as Figure 1 shown, the present invention includes the following steps 1-step 5.

[0040] Step 1: Use a deep learning network model to extract features from the input multi-modal data, and obtain two-dimensional pixel features and three-dimensional point features respectively.

[0041] In both the source domain / target domain, multi-modal data is used as the input, and the multi-modal data includes three-dimensional LiDAR point cloud data and two-dimensional image data. For the input multi-modal data, the Resnet network is used to process the image to obtain two-dimensional features, and the SparseConvNet network is used to process the point cloud to obtain three-dimensional features.

[0042] In one example, the dimensions of the two-dimensional pixel features and the three-dimensional point features are both 128 dimensions.

[0043] Step 2: By using the calibration data of the camera, align the data from different modalities and unify the number of features.

[0044] Since the input data comes from different modalities and they are not completely corresponding, and the number of pixels in the image is much larger than the number of points in the point cloud, it is necessary to align the data.

[0045] As Figure 2As shown in the figure, the camera calibration data can be used to map the point cloud to the image, and then find the pixel corresponding to each point, so as to obtain the mapping relationship from the point cloud to the image. Since the number of points in the point cloud is not equal to the number of pixels in the image (the number of pixels is much greater than the number of points), at this stage, this mapping relationship can be used to retain only the pixels with corresponding points in the image, thereby achieving the unification of the number of three-dimensional features and the number of two-dimensional features. After this stage of processing, the number of three-dimensional point features and the number of two-dimensional pixel features are consistent and have a one-to-one correspondence, which is convenient for the subsequent fusion and processing.

[0046] Step 3: In the shared feature space, the one-to-one correspondence and contrastive learning between cross-modal data are used to shorten the distance between corresponding features, thereby strengthening feature representation.

[0047] After completing the data mapping in step 2, the one-to-one correspondence between three-dimensional features and two-dimensional features is used to constrain these features in the shared feature space using the contrastive learning method, so that the corresponding points and pixels between different modalities have similar feature representations in the feature space, and at the same time, the feature representations of samples that do not have a corresponding relationship are as far apart as possible, thereby narrowing the domain gap between three-dimensional modalities and strengthening the learning and representation of features. The formula for this step is as follows:

[0048]

[0049]

[0050] Among them, F 2D and F 3D They represent the image features and point cloud features obtained after processing the input data in the first stage, and They represent the one-to-one correspondence between two-dimensional pixel features and three-dimensional point features, N is the number of features in each mode, τ is the preset parameter coefficient, is the contrastive learning loss calculated around the i-th two-dimensional feature, is the contrastive learning loss calculated with the ith 3D feature as the center, L ctr (F 2D ,F 3D ) is the cross-modal contrastive learning loss obtained by considering all features, and the distance metric function The cosine function is the cosine distance function, defined as cos(u,v)=u T v / ∥u∥∥v∥,φ 2D and φ 3D It is a feature projection function, which is responsible for projecting two-dimensional features and three-dimensional features into a shared feature space for measurement.

[0051] Step 4: Use the contrastive learning loss function on the source domain and the target domain and the supervised learning loss function on the source domain to perform gradient backpropagation and optimization on the model.

[0052] Since cross-modal contrastive learning does not require sample labels, it can be processed on both the labeled source domain and the unlabeled target domain. In addition, on the source domain, supervised training can also be performed using sample labels, and the cross-entropy loss between the prediction result and the true result is calculated. The calculated cross-entropy loss is added to the contrastive learning loss in Step 3 as the total loss function, and the Adam optimizer is used to optimize the parameters of the model during training until the loss value converges during the training process.

[0053] In one embodiment, the total loss function in the model training process is represented as and are the contrastive learning losses calculated according to Step 3 in the source domain and the target domain respectively, is the cross-entropy loss between the prediction result obtained using the segmentation head in the source domain and the true result:

[0054]

[0055] Step 5: After obtaining the trained model, perform label prediction on the LiDAR point cloud in the target domain to complete domain transfer.

[0056] Using the semantic segmentation model trained in the previous stage, for the input LiDAR point cloud in the target domain, it can be processed and the segmentation result can be predicted to obtain the category to which each point belongs, as Figure 3 shown.

[0057] Among them, the semantic segmentation model used in the present invention can be constructed, trained and used to complete three-dimensional scene segmentation using the Pytorch library in Python.

[0058] Based on the same concept, the present invention also discloses a three-dimensional scene segmentation domain transfer device based on multi-source heterogeneous data fusion, and the device includes:

[0059] A feature extraction module, configured to obtain multi-modal data of the source domain and the target domain and the segmentation labels of the three-dimensional point cloud data of the source domain, and extract the features of the multi-modal data; the multi-modal data includes: three-dimensional point cloud data and two-dimensional image data, and the features of the multi-modal data include: source domain point features, source domain pixel features, target domain point features, and target domain pixel data;

[0060] A feature alignment module that divides the source domain pixel features and the target domain pixel data into pixel features with a mapping relationship and pixel features without a mapping relationship respectively based on the mapping relationship between two-dimensional image data and three-dimensional point cloud data;

[0061] A model training module for training a semantic segmentation model through comprehensive cross-modal contrast learning and supervised learning to obtain a trained semantic segmentation model; wherein, the cross-modal contrast learning means increasing the feature representation similarity of point features and pixel features with a mapping relationship between different modalities in the source domain and the target domain in the same feature space, and reducing the feature representation similarity of point features and pixel features without a mapping relationship between different modalities in this feature space; the supervised learning means using the segmentation labels of the source domain three-dimensional point cloud data for supervised training;

[0062] A semantic segmentation module for inputting the three-dimensional point cloud data to be segmented in the target domain into the trained semantic segmentation model to obtain a three-dimensional scene segmentation domain migration result.

[0063] The exemplary device is a device embodiment corresponding to the above exemplary method. The specific operations of each module can be understood with reference to the description of the method embodiment and will not be elaborated here.

[0064] Based on the same concept, the present invention also discloses an electronic device. The electronic device can be a computer device, a laptop computer, a server or other types of electronic devices.

[0065] An electronic device may include at least one processor and a memory. The processor can execute instructions stored in the memory. The processor is communicatively connected to the memory via a data bus. In addition to the memory, the processor can also be communicatively connected to an input device, an output device, and a communication device via the data bus.

[0066] The processor can be any conventional processor. The processor can include, for example, a Central Processing Unit (CPU), a Graphic Processing Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit (ASIC), or a combination thereof.

[0067] The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0068] In an embodiment of the present disclosure, executable instructions are stored in the memory, and the processor can read the executable instructions from the memory and execute the instructions to implement all or part of the steps of the vehicle maneuver smoothness evaluation method in the above exemplary embodiment.

[0069] Based on the same concept, the present invention also discloses a computer program product or a computer-readable storage medium storing the computer program product. The computer program product includes computer program instructions, and the computer program instructions can be executed by the processor to implement all or part of the steps described in the above exemplary embodiment.

[0070] The computer program product can be written in any combination of one or more programming languages for writing program code for performing the operations of the embodiments of the present application. The programming languages include object-oriented programming languages, such as Java, C++, etc., and also include conventional procedural programming languages, such as the "C" language or similar programming languages, as well as scripting languages (such as Python). The program code can be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0071] The computer-readable storage medium can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the readable storage medium include: static random access memory (SRAM) with one or more wire electrical connections, electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk, or any suitable combination of the above.

[0072] It should be noted that the purpose of disclosing the embodiments is to help further understand the present invention. However, those skilled in the art can understand that various substitutions and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the embodiments, and the scope of protection claimed by the present invention shall be defined by the scope defined in the claims.

Claims

1. A three-dimensional scene segmentation domain migration method based on multi-source heterogeneous data fusion, characterized in that The method includes: Obtaining multi-modal data of the source domain and the target domain, as well as the segmentation labels of the three-dimensional point cloud data in the source domain and the true class labels of the two-dimensional pixels in the source domain, and extracting the features of the multi-modal data; the multi-modal data includes: three-dimensional point cloud data and two-dimensional image data, and the features of the multi-modal data include: source domain point features, source domain pixel features, target domain point features, and target domain pixel data; Based on the mapping relationship between the two-dimensional image data and the three-dimensional point cloud data, dividing the source domain pixel features and the target domain pixel data into pixel features with a mapping relationship and pixel features without a mapping relationship respectively; Training a semantic segmentation model by integrating cross-modal contrastive learning and supervised learning to obtain a trained semantic segmentation model; wherein, the cross-modal contrastive learning refers to increasing the feature representation similarity of point features and pixel features with mapping relationships between different modalities in the source domain and the target domain in the same feature space, and reducing the feature representation similarity of point features and pixel features without mapping relationships between different modalities in this feature space; the supervised learning refers to using the segmentation labels of the source domain three-dimensional point cloud data and the true class labels of the two-dimensional pixels in the source domain for supervised training; the loss of the cross-modal contrastive learning F 2D and F 3D respectively represent the image features and point cloud features obtained after processing the two-dimensional image data and the three-dimensional point cloud data, N is the number of features in each modality, and respectively represent the corresponding two-dimensional pixel features and three-dimensional point features, τ is a preset parameter coefficient, is the contrastive learning loss calculated with the i-th two-dimensional feature as the center, is the contrastive learning loss calculated with the i-th three-dimensional feature as the center, the distance metric function The cos function is the cosine distance function, defined as cos(u,v) = u T v / ||u||||v||, φ 2D and φ 3D are feature projection functions, responsible for projecting two-dimensional features and three-dimensional features into a shared feature space for measurement; Inputting the three-dimensional point cloud data to be segmented in the target domain into the trained semantic segmentation model to obtain the three-dimensional scene segmentation domain transfer result.

2. The method according to claim 1, characterized in that, The extracting of the features of the three-dimensional point cloud data in the source domain includes: Training the SparseConvNet network based on the first training dataset to obtain a three-dimensional feature extraction model; wherein, the first training dataset is composed of a number of three-dimensional point cloud sample data; Inputting the three-dimensional point cloud data in the source domain into the three-dimensional feature extraction model to obtain source domain point features.

3. The method according to claim 1, characterized in that, The extracting of the features of the two-dimensional image data in the source domain includes: Training the ResNet network based on the second training dataset to obtain a two-dimensional feature extraction model; wherein, the second training dataset is composed of a number of image sample data; Inputting the two-dimensional image data in the source domain into the two-dimensional feature extraction model to obtain source domain pixel features.

4. The method according to claim 1, characterized in that The dividing of the source domain two-dimensional image features into pixel features with a mapping relationship and pixel features without a mapping relationship based on the mapping relationship between the two-dimensional image data and the three-dimensional point cloud data includes: Using the calibration data of the camera to map the three-dimensional point cloud data in the source domain modality onto the two-dimensional image data in the source domain modality; Based on the mapping result of the three-dimensional point cloud data and the two-dimensional image data in the source domain modality, dividing the source domain two-dimensional image features into pixel features with a mapping relationship and pixel features without a mapping relationship.

5. The method according to claim 1, wherein The loss for supervised learning training of the semantic segmentation model wherein represents the true class label of two-dimensional pixels in the source domain, represents the true class label of three-dimensional points in the source domain, represents the class label of the predicted two-dimensional pixels in the source domain, represents the class label of the predicted three-dimensional points in the source domain.

6. The method according to claim 1, wherein The training of the semantic segmentation model by integrating cross-modal contrast learning and supervised learning to obtain the trained semantic segmentation model includes: Calculate the total training loss Among them, represents the supervised learning loss, represents the cross-modal contrastive learning loss in the source domain, represents the cross-modal contrastive learning loss in the target domain, and α represents the weight coefficient; The parameters of the semantic segmentation model are optimized using the Adam optimizer until the total training loss L in the training process total converges, thereby obtaining the trained semantic segmentation model.

7. A three-dimensional scene segmentation domain migration device based on multi-source heterogeneous data fusion, characterized in that, The device includes: A feature extraction module, configured to obtain multi-modal data of the source domain and the target domain, as well as the segmentation labels of the three-dimensional point cloud data in the source domain and the true class labels of the two-dimensional pixels in the source domain, and extract the features of the multi-modal data; the multi-modal data includes: three-dimensional point cloud data and two-dimensional image data, and the features of the multi-modal data include: source domain point features, source domain pixel features, target domain point features, and target domain pixel data; A feature alignment module, configured to divide the source domain pixel features and the target domain pixel data into pixel features with a mapping relationship and pixel features without a mapping relationship respectively based on the mapping relationship between the two-dimensional image data and the three-dimensional point cloud data; A model training module for training a semantic segmentation model by integrating cross-modal contrast learning and supervised learning to obtain a trained semantic segmentation model; wherein, the cross-modal contrast learning means that in the source domain and the target domain, the feature representation similarity between point features of different modalities and pixel features with a mapping relationship in the same feature space is increased, and the feature representation similarity between point features of different modalities and pixel features without a mapping relationship in this feature space is reduced; the supervised learning means using the segmentation labels of the source domain three-dimensional point cloud data and the true category labels of the two-dimensional pixels in the source domain for supervised training; the loss of the cross-modal contrast learning F 2D and F 3D respectively represent the image features and point cloud features obtained after processing the two-dimensional image data and the three-dimensional point cloud data, N is the number of features in each modality, and respectively represent the corresponding two-dimensional pixel features and three-dimensional point features, τ is a preset parameter coefficient, is the contrast learning loss calculated with the i-th two-dimensional feature as the center, is the contrast learning loss calculated with the i-th three-dimensional feature as the center, and the distance metric function The cos function is the cosine distance function, defined as cos(u,v) = u T v / ||u||||v||, φ 2D and φ 3D are feature projection functions responsible for projecting two-dimensional features and three-dimensional features into a shared feature space for measurement; A semantic segmentation module, configured to input the three-dimensional point cloud data to be segmented in the target domain into the trained semantic segmentation model to obtain the three-dimensional scene segmentation domain transfer result.

8. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the three-dimensional scene segmentation domain migration method based on multi-source heterogeneous data fusion as described in any one of claims 1-6 is implemented.

9. A computer-readable storage medium, characterized in that, Computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the three-dimensional scene segmentation domain migration method based on multi-source heterogeneous data fusion as described in any one of claims 1-6 is implemented.