Three-dimensional scene semantic segmentation and generation method based on semi-supervised diffusion model

By combining semi-supervised diffusion model with labeled and unlabeled point cloud data training, the problem of low model accuracy in three-dimensional semantic segmentation is solved, and higher semantic segmentation accuracy is achieved.

CN120339757AInactive Publication Date: 2025-07-18CHINA COAL RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510821007.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing semi-supervised learning methods have low model accuracy in three-dimensional semantic segmentation and insufficient semantic segmentation accuracy.

Method used

The semi-supervised diffusion model is adopted, and the supervised branch is trained with labeled point cloud data, and the unsupervised branch is trained with unlabeled point cloud data. The diffusion model is combined to generate pseudo-labels and enhance real data, and the classification model is optimized.

Benefits of technology

The model accuracy and semantic segmentation accuracy of the classification model are improved, and the semantic segmentation effect of three-dimensional scenes is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339757A_ABST
    Figure CN120339757A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional scene semantic segmentation and generation method based on a semi-supervised diffusion model, and relates to the technical field of artificial intelligence. In some embodiments of the present disclosure, a three-dimensional scene image is acquired; inputting the three-dimensional scene image into the trained semi-supervised learning classification model for image semantic segmentation to obtain an image semantic segmentation result; wherein the supervised branches are trained by using the marked point cloud data, and the unsupervised branches are trained by using the unmarked point cloud data; semi-supervised learning is adopted for model training, the model precision of the classification model is improved, and the semantic segmentation accuracy of the classification model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a three-dimensional scene semantic segmentation and generation method based on a semi-supervised diffusion model. Background Art

[0002] Existing semi-supervised learning (SSL) mainly focuses on image classification and image semantic segmentation. Semi-supervised three-dimensional semantic segmentation is to train a model using a small number of labeled dense point cloud frameworks and a large number of unlabeled point cloud frameworks, which can reduce the annotation burden to a certain extent.

[0003] Currently, the classical works in semi-supervised classification include generative methods based on VAE and GAN, as well as methods with confidence regularization, consistency regularization, etc.

[0004] Currently, for the classification model trained by a diffusion model with a small number of labels, the model accuracy is low and the semantic segmentation accuracy is low. Summary of the Invention

[0005] The present disclosure provides a three-dimensional scene semantic segmentation and generation method based on a semi-supervised diffusion model, so as to at least solve the problems of low model accuracy and low semantic segmentation accuracy of the existing classification model.

[0006] The technical solution of the present disclosure is as follows: An embodiment of the present disclosure provides a three-dimensional semantic segmentation method, including: Obtaining a three-dimensional scene image; Inputting the three-dimensional scene image into a trained semi-supervised learning classification model for image semantic segmentation to obtain an image semantic segmentation result; Wherein, the semi-supervised learning classification model is based on the obtained labeled point cloud data and unlabeled point cloud data; the supervised branch is trained according to the labeled point cloud data, and the unsupervised branch is trained according to the unlabeled point cloud data.

[0007] An embodiment of the present disclosure further provides a three-dimensional semantic segmentation device, including: An obtaining module, configured to obtain a three-dimensional scene image; A semantic segmentation module, configured to input the three-dimensional scene image into a trained semi-supervised learning classification model for image semantic segmentation to obtain an image semantic segmentation result; Wherein, the semi-supervised learning classification model is based on the obtained labeled point cloud data and unlabeled point cloud data; the supervised branch is trained according to the labeled point cloud data, and the unsupervised branch is trained according to the unlabeled point cloud data.

[0008] An embodiment of the present disclosure further provides an electronic device, including: A processor; A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the instructions to implement the steps in the above method.

[0009] An embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above method are implemented.

[0010] An embodiment of the present disclosure further provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps in the above method are implemented.

[0011] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: In some embodiments of the present disclosure, a three-dimensional scene image is acquired; the three-dimensional scene image is input into a trained semi-supervised learning classification model for image semantic segmentation to obtain an image semantic segmentation result; wherein, the supervised branch is trained using labeled point cloud data, and the unsupervised branch is trained using unlabeled point cloud data; the present disclosure uses semi-supervised learning for model training to improve the model accuracy of the classification model and the semantic segmentation accuracy of the classification model.

[0012] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation to the present disclosure.

[0014] Figure 1 A flowchart of a three-dimensional semantic segmentation method provided for an exemplary embodiment of the present disclosure; Figure 2 An architecture diagram of a semi-supervised learning classification model provided for an exemplary embodiment of the present disclosure; Figure 3 A structural diagram of a three-dimensional semantic segmentation device provided for an exemplary embodiment of the present disclosure; Figure 4 A structural diagram of an electronic device provided for an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0015] To enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0016] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure.

[0017] It should be noted that the user information involved in the present disclosure includes, but is not limited to: user device information and user personal information; the collection, storage, use, processing, transmission, provision, and disclosure of user information in the present disclosure and other processing are all in compliance with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0018] In view of the above technical problems, in some embodiments of the present disclosure, a three-dimensional scene image is obtained; the three-dimensional scene image is input into a trained semi-supervised learning classification model for image semantic segmentation to obtain an image semantic segmentation result; wherein, the supervised branch is trained using labeled point cloud data, and the unsupervised branch is trained using unlabeled point cloud data; the present disclosure uses semi-supervised learning for model training to improve the model accuracy of the classification model and the semantic segmentation accuracy of the classification model.

[0019] The technical solutions provided by the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0020] Figure 1 It is a schematic flowchart of a three-dimensional semantic segmentation method provided for an exemplary embodiment of the present disclosure. As Figure 1 shown, the method includes: S101: Obtain a three-dimensional scene image; S102: Input the three-dimensional scene image into a trained semi-supervised learning classification model for image semantic segmentation to obtain an image semantic segmentation result; Among them, the semi-supervised learning classification model is a model obtained according to the acquired labeled point cloud data and unlabeled point cloud data; the supervised branch is trained according to the labeled point cloud data, and the unsupervised branch is trained according to the unlabeled point cloud data.

[0021] In this embodiment, the execution subject of the above method may be a server or a terminal device.

[0022] Among them, the terminal device includes, but is not limited to, a mobile station (MS), a mobile terminal, a mobile telephone, a handset, and portable equipment, etc. The terminal device can communicate with one or more core networks via a radio access network (RAN). For example, the terminal device can be a mobile phone (or a "cellular" phone), a computer with wireless communication capabilities, etc. The terminal device can also be a computer with wireless transceiver functions, a virtual reality (VR) terminal device, an AR terminal device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, etc. Moreover, the operating systems installed on the terminal device include, but are not limited to, operating systems such as IOS, Android, windows, linux, Mac OS, etc. In different networks, the terminal can be called by different names. For example: user equipment, mobile station, user unit, station, cellular phone, personal digital assistant, wireless modem, wireless communication device, handheld device, laptop computer, cordless phone, wireless local loop station, TV, etc. For the convenience of description, it is simply referred to as the terminal device in this embodiment.

[0023] In this embodiment, the implementation form of the server is not limited. For example, the server can be a conventional server, a cloud server, a cloud host, a virtual center, and other server devices. Among them, the server mainly consists of a processor, a hard disk, a memory, a system bus, etc., and is of a general computer architecture type.

[0024] In this embodiment, a three-dimensional scene image is obtained; the three-dimensional scene image is input into a trained semi-supervised learning classification model for image semantic segmentation to obtain an image semantic segmentation result; among them, the supervised branch is trained using labeled point cloud data, and the unsupervised branch is trained using unlabeled point cloud data; the present disclosure uses semi-supervised learning for model training to improve the model accuracy of the classification model and the semantic segmentation accuracy of the classification model.

[0025] In some embodiments of the present disclosure, before using the semi-supervised learning classification model, it is necessary to first train to obtain the semi-supervised learning classification model. One achievable way is to obtain labeled point cloud data and unlabeled point cloud data; train the supervised branch according to the labeled point cloud data; and train the unsupervised branch according to the unlabeled point cloud data to obtain the semi-supervised learning classification model.

[0026] It should be noted that the supervised branch includes: a first backbone network and a first classifier; the unsupervised branch includes: a second backbone network, a second classifier, a first projector, a third backbone network, a third classifier, and a second projector.

[0027] In the above embodiments of the present disclosure, the supervised branch is trained according to the labeled point cloud data; and the unsupervised branch is trained according to the unlabeled point cloud data to obtain the semi-supervised learning classification model. One achievable way is to perform data augmentation on the labeled point cloud data to obtain first augmented point cloud data; input the first augmented point cloud data into the first backbone network to obtain first point-wise features; input the first point-wise features into the first classifier to obtain a first semantic prediction result; train the supervised branch according to the cross-entropy loss function and the first semantic prediction result; perform data augmentation on the unlabeled point cloud data to obtain second augmented point cloud data and third augmented point cloud data respectively; input the second augmented point cloud data into the second backbone network to obtain second point-wise features; input the second point-wise features into the second classifier to obtain a second semantic prediction result; and input the second point-wise features into the first projector to obtain a first feature embedding; input the third augmented point cloud data into the third backbone network to obtain third point-wise features; input the third point-wise features into the third classifier to obtain a third semantic prediction result; and input the second point-wise features into the second projector to obtain a second feature embedding; train the unsupervised branch according to the contrastive loss function, the second semantic prediction result, the third semantic prediction result, the first feature embedding, and the second feature embedding.

[0028] Figure 2 It is an architecture diagram of a semi-supervised learning classification model provided for an exemplary embodiment of the present disclosure. In the present disclosure, in the labeled point cloud data and the unlabeled point cloud data, a guidance point contrast learning framework for point cloud segmentation based on semi-supervised learning is adopted, and semantic prediction is used as pseudo guidance to improve the contrastive learning of the unlabeled point cloud. As Figure 2 shown, the semi-supervised learning classification model includes a supervised branch and an unsupervised branch. In the figure, the subscripts l and u respectively represent "labeled" and "unlabeled". P is the input point cloud. Pu is independently augmented to form Pu1 and Pu2 as inputs to the unsupervised branch. F is the output point-wise feature of the backbone U-Net (i.e., the backbone network), which is further fed to the classifier to predict the semantic score S. In the unsupervised branch, F is also fed to the projector to generate the feature embedding E. represents point - by - point class prediction, while is the ground - truth label. The weights of the backbone network, classifier, and projector are shared for all input point clouds. The cross - entropy loss constrains the supervised training using the labeled data, while our guided contrastive loss guides the feature learning in the unsupervised branch.

[0029] In some embodiments of the present disclosure, during the training of the supervised branch based on the labeled point - cloud data; and the training of the unsupervised branch based on the unlabeled point - cloud data to obtain a semi - supervised learning classification model, the pseudo - labels of the unlabeled point - cloud data are predicted using the semi - supervised learning classification model; based on the real images carrying the pseudo - labels, a conditional generation model based on a diffusion model is trained; according to the conditional generation model, pseudo - images corresponding to the given random labels are generated; the semi - supervised learning classification model is trained on the real data enhanced by the pseudo - images to obtain an updated semi - supervised learning classification model. The present disclosure adopts a diffusion model after training the model for the second - time data augmentation, and for the first time realizes the combination of semantic segmentation and generation through a diffusion model in a three - dimensional scene. During the training process of semi - supervised classification and the diffusion model, two opposite conditional models, namely the diffusion model and the classifier, provide complementary learning signals to each other, thereby training better classification and generation models. The present disclosure realizes a diffusion - based semi - supervised learning classification and generation framework by introducing semi - supervised learning to train a classification model and introducing a diffusion model to generate data.

[0030] In the above - mentioned embodiment, the pseudo - labels of the unlabeled point - cloud data are predicted using the semi - supervised learning classification model. One feasible way is that for a pair of perturbed point clouds in the unlabeled point - cloud data, according to the semantic scores predicted by the semi - supervised learning classification model, the pseudo - labels corresponding to the perturbed point clouds and the label confidence are generated, and the formula is as follows: where * is u1 or u2, represents the softmax function; For a given set of positive point pairs and a set of negative point pairs and , where represents the set of matching positive point pairs on the point clouds u1 and u2 perturbed from the same input. For the negative point set, sampling is performed separately to ensure that the negative samples can also come from non - matching regions. The negative point sets sampled from the point clouds u1 and u2 are denoted as and , where and represent the number of points in the associated point clouds; the guided contrastive loss of the positive point pair ​ is:

[0031] wherein, is the confidence threshold; is the feature embedding of point cloud u1, is the feature embedding of point cloud u2; is the pseudo-label guidance for filtering negative point pairs with the same pseudo-label.

[0032] For each side, the features from the other side are detached to stop the gradient and thus regarded as a constant reference to better optimize the features of the current side. is defined as: .

[0033] In an alternative embodiment, the diffusion model is a denoising diffusion probability model. The conditional generation model includes: a denoising diffusion probability model and classifier-free guidance.

[0034] In the above embodiment, based on the real images carrying pseudo-labels, a conditional generation model based on the diffusion model is trained. One possible way is that the denoising diffusion probability model adds noise to the data gradually from time to and removes the noise gradually from to recover the data in the reverse process; the predictor ϵθ is trained by the following objective to predict the noise ϵ:

[0035] wherein, represents the target generation class; Classifier-free guidance utilizes the conditional noise predictor and the unconditional noise predictor in inference to improve the sample quality and enhance the semantics; CFG iterates the following equation starting from :

[0036] wherein, , is the guidance strength, Z~N(0, 1), , , and are variables with respect to the time constant t.

[0037] In the above embodiments, the semi-supervised learning classification model is trained on the real data enhanced by the pseudo-image to obtain an updated semi-supervised learning classification model. Specifically, the semi-supervised learning classification model is trained or fine-tuned on the real data enhanced by the pseudo-image with labels. Based on this semi-supervised learning classification model, a new module is connected, and the real data enhanced by the pseudo-image is used to train the semi-supervised learning classification model to improve the classification performance. For simplicity and efficiency, the present disclosure trains linear probing. Linear probing is a method for evaluating the performance of a trained model by replacing the last layer of the model with a linear layer while keeping the rest unchanged. During this process, the linear layer is trained to optimize the model's representation learning ability while improving the model training accuracy.

[0038] In the above method embodiment, a three-dimensional scene image is acquired; the three-dimensional scene image is input into a trained semi-supervised learning classification model for image semantic segmentation to obtain an image semantic segmentation result; wherein, the supervised branch is trained using the labeled point cloud data, and the unsupervised branch is trained using the unlabeled point cloud data; the present disclosure uses semi-supervised learning for model training to improve the model accuracy of the classification model and the semantic segmentation accuracy of the classification model.

[0039] Figure 3 FIG. is a schematic structural diagram of a three-dimensional semantic segmentation device 30 provided by an exemplary embodiment of the present disclosure. As Figure 3 shown, the three-dimensional semantic segmentation device 30 includes: an acquisition module 31 and a semantic segmentation module 32.

[0040] Among them, the acquisition module 31 is used to acquire a three-dimensional scene image; The semantic segmentation module 32 is used to input the three-dimensional scene image into a trained semi-supervised learning classification model for image semantic segmentation to obtain an image semantic segmentation result; Among them, the semi-supervised learning classification model is a model obtained according to the acquired labeled point cloud data and unlabeled point cloud data; the supervised branch is trained according to the labeled point cloud data, and the unsupervised branch is trained according to the unlabeled point cloud data.

[0041] Optionally, before using the semi-supervised learning classification model, the semantic segmentation module 32 can also be used to: Acquire labeled point cloud data and unlabeled point cloud data; Train the supervised branch according to the labeled point cloud data; and train the unsupervised branch according to the unlabeled point cloud data to obtain a semi-supervised learning classification model.

[0042] Optionally, the supervised branch includes: a first backbone network and a first classifier; the unsupervised branch includes: a second backbone network, a second classifier, a first projector, a third backbone network, a third classifier, and a second projector; the semantic segmentation module 32 is used for training the supervised branch according to the labeled point cloud data and training the unsupervised branch according to the unlabeled point cloud data to obtain a semi-supervised learning classification model: Performing data augmentation on the labeled point cloud data to obtain first augmented point cloud data; inputting the first augmented point cloud data into the first backbone network to obtain first pointwise features; inputting the first pointwise features into the first classifier to obtain a first semantic prediction result; training the supervised branch according to the cross-entropy loss function and the first semantic prediction result; Performing data augmentation on the unlabeled point cloud data to obtain second augmented point cloud data and third augmented point cloud data respectively; inputting the second augmented point cloud data into the second backbone network to obtain second pointwise features; inputting the second pointwise features into the second classifier to obtain a second semantic prediction result; and inputting the second pointwise features into the first projector to obtain a first feature embedding; inputting the third augmented point cloud data into the third backbone network to obtain third pointwise features; inputting the third pointwise features into the third classifier to obtain a third semantic prediction result; and inputting the second pointwise features into the second projector to obtain a second feature embedding; training the unsupervised branch according to the contrastive loss function, the second semantic prediction result, the third semantic prediction result, the first feature embedding, and the second feature embedding.

[0043] Optionally, after the semantic segmentation module 32 trains the supervised branch according to the labeled point cloud data and trains the unsupervised branch according to the unlabeled point cloud data to obtain a semi-supervised learning classification model, it can also be used for: Predicting the pseudo-labels of the unlabeled point cloud data using the semi-supervised learning classification model; Training a conditional generation model based on a diffusion model according to the real images with pseudo-labels; Generating pseudo-images corresponding to given random labels according to the conditional generation model; Training the semi-supervised learning classification model on the real data enhanced by the pseudo-images to obtain an updated semi-supervised learning classification model.

[0044] Optionally, when the semantic segmentation module 32 predicts the pseudo-labels of the unlabeled point cloud data using the semi-supervised learning classification model, it is used for: For a pair of perturbed point clouds in the unlabeled point cloud data , generating a pseudo-label and a label confidence corresponding to the perturbed point cloud according to the semantic scores predicted by the semi-supervised learning classification model, and the formula is as follows: ​ where * is either u1 or u2, denotes the softmax function; For a given set of positive point pairs and a set of negative point pairs and , where denotes the set of matching positive point pairs on point clouds u1 and u2 perturbed from the same input; the set of negative points sampled from point clouds u1 and u2 is denoted as and , where and denote the number of points in the associated point clouds; the bootstrapped contrastive loss for the positive point pair is:

[0045] where is the confidence threshold; is the feature embedding of point cloud u1, is the feature embedding of point cloud u2; is the pseudo-label guidance for filtering negative point pairs with the same pseudo-label, is defined as: .

[0046] Optionally, the diffusion model is a denoising diffusion probability model, and the conditional generation model includes: a denoising diffusion probability model and classifier-free guidance; when training the conditional generation model based on the diffusion model according to the real image carrying pseudo-labels, the semantic segmentation module 32 is used for: The denoising diffusion probability model gradually adds noise to the data from time in the forward process and gradually removes the noise to recover the data starting from in the reverse process; Train the predictor ϵθ to predict the noise ϵ through the following objective:

[0047] where denotes the target generation class; Classifier-free guidance utilizes the conditional noise predictor and the unconditional noise predictor in inference to improve the sample quality and enhance the semantics; CFG starts iterating the following equation from :

[0048] ​Among them, , is the guidance intensity, Z~N(0, 1), , , and are variables with respect to the time constant t.

[0049] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0050] Figure 4 This is a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. As Figure 4 shown, the electronic device includes: a memory 41 and a processor 42. Additionally, the electronic device further includes a power supply component 43 and a communication component 44.

[0051] The memory 41 is used to store computer programs and can be configured to store various other data to support operations on the electronic device. Examples of such data include instructions for any application program or method operating on the electronic device.

[0052] The memory 41 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0053] The communication component 44 is used for data transmission with other devices.

[0054] The processor 42 can execute the computer instructions stored in the memory 41 to: acquire a three-dimensional scene image; input the three-dimensional scene image into a trained semi-supervised learning classification model for image semantic segmentation to obtain an image semantic segmentation result; wherein, the semi-supervised learning classification model is based on the acquired labeled point cloud data and unlabeled point cloud data; the supervised branch is trained according to the labeled point cloud data, and the unsupervised branch is trained according to the unlabeled point cloud data to obtain the model.

[0055] Correspondingly, an exemplary embodiment of the present disclosure also provides a computer-readable storage medium storing a computer program. When the computer-readable storage medium stores the computer program and the computer program is executed by one or more processors, it causes the one or more processors to execute Figure 1 each step in the method embodiments.

[0056] Correspondingly, an embodiment of the present disclosure further provides a computer program product, which includes computer programs / instructions that are executed by a processor Figure 1 for each step in the method embodiment.

[0057] The above Figure 4 The communication component is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on technologies such as Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0058] The above Figure 4 The power component provides power for various components of the device where the power component is located. The power component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power component is located.

[0059] The above electronic device further includes a display screen and an audio component.

[0060] The display screen includes a screen, and the screen may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations.

[0061] The audio component is configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), which is configured to receive external audio signals when the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory or sent via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.

[0062] In the above-mentioned device, equipment, storage medium, and computer program product embodiments of the present disclosure, a three-dimensional scene image is obtained; the three-dimensional scene image is input into a trained semi-supervised learning classification model for image semantic segmentation to obtain an image semantic segmentation result; among them, the supervised branch is trained using labeled point cloud data, and the unsupervised branch is trained using unlabeled point cloud data; the present disclosure uses semi-supervised learning for model training to improve the model accuracy of the classification model and the semantic segmentation accuracy of the classification model.

[0063] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, system, or computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0064] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0065] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0066] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0067] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0068] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0069] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0070] It should be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0071] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A three-dimensional scene semantic segmentation and generation method based on a semi-supervised diffusion model, characterized in that, Including: Obtain a three-dimensional scene image; Input the three-dimensional scene image into a trained semi-supervised learning classification model for image semantic segmentation to obtain an image semantic segmentation result; Among them, the semi-supervised learning classification model is based on the obtained labeled point cloud data and unlabeled point cloud data; the supervised branch is trained according to the labeled point cloud data, and the unsupervised branch is trained according to the unlabeled point cloud data to obtain the model.

2. The method according to claim 1, wherein Before using the semi-supervised learning classification model, the method further includes: Obtain labeled point cloud data and unlabeled point cloud data; Train the supervised branch according to the labeled point cloud data; and train the unsupervised branch according to the unlabeled point cloud data to obtain a semi-supervised learning classification model.

3. The method according to claim 2, wherein The supervised branch includes: a first backbone network and a first classifier; the unsupervised branch includes: a second backbone network, a second classifier, a first projector, a third backbone network, a third classifier, and a second projector; the training of the supervised branch according to the labeled point cloud data; and the training of the unsupervised branch according to the unlabeled point cloud data to obtain a semi-supervised learning classification model includes: Perform data augmentation on the labeled point cloud data to obtain first augmented point cloud data; input the first augmented point cloud data into the first backbone network to obtain first point-wise features; input the first point-wise features into the first classifier to obtain a first semantic prediction result; train the supervised branch according to the cross-entropy loss function and the first semantic prediction result; Perform data augmentation on the unlabeled point cloud data to obtain second augmented point cloud data and third augmented point cloud data respectively; input the second augmented point cloud data into the second backbone network to obtain second point-wise features; input the second point-wise features into the second classifier to obtain a second semantic prediction result; and input the second point-wise features into the first projector to obtain a first feature embedding; input the third augmented point cloud data into the third backbone network to obtain third point-wise features; input the third point-wise features into the third classifier to obtain a third semantic prediction result; and input the second point-wise features into the second projector to obtain a second feature embedding; train the unsupervised branch according to the contrastive loss function, the second semantic prediction result, the third semantic prediction result, the first feature embedding, and the second feature embedding.

4. The method according to claim 2, wherein After the training of the supervised branch according to the labeled point cloud data; and the training of the unsupervised branch according to the unlabeled point cloud data to obtain a semi-supervised learning classification model, the method further includes: Use the semi-supervised learning classification model to predict the pseudo-labels of the unlabeled point cloud data; Train a conditional generation model based on a diffusion model according to the real image carrying the pseudo-labels; Generate a pseudo-image corresponding to a given random label according to the conditional generation model; Train the semi-supervised learning classification model on the real data enhanced by the pseudo-image to obtain an updated semi-supervised learning classification model.

5. The method according to claim 4, characterized in that, Predicting the pseudo-labels of the unlabeled point cloud data using the semi-supervised learning classification model includes: For a pair of perturbed point clouds in the unlabeled point cloud data , according to the semantic scores predicted by the semi-supervised learning classification model , generate the pseudo-labels and label confidence for the perturbed point clouds, and the formula is as follows: where * is u1 or u2, represents the softmax function; For a given set of positive point pairs and a set of negative point pairs and , where represents the set of matching positive point pairs on point clouds u1 and u2 from the same input perturbation; the negative point sets sampled from point clouds u1 and u2 are denoted as and , where and represent the number of points in the associated point clouds; the bootstrapped contrastive loss for the positive point pair is given by: Among them, is the confidence threshold; is the feature embedding of point cloud u1, is the feature embedding of point cloud u2; is the pseudo-label guidance for filtering negative point pairs with the same pseudo-labels, is defined as: 。 6. The method according to claim 4, wherein The diffusion model is a denoising diffusion probability model, and the conditional generation model includes: the denoising diffusion probability model and classifier-free guidance; training the conditional generation model based on the diffusion model according to the real image carrying the pseudo-labels includes: The denoising diffusion probability model adds noise to the data gradually from time to during the forward process, and removes the noise gradually starting from to recover the data during the reverse process; ​ Training the predictor ϵθ to predict the noise ϵ through the following objective: Among them, represents the target generation class; Utilize conditional noise predictors in inference without classifier guidance and unconditional noise predictors to improve sample quality and enhance semantics; CFG starts iterating the following equation from and proceeds as follows: Among them, , is the guiding strength, Z~N(0, 1), , , and are variables regarding the time constant t.