Domain generalization semantic segmentation method and system, electronic equipment and storage medium

By using the domain generalized semantic segmentation method in the semantic segmentation task, the cross-attention fusion of virtual simulation data and visual basic model is solved, and efficient semantic segmentation and accuracy improvement in real-world scenarios are achieved.

CN120014270AActive Publication Date: 2025-05-16SUN YAT SEN UNIVERSITY SHENZHEN +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510089753.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

Existing deep neural networks have domain offset problems in semantic segmentation tasks, resulting in poor performance in different scenarios or data distributions, especially in real-world scenarios, which are difficult to achieve efficient semantic segmentation.

Method used

A domain generalized semantic segmentation method is proposed. By loading pre-training parameters of the visual basic model, using the virtual engine to generate a virtual simulation data set, freeze the parameters of the visual basic model, extract intermediate features, and perform cross-attention fusion with the word vector, randomly specify the category labels and do not participate in the loss calculation, loop iterative training until the loss converges, and obtain the trained target model.

Benefits of technology

It realizes efficient semantic segmentation in real-world scenarios, reduces data annotation costs, and improves the generalization ability and segmentation accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014270A_ABST
    Figure CN120014270A_ABST
Patent Text Reader

Abstract

The invention discloses a domain generalization semantic segmentation method and system, electronic equipment and a storage medium, and the method comprises the steps: loading pre-training parameters of a visual basic large model, and generating a virtual simulation data set; extracting an intermediate feature of the input image, carrying out cross attention fusion on the intermediate feature and the word vector, inputting a segmentation head to carry out semantic segmentation prediction, and obtaining a category prediction result corresponding to each pixel; randomly specifying a category label not to participate in loss calculation during each small-batch iterative training; calculating a loss value between the prediction category and the real category; setting a corresponding word vector gradient to be zero according to an index mode during gradient back propagation; iteratively training to obtain a target model; and performing domain generalization semantic segmentation processing on the target data to be segmented in the target scene of the real world to obtain a segmentation result. The method can reduce the data annotation cost, achieves the universal generalization segmentation performance of the model for the real world scene, and can be widely applied to the technical field of computer vision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a domain generalization semantic segmentation method, system, electronic device and storage medium. Background Art

[0002] Deep Neural Network (DNN) has achieved remarkable success in autonomous vehicles and related tasks in recent years. This is mainly due to the powerful feature learning ability of DNN, which enables it to automatically extract meaningful features from massive data. However, the training of DNN usually relies on the independent and identically distributed (IID) assumption, that is, the training data and test data come from the same distribution. Once the distribution of test data changes, the model performance may drop significantly. This phenomenon is called domain shift. The existence of domain shift limits the generalization ability of DNN, causing it to perform poorly in different scenarios or data distributions. This problem is particularly prominent in practical applications.

[0003] The problem of domain shift is particularly serious in the task of semantic segmentation. The goal of semantic segmentation is to accurately classify each pixel in the input image into a specific semantic category in order to achieve high-level visual understanding. This task requires not only a good feature extraction capability of the model, but also rich annotated data to support it. However, the annotation cost of the semantic segmentation task is extremely high. Unlike image classification or object detection tasks, semantic segmentation requires annotating each pixel one by one, which is a complex and time-consuming task. Especially when dealing with high-resolution images and complex scenes, the annotation process may take hours or even days. With the development of 3D modeling and simulation technology, this problem has been significantly alleviated. Modern 3D engines (such as Unreal Engine, Unity, etc.) can realistically simulate real-world scenes and support automatic generation of pixel-level semantic annotations. For example, by setting different materials and object categories in the virtual scene, the corresponding semantic label can be automatically assigned to each pixel without manual intervention. This approach can not only significantly reduce the annotation cost, but also flexibly generate diverse data covering different environments, lighting conditions, weather changes, etc., thereby enhancing the robustness and generalization ability of the model.

[0004] However, there are obvious differences between virtual data and real-world scenes. These differences are mainly reflected in the fact that the data simulated by the virtual engine is not realistic enough and has different lighting environments and shooting conditions from the real world. Therefore, if you want to train the model using only virtual data and achieve excellent semantic segmentation performance in real-world scenes, you must obtain robust image features with domain-independent characteristics.

[0005] With the development of large model technology, many basic visual models can obtain more robust image representations, and explicit visual pre-training and fine-tuning strategies in a variety of downstream tasks have demonstrated powerful image feature extraction capabilities. At the same time, efficient fine-tuning methods based on basic visual large models have shown significantly superior performance in out-of-distribution (OOD) generalization tasks. At present, most solutions focus on general domain generalization rather than specifically targeting out-of-distribution semantic segmentation tasks. Therefore, the key to the problem is how to use the basic visual large model in out-of-distribution semantic segmentation tasks and combine it with efficient fine-tuning strategies to ensure fewer parameters and better generalization performance. Summary of the invention

[0006] The main purpose of the embodiments of the present invention is to propose an efficient and high-precision domain generalization semantic segmentation method, system, electronic device and storage medium, which can reduce the cost of data annotation and achieve the general generalization segmentation performance of the model for real-world scenes.

[0007] To achieve the above objective, an embodiment of the present invention provides a domain generalization semantic segmentation method, comprising the following steps:

[0008] Load the pre-trained parameters of the visual basic model, and the virtual engine generates a virtual simulation data set;

[0009] Freezing the parameters of the visual basic model, and after inputting the simulation data set into the visual basic model, extracting the intermediate features of the input image through the encoder of the visual basic model;

[0010] Cross-attention fusion of the extracted intermediate features and word vectors;

[0011] The fused features are input into the segmentation head for semantic segmentation prediction to obtain the category prediction result corresponding to each pixel;

[0012] In each small batch iterative training, the randomly assigned category label is not involved in the loss calculation; according to the assigned category label, the loss value between the predicted category and the true category of each pixel is calculated;

[0013] During gradient backpropagation, obtain the category labels that are not involved in the loss calculation, and set the corresponding word vector gradients to zero by index;

[0014] The visual basic large model is trained iteratively until the loss converges, and the corresponding word vectors and segmentation heads are saved to obtain a trained target model;

[0015] According to the target model, domain generalization semantic segmentation processing is performed on the target data in the real-world target scene to be segmented to obtain segmentation results with different semantics.

[0016] In some embodiments, the cross-attention fusion of the extracted intermediate features and the word vectors comprises the following steps:

[0017] The cross-attention method is used to fuse and complement the deep semantic information of visual features and the category expression information guided by word vectors; the query matrix Q is obtained by linearly transforming the word vectors; the key matrix K is obtained by linearly transforming the features; and the value matrix V is obtained by linearly transforming the features.

[0018] The expression of cross attention fusion is: Among them, d k is the learnable temperature coefficient; A is the output result.

[0019] In some embodiments, in the step of randomly assigning a class label not to participate in the loss calculation, when randomly selecting a label, the label set of the small batch is {label i |i=1,...,N}, where N≤19 is the total number of small batch label categories. Each time, n category labels are randomly specified in the set, and the specified labels are set to not participate in the loss calculation.

[0020] In some embodiments, the limiting condition for the number of the specified n category labels is:

[0021]

[0022] Among them, n max Represents the maximum number of category labels, 1≤n≤n max .

[0023] In some embodiments, in the step of inputting the fused features into a segmentation head for semantic segmentation prediction to obtain a category prediction result corresponding to each pixel, mask2forme is used as the segmentation head;

[0024] The calculation expression of the total model loss of the visual basic large model is: loss=h_loss_cls+h_loss_dice+h_loss_focal, where h_loss_cls is the classification loss of the prediction result, h_loss_dice is the overlap loss between the predicted mask and the true mask, and h_loss_focal is the loss of the unbalanced sample ratio.

[0025] In some embodiments, the method further comprises: evaluating the processing result of the domain generalization semantic segmentation by using the pixel average intersection-over-union ratio as an evaluation index of the semantic segmentation;

[0026] The calculation formula of the pixel average intersection-over-union ratio is:

[0027]

[0028] Among them, mIoU represents the average pixel intersection over union; N represents the total number of categories; TP i is the number of pixels correctly predicted as class i; FP i is the number of pixels incorrectly predicted as class i; FN i is the number of pixels of class i that are missed.

[0029] Another aspect of an embodiment of the present invention further provides a domain generalization semantic segmentation system, including:

[0030] The first module is used to load the pre-trained parameters of the visual basic model, and the virtual engine generates a virtual simulation data set;

[0031] The second module is used to freeze the parameters of the visual basic model, and after the simulation data set is input into the visual basic model, the intermediate features of the input image are extracted through the encoder of the visual basic model;

[0032] The third module is used to perform cross-attention fusion of the extracted intermediate features and word vectors;

[0033] The fourth module is used to input the fused features into the segmentation head for semantic segmentation prediction to obtain the category prediction result corresponding to each pixel;

[0034] The fifth module is used to randomly specify a category label that does not participate in the loss calculation during each small batch iterative training; based on the specified category label, the loss value between the predicted category and the true category of each pixel is calculated;

[0035] The sixth module is used to obtain the category labels that are not involved in the loss calculation during gradient back propagation, and set the corresponding word vector gradient to zero by index;

[0036] The seventh module is used to iteratively train the visual basic large model until the loss converges, save the corresponding word vectors and segmentation heads, and obtain the trained target model;

[0037] The eighth module is used to perform domain generalization semantic segmentation processing on the target data in the real-world target scene to be segmented according to the target model to obtain segmentation results with different semantics.

[0038] To achieve the above objective, another aspect of an embodiment of the present invention provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the above-mentioned method when executing the computer program.

[0039] To achieve the above objective, another aspect of an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned method is implemented.

[0040] The embodiment of the present invention also discloses a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device can read the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the above method.

[0041] The embodiments of the present invention include at least the following beneficial effects: The present invention provides a domain generalization semantic segmentation method, system, electronic device and storage medium, wherein the scheme generates a virtual simulation data set by loading the pre-trained parameters of the visual basic large model; freezing the parameters of the visual basic large model, inputting the simulation data set into the visual basic large model, and then extracting the intermediate features of the input image through the encoder of the visual basic large model; cross-attention fusion of the extracted intermediate features and word vectors; inputting the fused features into the segmentation head for semantic segmentation prediction to obtain the category prediction result corresponding to each pixel; in each small batch iterative training, randomly specifying the category label not to participate in the loss calculation; according to the specified category label, calculating the loss value between the predicted category and the true category of each pixel; in the gradient back propagation, obtaining the category label not involved in the loss calculation, and setting the corresponding word vector gradient to zero in an indexed manner; cyclically iteratively training the visual basic large model until the loss converges, saving the corresponding word vector and segmentation head, and obtaining a trained target model; performing domain generalization semantic segmentation processing on the target data in the target scene of the real world to be segmented according to the target model, and obtaining segmentation results with different semantics. The embodiments of the present invention are efficient and highly accurate, can reduce data annotation costs, and achieve universal generalization segmentation performance of the model for real-world scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present invention;

[0043] Figure 2 is a flow chart of the overall steps provided by an embodiment of the present invention;

[0044] Figure 3It is a flowchart of the specific implementation process provided by the embodiment of the present invention;

[0045] Figure 4 is an example diagram of an end-to-end domain generalization semantic segmentation model provided by an embodiment of the present invention;

[0046] Figure 5 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present invention. They are only examples of devices and methods consistent with some aspects of the embodiments of the present invention as detailed in the attached claims.

[0048] It is understood that the terms "first", "second", "third", "fourth", etc. (if any) in the description of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0049] It should be understood that in the present invention, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can represent: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.

[0051] The domain generalization semantic segmentation method, system, electronic device and storage medium provided by the embodiments of the present invention relate to the field of computer vision technology. The domain generalization semantic segmentation method provided by the embodiments of the present invention can be applied to a terminal, can be applied to a server, or can be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, and can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements the domain generalization semantic segmentation method, etc., but is not limited to the above forms.

[0052] The present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0053] like Figure 1 FIG. 1 is a schematic diagram of an implementation environment provided by an embodiment of the present invention. Figure 1 The implementation environment includes at least one terminal 102 and a server 101. The terminal 102 and the server 101 can be connected to a network wirelessly or wired to complete data transmission and exchange.

[0054] Server 101 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.

[0055] In addition, the server 101 can also be a node server in the blockchain network. Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm.

[0056] The terminal 102 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc. The terminal 102 may also be a vehicle-mounted terminal of various device types described above, but is not limited thereto. The terminal 102 and the server 101 may be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment of the present invention.

[0057] Based on the example Figure 1 In the implementation environment shown, an embodiment of the present invention provides a domain generalization semantic segmentation method. The following is an example of applying the domain generalization semantic segmentation method in the server 101. It can be understood that the method can also be applied to the terminal 102.

[0058] Reference Figure 2 , Figure 2 The flowchart of the domain generalization semantic segmentation method applied to the server provided in the embodiment of the present invention, the execution subject of the method can be any of the aforementioned computer devices (including servers or terminals). Figure 2 , the method may include the following steps:

[0059] Load the pre-trained parameters of the visual basic model, and the virtual engine generates a virtual simulation data set;

[0060] Freezing the parameters of the visual basic model, and after inputting the simulation data set into the visual basic model, extracting the intermediate features of the input image through the encoder of the visual basic model;

[0061] Cross-attention fusion of the extracted intermediate features and word vectors;

[0062] The fused features are input into the segmentation head for semantic segmentation prediction to obtain the category prediction result corresponding to each pixel;

[0063] In each small batch iterative training, the randomly assigned category label is not involved in the loss calculation; according to the assigned category label, the loss value between the predicted category and the true category of each pixel is calculated;

[0064] During gradient backpropagation, obtain the category labels that are not involved in the loss calculation, and set the corresponding word vector gradients to zero by index;

[0065] The visual basic large model is trained iteratively until the loss converges, and the corresponding word vectors and segmentation heads are saved to obtain a trained target model;

[0066] According to the target model, domain generalization semantic segmentation processing is performed on the target data in the real-world target scene to be segmented to obtain segmentation results with different semantics.

[0067] In some embodiments, the cross-attention fusion of the extracted intermediate features and the word vectors comprises the following steps:

[0068] The cross-attention method is used to fuse and complement the deep semantic information of visual features and the category expression information guided by word vectors; the query matrix Q is obtained by linearly transforming the word vectors; the key matrix K is obtained by linearly transforming the features; and the value matrix V is obtained by linearly transforming the features.

[0069] The expression of cross attention fusion is: Among them, d k is the learnable temperature coefficient; A is the output result.

[0070] In some embodiments, in the step of randomly assigning a class label not to participate in the loss calculation, when randomly selecting a label, the label set of the small batch is {label i |i=1,...,N}, where N≤19 is the total number of small batch label categories. Each time, n category labels are randomly specified in the set, and the specified labels are set to not participate in the loss calculation.

[0071] In some embodiments, the limiting condition for the number of the specified n category labels is:

[0072]

[0073] Among them, n max Represents the maximum number of category labels, 1≤n≤n max .

[0074] In some embodiments, in the step of inputting the fused features into a segmentation head for semantic segmentation prediction to obtain a category prediction result corresponding to each pixel, mask2forme is used as the segmentation head;

[0075] The calculation expression of the total model loss of the visual basic large model is: loss=h_loss_cls+h_loss_dice+h_loss_focal, where h_loss_cls is the classification loss of the prediction result, h_loss_dice is the overlap loss between the predicted mask and the true mask, and h_loss_focal is the loss of the unbalanced sample ratio.

[0076] In some embodiments, the method further comprises: evaluating the processing result of the domain generalization semantic segmentation by using the pixel average intersection-over-union ratio as an evaluation index of the semantic segmentation;

[0077] The calculation formula of the pixel average intersection-over-union ratio is:

[0078]

[0079] Among them, mIoU represents the average pixel intersection over union; N represents the total number of categories; TP i is the number of pixels correctly predicted as class i; FP i is the number of pixels incorrectly predicted as class i; FN i is the number of pixels of class i that are missed.

[0080] The following describes the specific implementation process of the method of the present invention in detail by taking a specific application scenario as an example:

[0081] In view of the problems existing in the prior art, the present invention proposes a domain generalization semantic segmentation method for efficiently fine-tuning a visual basic large model, which is suitable for end-to-end training of actual driving scenes. In order to reduce the cost of data annotation, the present invention proposes to rely entirely on the data generated by the virtual engine simulation scene for training, extract discriminative features through the visual basic large model, and combine with efficient fine-tuning strategies to enable the model to quickly adapt to the real driving environment. Specifically, the present invention guides the feature expression by fusing the features of each layer of the visual large model with the randomly initialized word vector, thereby promoting the distinction of category information. Since the category information is domain-independent, it can improve the generalization ability of the model and enable it to maintain good segmentation performance in different driving scenarios. In order to ensure the effective binding of the word vector and the category information, the present invention proposes that in each training, the randomly specified category label does not participate in the loss calculation, and the gradient of the corresponding word vector is set to zero according to the index to ensure the relevance of the word vector to the specific category. Through virtual simulation data, combined with the visual basic large model to extract robust features, and further using the word vector to guide the expression of category information, the present invention can significantly improve the accuracy of domain generalization semantic segmentation and achieve the excellent generalization performance of the model in real driving scenes.

[0082] like Figure 3As shown, the specific implementation process of the method of the present invention includes the following steps:

[0083] Step 1: Prepare a virtual simulation dataset d; load the pre-trained parameters of the visual basic model m. The simulation dataset is generated by the virtual engine and simulates different driving scenarios to reduce the data annotation cost.

[0084] Step 2: Freeze the parameters of the basic visual model m and use the encoder of m to extract the intermediate features f of the input image;

[0085] 2. Step 3: Cross-attention fusion of the extracted intermediate feature f and the word vector t; In this embodiment, the intermediate feature f is fused with the word vector t by using the cross-attention method to fuse and complement the deep semantic information of the visual feature and the category expression information guided by the word vector. Q is the query matrix, which is obtained from the word vector t through linear transformation; K is the key matrix, which is obtained from the feature f through linear transformation; V is the value matrix, which is obtained from the feature f through linear transformation; d k is the learnable temperature coefficient; A is the output result obtained; as shown in formula (1).

[0086]

[0087] Step 4: Input the fused feature F into the segmentation head to h for semantic segmentation prediction, and each pixel gets the corresponding category prediction result;

[0088] Step 5: During each mini-batch iteration, the randomly assigned category label is not included in the loss calculation. When randomly selecting labels in step 5, the label set of the mini-batch is {label i |i=1,...,N}, where N≤19 is the total number of small batch label categories. Each time, n category labels are randomly specified in the set, and the specified labels are set to not participate in the loss calculation. Each time, n category labels are randomly specified not to participate in the loss calculation, n is an integer and 1≤n≤n max , where n max The definition of is as follows:

[0089]

[0090] Step 6: Based on the selection in step 5, calculate the loss value l between the predicted category and the true category of each pixel;

[0091] Step 7: During gradient back propagation, obtain the category labels that are not involved in the loss calculation, and set the corresponding word vector gradient to zero in an indexed manner; in step 7, select the labels that do not participate in the loss calculation, each label corresponds to some pixels of the sample, and the features corresponding to these pixels will not be used to predict the object category. Therefore, the gradient obtained by the network during back propagation will not act on the corresponding word vector. Assuming that the category label is 1 and does not participate in the loss calculation, the first word vector will not be updated by the gradient in an indexed manner. The present invention uses virtual simulation data to efficiently fine-tune the visual basic model, combines robust image features and uses word vectors to guide the expression of category information. The model can maintain a high generalization ability in real driving scenarios, effectively reduce the data annotation cost of semantic segmentation, and significantly improve the accuracy of domain generalization semantic segmentation tasks.

[0092] Step 8: Repeat steps 2-7 until the loss l converges, and save the word vector t and the segmentation head h. The driving scene domain generalization model trained by the present invention is an end-to-end semantic segmentation model, that is, objects (roads, pedestrians, cars, plants, sky, etc.) in real-world driving scenes are used as categories to be segmented, and different objects with semantics in the content of the pictures taken of the driving scenes are identified through the semantic segmentation method.

[0093] The semantic segmentation framework combined with the basic visual big model can select any appropriate end-to-end basic visual big model according to the actual project requirements (inference speed, segmentation accuracy, stability, etc.). Figure 4 The DINOV2 basic visual large model shown can be used as a feature encoding model for semantic segmentation. It can extract image features with strong discriminability and generalization, and achieve the fusion and complementarity of high-level visual features and category boundary information by combining word vectors, so as to realize efficient fine-tuning under semantic segmentation tasks. The so-called end-to-end means that the semantic segmentation model is generally composed of multiple parts, and different parts realize different functions. This embodiment involves three parts: the basic visual large model is used as the feature extraction part of the image, and the network parameters of this part are fixed and do not participate in the gradient update; a series of learnable word vectors are fused with the intermediate layer image features, and the parameters of this part need to be updated for efficient fine-tuning of the large model to adapt to the semantic segmentation task; the segmentation head decodes the fused features and realizes end-to-end loss training, and the parameters of this part need to be updated to ensure the accuracy of the segmentation function; these parts are serially connected to achieve the overall prediction function of the network. During training, the training method in which the total loss is simultaneously updated by gradient descent for the parameters that need to be updated in the model is called end-to-end training. Figure 4In this embodiment, mask2forme is used as the segmentation head. Its module loss includes the classification loss h_loss_cls of the prediction result, the overlap loss h_loss_dice between the predicted mask and the true mask, and the loss h_loss_focal to solve the imbalanced sample ratio. Its total loss is shown in formula (1):

[0094] loss=h_loss_cls+h_loss_dice+h_loss_focal (1)

[0095] The loss calculated in step 6 is the total loss of the model here.

[0096] The fusion of the intermediate feature f and the word vector t mentioned in step 3 is to use the cross-attention method to fuse and complement the deep semantic information of the visual feature and the category expression information guided by the word vector. Where Q is the query matrix, which is obtained from the word vector t through linear transformation; K is the key matrix, which is obtained from the feature f through linear transformation; V is the value matrix, which is obtained from the feature f through linear transformation; d k is the learnable temperature coefficient; A is the output result obtained; as shown in formula (2).

[0097]

[0098] Step 5: When randomly selecting labels, let the label set of the mini-batch be {label i |i=1,...,N}, where N≤19 is the total number of small batch label categories. Each time, n category labels are randomly specified in the set, and the specified category labels are set to not participate in the loss calculation. n is an integer and 1≤n≤n max , where n max The definition of is as follows:

[0099]

[0100] In addition, the labels that are not involved in the loss calculation mentioned in step 7 are selected. Each label corresponds to some pixels of the sample, and the features corresponding to these pixels will not be used to predict the object category. Therefore, the gradient obtained by the network during back propagation will not act on the corresponding word vector. Assuming that the category label is 1 and does not participate in the loss calculation, according to the index method, the first word vector will not be updated by the gradient.

[0101] By using virtual simulation data to efficiently fine-tune the visual basic model, combining robust image features and using word vectors to guide the expression of category information, the model can maintain a high generalization ability in real driving scenarios, effectively reduce the data annotation cost of semantic segmentation tasks, and significantly improve the accuracy of domain generalization semantic segmentation tasks.

[0102] The commonly used evaluation index for semantic segmentation is the mean intersection over union (mIoU). mIoU is an indicator that measures the degree of overlap between the predicted result and the true label. Formula (4) gives a specific definition:

[0103]

[0104] N represents the total number of categories, where TP i is the number of pixels correctly predicted as class i; FP i is the number of pixels incorrectly predicted as class i; FN i is the number of pixels of class i that are missed.

[0105] In summary, the present invention starts from the perspective of efficient fine-tuning on the visual basic large model, and can perform good semantic segmentation of real-world driving scenes when only virtual data is used for training; robust image features are extracted through the visual basic large model, and a series of learnable word vectors are used to combine with image features for cross-attention, and the word vectors are further bound to the category information by erasing the specified category labels and setting the word vector gradients to zero. Since the category information has the characteristic of domain independence, the image features guided by the word vectors will also have the characteristic of domain independence, which significantly reduces the cost of data annotation while achieving the model's general generalization segmentation performance for real-world scenes, and can be widely used in the fields of image processing and computer vision technology.

[0106] Another aspect of an embodiment of the present invention further provides a domain generalization semantic segmentation system, including:

[0107] The first module is used to load the pre-trained parameters of the visual basic model, and the virtual engine generates a virtual simulation data set;

[0108] The second module is used to freeze the parameters of the visual basic model, and after the simulation data set is input into the visual basic model, the intermediate features of the input image are extracted through the encoder of the visual basic model;

[0109] The third module is used to perform cross-attention fusion of the extracted intermediate features and word vectors;

[0110] The fourth module is used to input the fused features into the segmentation head for semantic segmentation prediction to obtain the category prediction result corresponding to each pixel;

[0111] The fifth module is used to randomly specify a category label that does not participate in the loss calculation during each small batch iterative training; based on the specified category label, the loss value between the predicted category and the true category of each pixel is calculated;

[0112] The sixth module is used to obtain the category labels that are not involved in the loss calculation during gradient back propagation, and set the corresponding word vector gradient to zero by index;

[0113] The seventh module is used to iteratively train the visual basic large model until the loss converges, save the corresponding word vectors and segmentation heads, and obtain the trained target model;

[0114] The eighth module is used to perform domain generalization semantic segmentation processing on the target data in the real-world target scene to be segmented according to the target model to obtain segmentation results with different semantics.

[0115] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0116] The embodiment of the present invention further provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the above domain generalization semantic segmentation method when executing the computer program. The electronic device can be any intelligent terminal including a tablet computer, a car computer, etc.

[0117] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0118] See also Figure 5 , Figure 5 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:

[0119] The processor 501 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention;

[0120] The memory 502 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 502 can store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 502, and the processor 501 calls and executes the domain generalization semantic segmentation method of the embodiment of the present invention;

[0121] Input / output interface 503, used to implement information input and output;

[0122] Communication interface 504, used to realize communication interaction between the device and other devices, which can be realized by wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.);

[0123] A bus 505 that transmits information between the various components of the device (e.g., the processor 501, the memory 502, the input / output interface 503, and the communication interface 504);

[0124] The processor 501 , the memory 502 , the input / output interface 503 and the communication interface 504 are connected to each other in communication within the device via the bus 505 .

[0125] An embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned domain generalization semantic segmentation method is implemented.

[0126] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0127] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0128] It should be noted that in various specific embodiments of the present invention, when it comes to the need to perform relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present invention needs to obtain the user's sensitive personal information, it will obtain the user's separate permission or consent through a pop-up window or jump to a confirmation page, and after clearly obtaining the user's separate permission or consent, it will obtain the necessary user-related data for the normal operation of the embodiment of the present invention.

[0129] The embodiments described in the embodiments of the present invention are intended to more clearly illustrate the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those skilled in the art can appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are also applicable to similar technical problems.

[0130] Those skilled in the art will appreciate that the technical solutions shown in the figures do not limit the embodiments of the present invention and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0131] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0132] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0133] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0134] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0135] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0136] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including multiple instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods of various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store programs.

[0137] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the embodiments of the present invention is not limited thereby. Any modification, equivalent substitution and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present invention shall be within the scope of the rights of the embodiments of the present invention.

Claims

1. A domain generalization semantic segmentation method, characterized in that: The following steps are involved: Load the pre-trained parameters of the visual basic model, and the virtual engine generates a virtual simulation data set; Freezing the parameters of the visual basic model, and after inputting the simulation data set into the visual basic model, extracting the intermediate features of the input image through the encoder of the visual basic model; Cross-attention fusion of the extracted intermediate features and word vectors; The fused features are input into the segmentation head for semantic segmentation prediction to obtain the category prediction result corresponding to each pixel; In each mini-batch iterative training, the randomly assigned category labels are not involved in the loss calculation; According to the specified category label, calculate the loss value between the predicted category and the true category of each pixel; During gradient backpropagation, obtain the category labels that are not involved in the loss calculation, and set the corresponding word vector gradients to zero by index; The visual basic large model is trained iteratively until the loss converges, and the corresponding word vectors and segmentation heads are saved to obtain a trained target model; According to the target model, domain generalization semantic segmentation processing is performed on the target data in the real-world target scene to be segmented to obtain segmentation results with different semantics.

2. A domain generalization semantic segmentation method according to claim 1, characterized in that: The step of cross-attentionally fusing the extracted intermediate features with the word vector comprises the following steps: The cross-attention method is used to fuse and complement the deep semantic information of visual features and the category expression information guided by word vectors; the query matrix Q is obtained by linearly transforming the word vectors; the key matrix K is obtained by linearly transforming the features; and the value matrix V is obtained by linearly transforming the features. The expression of cross attention fusion is: Among them, d k is the learnable temperature coefficient; A is the output result.

3. A domain generalization semantic segmentation method according to claim 1, characterized in that: In the step of randomly assigning category labels not to participate in the loss calculation, when randomly selecting labels, the label set of the small batch is {label i |i=1,...,N}, where N≤19 is the total number of small batch label categories. Each time, n category labels are randomly specified in the set, and the specified labels are set to not participate in the loss calculation.

4. A domain generalization semantic segmentation method according to claim 3, characterized in that: The limiting condition for the number of the specified n category labels is: Among them, n max Represents the maximum number of category labels, 1≤n≤n max .

5. The domain generalization semantic segmentation method according to claim 1, characterized in that: In the step of inputting the fused features into the segmentation head for semantic segmentation prediction to obtain the category prediction result corresponding to each pixel, mask2forme is used as the segmentation head; The calculation expression of the total model loss of the visual basic large model is: loss=h_loss_cls+h_loss_dice+h_loss_focal, where h_loss_cls is the classification loss of the prediction result, h_loss_dice is the overlap loss between the predicted mask and the true mask, and h_loss_focal is the loss of the unbalanced sample ratio.

6. The domain generalization semantic segmentation method according to claim 1, characterized in that: The method further includes: evaluating the processing result of the domain generalization semantic segmentation by using the pixel average intersection-over-union ratio as an evaluation index of the semantic segmentation; The calculation formula of the pixel average intersection-over-union ratio is: Among them, mIoU represents the average pixel intersection over union; N represents the total number of categories; TP i is the number of pixels correctly predicted as class i; FP i is the number of pixels incorrectly predicted as class i; FN i is the number of pixels of class i that are missed.

7. A domain generalization semantic segmentation system, characterized in that include: The first module is used to load the pre-trained parameters of the visual basic model, and the virtual engine generates a virtual simulation data set; The second module is used to freeze the parameters of the visual basic model, and after the simulation data set is input into the visual basic model, the intermediate features of the input image are extracted through the encoder of the visual basic model; The third module is used to perform cross-attention fusion of the extracted intermediate features and word vectors; The fourth module is used to input the fused features into the segmentation head for semantic segmentation prediction to obtain the category prediction result corresponding to each pixel; The fifth module is used to randomly assign category labels not to participate in loss calculation in each small batch iterative training; According to the specified category label, calculate the loss value between the predicted category and the true category of each pixel; The sixth module is used to obtain the category labels that are not involved in the loss calculation during gradient back propagation, and set the corresponding word vector gradient to zero by index; The seventh module is used to iteratively train the visual basic large model until the loss converges, save the corresponding word vectors and segmentation heads, and obtain the trained target model; The eighth module is used to perform domain generalization semantic segmentation processing on the target data in the real-world target scene to be segmented according to the target model to obtain segmentation results with different semantics.

8. An electronic device, characterized in that: including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Data processing method, neural network and related equipment

    CN117392488A

  • Semantic segmentation method and device, electronic equipment and storage medium

    CN118608781A