Comprehensive Calculation Method, Device and Equipment for Loss Function

By combining multiple loss functions, the problems of coherence and false positives in the segmentation of slender structures are solved, and the efficiency and accuracy of the segmentation of slender lines are achieved, and the safety and reliability of the autonomous driving system are improved.

CN119762788BActive Publication Date: 2025-06-10ZHEJIANG XINMAI SILICON CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510258224.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-10
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

The prior art is difficult to maintain its coherence when dealing with elongated structures, resulting in fractures or discontinuities in the segmentation results, which in turn affects the accurate definition of the safe area and generates a large number of false positives and the output values ​​of the segmentation model are too large or too small.

Method used

By combining multiple loss functions, including coherence loss, false positive penalty term, and regularized term loss function constraints, we comprehensively consider segmentation accuracy, coherence and false positive control, and improve the overall performance of slender line segmentation in road scenarios.

Benefits of technology

The coherence and accuracy of slender line segmentation is achieved, false positives are reduced, the accurate identification of safe areas is ensured, and the safety and reliability of the autonomous driving system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119762788B_ABST
    Figure CN119762788B_ABST
Patent Text Reader

Abstract

The present invention discloses a comprehensive calculation method, device and equipment for loss functions, which relates to the field of loss function calculation, and includes the steps of: obtaining the network input prediction results, true labels and features of each layer of the model; respectively obtaining the output results of four loss functions through the focal loss function, connectivity loss function, false positive penalty term loss function and regularization term loss function; respectively weighting the output results of the four loss functions to obtain the total loss result. Main technical solutions and effects: Through the synergistic effect of multiple loss functions, comprehensively considering segmentation accuracy, coherence, false positive control and quantization error, the model is significantly improved in multiple dimensions, has stronger adaptability, and has obvious improvement in the coherent segmentation of slender lines.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of loss function calculation, and particularly to a comprehensive loss function calculation method, device and equipment. Background Art

[0002] In autonomous driving and intelligent transportation systems, environmental perception is one of the key technologies for achieving safe driving. Among them, road scene segmentation is an important part of environmental perception, which differentiates different regions in a road image through the dividing lines between road surfaces to identify key structures such as road surfaces, curbs, green belts, fences, etc. These slender line structures play an important role in defining safe areas. For example, pedestrians and vehicles outside fences, outside green belts or on curbs are considered to be in safe areas and do not require an alarm to be triggered.

[0003] The existing loss function calculation methods mainly have the following defects:

[0004] 1. Traditional loss functions such as cross-entropy loss and Dice loss are difficult to maintain the coherence of slender structures during processing, resulting in breaks or discontinuities in the segmentation results, thereby affecting the accurate definition of safe areas.

[0005] 2. In complex road environments, the background is complex and diverse, which easily causes the model to generate a large number of false positives (i.e., misjudging non-safe areas as safe areas), which may lead to missed reports of dangerous vehicles or driving.

[0006] 3. Existing methods usually lack effective constraints on the output range of the model. The training method of the segmentation model easily makes the output values too large or too small, thereby affecting the quantization and use of the model at the side end.

[0007] 4. Most methods only focus on the accuracy of segmentation, and fail to comprehensively consider coherence, false positive control and the stability of model quantization, resulting in limited overall performance in specific application scenarios. Summary of the Invention

[0008] Object of the Invention: The object of the present invention is to solve the technical problems in the prior art that it is difficult to maintain the coherence of slender structures during processing, resulting in breaks or discontinuities in the segmentation results, thereby affecting the accurate definition of safe areas, and also generating a large number of false positives and the training method of the segmentation model easily making the output values too large or too small. The present invention provides a comprehensive loss function calculation method, device and equipment, which realizes the comprehensive consideration of segmentation accuracy, coherence and false positive control by combining multiple loss functions (including coherence loss, false positive penalty term, regularization term loss function constraint, etc.), improves the overall performance of slender line segmentation in road scenes, ensures the accurate identification of safe areas, and further improves the safety and reliability of the autonomous driving system.

[0009] One or more embodiments of this specification simultaneously relate to a comprehensive loss function calculation device, an electronic device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.

[0010] Technical solution:

[0011] In the first aspect, this application proposes a comprehensive loss function calculation method, including the steps of:

[0012] Obtain the network input prediction result, the true label, and the features of each layer of the model;

[0013] Obtain the output results of four loss functions through the focal loss function, the connectivity loss function, the false positive penalty term loss function, and the regularization term loss function respectively;

[0014] Weight the output results of the four loss functions respectively to obtain the total loss result.

[0015] Preferably, obtaining the output results of four loss functions through the focal loss function, the connectivity loss function, the false positive penalty term loss function, and the regularization term loss function respectively includes:

[0016] Activate the network input prediction result through the Sigmoid function to obtain the prediction probability value;

[0017] Flatten the prediction probability value and the true label;

[0018] Calculate the number of true positives, false positives, and false negatives;

[0019] Calculate the output result of the focal loss function.

[0020] Preferably, calculating the number of true positives, false positives, and false negatives includes the following formula;

[0021]

[0022]

[0023] ;

[0024] Wherein, is the value after flattening the prediction probability value, is the value after flattening the true label, TP is the number of true positives calculated, FP is the number of false positives calculated, and FN is the number of false negatives calculated.

[0025] Preferably, calculating the output result of the focal loss function includes:

[0026] Calculate the loss index, and the formula is as follows;

[0027] ;

[0028] The output result of the intersection loss function is calculated through the loss index, and the formula is as follows;

[0029] ;

[0030] Among them, and are weight parameters that adjust the influence of FP and FN on the Tversky index, is a very small constant used to avoid the case where the denominator is 0, controls the modulation intensity of the loss.

[0031] Preferably, the output results of the four loss functions are obtained through the focal loss function, the connectivity loss function, the false positive penalty term loss function, and the regularization term loss function respectively, including:

[0032] An average pooling operation is performed on the network input prediction result to obtain a local average value, and the formula is as follows:

[0033] ;

[0034] Calculate the difference between the prediction result and its local average value, and the formula is as follows:

[0035] ;

[0036] Among them, p is the predicted probability value obtained by activating the network input prediction result through the Sigmoid function.

[0037] Preferably, the output results of the four loss functions are obtained through the focal loss function, the connectivity loss function, the false positive penalty term loss function, and the regularization term loss function respectively, including the following formulas:

[0038] ;

[0039] Among them: N 0 is the number of background pixels, and p i is the probability that the i-th pixel in the prediction result belongs to the foreground, and i ∈ background means that i is a background pixel.

[0040] Preferably, the output results of the four loss functions are obtained through the focal loss function, the connectivity loss function, the false positive penalty term loss function, and the regularization term loss function respectively, including the following formulas:

[0041] ;

[0042] Among them, Ojk is the feature at the output end of each layer of the model, j represents the output layer, k represents the pixel index, and N j is the total number of pixels output by the j-th layer, and m represents the size that restricts the model output.

[0043] Preferably, the output results of the four loss functions are weighted respectively to obtain the total loss result, including:

[0044] ;

[0045] where λ connectivity 、λ fp 、λ reg represent the weights of the three auxiliary losses respectively.

[0046] The second unit, an embodiment of the present invention provides a loss function comprehensive calculation device, including:

[0047] An acquisition unit, configured to acquire the network input prediction result, the true label, and the features of each layer of the model;

[0048] A calculation unit, configured to obtain the output results of the four loss functions respectively through the focal loss function, the connectivity loss function, the false positive penalty term loss function, and the regularization term loss function;

[0049] A merging unit, configured to weight the output results of the four loss functions respectively to obtain the total loss result.

[0050] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor and a memory. Among them, the memory is used to store one or more computer programs; when one or more computer programs stored in the memory are executed by the processor, the electronic device can implement the method of any possible design in the first aspect above.

[0051] In a fourth aspect, the present invention provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method in any one of the above embodiments is implemented.

[0052] In a fifth aspect, an embodiment of the present invention further provides a computer program product, when the computer program product runs on an electronic device, the electronic device is enabled to execute the method of any possible design in any one of the above aspects.

[0053] Beneficial effects: Provide a comprehensive loss function, combine four separate loss functions including the focal loss function, the connectivity loss function, the false positive penalty term loss function, and the regularization term loss function, and achieve the following effects:

[0054] 1. Improve segmentation coherence: Introduce connectivity loss. Through local consistency constraints, ensure the coherence of slender lines, avoid breaks or discontinuities in the segmentation results, and thus accurately define the boundaries of the safety area;

[0055] 2. Reduce false positives: The false positive penalty term effectively suppresses the misdetection of background areas, reduces the false positive rate, improves the accuracy of the segmentation results, reduces the possibility of missed reports of dangerous targets, and enhances the reliability of the system;

[0056] 3. Enhance model stability: The regularization term restricts the model output values within a predetermined range, prevents the values from being too large or too small, and enhances the stability and consistency of the model during the side quantization process;

[0057] 4. Comprehensive optimization objective: Through the synergistic effect of multiple loss functions, comprehensively consider segmentation accuracy, false positive control, and quantization error, so that the model has significant improvements in multiple dimensions, is more adaptable, and has an obvious improvement in the coherent segmentation of slender lines. Description of the Drawings

[0058] Figure 1 It is a schematic diagram of the method framework provided by the present invention;

[0059] Figure 2 It is a schematic structural diagram of the device provided by an embodiment of the present application;

[0060] Figure 3 It is a block diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0061] To make the technical solutions of the present invention clearer, the following further describes the present invention in detail with specific embodiments in conjunction with the accompanying drawings.

[0062] Embodiment 1

[0063] To make the purpose, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meaning as understood by those of ordinary skill in the art in the field to which the present invention belongs. The words such as "including" used herein are intended to mean that the elements or items appearing before the word cover the elements or items listed after the word and their equivalents, without excluding other elements or items.

[0064] Regarding the problems existing in the prior art, such asFigure 1 As shown in the figure, a comprehensive calculation method for loss functions includes the following steps:

[0065] Step S101: Obtain the network input prediction result, the ground truth label, and the features of each layer of the model;

[0066] Step S102: Obtain the output results of four loss functions through the focal loss function, the connectivity loss function, the false positive penalty term loss function, and the regularization term loss function respectively;

[0067] Step S103: Weight the output results of the four loss functions respectively to obtain the total loss result.

[0068] Specifically, through the synergistic effect of multiple loss functions, comprehensively considering segmentation accuracy, coherence, false positive control, and quantization error, the model is significantly improved in multiple dimensions and has stronger adaptability;

[0069] When calculating the loss function, only the network input prediction result, the ground truth label, and the parameters of the features of each layer of the model need to be obtained. The prediction result is calculated on the loss function, and the output of the model is directly optimized through the post-processing mechanism, reducing the over-reliance on the intermediate representation of the feature map and being able to more directly and effectively improve the local consistency of the final result;

[0070] The network input prediction result refers to the predicted output result obtained through the forward propagation of the model, which is the road scene segmentation result predicted by the model according to the input data (for example, the segmented areas include the safe area and the non-safe area, etc.);

[0071] The ground truth label refers to the correct segmentation label given in the manual annotation or reference data, which is used to compare with the model prediction result to calculate the error;

[0072] Features of each layer: refer to the feature maps or activation values output by each layer of the network (especially the convolutional layer). These feature maps contain different levels of feature information extracted from the input image and are used to capture the spatial, edge, texture, etc. of the image. By analyzing the output of each layer, it can help the model to be further optimized.

[0073] In some preferred embodiments, obtaining the output results of four loss functions through the focal loss function, the connectivity loss function, the false positive penalty term loss function, and the regularization term loss function respectively includes:

[0074] Activating the network input prediction result through the Sigmoid function to obtain the prediction probability value;

[0075] Flatten the prediction probability value and the ground truth label;

[0076] Calculate the number of true positives, false positives, and false negatives;

[0077] Calculate the output result of the focal loss function.

[0078] Specifically, using the Sigmoid activation function, the predicted value (logits) of each pixel is converted into a probability value between 0 and 1. After converting the prediction results into probabilities, subsequent loss calculations will be based on these probabilities rather than the original classification outputs. In deep learning models, the output of an image segmentation task is usually a multi-dimensional probability map (e.g., a tensor of shape H×W×C, where H and W are the height and width of the image, and C is the number of classes). To calculate the loss function, these high-dimensional prediction results and true labels must be flattened into one-dimensional vectors for element-wise comparison. The flattening operation converts the two-dimensional (or multi-dimensional) output predictions and labels into one-dimensional vectors, enabling subsequent loss calculations to directly perform point-wise comparisons. By calculating true positives, false positives, and false negatives, the accuracy of the model's predictions can be measured, especially in cases of class imbalance. The focal loss function optimizes the model based on these metrics, particularly by adjusting the weights of false positives and false negatives;

[0079] Among them, true positives (TP) refer to the pixels predicted as the positive class and actually being the positive class, that is, the pixels predicted as the "safe area" and labeled as the "safe area";

[0080] False positives (FP): refer to the pixels predicted as the positive class but actually being the negative class, that is, the pixels predicted as the "safe area" but labeled as the "non-safe area";

[0081] False negatives (FN): refer to the pixels predicted as the negative class but actually being the positive class, that is, the pixels predicted as the "non-safe area" but labeled as the "safe area";

[0082] The focal loss function is an improvement over the traditional cross-entropy loss and is specifically designed to handle class imbalance problems. It adjusts the loss weights of difficult-to-classify samples to increase the model's attention to difficult samples.

[0083] In some preferred embodiments, calculating the number of true positives, false positives, and false negatives includes the following formula;

[0084]

[0085]

[0086] ;

[0087] Among them, is the value after the flattening operation of the predicted probability value, The value after flattening for the true label, TP is the number of true positives calculated, FP is the number of false positives calculated, and FN is the number of false negatives calculated.

[0088] In some preferred embodiments, calculating the output result of the focal loss function includes:

[0089] Calculating the loss index, with the formula as follows;

[0090] ;

[0091] Calculating the output result of the intersection loss function through the loss index, with the formula as follows;

[0092] ;

[0093] Wherein, and are weight parameters that adjust the influence of FP and FN on the Tversky index, is a very small constant used to avoid the case of a denominator of 0, controls the modulation intensity of the loss.

[0094] Specifically, the Tversky index is used to handle the class imbalance problem. Especially when the number of positive and negative classes differs greatly, by adjusting α and β, the influence of false positives and false negatives on the final loss can be flexibly controlled;

[0095] Focal Tversky Loss, based on the Tversky index, further focuses more strongly on difficult samples by adjusting the γ parameter, avoiding the model from overly focusing on easily classified samples, thereby improving the model's recognition ability for difficult samples.

[0096] In some preferred embodiments, obtaining the output results of four loss functions respectively through the focal loss function, connectivity loss function, false positive penalty term loss function, and regularization term loss function includes:

[0097] Performing average pooling operation on the network input prediction result to obtain the local average value, with the formula as follows:

[0098] ;

[0099] Calculating the difference between the prediction result and its local average value, with the formula as follows:

[0100] ;

[0101] Wherein, p is the predicted probability value obtained by activating the network input prediction result through the Sigmoid function.

[0102] Specifically, the above method is a connectivity loss function method. The current methods all perform local average pooling on the feature map in the model to obtain local features at the channel level. However, the pooling operation mainly targets the spatial level of the feature map, rather than directly optimizing the local consistency of the prediction results. By performing local average pooling on the prediction results in the loss function, the local consistency of the prediction results is strengthened. This method does not rely on the pooling layer inside the network, but directly optimizes the output of the model through a post-processing mechanism, reducing the over-reliance on the intermediate representation of the feature map and being able to more directly and effectively improve the local consistency of the final result;

[0103] AvgPool2D performs two-dimensional average pooling on the prediction results to generate the average value of the local region. This helps to calculate the mean of each local region, thus ensuring the coherence of image segmentation and avoiding breaks in the prediction results;

[0104] By calculating the difference between the predicted value of each pixel and its local average value, the connectivity loss function encourages the model to maintain the coherence of the segmentation results. This loss function helps to maintain the continuity of slender structures (such as road dividers) and avoid breaks or discontinuities in image segmentation.

[0105] In some preferred embodiments, the output results of four loss functions are obtained through the focal loss function, the connectivity loss function, the false positive penalty term loss function, and the regularization term loss function respectively, including the following formulas:

[0106] ;

[0107] Where: N 0 is the number of background pixels, p i is the probability that the i-th pixel in the prediction result belongs to the foreground, and i ∈ background means that i is a background pixel.

[0108] Specifically, the false positive penalty term effectively suppresses the false detection in the background region, reduces the false positive rate, improves the accuracy of the segmentation result, reduces the possibility of missing dangerous targets, and enhances the reliability of the system;

[0109] The purpose of this loss function is to reduce false positives, that is, to reduce the probability that the background region is mispredicted as the foreground region. For each pixel point in the background region, the model calculates the probability p i that it belongs to the background, and then sums up the overall loss according to the number N 0 of background pixels;

[0110] Calculation process: By calculating the predicted probability of background pixels and averaging it, situations where the background is misclassified as the foreground are penalized. This penalty helps improve the accuracy of the model in complex environments and avoids incorrectly detecting background regions as target regions.

[0111] In some preferred embodiments, the output results of four loss functions are obtained respectively through a focal loss function, a connectivity loss function, a false positive penalty term loss function, and a regularization term loss function, including the following formulas:

[0112] ;

[0113] Among them, O jk is the feature at the output end of each layer of the model, j represents the output layer, k represents the pixel index, N j is the total number of pixels output by the j-th layer, and m represents the size that limits the model output.

[0114] Specifically, the above scheme calculates the loss function for the regularization term loss function. The purpose of this loss function is to limit the output of each layer through regularization means, keeping its value within a reasonable range, avoiding the situation where the model output value is too large or too small, which is crucial for the model during side-end quantization deployment because extreme output values may lead to loss of accuracy;

[0115] Calculation process: For the output of each layer, by taking the absolute value and subtracting the maximum value m, if it exceeds the range (i.e., the part greater than m), the loss is calculated. This loss helps ensure that the model's output does not cause unstable prediction results due to over-activating certain neurons;

[0116] The regularization term of the present invention regularizes the output of each layer of the network structure, rather than calculating the regularization loss for the output result of the network and the binary label. The reason for adopting this regularization is that in the segmentation task, due to the need to distinguish the differences between the foreground and background in the image, the output values of the network often vary greatly among different pixels, which may lead to loss of accuracy after the model is deployed for side-end quantization. This scheme applies regularization to the output of each layer, ensuring that the entire network can maintain high accuracy during the quantization process. The regularization term limits the model output value within a predetermined range, preventing it from being too large or too small, and enhancing the stability and consistency of the model during the side-end quantization process.

[0117] In some preferred embodiments, the output results of the four loss functions are weighted respectively to obtain the total loss result, including:

[0118] ;

[0119] Among them, λ connectivity 、λ fp 、λreg They respectively represent the weights of three auxiliary losses.

[0120] Specifically, weight 1 (λ connectivity ) is used to control the intensity of local consistency of the prediction result;

[0121] Increasing weight 1: will strengthen the local consistency of the model output and reduce the breaks in the middle of the slender lines (such as the safety area segmentation line) in the prediction result; decreasing weight 1: will make the local consistency weaker, allowing more detailed information to be retained in the segmentation result, thereby reducing the risk of overfitting;

[0122] Weight 2 (λ fp ) is used to control the intensity of false positive (false detection) suppression;

[0123] Increasing weight 2: can effectively suppress the appearance of false positives, especially in areas where it is difficult to distinguish (such as the dry-wet boundary on a rainy day or lane lines under a complex background); decreasing weight 2: will reduce the intensity of false positive suppression, which can increase the detection rate and reduce the risk of missed detection.

[0124] Weight 3 (λ reg ) is used to control the intensity of the regularization term.

[0125] Increasing weight 3: can more strictly limit the output range of the model to ensure that the output value does not exceed the specified interval. This helps to avoid large errors during the quantization process and ensure the accuracy of the model output. Decreasing weight 3: will weaken the effect of regularization, allowing the model to output a larger range, thereby possibly improving the flexibility and adaptability of the model. Therefore, reasonable adjustment of each weight is required to find the best balance among accuracy, robustness, and efficiency according to the requirements of the task, ensuring that the model can exhibit optimal performance in specific application scenarios.

[0126] In some preferred embodiments, in combination with Figure 2 , the embodiment of the present invention provides a comprehensive loss function calculation device, including:

[0127] An acquisition unit 301, configured to acquire the network input prediction result, the true label, and the features of each layer of the model;

[0128] A calculation unit 302, configured to obtain the output results of four loss functions respectively through a focal loss function, a connectivity loss function, a false positive penalty term loss function, and a regularization term loss function;

[0129] A merging unit 303, configured to weight the output results of the four loss functions respectively to obtain a total loss result.

[0130] All relevant content of each step involved in the above method embodiments can be cited in the function descriptions of the corresponding functional modules, and will not be elaborated here.

[0131] In some other embodiments of the present invention, embodiments of the present invention disclose an electronic device 400, as Figure 3 shown, the electronic device may include: one or more processors 401; a memory 402; a display 403; one or more applications (not shown); and one or more computer programs 404. The above components may be connected through one or more communication buses 405. Wherein the one or more computer programs 404 are stored in the above memory 402 and configured to be executed by the one or more processors 401. The one or more computer programs 404 include instructions, and the above instructions can be used to execute the steps in Figures 1 to 2 and the corresponding embodiments.

[0132] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working processes of the above-described system, device, and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here.

[0133] In each embodiment of the present invention, each functional unit may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0134] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk, or optical disc that can store program codes.

[0135] As described above, it is only the specific implementation manner of the embodiments of the present invention, but the protection scope of the embodiments of the present invention is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of the present invention should be covered within the protection scope of the embodiments of the present invention. Therefore, the protection scope of the embodiments of the present invention shall be subject to the protection scope of the claims.

Claims

1. A loss function comprehensive calculation method, applied to image segmentation, characterized in that: Includes steps: Get the network input prediction results, true labels, and features of each layer of the model; The output results of four loss functions are obtained through the focus loss function, connectivity loss function, false positive penalty loss function, and regularization loss function, including: Get the output result of the connectivity loss function, perform average pooling on the network input prediction result, and get the local average value. The formula is as follows: ; Calculate the difference between the prediction result and its local mean, the formula is as follows: ; Among them, p is the predicted probability value obtained by activating the network input prediction result through the Sigmoid function; The output result of the false positive penalty loss function is obtained, and the formula is as follows: ; Where: N0 is the number of background pixels, p i is the probability that the i-th pixel in the prediction result belongs to the foreground, i∈background means that i is a background pixel; The output result of the regularization loss function is obtained, and the formula is as follows: ; Among them, O jk is the feature of each layer output of the model, j represents the output layer, k represents the pixel index, N j is the total number of pixels output by the jth layer, and m represents the size of the model output; The output results of the four loss functions are weighted respectively to obtain the total loss result.

2. The method according to claim 1, characterized in that The output results of four loss functions are obtained through the focus loss function, connectivity loss function, false positive penalty loss function, and regularization loss function, including: The network input prediction result is activated by the Sigmoid function to obtain the predicted probability value; Flatten the predicted probability values ​​and true labels; Calculate the number of true positives, false positives, and false negatives; The output of the focal loss function is calculated.

3. The method according to claim 2, characterized in that Calculating the number of true positives, false positives, and false negatives involves the following formula; ; ; ; in, is the value after flattening the predicted probability value. It is the value after flattening the real label, TP is the number of true positives, FP is the number of false positives, and FN is the number of false negatives.

4. The method according to claim 3, characterized in that The output results of the focal loss function are calculated, including: Calculate the loss index, the formula is as follows; ; The output of the focal loss function is calculated by the loss index. The formula is as follows: ; in, and is the weight parameter that adjusts the impact of FP and FN on the Tversky index, is a constant used to avoid the denominator being 0. Controls the modulation intensity of the loss.

5. The method according to claim 4, characterized in that The output results of the four loss functions are weighted to obtain the total loss result, including: ; Among them, λ connectivity , fp , reg Represent the weights of the three auxiliary losses respectively.

6. A loss function comprehensive calculation device used in the loss function comprehensive calculation method according to any one of claims 1 to 5, characterized in that: include: The acquisition unit is used to obtain the network input prediction results, true labels, and each layer feature of the model; A computing unit, used to obtain output results of four loss functions through a focal loss function, a connectivity loss function, a false positive penalty loss function, and a regularization loss function; The merging unit is used to weight the output results of the four loss functions respectively to obtain the total loss result.

7. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program executable on the processor, and when the computer program is executed by the processor, the processor implements the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • U-Net image segmentation method based on FTL loss function and attention

    CN113538458A

  • Multi-sensor fusion vehicle-road collaborative sensing method for automatic driving

    CN114821507A