Lane detection model post-training quantification method and device based on semantic sensitivity

By introducing semantic sensitivity into the lane detection model and dynamically adjusting the weight coefficient and loss function, the problem of semantic information being ignored in the prior art is solved, and the effect of improving detection accuracy while reducing resource requirements is achieved.

CN120339977APending Publication Date: 2025-07-18BEIHANG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510044128.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-22
Filing Date
2025-01-10
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing post-training quantization method ignores semantic information in the lane detection model, resulting in a decrease in detection accuracy and unable to effectively reduce the demand for computing resources and storage space.

Method used

By collecting the label-free training data set, the semantic sensitivity of different semantic output heads is calculated, the weight coefficient is dynamically adjusted, and the area sensitivity loss function is used for post-training quantization until the model converges, and the simulated quantization weight is converted into fixed-point numbers.

Benefits of technology

While reducing computing resources and storage space, the accuracy and efficiency of the lane detection model are improved, and the calculation amount and training time are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339977A_ABST
    Figure CN120339977A_ABST
Patent Text Reader

Abstract

The invention discloses a lane detection model post-training quantification method and device based on semantic sensitivity. The post-training quantification method for the lane detection model comprises the following steps: firstly, collecting an unlabeled training data set of the lane detection model; secondly, simulating and calculating semantic sensitivities of different semantic output heads in the same lane detection model by utilizing the lane deformation score and the noise level; based on the semantic sensitivities of the different semantic output heads, dynamically adjusting the weight coefficients of the different semantic output heads in a post-training quantization process, and performing post-training quantization through a region sensitivity loss function until the lane detection model converges; and finally, deploying the lane detection model after post-training quantization, and converting the weight of analog quantization into a fixed point number so as to reduce required computing resources and storage space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for post-training quantization of a lane detection model based on semantic sensitivity, and also relates to a corresponding device for post-training quantization of a lane detection model, belonging to the technical field of autonomous driving. Background Art

[0002] In the technical field of autonomous driving, lane detection is a core function. It identifies the lane lines on the road surface through a camera installed in front of the vehicle, helps the vehicle determine its own position on the road, and thus controls the direction of the vehicle to ensure that the vehicle always stays in the center of the lane during driving. The accuracy and real-time performance of lane detection directly affect the driving safety of autonomous vehicles. Therefore, people have begun to widely apply deep neural networks (DNNs) in lane detection technology.

[0003] The application of deep neural networks in lane detection technology is diverse, including feature extraction, spatial information processing, improvement of model generalization ability, application of attention mechanism, optimization of real-time performance, and multi-task learning. With the excellent performance of deep neural networks in image recognition and classification tasks, more and more lane detection models are constructed based on deep neural networks. However, deep neural networks usually require a large amount of computing resources and storage space, which poses a huge challenge to the edge computing devices or in-vehicle systems commonly used in autonomous vehicles. To solve this problem, the neural network model quantization technology has emerged, which reduces the demand for computing resources by reducing the number of bits of weights and activation functions in the neural network model.

[0004] Currently, the neural network model quantization methods are mainly divided into two categories: quantization-aware training (QAT) methods and post-training quantization (PTQ) methods. QAT methods introduce quantization during the training process of the neural network model, while PTQ methods perform quantization after the neural network model is trained. Although QAT methods can retain the performance of the neural network model to a certain extent, they require re-training the neural network model, consuming a large amount of computing resources and time. In contrast, PTQ methods have received increasing attention in the industry due to their simple operation and low cost.

[0005] However, when the existing PTQ methods are applied to lane detection models, they face an important problem: they usually ignore the semantic information contained in the output of the lane detection model, especially in complex post-processing steps. The lane detection model not only needs to identify the lane lines in the image, but also needs to understand the semantic attributes of the lane lines, such as position, shape, and inclination. These semantic information are extremely vulnerable to damage during the quantization process, thus affecting the accuracy of lane detection. Summary of the Invention

[0006] The primary technical problem to be solved by the present invention is to provide a post-training quantization method for a lane detection model based on semantic sensitivity.

[0007] Another technical problem to be solved by the present invention is to provide a post-training quantization device for a lane detection model based on semantic sensitivity.

[0008] To achieve the above-mentioned invention purpose, the present invention adopts the following technical solutions:

[0009] According to the first aspect of the embodiments of the present invention, a post-training quantization method for a lane detection model based on semantic sensitivity is provided, including the following steps:

[0010] S1, collecting an unlabeled training data set of the lane detection model;

[0011] S2, using the lane deformation score and the noise level to simulate and calculate the semantic sensitivity of different semantic output heads in the same lane detection model;

[0012] S3, based on the semantic sensitivity of different semantic output heads, dynamically adjusting the weight coefficients of different semantic output heads in the post-training quantization process, and performing post-training quantization through the regional sensitivity loss function until the lane detection model converges;

[0013] S4, deploying the lane detection model after post-training quantization, and converting the simulated quantized weights into fixed-point numbers to reduce the required computing resources and storage space.

[0014] Preferably, in the step S1, the unlabeled training data set includes lane images under various road environments and the corresponding outputs X1,..., X of the lane detection model p ; where X i represents the concatenated vector corresponding to the outputs of all semantic output heads of the lane detection model.

[0015] Preferably, in the step S2, different sizes of random noise are respectively applied to the concatenated vectors corresponding to the outputs of different semantic output heads, and repeated multiple times at each noise level, and the semantic sensitivity of different semantic output heads to semantic information under different noise levels is obtained by combining the lane deformation scores obtained on the unlabeled training data set.

[0016] Preferably, the lane deformation score is calculated by the following formula:

[0017]

[0018] Among them, Score is the lane deformation score; M represents the set of matching points after the two lane lines are interpolated to the size of the input image; d(p0, p1) is the distance between the matching points; b(M) is used to represent the boundary length normalization to adapt to lanes of different lengths; n is the number of mismatched points; v is a fixed penalty value.

[0019] Preferably, the semantic sensitivity of different semantic output heads to semantic information is calculated by the following formula:

[0020]

[0021] Among them, w i is a coefficient calculated according to the semantic sensitivity of different semantic output heads, m is the equivalent number of bits corresponding to the noise level, and L is a function for calculating the deformation coefficient of the lane lines decoded by different semantic output heads.

[0022] Preferably, in the step S3, post-training quantization is performed on the lane detection model. According to the output result of each quantization step, the current noise level is estimated, and then a look-up table is made using the semantic sensitivity obtained in the step S2, thereby obtaining the relative relationship between the semantic sensitivities of different semantic output heads.

[0023] Preferably, in the look-up table obtained in the step S2, the semantic sensitivity corresponding to the corresponding semantic output head and noise level is found through the table, and the weight coefficient assigned to each semantic output head is calculated through normalization.

[0024] Preferably, after calculating the quantization output of each semantic output head, the region sensitivity loss is calculated according to the quantization output and the features output by the original lane detection model. After being weighted by the weight coefficients of each semantic output head, the overall loss is obtained; the parameters of the lane detection model are updated through gradient descent and the set learning rate until the lane detection model converges.

[0025] Preferably, in the step S4, the simulated quantization parameters and weight coefficients in the post-training quantization are converted into fixed-point numbers. After comparing the parameters and activation values layer by layer, after confirming that the quantized lane detection model is correct, the deployment of the lane detection model is completed.

[0026] According to the second aspect of the embodiments of the present invention, a post-training quantization device for a lane detection model based on semantic sensitivity is provided, including a processor and a memory. The processor reads the computer program in the memory and is used to execute the above-mentioned post-training quantization method for the lane detection model.

[0027] Compared with the prior art, for the first time, the present invention quantifies and models the lane detection model from the perspective of semantic sensitivity, utilizes the knowledge of the lane detection model itself in a label-free manner, and guides the effective quantization of the neural network model. It effectively captures the information about lane semantics in the neural network model. During the inference process, compared with the baseline model using the block reconstruction algorithm, it can effectively improve the accuracy of the lane detection model after quantization while significantly reducing the computational load and training time. Description of the Drawings

[0028] Figure 1 It is a logical framework diagram of the post-training quantization method for the lane detection model based on semantic sensitivity in an embodiment of the present invention;

[0029] Figure 2 It is a flowchart of the post-training quantization method for the lane detection model based on semantic sensitivity in an embodiment of the present invention;

[0030] Figure 3 It is a schematic diagram of common lane deformation types caused by slight perturbations in an embodiment of the present invention. Among them, (a) Bending: The terminal is unexpectedly offset; (b) Spike: The middle is offset; (c) Dislocation: Lane points are missing or redundant. Any lane deformation can be expressed as a combination of these three deformations; (d) Measuring the distortion between lanes using two types of point relationships: Matching and non-matching.

[0031] Figure 4 It is a schematic diagram of the post-training quantization device for the lane detection model based on semantic sensitivity in an embodiment of the present invention. Detailed Embodiment

[0032] The technical content of the present invention will be described in detail below in conjunction with the drawings and specific embodiments.

[0033] First of all, it should be noted that in the lane detection model, each semantic output head i contains two different functions: S i (·) and C i (·). The former function generates an output with physical meaning, which we call semantics. These semantics are linked to physical properties such as distance and angle. The latter function C i (·) generates a corresponding confidence output for the semantics of each semantic output head.

[0034] Such as Figure 1As shown in the figure, the main technical concept of the embodiment of the present invention lies in introducing the concept of Semantic Sensitivity in the post-training quantization process of the lane detection model. Among them, the SemanticGuided Focus algorithm is used to combine semantics to generate a practical proxy mask, thereby guiding the existing PTQ method to optimize the key areas therein to solve the semantic sensitivity within each semantic output head (i.e., the sensitivity within the head). The sensitivity within the head here reflects the high sensitivity of a limited number of foreground (lane) areas to quantization noise during the post-processing process. On the other hand, considering the semantic sensitivity between different semantic output heads (i.e., the sensitivity between heads), the Sensitivity Aware Selection algorithm is used to perceive the real-time sensitivity of each semantic output head according to the lane deformation score, thereby effectively adjusting the optimization target.

[0035] As Figure 2 shown, the post-training quantization method for the lane detection model provided by the embodiment of the present invention at least includes the following steps: S1, collecting an unlabeled training data set of the lane detection model; S2, using the lane deformation score and the noise level to simulate and calculate the semantic sensitivity of different semantic output heads in the same lane detection model; S3, based on the semantic sensitivity of different semantic output heads, dynamically adjusting the weight coefficients of different semantic output heads during the post-training quantization process, and performing post-training quantization through the regional sensitivity loss function until the lane detection model converges; S4, deploying the lane detection model after post-training quantization to ensure that while maintaining the original lane detection performance, converting the simulated quantization weights into fixed-point numbers to reduce the required computing resources and storage space.

[0036] Next, the specific implementation process of each step will be described separately:

[0037] First, in step S1, an unlabeled training data set of the lane detection model is collected. These training data sets do not need to be labeled to avoid affecting the subsequent calculation of semantic sensitivity. In an embodiment of the present invention, the unlabeled training data set may only require about 200 pictures containing lane lines, including lane images in various road environments and the corresponding concatenated vectors X1,…,X p . Among them, X i represents the concatenated vector corresponding to the outputs of all semantic output heads of the lane detection model, and p is a positive integer.

[0038] Next, in step S2, different magnitudes of random noise are applied to the concatenated vectors corresponding to different semantic output heads respectively, and repeated multiple times at each noise level. Combining the lane deformation scores obtained on the unlabeled training dataset, the semantic sensitivities of different semantic output heads to semantic information are obtained at different noise levels. In the lane detection task, the semantic information here mainly refers to the description of the lane, such as the positions of key points on the lane, the parameters of the fitted curve, etc., which are decoded and output by the post-processing module in the lane detection model. Quantifying the impact on it is ultimately reflected in the degree of lane deformation.

[0039] Figure 3 In the embodiment of the present invention, it is a schematic diagram of common lane deformation types caused by slight perturbation. Among them, (a) bending: the terminal is accidentally offset; (b) spike: the middle is offset; (c) dislocation: lane points are missing or redundant. Any lane deformation can be expressed as a combination of these three deformations; (d) measuring the distortion between lanes with two types of point relationships: matching and non-matching. In step S2, the formula for calculating the lane deformation score Score is as follows:

[0040]

[0041] Among them, the lane deformation score Score is used to reflect the moving distance of all matching points from the perturbed lane to the original lane, and the score of non-matching points is a fixed penalty value v. M represents the set of matching points after the two lane lines are interpolated to the input image size, and d(p0, p1) is the distance between matching points; b(M) is used to represent the boundary length normalization to adapt to lanes of different lengths. n is the number of non-matching points.

[0042] Furthermore, the semantic sensitivities of different semantic output heads to semantic information are calculated by the following formula:

[0043]

[0044] Among them, w i is the coefficient calculated according to the semantic sensitivities of different semantic output heads, m is the equivalent number of bits of the corresponding noise level, and L is a function for calculating the deformation coefficient of the lane lines decoded by different semantic output heads. The deformation coefficient here statistically represents the average distance between matching points of two lane lines X and Y at the same height:

[0045]

[0046] Among them, M represents the set of matching points after the two lane lines are interpolated to the input image size, M’ represents the set of single points where matching points cannot be found, and v is a fixed penalty value, meaning that points exceeding this distance are considered not to be on the same lane line.

[0047] Further, in step S3, post-training quantization is performed on the lane detection model. According to the output result of each quantization step, the current noise level is estimated. Then, a look-up table is made using the semantic sensitivity obtained in step S2, and the relative relationship between the semantic sensitivities of different semantic output heads is obtained therefrom. In an embodiment of the present invention, given the original output X and the quantized output X' of a certain semantic output head, its noise level can be estimated as:

[0048]

[0049] In the look-up table obtained in step S2, the semantic sensitivity corresponding to the semantic output head and the noise level is found, and the weight coefficients assigned to each semantic output head are calculated through normalization. After calculating the quantized output of each semantic output head, the regional sensitivity loss is calculated based on the quantized output and the features output by the original lane detection model. After being weighted by the weight coefficients of each semantic output head, the overall loss is obtained. Further, the parameters of the lane detection model are updated through gradient descent and the set learning rate until the lane detection model converges.

[0050] Finally, in step S4, the simulated quantization parameters and weight coefficients in the post-training quantization are converted into fixed-point numbers. After layer-by-layer comparison of the parameters and activation values, after confirming that the quantized lane detection model is correct, the deployment of the lane detection model is completed. In the above process, if there is a large quantization loss, the learning rate parameters can be traversed from 1e-1 to 1e-6 in a magnitude of 1e-0.3, and step S3 is re-executed to select the result with the optimal performance.

[0051] In an embodiment of the present invention, according to the total loss function l sem the parameters of the lane detection model are updated to make the lane detection model converge preliminarily; then, steps S3 and S4 are iteratively executed until the quantized lane detection model with the highest accuracy is selected. Among them, the total loss function l sem The calculation formula is as follows:

[0052]

[0053] Among them, w i is the coefficient calculated according to the semantic sensitivity of different semantic output heads, and l i is the regional sensitivity loss function.

[0054] In an embodiment of the present invention, the above-mentioned regional sensitivity loss l i is calculated using the following formula:

[0055]

[0056] Among them, C nis the lane line confidence corresponding to the nth image; X n is the output feature of the original lane detection model corresponding to the nth image; X′ n is the output feature of the quantized lane detection model corresponding to the nth image.

[0057] Next, the Semantic Guided Focus algorithm and the Sensitivity Aware Selection algorithm used in the embodiments of the present invention will be specifically described.

[0058] In the Semantic Guided Focus algorithm, the inventors considered the relationship between semantics (head output) and post - processing. Some of the semantics are related to the foreground region in post - processing, while others correspond to the background. Therefore, we are committed to enhancing the accurate expression of pixels related to the foreground region and suppressing the expression of background pixels, so as to distinguish whether each element within each semantic output head is for the foreground region (lane) or the background.

[0059] On this basis, we first focus on the semantic term in the following formula and introduce a masking function M to achieve the distinction:

[0060]

[0061] Among them, the masking function M removes the error terms related to the background and retains the elements of the foreground region. Then, we multiply them element - by - element, and the semantic optimization objective becomes:

[0062]

[0063] We incorporate the reconstruction loss into the confidence value of each semantic output head. To prevent elements related to the background region from becoming foreground, especially in the presence of a large amount of noise, we enhance the alignment of the confidence value output by adopting a new parameter λ>1, as shown in the following formula:

[0064]

[0065] On the other hand, the Sensitivity Aware Selection algorithm is shown in the following table:

[0066]

[0067] Specifically, the Sensitivity Aware Selection algorithm is used to select semantic output heads in the floating - point model (FP model). It requires inputting the semantic output head set H, the hyperparameter k, and the calibration dataset D. For each semantic output head i, first calculate the full - precision lane L, and then use the quantized version S i(x) replaces the original semantic function S ^i (x) to perturb the model. Then, it calculates the perturbed lane L^, and uses the Score calculation formula to calculate the similarity s of (L, L′). Finally, it updates the score Score of the i-th semantic output head according to the calculated similarity s i . After all semantic output heads are processed, the semantic output heads are sorted according to the scores, and the top k semantic output heads are returned as the results

[0068] Based on the above post-training quantization method of the lane detection model based on semantic sensitivity, the present invention further provides a post-training quantization device for the lane detection model based on semantic sensitivity. As Figure 4 shown, the post-training quantization device for the lane detection model includes one or more processors 41 and a memory 42. Among them, the memory 42 is coupled to the processor 41 and is used to store one or more programs. When the one or more programs are executed by the one or more processors 41, the one or more processors 41 implement the post-training quantization method of the lane detection model based on semantic sensitivity in the above embodiments

[0069] Among them, the processor 41 is used to control the overall operation of the post-training quantization device for the lane detection model based on semantic sensitivity to complete all or part of the steps of the post-training quantization method of the lane detection model based on semantic sensitivity in the above. In the embodiments of the present invention, the processor 41 is preferably a GPU (Graphics Processing Unit), but it can also be an FPGA (Field Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), a DSP (Digital Signal Processor), etc. The memory 42 is used to store various types of data to support the operation of the post-training quantization method of the lane detection model based on semantic sensitivity. These data may include, for example, instructions for any application program or method operating on the post-training quantization device for the lane detection model based on semantic sensitivity, and application program-related data

[0070] The memory 42 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, etc

[0071] In an exemplary embodiment, the post-training quantization device for the lane detection model based on semantic sensitivity can be specifically implemented by a computer chip or an entity, or by a product with certain functions, and is used to execute the above post-training quantization method for the lane detection model based on semantic sensitivity, and achieve the same technical effects as those of the above method. A typical embodiment is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, an in-vehicle human-machine interaction device, a checkpoint inspection device, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0072] In another exemplary embodiment, the present invention also provides a computer-readable storage medium including program instructions. When the program instructions are executed by a processor, the steps of the post-training quantization method for the lane detection model based on semantic sensitivity in any one of the above embodiments are implemented. For example, the computer-readable storage medium can be a memory including program instructions, and the above program instructions can be executed by the processor of the post-training quantization device for the lane detection model based on semantic sensitivity to complete the above post-training quantization method for the lane detection model based on semantic sensitivity, and achieve the same technical effects as those of the above method.

[0073] Compared with the prior art, the present invention first quantizes and models the lane detection model from the perspective of semantic sensitivity, utilizes the knowledge of the lane detection model itself in a label-free manner, and guides the effective quantization of the neural network model. It effectively captures the information about lane semantics in the neural network model. Compared with the baseline model using the block reconstruction algorithm during the inference process, it can effectively improve the accuracy of the quantized lane detection model while significantly reducing the computational amount and training time.

[0074] The above provides a detailed description of the post-training quantization method and device for the lane detection model based on semantic sensitivity of the present invention. For those of ordinary skill in the art, any obvious changes made without departing from the substantial content of the present invention will constitute an infringement of the patent right of the present invention and will bear corresponding legal responsibilities.

Claims

1. A post-training quantization method for a lane detection model based on semantic sensitivity, characterized in that It includes the following steps: S1. Collect the unlabeled training dataset of the lane detection model; S2. Use the lane deformation score and the noise level to simulate and calculate the semantic sensitivities of different semantic output heads in the same lane detection model; S3. Based on the semantic sensitivities of different semantic output heads, dynamically adjust the weight coefficients of different semantic output heads during the post-training quantization process, and perform post-training quantization through the regional sensitivity loss function until the lane detection model converges; S4. Deploy the lane detection model after post-training quantization, and convert the simulated quantization weights into fixed-point numbers to reduce the required computing resources and storage space.

2. The post-training quantization method for the lane detection model according to claim 1, wherein In the step S1, the tagless training data set includes lane images in various road environments and their corresponding concatenated vectors X1, …, X output by the lane detection model. p ; where X i represents the concatenated vector corresponding to the outputs of all semantic output heads of the lane detection model.

3. The post-training quantization method for the lane detection model according to claim 2, wherein In the step S2, different magnitudes of random noise are respectively applied to the concatenated vectors corresponding to different semantic output heads, and repeated multiple times at each noise level. Combining the lane deformation scores obtained on the unlabeled training dataset, the semantic sensitivities of different semantic output heads to semantic information are obtained at different noise levels.

4. The lane detection model post-training quantization method according to claim 3, characterized in that The lane deformation score is calculated by the following formula: where Score is the lane deformation score; M represents the set of matching points after interpolating two lane lines to the input image size; d(p0, p1) is the distance between the matching points; b(M) is used to represent the boundary length normalization to adapt to lanes of different lengths; n is the number of mismatched points; and v is a fixed penalty value.

5. The post-training quantization method for the lane detection model according to claim 3, wherein The semantic sensitivities of different semantic output heads to semantic information are calculated by the following formula: where w i is a coefficient calculated based on the semantic sensitivity of the output head for different semantics, m is the equivalent number of bits for the corresponding noise level, and L is a function for calculating the deformation coefficient of the lane line decoded by different semantic output heads.

6. The post-training quantization method for the lane detection model according to claim 1, wherein In the step S3, perform post-training quantization on the lane detection model, estimate the current noise level according to the output result of each quantization step, and then use the semantic sensitivities obtained in the step S2 to create a lookup table, thereby obtaining the relative relationship between the semantic sensitivities of different semantic output heads.

7. The method for post-training quantization of the lane detection model according to claim 6, wherein In the lookup table obtained in the step S2, look up the semantic sensitivities at the corresponding semantic output heads and noise levels, and calculate the weight coefficients assigned to each semantic output head through normalization.

8. The post-training quantization method for the lane detection model according to claim 7, characterized in that After calculating the quantization outputs of each semantic output head, calculate the regional sensitivity loss according to the quantization output and the features output by the original lane detection model. After weighting by the weight coefficients of each semantic output head, obtain the overall loss; update the parameters of the lane detection model through gradient descent and the set learning rate until the lane detection model converges.

9. The post-training quantization method for the lane detection model according to claim 1, wherein In the step S4, convert the simulated quantization parameters and weight coefficients in the post-training quantization into fixed-point numbers. After comparing the parameters and activation values layer by layer, and confirming that the lane detection model after quantization is correct, complete the deployment of the lane detection model.

10. A post-training quantization device for a lane detection model based on semantic sensitivity, characterized in that It includes a processor and a memory. The processor reads the computer program in the memory and is used to execute the post-training quantization method of the lane detection model according to any one of claims 1 to 9.

Citation Information

Cited By

  • Lane line real-time identification method and system, electronic equipment and storage medium

    CN121482736A