Semantic information prompt-based front scene display method and device

By integrating semantic segmentation into the Stixel-world algorithm, the method accurately differentiates road and obstacle intersections, enhancing scene representation and driving safety through improved Stixel generation.

CN120318785APending Publication Date: 2025-07-15ZHONGKE HUIYAN (TIANJIN) ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510243702.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The multi-layer Stixel-world method causes erroneous identification at the junction of roads and obstacles due to the same depth.

Method used

By acquiring the original camera stereoscopic image, a neural network is used to generate a parallax map and a three-class semantic segmentation map, and semantic segmentation information is added as prompt information in the multi-layer Stixel-world dynamic programming process to improve the accuracy of multi-layer Stixel.

Benefits of technology

The identification accuracy of multi-layer Stixel-world at the junction of roads and obstacles has been improved, and the display accuracy of front-vehicle scenes has been improved through the introduction of semantic information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318785A_ABST
    Figure CN120318785A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for displaying a scene in front of a vehicle based on semantic information prompt, which are used for improving the dividing accuracy of obstacles and roads in an original multi-layer stixel-word calculation process. The method for displaying the scene in front of the vehicle based on the semantic information prompt comprises the steps of obtaining an original camera stereo image, and obtaining a disparity map and a three-classification semantic segmentation map through a neural network; using the disparity map to generate multilayer Stixles according to a multilayer Stixle-word algorithm, and using a three-classification semantic segmentation map to obtain semantic segmentation information; semantic segmentation information is added to serve as prompt information in the multi-layer Stixle-word dynamic programming solving process, and the multi-layer Stixles with the higher accuracy rate is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of assisted driving, and particularly to a method and device for displaying a front vehicle scene based on semantic information prompt. Background Art

[0002] Stixel-world is a method for comprehensively and quickly representing the front vehicle scene while retaining depth information. The multi-layer Stixel-world method represents the front vehicle scene as three categories: road, obstacle, and sky according to depth information. Since the depths of the road and the obstacle are the same at the junction, the multi-layer Stixel-world will make mistakes at the junction of the road and the obstacle.

[0003] In view of this, the present invention is proposed. Summary of the Invention

[0004] The main purpose of the present invention is to disclose a method and device for displaying a front vehicle scene based on semantic information prompt, which is used to solve the problem that the multi-layer Stixel-world makes mistakes at the junction of the road and the obstacle because the depths of the road and the obstacle are the same in the prior art.

[0005] To achieve the above object, according to one aspect of the present invention, a method for displaying a front vehicle scene based on semantic information prompt is provided, and the following technical solutions are adopted:

[0006] The method for displaying a front vehicle scene based on semantic information prompt includes: obtaining an original camera stereo image, obtaining a disparity map and a three-class semantic segmentation map through a neural network; using the disparity map to generate multi-layer Stixles according to the multi-layer Stixle-world algorithm, and obtaining semantic segmentation information using the three-class semantic segmentation map; adding the semantic segmentation information as prompt information in the process of solving the multi-layer Stixle-world dynamic programming to obtain multi-layer Stixles with higher accuracy.

[0007] Further, the generating multi-layer Stixles using the disparity map according to the multi-layer Stixle-world algorithm includes: given the calibrated left image I with size w*h in the stereo image pair and the possible annotation set , the multi-layer Stixel world corresponds to splitting the annotation L belonging to into the category set column by column, that is, the road g, the obstacle o, and the sky s;

[0008]

[0009] represents the Annotations of columns and contain multiple segments , the total number of segments in each column is given by , the maximum value of the segment is implicitly limited by the height of the image, i.e., , h is the height of the image, w is the width of the image, and in the formula The bottom point and The vertex mark the starting and ending points of the segment , each segment is assigned to a category , is an arbitrary function for calculating the disparity of this segment; , and is only determined by the row number v, , all segments and are vertically adjacent, ensuring that each pixel in the image will only have one label; in the annotation step, all segments are designated as road, obstacle, and sky, so it is assumed that all segments are approximately piecewise planes in three-dimensional space. The selection of the function is simplified to a set of linear functions. The segments assigned to obstacles are assumed to have a constant disparity, and the corresponding function is , is the representative disparity within the segment. The surface of the segment labeled as road is modeled as a linear function; , is the expected road disparity gradient, is the row coordinate of the horizon, and can both be extracted from the existing camera geometry. The disparity function of the segment labeled as sky is .

[0010] Furthermore, the obtaining of semantic segmentation information using the three-class semantic segmentation map includes:

[0011] The semantic segmentation map obtained by the semantic segmentation neural network contains three values: 0 for road, 1 for obstacle, and 2 for sky. For a column Se of the image, the semantic segmentation information of this column is a permutation of the three values 0, 1, and 2, i.e.,

[0012]

[0013] Construct a semantic transition table T according to whether the semantic value of the current pixel is equal to that of the previous pixel. If the semantic values are equal, the corresponding position in the transition table is 0; if they are not equal, it is 1, i.e.,

[0014]

[0015] Then calculate the cumulative value of the semantic transition table from the bottom to this pixel at each pixel of the semantic transition table to construct a semantic transition cumulative table A, Indicates from this column row to row, the number of semantic transitions experienced by the image:

[0016]

[0017]

[0018] If the value of the semantic transition accumulation table at a certain pixel is 0, that is = 0, it means that from this column row to row, this fragment should be an entire stixel.

[0019] Furthermore, adding semantic segmentation information as hint information in the multi-layer Stixle-world dynamic programming solution process to obtain multi-layer Stixles with higher accuracy includes:

[0020] Finding the label with the highest probability under the condition of a given disparity input P :

[0021]

[0022] Through Bayes' formula, it can be expressed as:

[0023]

[0024] is ignored because this term is just a normalization factor, represents the probability of the input disparity P given the label L, while represents the probability of the label L occurring, under the condition of assuming that the labels between columns and the disparities between columns of the image are independent, and can be further decomposed:

[0025]

[0026]

[0027] .

[0028] Assume that the width of the image I is w, P is the disparity map of I, and L is the label map of I, is the disparity of the u-th column, is the label of the u-th column.

[0029] Further, the method of adding semantic segmentation information as a hint during the multi-layer Stixle-world dynamic programming solution to obtain multi-layer Stixles with higher accuracy further includes:

[0030] Define a scaling function , where v is the number of rows and c is the semantic category. If there is no semantic change in a certain column of the image from the lowest row to the v-th row, the function value is a constant less than 1; otherwise, the function value is 1.

[0031] The dynamic programming for each column of the image can be formulated as: The cost for a length of 1 is:

[0032]

[0033] The cost for a length of 2 is:

[0034]

[0035] As for can be expressed as

[0036]

[0037] The definition of can be given recursively:

[0038]

[0039] With the hint of the semantic transition table, during the dynamic programming solution process, if there is no semantic change before a certain row, dynamic programming will be more inclined to regard this section as a whole.

[0040] According to another aspect of the present invention, a front-vehicle scene display device based on semantic information hint is provided, and the following technical solutions are adopted:

[0041] The front-vehicle scene display device based on semantic information hint includes: an acquisition module, configured to acquire an original camera stereo image and obtain a disparity map and a three-class semantic segmentation map through a neural network; a generation module, configured to generate multi-layer Stixles according to the multi-layer Stixle-world algorithm using the disparity map, and obtain semantic segmentation information using the three-class semantic segmentation map; a solution module, configured to add semantic segmentation information as a hint during the multi-layer Stixle-world dynamic programming solution process to obtain multi-layer Stixles with higher accuracy.

[0042] Further, the generation module is further configured to: Given the calibrated left image I with a size of w*h in the stereo image pair and a possible annotation set , the multi-layer Stixel world corresponds to belonging to The annotation L of is divided into sets of categories column by column , namely road g, obstacle o, and sky s;

[0043]

[0044] Indicates the annotation of the column of the image and contains multiple segments , the total number of segments in each column is given by , and the maximum value of the segments is implicitly limited by the height of the image, that is , in the formula the bottom point and the vertex mark the starting point and the ending point of the segment , and each segment is assigned to a category , is an arbitrary function for calculating the disparity of the segment;

[0045] , and is only determined by the number of rows v , all segments and are vertically adjacent, ensuring that each pixel in the image has only one label; in the annotation step, all segments are designated as roads, obstacles, and skies, so it is assumed that all segments are approximately piecewise planes in three-dimensional space, and the selection of the function is simplified to a set of linear functions. The segments assigned to obstacles are assumed to have a constant disparity, and the corresponding function is , is the representative disparity within the segment, and the surface of the segment labeled as a road is modeled as a linear function;

[0046] , is the expected road disparity gradient is the row coordinate of the horizon and can both be extracted from the existing camera geometry. The disparity function of the segment labeled as the sky is .

[0047] Furthermore, the obtaining module is further configured to: The semantic segmentation map obtained through the semantic segmentation neural network contains three values: 0 for road, 1 for obstacle, and 2 for sky. For a column Se of the image, the semantic segmentation information of this column is an arrangement of the three values 0, 1, and 2, that is,

[0048]

[0049] Construct a semantic transition table T based on whether the semantic value of the current pixel is equal to that of the previous pixel. If the semantic values are equal, the corresponding position in the transition table is 0; if not, it is 1. That is,

[0050]

[0051] Then, calculate the cumulative semantic transition value from the bottom to each pixel in the semantic transition table to construct a cumulative semantic transition table A, indicating the number of semantic transitions experienced by the image from the row to the row of this column:

[0052]

[0053]

[0054] If the value of the cumulative semantic transition table at a certain pixel is 0, that is, = 0, it means that this segment from the row to the row of this column should be an integral stixel.

[0055] Furthermore, the solving module is further configured to: find the label with the highest probability under the condition of a given disparity input P :

[0056]

[0057] Through Bayes' formula, it can be expressed as:

[0058]

[0059] is ignored because this term is just a normalization factor, represents the probability of the input disparity P given the label L, while represents the probability of the label L appearing. Under the assumption that the labels between columns and the disparities between columns of the image are independent, and can be further decomposed:

[0060]

[0061]

[0062] 。

[0063] Furthermore, the solving module is further configured to: define a scaling function , where v is the number of rows and c is the semantic category. If there is no semantic change in a certain column of the image from the lowest row to row v, the function value is a constant less than 1; otherwise, the function value is 1. The dynamic programming for each column of the image can be formulated as follows: The cost for a length of 1 is:

[0064]

[0065] The cost for a length of 2 is:

[0066]

[0067] As for It can be expressed as

[0068]

[0069] The definition of can be given recursively:

[0070]

[0071] With the hint of the semantic transition table, during the process of solving by dynamic programming, if there is no semantic change before a certain row, dynamic programming is more inclined to take this section as a whole.

[0072] In summary, different from the existing multi-layer stixels and semantic stixels that regard semantic information as equally important as disparity information, we only use semantic information as hint information. During the calculation process of multi-layer stixels, through the hint of semantic information, the calculation accuracy of multi-layer stixels is improved, and a neural network is used to obtain a disparity map with better quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings.

[0074] Figure 1 It is a flowchart of a method for displaying a front vehicle scene based on semantic information hint according to an embodiment of the present invention;

[0075] Figure 2 It is a flowchart of another method for displaying a front vehicle scene based on semantic information hint according to an embodiment of the present invention;

[0076] Figure 3 It is a schematic diagram of semantic transition and semantic transition accumulation according to an embodiment of the present invention;

[0077] Figure 4 The rendering effect diagram corrected using semantic segmentation information according to the embodiments of the present invention; and

[0078] Figure 5 The structural diagram of the vehicle front scene display device based on semantic information prompt according to the embodiments of the present invention. Detailed implementation manners

[0079] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings, but the present invention can be implemented in many different ways defined and covered by the claims.

[0080] Figure 1 The flowchart of a vehicle front scene display method based on semantic information prompt according to the embodiments of the present invention.

[0081] Refer to Figure 1 As shown, a vehicle front scene display method based on semantic information prompt includes:

[0082] S101: Obtain the original camera stereo image, and obtain the disparity map and the three-class semantic segmentation map through the neural network;

[0083] S103: Use the disparity map to generate multi-layer Stixles according to the multi-layer Stixle-world algorithm, and use the three-class semantic segmentation map to obtain semantic segmentation information;

[0084] S105: Add the semantic segmentation information as prompt information during the multi-layer Stixle-world dynamic programming solution process to obtain multi-layer Stixles with higher accuracy.

[0085] Figure 2 The specific algorithm flowchart of another intelligent chassis control method based on binocular cameras according to the embodiments of the present invention.

[0086] Figure 2 The overall flowchart of the algorithm is given. The original multi-layer Stixel-world is for a labeling task. Given the calibrated left image I of size w*h in the stereo image pair and the possible labeling set , the multi-layer Stixel world corresponds to splitting the label L belonging to into the category set , that is, road, obstacle, and sky, column by column.

[0087]

[0088] represents the label of the column of the image and contains multiple segments . The total number of segments in each column is determined by Given that the maximum value of a stixel is implicitly limited by the height of the image, i.e., in the formula, (bottom point) and (top point) mark the start and end points of the stixel . Each stixel is assigned to a class , which is an arbitrary function for computing the disparity of the stixel and is only determined by the row number v, .

[0089] All stixels and are vertically adjacent, ensuring that each pixel in the image has only one label.

[0090] In the labeling step, all stixels are assigned as road, obstacle, and sky. Thus, it is assumed that all stixels are approximately piecewise planes in 3D space, and the selection of the function is simplified to a set of linear functions. Considering that the world geometry in a realistic man-made environment mainly consists of vertical and horizontal planes, the function set can be further simplified. The stixels assigned to obstacles are assumed to have a constant disparity, and the corresponding function is , where is the representative disparity within the stixel. The surface of the stixels labeled as road is modeled as a linear function , where is the expected road disparity gradient, is the row coordinate of the horizon, and can both be extracted from the existing camera geometry. The disparity function of the stixels labeled as sky is because the sky is at an infinite distance.

[0091] Finding the optimal Stixel representation is a typical MAP problem. The goal is to find the label with the maximum probability given the disparity input P :

[0092]

[0093] Through Bayes' formula, can be expressed as,

[0094]

[0095] is ignored because this term is just a normalization factor. represents the probability of the input disparity P given the label L, while Denotes the probability of the label L occurring, under the condition that the labels between columns of the assumed image and the disparity between columns are independent and can be further decomposed

[0096]

[0097]

[0098]

[0099] Likelihood probability and prior probability also need to be further decomposed and defined. This will not be elaborated here

[0100] Dynamic programming is a solution scheme that has been successfully applied to a large number of optimization problems. Its main advantage is that it can obtain the global optimal solution of the problem non-iteratively in polynomial time without the risk of getting stuck in a local optimal solution. During the dynamic programming process, a cost minimization is performed. This cost is derived from the previous data terms and prior terms. Since the natural logarithm is a strictly increasing continuous function within the likelihood probability range, maximizing the posterior probability is equal to minimizing the log-likelihood. Taking the negative logarithm of the posterior probability also has the advantage of converting the product of probability values into the summation of the corresponding negative logarithms

[0101] To use dynamic programming to solve an optimization problem, this optimization problem needs to have two criteria. First, the optimization problem must have a discrete nature. Second, the optimization problem must exhibit an optimal substructure. This means that the optimization problem can be recursively represented as a combination of a set of smaller sub-problems, each of which can be further decomposed or obtain an optimal solution. The columns of the image are solved independently, and dynamic programming is used to solve the optimal segmentation of each column

[0102] In our optimization problem, the object of optimization is the segment , if all elements of the segment have a discrete nature, then the segment as a whole also has a discrete nature. But only , and are discrete is not discrete, but can be regarded as an attribute of the segment and searched for in the candidate functions during dynamic programming. Dynamic programming only acts on , and the three variables, then the optimization problem has a discrete nature

[0103] Dynamic programming requires that the optimization problem must have an optimal substructure. To prove that the optimization problem has the property of an optimal substructure, the following recursive definition of the optimization problem is given.

[0104] Introduce the following variables: variable , , represents the cost generated by allocating the segment with the bottom row b and the top row t to the road, obstacle, and sky. , , represents the minimum cost of allocating the segment from row 0 to row t to the road, obstacle, and sky. This segment can be composed of multiple sub - segments, but the last sub - segment, i.e., the sub - segment with the top row t, must be the sub - segment corresponding to the capital letter. The function c represents the prior cost between different sub - segments. For example, the binary term represents the prior cost value of connecting the sub - segment with the top row 10 ending with the road type and the sub - segment with obstacles from row 11 to row 20, while the unary term represents the prior cost value of the first sub - segment with the top row 10 being an obstacle.

[0105] First, determine the cost of length 1:

[0106]

[0107] Subsequently, for segments of length 2 from 0 to 1, they can be a single segment or multiple segments composed of the previous segments.

[0108]

[0109] As for it can be expressed as,

[0110]

[0111] The definition of can be given recursively:

[0112]

[0113] The minimum cost Mark the path obtained by reverse solving of dynamic programming. Calculated in this way, the obtained multi-layer stixels can get a good result, but there are still some deficiencies. Analyzing the reasons, the parallax of the sky part is basically 0, and the sky and obstacles can be clearly distinguished. However, since the parallax of both obstacles and the road surface is not 0, and the parallax at the junction of obstacles and the road is equal, there is no definite difference in numerical values between the two, making it difficult to clearly distinguish between the two. Therefore, we introduce semantic segmentation information, aiming to use the information of semantic segmentation to help multi-layer stixels clearly distinguish obstacles and the road.

[0114] The semantic segmentation map obtained through the semantic segmentation neural network only contains three values: 0 (road), 1 (obstacle), and 2 (sky). For a column Se of the image, the semantic segmentation information of this column is an arrangement of the three values 0, 1, and 2, that is,

[0115]

[0116] A semantic transition table T can be constructed according to whether the semantic value of the current pixel is equal to that of the previous pixel. If the semantic values are equal, the corresponding position in the transition table is 0, and if they are not equal, it is 1, that is,

[0117]

[0118] Then, calculate the cumulative value of the semantic transition table from the bottom to this pixel at each pixel of the semantic transition table to construct a semantic transition cumulative table A, as Figure 3 , where the cube in the segmantic column represents the sky, the circle represents the obstacle, and the ellipse represents the road. Represents from the row to the row of this column, the number of semantic transitions experienced by the image.

[0119]

[0120]

[0121] If the value of the semantic transition cumulative table at a certain pixel is 0, that is, =0, it means that this segment from the row to the row of this column should be an integral stixel.

[0122] Define a scaling function , where v is the row number and c is the semantic category. If there is no semantic change in a column of the image from the lowest row to the v-th row, the function value is a constant less than 1, otherwise the function value is 1.

[0123] The dynamic programming for each column of the image can be formulated as follows: The cost of length 1 is

[0124]

[0125] The cost of length 2 is

[0126]

[0127] As for It can be expressed as,

[0128]

[0129] It can be recursively given Definition:

[0130]

[0131] No need for Multiply by the scaling function, because the bottom of the image cannot be the sky. Minimum cost Mark the reverse solution path of dynamic programming. With the help of the semantic transition table, during the dynamic programming solution process, if there is no semantic change before a certain line, dynamic programming will tend to treat this section as a whole. Due to the characteristics of dynamic programming, the improvement of the accuracy of the previous steps will also improve the accuracy of the subsequent steps, thereby improving the overall effect. Figure 4 In the middle, the right image uses semantic segmentation information to correct the segmentation error pointed by the arrow relative to the left image.

[0132] Figure 5 This is a structural diagram of a vehicle front scene display device based on semantic information prompts according to an embodiment of the present invention.

[0133] The vehicle front scene display device based on semantic information prompts includes: an acquisition module 50, which is used to acquire the original camera stereo image, and obtain a disparity map and a three-category semantic segmentation map through a neural network; a generation module 52, which is used to use the disparity map to generate multi-layer Stixles according to the multi-layer Stixle-world algorithm, and use the three-category semantic segmentation map to obtain semantic segmentation information; a solution module 54, which is used to add semantic segmentation information as prompt information in the multi-layer Stixle-world dynamic planning solution process to obtain multi-layer Stixles with higher accuracy.

[0134] Different from the existing multi-layer stixel and semantic stixel that regards semantic information as equally important as disparity information, we only use semantic information as suggestive information. During the calculation process of multi-layer stixel, the calculation accuracy of multi-layer stixel is improved through the hint of semantic information, and a neural network is used to obtain a disparity map with better quality.

[0135] Only some exemplary embodiments of the present embodiment are described above by way of illustration. Without doubt, for those of ordinary skill in the art, the described embodiments can be modified in various different ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

Claims

1. A method for displaying a vehicle front scene based on semantic information prompt, characterized in that Including: Obtain the original camera stereo image, and obtain the disparity map and the three-class semantic segmentation map through a neural network; Use the disparity map to generate multi-layer Stixles according to the multi-layer Stixle-world algorithm, and use the three-class semantic segmentation map to obtain semantic segmentation information; During the dynamic programming solution process of the multi-layer Stixle-world, add the semantic segmentation information as prompt information to obtain multi-layer Stixles with higher accuracy.

2. The vehicle front scene display method according to claim 1, wherein, The step of using the disparity map to generate multi-layer Stixles according to the multi-layer Stixle-world algorithm includes: Given a calibrated left image \(I\) of size \(w\times h\) in a stereo image pair and a set of possible annotations , a multi-layer Stixelworld corresponds to splitting the annotation \(L\) belonging to into a set of categories column by column, namely road \(g\), obstacle \(o\) and sky \(s\); Indicates the annotation of the columns of the image and contains multiple segments , and the total number of segments per column is given by . The maximum value of the segments is implicitly limited by the height of the image, i.e., , where h is the height of the image and w is the width of the image. In the formula, the bottom point and the vertex mark the starting and ending points of the segment . Each segment is assigned to a class , is an arbitrary function for calculating the disparity of the segment; and is determined only by the line number v, , all the segments and are vertically adjacent, ensuring that each pixel of the image will have only one label; in the annotation step, all the segments are designated as roads, obstacles, and sky, so it is assumed that all segments are approximately piecewise planes in three-dimensional space. The selection of the function is reduced to a set of linear functions. The segments assigned to obstacles are assumed to have a constant parallax, and the corresponding function is , is the representative parallax within the segment, and the surface of the segment labeled as a road is modeled as a linear function; , is the expected road parallax gradient, is the row coordinate of the horizon, and both can be extracted from the existing camera geometry, and the parallax function of the segment marked as the sky is .

3. The vehicle front scene display method according to claim 2, characterized in that, The step of using the three-class semantic segmentation map to obtain semantic segmentation information includes: The semantic segmentation map obtained through the semantic segmentation neural network contains three values: 0 for road, 1 for obstacle, and 2 for sky. For a column Se of the image, the semantic segmentation information of this column is an arrangement of the three values 0, 1, and 2, that is, Construct a semantic transition table T according to whether the semantic value of the current pixel is equal to that of the previous pixel. If the semantic values are equal, the corresponding position in the transition table is 0; if they are not equal, it is 1, that is, Then, at each pixel of the semantic transition table, calculate the semantic transition table cumulative value from the bottom to this pixel to construct the semantic transition cumulative table A. Indicates from this column Row to The number of semantic transitions experienced by the image in the row: If the value of the semantic transition accumulation table at a certain pixel is 0, that is = 0, it means that the segment from the row to the row of this column should be an integral stixel.

4. The method for displaying a front vehicle scene according to claim 3, wherein, The step of adding the semantic segmentation information as prompt information during the dynamic programming solution process of the multi-layer Stixle-world to obtain multi-layer Stixles with higher accuracy includes: Find the label with the highest probability given the parallax input P : Through Bayes' formula, it can be expressed as: is ignored because this is just a normalization factor, represents the probability of the input disparity P given the label L, while represents the probability of the label L occurring, under the assumption that the labels between columns of the image and the disparities between columns are independent, and can be further decomposed: Assume that the width of image I is w, P is the disparity map of I, and L is the label map of I. is the disparity of the u-th column, is the label of the u-th column.

5. The method for displaying a front vehicle scene according to claim 4, characterized in that, The step of adding the semantic segmentation information as prompt information during the dynamic programming solution process of the multi-layer Stixle-world to obtain multi-layer Stixles with higher accuracy further includes: Define the scaling function , where v is the number of rows and c is the semantic category. If there is no semantic change in a certain column of the image from the lowest row to row v, the function value is a constant less than 1; otherwise, the function value is 1. The dynamic programming for each column of the image can be formulated as: The cost for a length of 1 is: The cost for a length of 2 is: As for It can be expressed as It can be recursively given the definition of With the prompt of the semantic transition table, during the dynamic programming solution process, if there is no semantic change before a certain row, dynamic programming will be more inclined to regard this section as a whole.

6. A vehicle front scene display device based on semantic information prompt, characterized in that Including: An acquisition module, used to obtain the original camera stereo image, and obtain the disparity map and the three-class semantic segmentation map through a neural network; A generation module, used to generate multi-layer Stixles according to the multi-layer Stixle-world algorithm using the disparity map, and obtain semantic segmentation information using the three-class semantic segmentation map; A solution module, used to add the semantic segmentation information as prompt information during the dynamic programming solution process of the multi-layer Stixle-world to obtain multi-layer Stixles with higher accuracy.

7. The vehicle front scene display device according to claim 6, characterized in that, The generation module is further used for: The calibrated left image I of size w*h in a given stereo image pair and a set of possible annotations , where the multi-layer Stixel world corresponds to splitting the annotation L belonging to into a set of categories column by column, namely road g, obstacle o, and sky s; Indicates the annotation of the column of the image and contains multiple segments . The total number of segments per column is given by , and the maximum value of the segment is implicitly limited by the height of the image, i.e., . In the formula, the bottom point and the vertex mark the starting point and the ending point of the segment . Each segment is assigned to a category . is an arbitrary function for calculating the disparity of the segment; and is determined only by the line number v, , all segments and are vertically adjacent, ensuring that each pixel of the image will have only one label; In the annotation step, all segments are designated as roads, obstacles, and sky, so it is assumed that all segments are approximately piecewise planes in three-dimensional space. The selection of the function is reduced to a set of linear functions. The segments assigned to obstacles are assumed to have a constant parallax, and the corresponding function is , where is the representative parallax within the segment. The surface of the segments labeled as roads is modeled as a linear function; , is the expected road parallax gradient, is the row coordinate of the horizon, and both can be extracted from the existing camera geometry. The parallax function of the segment marked as sky is .

8. The vehicle front scene display device according to claim 7, characterized in that, The acquisition module is further used for: The semantic segmentation map obtained through the semantic segmentation neural network contains three values: 0 for road, 1 for obstacle, and 2 for sky. For a column Se of the image, the semantic segmentation information of this column is an arrangement of the three values 0, 1, and 2, that is, Construct a semantic transition table T according to whether the semantic value of the current pixel is equal to that of the previous pixel. If the semantic values are equal, the corresponding position in the transition table is 0; if they are not equal, it is 1, that is, Then, at each pixel of the semantic transition table, calculate the semantic transition table accumulation value from the bottom to this pixel to construct the semantic transition accumulation table A. Indicates from this column Row to The number of semantic transitions experienced by the image in the row: If the value of the semantic transition accumulation table at a certain pixel is 0, that is = 0, it means that the segment from the row to the row of this column should be an integral stixel.

9. The vehicle front scene display device according to claim 8, characterized in that, The solution module is further used for: Find the label with the highest probability given the parallax input P : Through Bayes' formula, it can be expressed as: is ignored because this term is just a normalization factor, represents the probability of the input disparity P given the label L, while represents the probability of the label L occurring, under the assumption that the labels between columns of the image and the disparities between columns are independent, and can be further decomposed: 。 10. The vehicle front scene display device according to claim 9, wherein, The solution module is further used for: Define the scaling function , where v is the number of rows and c is the semantic category. If there is no semantic change in a certain column of the image from the lowest row to row v, the function value is a constant less than 1; otherwise, the function value is 1. The dynamic programming for each column of the image can be formulated as: The cost for a length of 1 is: The cost for a length of 2 is: As for It can be expressed as The definition of can be given recursively as follows: is defined as: With the hint of the semantic transition table, during the process of solving by dynamic programming, if there is no semantic change before a certain row, dynamic programming will be more inclined to take this section as a whole.