A semantic segmentation method, device and equipment for an in-vehicle camera

By aligning and fusing images from different camera angles using optical flow compensation, the method addresses the challenge of integrating multi-camera data for improved semantic segmentation accuracy and real-time performance in automatic driving vehicles.

CN115965779BActive Publication Date: 2025-07-15SAIC MOTOR
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111187724.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-12
Publication Date
2025-07-15
Estimated Expiration
2041-10-12

AI Technical Summary

Technical Problem

In autonomous driving vehicles, when equipped with a multi-eye camera system, it is difficult for the prior art to achieve the accuracy and real-timeness of semantic segmentation results of multiple cameras in scenarios with high real-time requirements, and the interpretability information obtained by different cameras cannot be effectively correlated.

Method used

The target image is obtained through on-board cameras of different perspectives, and the view angle conversion and optical flow transformation are performed to compensate for the distortion loss. Then the feature map and semantic segmentation results are fusion to achieve semantic sharing and fusion between different cameras.

Benefits of technology

The accuracy and real-time nature of semantic segmentation results are improved, and the semantic segmentation effect of multi-eye camera systems is improved through the complementary advantages between different cameras.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115965779B_ABST
    Figure CN115965779B_ABST
Patent Text Reader

Abstract

The present application discloses a semantic segmentation method, device and equipment for an in-vehicle camera. The method includes: First, obtain a first target image and a second target image respectively through the in-vehicle cameras from the first perspective and the second perspective; then perform perspective transformation on both of them respectively to obtain the transformed first target image from the second perspective and the transformed second target image from the first perspective; then use optical flow transformation to compensate for the distortion loss to obtain the compensated first target image and second target image; further perform semantic fusion on the feature maps of the compensated first target image, the feature map of the first target image, and the feature map of the second target image to determine the semantic segmentation result of the second target image, and then perform semantic fusion on the feature map of the compensated second target image, the semantic segmentation result of the second target image, and the feature map of the first target image to determine the semantic segmentation result of the first target image. Thereby, the accuracy and real-time performance of the semantic segmentation result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of vehicles, and particularly to a semantic segmentation method, device and equipment for an in-vehicle camera. Background Art

[0002] With the improvement of people's living standards and the rapid development of social economy, the utilization rate of automobiles has gradually increased, and more and more automobiles have entered people's lives, bringing great convenience to all aspects of people's lives. Among them, with the rapid development of intelligent driving technology, people's acceptance and demand for autonomous vehicles are gradually increasing.

[0003] At present, the semantic segmentation method in visual perception has received more and more attention as an important function of autonomous vehicles. Since the semantic segmentation method is a dense classification problem, how to achieve better classification accuracy in scenarios with high real-time requirements is still a difficult problem that needs to be solved urgently, especially for autonomous vehicles equipped with a multi-camera system and limited computing resources. Currently, mainstream autonomous vehicles usually carry a multi-camera vision system composed of multiple cameras with different perspectives. In order to obtain the semantic segmentation results of these multiple cameras, traditional methods usually independently process the video sequences collected by multiple cameras, which will cause the interpretable information obtained by multiple cameras to be uncorrelated, and then lead to insufficient semantic segmentation results of a single camera and poor real-time performance of segmentation. Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to provide a semantic segmentation method, device and equipment for an in-vehicle camera, which can utilize the semantic sharing and fusion between different cameras to achieve the complementary advantages between the semantic segmentation results of different cameras, thereby improving the accuracy and real-time performance of the semantic segmentation results.

[0005] The embodiments of the present application provide a semantic segmentation method for an in-vehicle camera, including:

[0006] Obtaining a first target image through an in-vehicle camera with a first perspective on a target vehicle; and obtaining a second target image through an in-vehicle camera with a second perspective on the target vehicle; the first perspective and the second perspective are different;

[0007] Performing perspective conversion on the first target image to obtain a converted first target image with the second perspective; and performing perspective conversion on the second target image to obtain a converted second target image with the first perspective;

[0008] Using optical flow transformation to compensate for the distortion loss of the converted first target image with the second perspective to obtain a compensated first target image; and using optical flow transformation to compensate for the distortion loss of the converted second target image with the first perspective to obtain a compensated second target image;

[0009] Semantically fuse the feature map of the compensated first target image, the feature map of the first target image, and the feature map of the second target image, and determine the semantic segmentation result of the second target image according to the obtained fusion result;

[0010] Semantically fuse the feature map of the compensated second target image, the semantic segmentation result of the second target image, and the feature map of the first target image, and determine the semantic segmentation result of the first target image according to the obtained fusion result.

[0011] In an optional implementation, the vehicle-mounted camera with the first view angle is a vehicle-mounted camera with a 120-degree horizontal view angle; the vehicle-mounted camera with the second view angle is a vehicle-mounted camera with a 60-degree horizontal view angle.

[0012] In an optional implementation, the step of semantically fusing the feature map of the compensated first target image, the feature map of the first target image, and the feature map of the second target image, and determining the semantic segmentation result of the second target image according to the obtained fusion result includes:

[0013] Multiply the feature map of the compensated first target image and the feature map of the first target image, and splice the feature map obtained after multiplication with the feature map of the second target image to obtain a spliced feature map;

[0014] Process the spliced feature map through a 1×1 convolutional layer to output a fused feature as the semantic fusion result, and determine the semantic segmentation result of the second target image according to the semantic fusion result; or, first process the spliced feature map through a standard residual structure with a 3×3 convolution, and then process it through a 1×1 convolutional layer to output a fused feature as the semantic fusion result, and determine the semantic segmentation result of the second target image according to the semantic fusion result.

[0015] In an optional implementation, the step of semantically fusing the feature map of the compensated second target image, the semantic segmentation result of the second target image, and the feature map of the first target image, and determining the semantic segmentation result of the first target image according to the obtained fusion result includes:

[0016] Multiply the feature map of the compensated second target image and the semantic segmentation result of the second target image, and splice the feature map obtained after multiplication with the feature map of the first target image to obtain a spliced feature map;

[0017] The concatenated feature map is processed through a 1×1 convolutional layer to output a fused feature as the semantic fusion result, and the semantic segmentation result of the first target image is determined according to the semantic fusion result; or, the concatenated feature map is first passed through a standard residual structure with 3×3 convolution and then through a 1×1 convolutional layer to output a fused feature as the semantic fusion result, and the semantic segmentation result of the first target image is determined according to the semantic fusion result.

[0018] Corresponding to the above semantic segmentation method of an in-vehicle camera, the present application proposes a semantic segmentation device for an in-vehicle camera, including:

[0019] An acquisition unit, configured to acquire a first target image through an in-vehicle camera with a first viewing angle on a target vehicle; and acquire a second target image through an in-vehicle camera with a second viewing angle on the target vehicle; the first viewing angle and the second viewing angle are different;

[0020] A conversion unit, configured to perform a viewing angle conversion on the first target image to obtain a converted first target image with the second viewing angle; and perform a viewing angle conversion on the second target image to obtain a converted second target image with the first viewing angle;

[0021] A compensation unit, configured to utilize optical flow transformation to compensate for the distortion loss of the converted first target image with the second viewing angle to obtain a compensated first target image; and utilize optical flow transformation to compensate for the distortion loss of the converted second target image with the first viewing angle to obtain a compensated second target image;

[0022] A first determination unit, configured to semantically fuse the feature map of the compensated first target image, the feature map of the first target image, and the feature map of the second target image, and determine the semantic segmentation result of the second target image according to the obtained fusion result;

[0023] A second determination unit, configured to semantically fuse the feature map of the compensated second target image, the semantic segmentation result of the second target image, and the feature map of the first target image, and determine the semantic segmentation result of the first target image according to the obtained fusion result.

[0024] In an optional implementation manner, the in-vehicle camera with the first viewing angle is an in-vehicle camera with a 120-degree horizontal viewing angle; the in-vehicle camera with the second viewing angle is an in-vehicle camera with a 60-degree horizontal viewing angle.

[0025] In an optional implementation manner, the first determination unit includes:

[0026] A first multiplication subunit, configured to multiply the feature map of the compensated first target image and the feature map of the first target image, and splice the feature map obtained after the multiplication with the feature map of the second target image to obtain a spliced feature map;

[0027] A first determination subunit, configured to output the fused feature after processing the spliced feature map through a convolutional layer as the semantic fusion result, and determine the semantic segmentation result of the second target image according to the semantic fusion result; or, first process the spliced feature map through a standard residual structure with convolution, and then output the fused feature after processing through the convolutional layer as the semantic fusion result, and determine the semantic segmentation result of the second target image according to the semantic fusion result.

[0028] In an optional implementation manner, the second determination unit includes:

[0029] A second multiplication subunit, configured to multiply the feature map of the compensated second target image and the semantic segmentation result of the second target image, and splice the feature map obtained after the multiplication with the feature map of the first target image to obtain a spliced feature map;

[0030] A second determination subunit, configured to output the fused feature after processing the spliced feature map through a 1×1 convolutional layer as the semantic fusion result, and determine the semantic segmentation result of the first target image according to the semantic fusion result; or, first process the spliced feature map through a standard residual structure with 3×3 convolution, and then output the fused feature after processing through a 1×1 convolutional layer as the semantic fusion result, and determine the semantic segmentation result of the first target image according to the semantic fusion result.

[0031] An embodiment of the present application further provides a semantic segmentation device for a vehicle-mounted camera, including: a processor, a memory, and a system bus;

[0032] The processor and the memory are connected through the system bus;

[0033] The memory is used to store one or more programs, and the one or more programs include instructions, and when the instructions are executed by the processor, the processor is caused to execute any one of the implementation manners of the semantic segmentation method for the vehicle-mounted camera described above.

[0034] An embodiment of the present application further provides a computer-readable storage medium, in which instructions are stored, and when the instructions are run on a terminal device, the terminal device is caused to execute any one of the implementation manners of the semantic segmentation method for the vehicle-mounted camera described above.

[0035] As can be seen, the embodiments of the present application have the following beneficial effects:

[0036] The embodiments of the present application provide a semantic segmentation method, device and equipment for an in-vehicle camera. First, a first target image is obtained through an in-vehicle camera with a first perspective on a target vehicle; and a second target image is obtained through an in-vehicle camera with a second perspective on the target vehicle; wherein, the first perspective and the second perspective are different. Then, the first target image is subjected to perspective transformation to obtain the transformed first target image with the second perspective; and the second target image is subjected to perspective transformation to obtain the transformed second target image with the first perspective. Next, optical flow transformation is used to compensate for the distortion loss of the transformed first target image with the second perspective to obtain the compensated first target image; and optical flow transformation is used to compensate for the distortion loss of the transformed second target image with the first perspective to obtain the compensated second target image. Furthermore, the feature maps of the compensated first target image, the feature map of the first target image, and the feature map of the second target image can be semantically fused, and the semantic segmentation result of the second target image can be determined according to the obtained fusion result. Then, the feature map of the compensated second target image, the semantic segmentation result of the second target image, and the feature map of the first target image are semantically fused, and the semantic segmentation result of the first target image is determined according to the obtained fusion result. Thus, it is possible to utilize the semantic sharing and fusion between in-vehicle cameras with different perspectives to achieve the complementary advantages among the semantic segmentation results of each camera, thereby improving the accuracy and real-time performance of the semantic segmentation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 It is a flowchart of a semantic segmentation method for an in-vehicle camera provided by an embodiment of the present application;

[0039] Figure 2 It is an overall schematic diagram of the semantic segmentation of an in-vehicle camera provided by an embodiment of the present application;

[0040] Figure 3 It is a composition schematic diagram of a semantic segmentation device for an in-vehicle camera provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.

[0042] As is well known, with the continuous progress of high-tech such as cloud computing, artificial intelligence, modern sensing, information fusion, communication, and automatic control, the future development speed of autonomous vehicles will accelerate, and at the same time, people's acceptance and demand for autonomous vehicles are gradually increasing.

[0043] Currently, mainstream autonomous vehicles usually carry a multi-camera vision system composed of cameras with different perspectives. In order to obtain semantic segmentation results of multiple cameras, traditional methods usually independently process the video sequences collected by multiple cameras. This results in the interpretability of the network still mainly depending on the design of a single network, and the interpretable information obtained by multiple cameras cannot be associated, thus leading to inaccurate semantic segmentation results for a single camera.

[0044] Based on this, this application proposes a semantic segmentation method, device, and equipment for on-vehicle cameras, which can utilize semantic sharing and fusion between different cameras to achieve complementary advantages between the semantic segmentation results of different cameras, thereby improving the accuracy and real-time performance of the semantic segmentation results.

[0045] The following will detail the semantic segmentation method for on-vehicle cameras provided in the embodiments of this application with reference to the accompanying drawings. Refer to Figure 1 As shown, it shows a flowchart of an embodiment of a semantic segmentation method for on-vehicle cameras provided in the embodiments of this application. This embodiment may include the following steps:

[0046] S101: Obtain a first target image through an on-vehicle camera with a first perspective on a target vehicle; and obtain a second target image through an on-vehicle camera with a second perspective on the target vehicle; wherein, the first perspective and the second perspective are different.

[0047] In this embodiment, any autonomous vehicle that uses the method of the embodiments of this application to implement semantic segmentation of on-vehicle cameras is defined as a target vehicle. In order to implement semantic segmentation of on-vehicle cameras with different perspectives on the target vehicle and improve the accuracy and real-time performance of each semantic segmentation result, this application proposes that it is first necessary to obtain a first target image through an on-vehicle camera with a first perspective on the target vehicle; and obtain a second target image through an on-vehicle camera with a second perspective on the target vehicle for subsequent step S102.

[0048] Among them, the first perspective and the second perspective are different. An optional implementation is that the in-vehicle camera with the first perspective can be an in-vehicle camera with a 120-degree horizontal viewing angle; the in-vehicle camera with the second perspective can be an in-vehicle camera with a 60-degree horizontal viewing angle. Then the first target image can be as Figure 2 shown as I 120 , and the second target image can be as Figure 2 shown as I 60 .

[0049] It should be noted that hereinafter, the present application will take the binocular vision system on the target vehicle as an example to introduce the semantic segmentation method of the in-vehicle camera provided by the present application. Specifically, hereinafter, the in-vehicle camera with the first perspective is an in-vehicle camera with a 120° horizontal viewing angle (i.e., cam-120), and the in-vehicle camera with the second perspective is an in-vehicle camera with a 60° horizontal viewing angle (i.e., am-60) as an example to introduce the semantic segmentation method of the in-vehicle camera provided by the present application. The semantic segmentation methods of in-vehicle cameras with other structures and viewing angles will not be elaborated one by one, and reference implementation can be made.

[0050] S102: Perform perspective transformation on the first target image to obtain the transformed first target image with the second perspective; and perform perspective transformation on the second target image to obtain the transformed second target image with the first perspective.

[0051] In this embodiment, in order to implement semantic segmentation of in-vehicle cameras with different perspectives on the target vehicle and improve the accuracy and real-time performance of each semantic segmentation result, after obtaining the first target image and the second target image through step S101, the first target image can be further subjected to perspective transformation to obtain the transformed first target image with the second perspective; and the second target image is subjected to perspective transformation to obtain the transformed second target image with the first perspective.

[0052] Specifically, as Figure 2 shown, after obtaining the first target image (I 120 ) using the Cam-120 camera, it can be subjected to perspective transformation to obtain an image approximately with a 60° camera viewing angle. The homography matrix used in this transformation process can be obtained from the internal and external parameters of the two cameras. The specific calculation formula is as follows:

[0053]

[0054] Among them, H represents the homography matrix for mapping the cam-120 image to the cam-60 image, K 60 and K 120 respectively represent the internal parameter matrices of the two cameras, and R represents the rotation matrix from the cam-120 camera to the cam-60 camera.

[0055] S103: Use optical flow transformation to compensate for the distortion loss of the first target image after the conversion of the second perspective, and obtain the compensated first target image; and use optical flow transformation to compensate for the distortion loss of the second target image after the conversion of the first perspective, and obtain the compensated second target image.

[0056] In this embodiment, after determining the first target image after the conversion of the second perspective and the second target image after the conversion of the first perspective through step S102, further use optical flow transformation to compensate for the distortion loss of the first target image after the conversion of the second perspective, and obtain the compensated first target image; and use optical flow transformation to compensate for the distortion loss of the second target image after the conversion of the first perspective, and obtain the compensated second target image, so as to execute the subsequent step S104.

[0057] For example, as Figure 2 shown, the first target image (I 120 ) obtains an image approximately with a 60° camera perspective after perspective transformation, but there is obvious distortion in the transformation of nearby objects. Therefore, it is necessary to further use optical flow transformation to compensate for the original distortion loss to obtain the compensated first target image, as Figure 2 shown in Warp 120->60 . Similarly, the first target image (I 60 ) obtains an image approximately with a 120° camera perspective after perspective transformation, and there is also obvious distortion. Therefore, it is still necessary to further use optical flow transformation to compensate for the original distortion loss to obtain the compensated second target image, as Figure 2 shown in Warp 60->120 .

[0058] In this way, a mapping and matching relationship is established between the first target image and the second target image from different perspectives, so as to execute the subsequent step S104.

[0059] S104: Semantically fuse the feature maps of the compensated first target image, the feature maps of the first target image, and the feature maps of the second target image, and determine the semantic segmentation result of the second target image according to the obtained fusion result.

[0060] In this embodiment, after obtaining the compensated first target image and the compensated second target image through step S103, further, the feature map of the compensated first target image (such as Figure 2 shown in Warp 120->60 ), the feature map of the first target image (such as Figure 2 F of 120 ), and the feature map of the second target image (such as Figure 2 F of 60)Perform semantic fusion and determine the semantic segmentation result of the second target image based on the obtained fusion result.

[0061] Among them, the feature map of the first target image (such as Figure 2 I of 120 ) (such as Figure 2 F of 120 ) is the high-level semantic feature obtained by passing the first target image (such as Figure 2 I of 120 ) through a semantic segmentation network similar to the traditional method (i.e., a full-functional segmentation network, which can be implemented by any network with this function, such as Figure 2 Semantic Segmentation Network shown in Figure 2 I of 60 ) (such as Figure 2 F of 60 ) is the high-level semantic feature obtained by passing the second target image (such as Figure 2 I of 60 ) through a lightweight CNN network (such as Figure 2 Lightweight CNN shown in Figure 2 Semantic SegmentationNetwork). Among them, the lightweight CNN network is responsible for extracting detailed and complementary features to refine the semantic segmentation result, and can be simply designed as a cascade of multiple convolutional layers or share the same structure with the feature extraction part of the aforementioned semantic segmentation network (such as

[0062] Specifically, an optional implementation method is that the implementation process of this step S104 may include the following steps A1 - A2:

[0063] Step A1: Multiply the feature map of the compensated first target image by the feature map of the first target image, and splice the obtained feature map after multiplication with the feature map of the second target image to obtain the spliced feature map.

[0064] Specifically, in this implementation method, as Figure 2 shown, the feature map of the compensated first target image (such as Figure 2 Warp in 120->60 ) can be multiplied by the feature map of the first target image (such as Figure 2 F of 120 ) to obtain the feature map after multiplication (such as Figure 2 in ), and then the feature map after multiplication (such as Figure 2 in ) is spliced with the feature map of the second target image (such as Figure 2 F of60 ) are concatenated to obtain the concatenated feature map for performing subsequent step A2.

[0065] Step A2: The concatenated feature map is processed through a 1×1 convolutional layer to output the fused feature as the semantic fusion result, and based on the semantic fusion result, the semantic segmentation result of the second target image is determined; or, the concatenated feature map is first passed through a standard residual structure with 3×3 convolution and then through a 1×1 convolutional layer to output the fused feature as the semantic fusion result, and based on the semantic fusion result, the semantic segmentation result of the second target image is determined.

[0066] In this implementation, as Figure 2 shown, through step A1, the feature map of the compensated first target image (such as Figure 2 Warp in 120->60 ) and the feature map of the first target image (such as Figure 2 F in 120 ) are multiplied to obtain the multiplied feature map (such as Figure 2 in ), and then the multiplied feature map (such as Figure 2 in ) is concatenated with the feature map of the second target image (such as Figure 2 F in 60 ) to obtain the concatenated feature map. Further, the concatenated feature map can be processed through a 1×1 convolutional layer to output the fused feature as the semantic fusion result, and based on the semantic fusion result, the semantic segmentation result of the second target image is determined (such as Figure 2 S in 60 ); or, the concatenated feature map is first passed through a standard residual structure with 3×3 convolution and then through a 1×1 convolutional layer to output the fused feature as the semantic fusion result, and based on the semantic fusion result, the semantic segmentation result of the second target image is determined (such as Figure 2 S in 60 ). Among them, the residual structure is used to solve the degradation problem and the gradient problem to improve the performance of the semantic fusion (feature fusion) network.

[0067] S105: Semantically fuse the feature map of the compensated second target image, the semantic segmentation result of the second target image, and the feature map of the first target image, and determine the semantic segmentation result of the first target image based on the obtained fusion result.

[0068] In this embodiment, after obtaining the semantic segmentation result of the second target image through step S104, further, the feature map of the compensated second target image (such as Figure 2 Warp in 60->120)), the semantic segmentation result of the second target image, and the feature map of the first target image (such as Figure 2 's F 120 ) are semantically fused, and the semantic segmentation result of the first target image is determined according to the obtained fusion result

[0069] Specifically, an optional implementation manner is that the implementation process of this step S105 may include the following steps B1 - B2:

[0070] Step B1: Multiply the feature map of the compensated second target image by the semantic segmentation result of the second target image, and splice the obtained feature map with the feature map of the first target image to obtain a spliced feature map.

[0071] Specifically, in this implementation manner, as Figure 2 shown, the feature map of the compensated second target image (such as Figure 2 Warp in 60->120 ) can be multiplied by the semantic segmentation result of the second target image to obtain the obtained feature map (such as Figure 2 in ), and then the obtained feature map (such as Figure 2 in ) is spliced with the feature map of the first target image (such as Figure 2 's F 120 ) to obtain a spliced feature map for performing the subsequent step B2.

[0072] Step B2: Process the spliced feature map through a 1×1 convolutional layer to output a fused feature as the semantic fusion result, and determine the semantic segmentation result of the first target image according to the semantic fusion result; or, first process the spliced feature map through a standard residual structure with 3×3 convolution, and then process it through a 1×1 convolutional layer to output a fused feature as the semantic fusion result, and determine the semantic segmentation result of the first target image according to the semantic fusion result.

[0073] In this implementation manner, as Figure 2 shown, through step B1, the feature map of the compensated second target image (such as Figure 2 Warp in 60->120 ) is multiplied by the semantic segmentation result of the second target image (specifically the fusion result obtained through step S104) to obtain the obtained feature map (such as Figure 2 in ), and then the obtained feature map (such as Figure 2 in ) is spliced with the feature map of the first target image (such as Figure 2 's F 120) After splicing to obtain the spliced feature map, the spliced feature map can be further processed through a 1×1 convolutional layer to output the fused feature as the semantic fusion result, and based on the semantic fusion result, determine the semantic segmentation result of the first target image (as shown in Figure 2 S in 120 ); alternatively, the spliced feature map can be first passed through a standard residual structure with a 3×3 convolution, and then processed through a 1×1 convolutional layer to output the fused feature as the semantic fusion result, and based on the semantic fusion result, determine the semantic segmentation result of the first target image (as shown in Figure 2 S in 120 ). Among them, similarly, the residual structure is used to solve the degradation problem and the gradient problem to improve the performance of the semantic fusion (feature fusion) network.

[0074] In this way, by performing the above steps S101-105, image data association is established between in-vehicle cameras with a common field of view but different perspectives. By using semantic sharing and fusion between different in-vehicle cameras, the interpretability of the existing multi-object semantic segmentation model in terms of structure and results is improved, and the complementary advantages between the semantic segmentation results of different cameras are realized, thereby improving the accuracy and real-time performance of the semantic segmentation result.

[0075] In summary, a semantic segmentation method for an in-vehicle camera provided in this embodiment first obtains a first target image through an in-vehicle camera with a first perspective on a target vehicle; and obtains a second target image through an in-vehicle camera with a second perspective on the target vehicle; where the first perspective and the second perspective are different. Then, perform perspective transformation on the first target image to obtain the transformed first target image with the second perspective; and perform perspective transformation on the second target image to obtain the transformed second target image with the first perspective; then, use optical flow transformation to compensate for the distortion loss of the transformed first target image with the second perspective to obtain the compensated first target image; and use optical flow transformation to compensate for the distortion loss of the transformed second target image with the first perspective to obtain the compensated second target image; furthermore, the feature map of the compensated first target image, the feature map of the first target image, and the feature map of the second target image can be semantically fused, and based on the obtained fusion result, determine the semantic segmentation result of the second target image. Then, semantically fuse the feature map of the compensated second target image, the semantic segmentation result of the second target image, and the feature map of the first target image, and based on the obtained fusion result, determine the semantic segmentation result of the first target image. Thus, it is possible to realize the complementary advantages between the semantic segmentation results of in-vehicle cameras with different perspectives by using semantic sharing and fusion between them, and further improve the accuracy and real-time performance of the semantic segmentation result.

[0076] See Figure 3As shown in the figure, the present application also provides an embodiment of a semantic segmentation device for an in-vehicle camera, which may include:

[0077] An acquisition unit 301, configured to acquire a first target image through an in-vehicle camera with a first viewing angle on a target vehicle; and acquire a second target image through an in-vehicle camera with a second viewing angle on the target vehicle; the first viewing angle and the second viewing angle are different;

[0078] A conversion unit 302, configured to perform viewing angle conversion on the first target image to obtain a converted first target image with the second viewing angle; and perform viewing angle conversion on the second target image to obtain a converted second target image with the first viewing angle;

[0079] A compensation unit 303, configured to utilize optical flow transformation to compensate for the distortion loss of the converted first target image with the second viewing angle to obtain a compensated first target image; and utilize optical flow transformation to compensate for the distortion loss of the converted second target image with the first viewing angle to obtain a compensated second target image;

[0080] A first determination unit 304, configured to perform semantic fusion on the feature map of the compensated first target image, the feature map of the first target image, and the feature map of the second target image, and determine the semantic segmentation result of the second target image according to the obtained fusion result;

[0081] A second determination unit 305, configured to perform semantic fusion on the feature map of the compensated second target image, the semantic segmentation result of the second target image, and the feature map of the first target image, and determine the semantic segmentation result of the first target image according to the obtained fusion result.

[0082] In some possible implementation manners of the present application, the in-vehicle camera with the first viewing angle is an in-vehicle camera with a 120-degree horizontal viewing angle; the in-vehicle camera with the second viewing angle is an in-vehicle camera with a 60-degree horizontal viewing angle.

[0083] In some possible implementation manners of the present application, the first determination unit 304 includes:

[0084] A first multiplication sub-unit, configured to multiply the feature map of the compensated first target image and the feature map of the first target image, and splice the obtained feature map with the feature map of the second target image to obtain a spliced feature map;

[0085] A first determination subunit, configured to output, as a semantic fusion result, the fused features obtained by processing the concatenated feature map through a convolutional layer, and determine the semantic segmentation result of the second target image according to the semantic fusion result; or, first pass the concatenated feature map through a standard residual structure with a convolution, and then output the fused features obtained by processing through the convolutional layer as the semantic fusion result, and determine the semantic segmentation result of the second target image according to the semantic fusion result.

[0086] In some possible implementation manners of this application, the second determination unit 305 includes:

[0087] A second multiplication subunit, configured to multiply the feature map of the compensated second target image and the semantic segmentation result of the second target image, and splice the feature map obtained after the multiplication with the feature map of the first target image to obtain a concatenated feature map;

[0088] A second determination subunit, configured to output, as a semantic fusion result, the fused features obtained by processing the concatenated feature map through a 1×1 convolutional layer, and determine the semantic segmentation result of the first target image according to the semantic fusion result; or, first pass the concatenated feature map through a standard residual structure with a 3×3 convolution, and then output the fused features obtained by processing through a 1×1 convolutional layer as the semantic fusion result, and determine the semantic segmentation result of the first target image according to the semantic fusion result.

[0089] As can be seen from the above embodiments, the semantic segmentation device of the in-vehicle camera provided by the embodiments of the present application first obtains a first target image through the in-vehicle camera with the first view angle on the target vehicle; and obtains a second target image through the in-vehicle camera with the second view angle on the target vehicle; wherein, the first view angle and the second view angle are different. Then, the first target image is subjected to view angle conversion to obtain the converted first target image with the second view angle; and the second target image is subjected to view angle conversion to obtain the converted second target image with the first view angle; Next, using optical flow transformation to compensate for the distortion loss of the converted first target image with the second view angle to obtain the compensated first target image; and using optical flow transformation to compensate for the distortion loss of the converted second target image with the first view angle to obtain the compensated second target image; Furthermore, the feature map of the compensated first target image, the feature map of the first target image, and the feature map of the second target image can be semantically fused, and the semantic segmentation result of the second target image can be determined according to the obtained fusion result. Then, the feature map of the compensated second target image, the semantic segmentation result of the second target image, and the feature map of the first target image are semantically fused, and the semantic segmentation result of the first target image is determined according to the obtained fusion result. Thus, it is possible to utilize the semantic sharing and fusion between in-vehicle cameras with different view angles to achieve complementary advantages among the semantic segmentation results of each camera, thereby improving the accuracy and real-time performance of the semantic segmentation result.

[0090] Furthermore, the embodiments of the present application also provide a semantic segmentation device for an in-vehicle camera, including: a processor, a memory, and a system bus;

[0091] The processor and the memory are connected through the system bus;

[0092] The memory is used to store one or more programs, and the one or more programs include instructions, and when the instructions are executed by the processor, the processor is caused to execute any implementation method of the above-mentioned semantic segmentation method for the in-vehicle camera.

[0093] Furthermore, the embodiments of the present application also provide a computer-readable storage medium, in which instructions are stored, and when the instructions are run on a terminal device, the terminal device is caused to execute any implementation method of the above-mentioned semantic segmentation method for the in-vehicle camera.

[0094] As can be seen from the description of the above embodiments, those skilled in the art can clearly understand that all or part of the steps in the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network communication device such as a media gateway, etc.) to execute the methods described in various embodiments or certain parts of the embodiments of the present application.

[0095] It should be noted that the various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part.

[0096] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0097] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A semantic segmentation method for an in-vehicle camera, characterized in that, Including: Obtaining a first target image through an in-vehicle camera at a first perspective on a target vehicle; And obtaining a second target image through an in-vehicle camera at a second perspective on the target vehicle; the first perspective and the second perspective are different; Performing perspective transformation on the first target image to obtain a transformed first target image at the second perspective; And performing perspective transformation on the second target image to obtain a transformed second target image at the first perspective; Using optical flow transformation to compensate for the distortion loss of the transformed first target image at the second perspective, obtaining a compensated first target image; And using optical flow transformation to compensate for the distortion loss of the transformed second target image at the first perspective, obtaining a compensated second target image; Semantically fusing the feature maps of the compensated first target image, the feature map of the first target image, and the feature map of the second target image, and determining the semantic segmentation result of the second target image according to the obtained fusion result; Semantically fusing the feature map of the compensated second target image, the semantic segmentation result of the second target image, and the feature map of the first target image, and determining the semantic segmentation result of the first target image according to the obtained fusion result.

2. The method according to claim 1, wherein The in-vehicle camera at the first perspective is an in-vehicle camera with a 120-degree horizontal viewing angle; the in-vehicle camera at the second perspective is an in-vehicle camera with a 60-degree horizontal viewing angle.

3. The method according to claim 1, characterized in that, The semantically fusing the feature maps of the compensated first target image, the feature map of the first target image, and the feature map of the second target image, and determining the semantic segmentation result of the second target image according to the obtained fusion result includes: Multiplying the feature map of the compensated first target image and the feature map of the first target image, and splicing the obtained feature map after multiplication with the feature map of the second target image to obtain a spliced feature map; Processing the spliced feature map through a 1×1 convolutional layer to output a fused feature as the semantic fusion result, and determining the semantic segmentation result of the second target image according to the semantic fusion result; or, first processing the spliced feature map through a standard residual structure with 3×3 convolution, and then processing it through a 1×1 convolutional layer to output a fused feature as the semantic fusion result, and determining the semantic segmentation result of the second target image according to the semantic fusion result.

4. The method according to claim 1, characterized in that, The semantically fusing the feature map of the compensated second target image, the semantic segmentation result of the second target image, and the feature map of the first target image, and determining the semantic segmentation result of the first target image according to the obtained fusion result includes: Multiplying the feature map of the compensated second target image and the semantic segmentation result of the second target image, and splicing the obtained feature map after multiplication with the feature map of the first target image to obtain a spliced feature map; The fused features are output after processing the concatenated feature maps through a 1×1 convolutional layer as the semantic fusion result, and the semantic segmentation result of the first target image is determined according to the semantic fusion result; or, the concatenated feature maps are first passed through a standard residual structure with 3×3 convolutions and then through a 1×1 convolutional layer to output the fused features as the semantic fusion result, and the semantic segmentation result of the first target image is determined according to the semantic fusion result.

5. A semantic segmentation device for an in-vehicle camera, characterized in that, including: an acquisition unit, configured to acquire a first target image through an in-vehicle camera at a first viewing angle on a target vehicle; and acquire a second target image through an in-vehicle camera at a second viewing angle on the target vehicle; the first viewing angle and the second viewing angle are different; a conversion unit, configured to perform a viewing angle conversion on the first target image to obtain a converted first target image at the second viewing angle; and perform a viewing angle conversion on the second target image to obtain a converted second target image at the first viewing angle; a compensation unit, configured to utilize optical flow transformation to compensate for the distortion loss of the converted first target image at the second viewing angle to obtain a compensated first target image; and utilize optical flow transformation to compensate for the distortion loss of the converted second target image at the first viewing angle to obtain a compensated second target image; a first determination unit, configured to perform semantic fusion on the feature maps of the compensated first target image, the feature maps of the first target image, and the feature maps of the second target image, and determine the semantic segmentation result of the second target image according to the obtained fusion result; a second determination unit, configured to perform semantic fusion on the feature maps of the compensated second target image, the semantic segmentation result of the second target image, and the feature maps of the first target image, and determine the semantic segmentation result of the first target image according to the obtained fusion result.

6. The device according to claim 5, characterized in that, The in-vehicle camera at the first viewing angle is an in-vehicle camera with a 120-degree horizontal viewing angle; the in-vehicle camera at the second viewing angle is an in-vehicle camera with a 60-degree horizontal viewing angle.

7. The device according to claim 5, characterized in that, The first determination unit includes: a first multiplication sub-unit, configured to multiply the feature maps of the compensated first target image and the feature maps of the first target image, and concatenate the obtained feature maps after multiplication with the feature maps of the second target image to obtain concatenated feature maps; a first determination sub-unit, configured to output the fused features after processing the concatenated feature maps through a convolutional layer as the semantic fusion result, and determine the semantic segmentation result of the second target image according to the semantic fusion result; or, the concatenated feature maps are first passed through a standard residual structure with convolutions and then through a convolutional layer to output the fused features as the semantic fusion result, and the semantic segmentation result of the second target image is determined according to the semantic fusion result.

8. The device according to claim 5, characterized in that, The second determination unit includes: a second multiplication sub-unit, configured to multiply the feature maps of the compensated second target image and the semantic segmentation result of the second target image, and concatenate the obtained feature maps after multiplication with the feature maps of the first target image to obtain concatenated feature maps; A second determination subunit, configured to output the fused features after processing the spliced feature map through a 1×1 convolutional layer as the semantic fusion result, and determine the semantic segmentation result of the first target image according to the semantic fusion result; or, first process the spliced feature map through a standard residual structure with 3×3 convolution, and then output the fused features after processing through a 1×1 convolutional layer as the semantic fusion result, and determine the semantic segmentation result of the first target image according to the semantic fusion result.

9. A semantic segmentation device for an in-vehicle camera, characterized in that, Comprising: A processor, a memory, and a system bus; The processor and the memory are connected through the system bus; The memory is configured to store one or more programs, the one or more programs include instructions, and the instructions, when executed by the processor, cause the processor to execute the method according to any one of claims 1-4.

10. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium, and when the instructions are run on a terminal device, the terminal device is caused to execute the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Image semantic segmentation method, terminal and readable storage medium

    CN110348351A

  • Lung lobe segmentation method and device based on multiple view angles

    CN111292343A