Video Image Stitching Method, Device and Equipment
The screen ratio and display position of the video image frame are calculated through the vehicle-mounted radar ranging data and preset models, which solves the depth of field problem under parallax and improves the effect and real-timeness of video image stitching.
Patent Information
- Application Number
- CN202111677340.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-31
AI Technical Summary
The prior art cannot effectively solve the depth of field problem in video image stitching under parallax, resulting in poor stitching effect.
By using the vehicle-mounted radar ranging data and preset models, the screen ratio and display position information of the area of interest (ROI) and the original video image frame are calculated, thereby cropping and splicing the target ROI to improve the splicing effect of the video image.
This method improves the continuity and display effect after stitching by reflecting the different depth of field of the video image, and meets the high requirements for real-timeness of video stitching.
Smart Images

Figure CN114331848B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving technology, and in particular, to a method, device and equipment for video image stitching. Background Art
[0002] With the continuous development of autonomous driving technology, vehicles are equipped with more and more sensors with different functions. For example, cameras and / or various types of radars are deployed around the vehicle. Among them, the radar is mainly used to measure the distance between the vehicle and obstacles, while the camera is used to collect the scene images around the vehicle. In remote driving, the vehicle can send the video images collected by the sensors to the remote cockpit, and the remote cockpit can perform stitching processing on the video images.
[0003] In the related art, when stitching images, a traditional pure vision image stitching algorithm is adopted, and the problem of low real-time performance in the video image stitching method is solved by asynchronously updating the stitching model. However, the video image stitching processing method in the related art cannot solve the depth of field problem under parallax, and the stitching effect is poor. Summary of the Invention
[0004] To solve or partially solve the problems existing in the related art, the present application provides a method, device and equipment for video image stitching, which can improve the stitching effect of video images.
[0005] The first aspect of the present application provides a method for video image stitching, including:
[0006] Obtaining the aspect ratio of the region of interest (ROI) to the original video image frame according to the on-vehicle radar ranging data and a first preset model;
[0007] Obtaining the display position information of the ROI on the original video image frame according to the aspect ratio and a second preset model;
[0008] Cropping the target ROI from the original video image frame according to the aspect ratio and the display position information;
[0009] Stitching the cropped target ROIs into a target video image.
[0010] In an embodiment, the step of obtaining the aspect ratio of the region of interest (ROI) to the original video image frame according to the on-vehicle radar ranging data and a first preset model includes: inputting the on-vehicle radar ranging data into a pre-trained shallow neural network model, and outputting the aspect ratio of the region of interest (ROI) to the original video image frame.
[0011] Obtaining the display position information of the ROI on the original video image frame according to the screen ratio and the second preset model includes: inputting the screen ratio into the fitting model and outputting the display position information of the ROI on the original video image frame.
[0012] In one embodiment, the shallow neural network model is trained in the following manner:
[0013] The shallow neural network model is trained using a training set to obtain the pre-trained shallow neural network model, where the training set includes labeled screen ratios and training radar ranging data, and the labeled screen ratio is the screen ratio of the labeled training ROI to the training video image frame.
[0014] In one embodiment, the training set is obtained in the following manner:
[0015] Select a set number of data from the collected radar ranging data as the training radar ranging data;
[0016] According to the principle of continuous frames, label the ROI that is continuous with the previous frame of the training video image frame from the training video image frames to obtain the labeled training ROI;
[0017] Compare the labeled training ROI with the training video image frame to obtain the labeled screen ratio;
[0018] Save the training radar ranging data and the labeled screen ratio as the training set.
[0019] In one embodiment, the fitting model is obtained in the following manner:
[0020] Perform fitting processing on the labeled display position of the labeled training ROI on the training video image frame and the labeled screen ratio to obtain the fitting model.
[0021] In one embodiment, performing the fitting of the labeled display position of the labeled training ROI on the training video image frame and the labeled screen ratio to obtain the fitting model includes:
[0022] Input each of the labeled screen ratios as known quantities into a target fitting equation with the target display position as the unknown quantity, and iteratively solve for the target display position, where the target fitting equation includes polynomial fitting coefficients;
[0023] When the deviation between the target display position and the labeled display position of the labeled training ROI on the training video image frame is less than a preset threshold, determine that the corresponding polynomial fitting coefficient value is the target fitting coefficient value, and use the target fitting equation determined by the target fitting coefficient value as the fitting model.
[0024] The second aspect of the present application provides a video image stitching device, including:
[0025] A first output module, configured to obtain the screen ratio of the region of interest (ROI) to the original video image frame according to the vehicle-mounted radar ranging data and a first preset model;
[0026] A second output module, configured to obtain the display position information of the ROI on the original video image frame according to the screen ratio and a second preset model;
[0027] A target area module, configured to crop the target ROI from the original video image frame according to the screen ratio obtained by the first output module and the display position information obtained by the second output module;
[0028] A stitching module, configured to stitch the target ROIs cropped by the target area module into a target video image.
[0029] In an embodiment, the first output module inputs the vehicle-mounted radar ranging data into a pre-trained shallow neural network model and outputs the screen ratio of the region of interest (ROI) to the original video image frame;
[0030] The second output module inputs the screen ratio into the fitting model and outputs the display position information of the ROI on the original video image frame.
[0031] In an embodiment, the device further includes:
[0032] A model training module, configured to train the shallow neural network model by using a training set to obtain the pre-trained shallow neural network model, where the training set includes labeled screen ratios and training radar ranging data, and the labeled screen ratio is the screen ratio of the labeled training ROI to the training video image frame.
[0033] The third aspect of the present application provides an electronic device, including:
[0034] A processor; and
[0035] A memory, on which executable code is stored, and when the executable code is executed by the processor, the processor is caused to execute the method as described above.
[0036] The fourth aspect of the present application provides a computer-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor is caused to execute the method as described above.
[0037] The technical solution provided by the present application may include the following beneficial effects:
[0038] In the present application, vehicle-mounted radar ranging data is used as the data input volume. The data volume of the vehicle-mounted radar ranging data itself is small, and the calculation amount input into the preset model is also relatively small. In addition, the vehicle-mounted radar ranging data can also reflect different depths of field of the video image. Therefore, the target ROI obtained according to the screen ratio of the region of interest ROI to the original video image frame and the display position information of the ROI on the original video image frame can make the subsequent spliced screen more continuous, and make the screen display effect of the target video image spliced according to the target ROI better, thereby improving the splicing effect of the video image.
[0039] Further, in the present application, the vehicle-mounted radar ranging data is input into a pre-trained shallow neural network model to output the screen ratio of the region of interest ROI to the original video image frame, and the screen ratio is input into a fitting model to output the display position information of the ROI on the original video image frame. Since the number of layers of the shallow neural network model itself is small, compared with the deep neural network, the operation process of the shallow neural network model is much simpler. Therefore, the calculation speed is relatively fast; and the fitting model belongs to a traditional mathematical model, and it can also quickly complete the output of the display position information of the ROI on the original video image frame. Therefore, the technical solution of the present application can meet the high real-time requirements of video splicing.
[0040] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] By describing the exemplary embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features and advantages of the present application will become more obvious. Among them, in the exemplary embodiments of the present application, the same reference numerals generally represent the same components.
[0042] Figure 1 is a schematic flowchart of a video image splicing method shown in an embodiment of the present application;
[0043] Figure 2 is another schematic flowchart of a video image splicing method shown in an embodiment of the present application;
[0044] Figure 3 is another schematic flowchart of a video image splicing method shown in an embodiment of the present application;
[0045] Figure 4 It is a schematic flowchart of training a model in the video image stitching method shown in the embodiments of the present application;
[0046] Figure 5 It is a schematic flowchart of applying a model in the video image stitching method shown in the embodiments of the present application;
[0047] Figure 6 It is a schematic structural diagram of a video image stitching device shown in the embodiments of the present application;
[0048] Figure 7 It is another schematic structural diagram of a video image stitching device shown in the embodiments of the present application;
[0049] Figure 8 It is a schematic structural diagram of an electronic device shown in the embodiments of the present application. Detailed implementation manners
[0050] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.
[0051] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the" and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0052] It should be understood that although the terms "first", "second", "third", etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, "a plurality" means two or more unless otherwise specifically defined.
[0053] The video image stitching processing method in the related art cannot solve the depth of field problem under parallax, and the stitching effect is poor. In view of the above problems, the embodiments of the present application provide a video image stitching method, which can improve the stitching effect of video images.
[0054] The technical solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0055] Figure 1 It is a schematic flowchart of the video image stitching method shown in the embodiments of the present application.
[0056] See Figure 1 , the method includes:
[0057] S101. Obtain the screen ratio of the region of interest (ROI) to the original video image frame according to the on-vehicle radar ranging data and the first preset model.
[0058] In this step, the on-vehicle radar ranging data can be input into a pre-trained shallow neural network model, and the screen ratio of the ROI (Region Of Interest) to the original video image frame is output.
[0059] Among them, the shallow neural network model is trained in the following manner: The shallow neural network model is trained with a training set to obtain a pre-trained shallow neural network model, where the training set includes the labeled screen ratio and the training radar ranging data, and the labeled screen ratio is the screen ratio of the labeled training ROI to the training video image frame.
[0060] Among them, the training set can be obtained in the following manner: Select a set number of data from the collected radar ranging data as the training radar ranging data; According to the principle of continuous frames, the ROI that is continuous with the previous frame of the training video image frame is labeled from the training video image frame to obtain the labeled training ROI; The labeled screen ratio is obtained by comparing the labeled training ROI with the training video image frame; The training radar ranging data and the labeled screen ratio are saved as the training set.
[0061] S102. Obtain the display position information of the ROI on the original video image frame according to the screen ratio and the second preset model.
[0062] In this step, the screen ratio can be input into the fitting model, and the display position information of the ROI on the original video image frame is output.
[0063] Among them, the fitting model is obtained in the following manner: The labeled display position and the labeled screen ratio of the labeled training ROI on the training video image frame are subjected to fitting processing to obtain the fitting model.
[0064] For example, each marked screen ratio can be input as a known quantity into a target fitting equation with the target display position as an unknown quantity, and the target display position can be iteratively solved. The target fitting equation includes polynomial fitting coefficients; when the deviation between the target display position and the marked display position of the marked training ROI on the training video image frame is less than a preset threshold, the corresponding polynomial fitting coefficient value is determined as the target fitting coefficient value, and the target fitting equation determined by the target fitting coefficient value is used as the fitting model.
[0065] S103. Crop the target ROI from the original video image frame according to the screen ratio and the display position information.
[0066] Since the screen size of the original video image frame is known or fixed, when the screen ratio of the ROI to the original video image frame and the display position information of the ROI on the original video image frame are determined, the target ROR, that is, the final ROI, can be conveniently cropped from the original video image frame.
[0067] S104. Stitch the cropped target ROIs into a target video image.
[0068] After obtaining the target ROR, that is, the final ROI, when video image stitching needs to be performed subsequently, stitching using the target ROI can make the screen of the stitched video image more continuous and the stitching effect better.
[0069] It can be seen from this embodiment that the present application uses the ranging data of the vehicle-mounted radar as the data input volume. The amount of data of the vehicle-mounted radar ranging data itself is small, and the calculation amount input to the preset model is also relatively small. In addition, the vehicle-mounted radar ranging data can also reflect the different depths of field of the video image. Therefore, the target ROI obtained according to the screen ratio of the region of interest ROI to the original video image frame and the display position information of the ROI on the original video image frame can make the subsequent stitched screen more continuous, and make the screen display effect of the target video image stitched according to the target ROI better, thereby improving the stitching effect of the video image.
[0070] Figure 2 It is another schematic flowchart of the video image stitching method shown in the embodiment of the present application.
[0071] See Figure 2 , the method includes:
[0072] S201. Input the ranging data of the vehicle-mounted radar into a pre-trained shallow neural network model, and output the screen ratio of the region of interest ROI to the original video image frame.
[0073] Wherein, the original video image frame is a video image frame of the vehicle driving environment collected by the vehicle-mounted camera.
[0074] In the embodiments of the present application, the vehicle-mounted radar may be an ultrasonic radar, a millimeter-wave radar, etc., which can be deployed around the vehicle, and the number can be deployed according to actual needs. For example, 12 ultrasonic radars can be deployed. The ranging data of the vehicle-mounted radar mainly includes the distance information between the target or obstacle measured by the vehicle-mounted radar in real time and the vehicle. This distance information can correspond to the image depth of field of the video image frame of the vehicle driving environment captured by the vehicle-mounted camera at the same time when the vehicle-mounted radar collects the ranging data. Therefore, the ranging data of the vehicle-mounted radar can also reflect the different depths of field of the video image.
[0075] The Region Of Interest (ROI) is the area on the image that is focused on when analyzing and processing the image. This area may be the position of the target or obstacle on the image. In machine vision and image processing, the area that needs to be processed outlined in the processed image in the form of a square, circle, ellipse, irregular polygon, etc. is called the Region Of Interest. The typical shape of the ROI can be a rectangle or other shapes. When the ROI is a rectangle, the aspect ratio of the ROI to the original video image frame includes the ratio of the length of the ROI to the length of the original video image frame and the ratio of the width of the ROI to the width of the original video image frame.
[0076] The shallow neural network model can be a basic neural network model that only includes an input layer, a hidden layer, and an output layer. Each layer uses the sigmoid (S-shaped function) as the activation function. The sigmoid function is used for the output of the hidden layer neurons, and its value range is (0, 1). It can map a real number to the interval (0, 1) and can be used for binary classification. It has a better effect when the features are relatively complex or not very different. Since the number of layers of the shallow neural network model itself is small, compared with the deep neural network, the operation process of the shallow neural network model is much simpler. Therefore, the calculation speed is fast, and it can better meet the requirements of high real-time performance.
[0077] This application can train a shallow neural network model using a training set to obtain a pre-trained shallow neural network model. Among them, the training set includes the labeled frame ratio and the radar ranging data for training. The labeled frame ratio is the ratio of the labeled training ROI to the training video image frame, and the labeled display position is the labeled display position of the labeled training ROI on the training video image frame. The training set can be obtained in the following way: select a set number of data from the collected radar ranging data as the radar ranging data for training; according to the principle of continuous frames, label the ROI that is continuous with the previous training video image frame from the training video image frames to obtain the labeled training ROI; compare the labeled training ROI with the training video image frame to obtain the labeled frame ratio; save the radar ranging data for training and the labeled frame ratio as the training set.
[0078] S202. Input the ratio of the ROI to the original video image frame into the fitting model, and output the display position information of the ROI on the original video image frame.
[0079] In the embodiment of this application, the role of the fitting model is that when the ratio of the ROI to the original video image frame is input, it can output the display position information of the ROI on the original video image frame.
[0080] Similar to the pre-training of the shallow neural network model before application, the fitting model here can also be a mathematical model obtained through fitting processing. This application can perform fitting processing on the labeled display position and the labeled frame ratio of the labeled training ROI on the training video image frame to obtain the fitting model.
[0081] For example, input each labeled frame ratio as a known quantity into the target fitting equation with the target display position as the unknown quantity, and iteratively solve the target display position. The target fitting equation includes polynomial fitting coefficients; when the deviation between the target display position and the labeled display position of the labeled training ROI on the training video image frame is less than the preset threshold, determine the corresponding polynomial fitting coefficient value as the target fitting coefficient value, and use the target fitting equation determined by the target fitting coefficient value as the fitting model.
[0082] S203. Crop the target ROI from the original video image frame according to the ratio of the ROI to the original video image frame and the display position information of the ROI on the original video image frame.
[0083] Since the frame size of the original video image frame is known or fixed, when the ratio of the ROI to the original video image frame and the display position information of the ROI on the original video image frame are determined, it is convenient to crop the target ROR, that is, the final ROI, from the original video image frame.
[0084] S204. Stitch the obtained target ROIs into a target video image frame.
[0085] The target ROIs obtained in the above process may be from the original video image frames of the same camera or different cameras (for example, the front camera and the left or right camera). When stitching these target ROIs, one of the cameras can be used as a reference, such as the original video image frame collected by the front - facing camera in front of the vehicle or the target ROI of the video image frame with better quality.
[0086] Among them, stitching the obtained target ROIs into a target video image frame can be realized by using existing stitching processing methods. For example, stitching can be performed according to reference factors such as the pose or shooting time corresponding to the video image frame. This application does not limit it here.
[0087] It can be seen from this embodiment that this application inputs the ranging data of the vehicle - mounted radar into a pre - trained shallow neural network model, outputs the aspect ratio of the region of interest (ROI) to the original video image frame, and inputs the aspect ratio into a fitting model to output the display position information of the ROI on the original video image frame. Since the number of layers of the shallow neural network model is relatively small, compared with the deep neural network, the operation process of the shallow neural network model is much simpler. Therefore, the calculation speed is relatively fast; and the fitting model belongs to a traditional mathematical model, and it can also quickly complete the output of the display position information of the ROI on the original video image frame. Therefore, the technical solution of this application can meet the high real - time requirements of video stitching.
[0088] Figure 3 It is another schematic flowchart of the video image stitching method shown in the embodiment of this application. Figure 3 Relative to Figure 1 and Figure 2 This application solution is described in more detail.
[0089] See Figure 3 , this method includes:
[0090] S301. Collect the ranging data of the vehicle - mounted radar and the video image frames of the camera and perform pre - processing.
[0091] In this application, when an object enters the effective range of the ultrasonic radar around the driving vehicle, the video image frame data of all current cameras and the ranging data of the vehicle - mounted radar can be recorded every set time, such as 5 s, and saved.
[0092] It should be noted that, for the convenience of training a shallow neural network model and further reducing the computational complexity, in the embodiments of the present application, the ranging data of the vehicle-mounted radar can be preprocessed. Such preprocessing includes, for example, data cleaning (e.g., removing dirty point data, i.e., ranging data that obviously does not meet the requirements) and normalization of the data, etc.; the normalized ranging data of the vehicle-mounted radar is more intuitive and convenient for researchers.
[0093] S302. Pre-train the shallow neural network model to obtain a trained shallow neural network model, and perform fitting processing on the fitting model to obtain a processed fitting model.
[0094] The processing procedure of this step can be referred to Figure 4 as shown in Figure 4 FIG. 10 is a schematic flowchart of training a model in the video image stitching method shown in the embodiments of the present application.
[0095] In the present application, a training set can be used to train the shallow neural network model to obtain a pre-trained shallow neural network model. The training set can be obtained in the following manner: select a set number of data from the collected radar ranging data as the training radar ranging data; according to the principle of continuous frames, mark the ROI that is continuous with the previous frame of the training video image frame from the training video image frames to obtain the marked training ROI; compare the marked training ROI with the training video image frames to obtain the marked picture ratio; save the training radar ranging data and the marked picture ratio as the training set.
[0096] Specifically, the process of obtaining the training set is described below through steps S1 to S4.
[0097] Step S1: Select a set number of data from the collected radar ranging data as the training radar ranging data.
[0098] From the collected radar ranging data, the first set number of data can be selected as the training radar ranging data, the second set number of data can be selected as the test radar ranging data, and the third set number of data can be selected as the verification radar ranging data. For example, 50% of the randomly selected data from the collected radar ranging data is used as the training radar ranging data, and the rest is used to test or verify the trained shallow neural network model. Among them, the selected training radar ranging data can be, for example, the radar ranging data of 6 ultrasonic radars in the front of the vehicle.
[0099] The radar ranging data measures the distance information between the target or obstacle and the vehicle in real time, and this distance information can correspond to the image depth of field of the vehicle driving environment video image frame captured by the vehicle-mounted camera at the same time when the vehicle-mounted radar collects the ranging data.
[0100] Step S2: According to the principle of continuous frames, mark the ROI that is continuous with the previous training video image frame from the training video image frames to obtain the marked training ROI.
[0101] Specifically, a region can be selected in the training video image frame for stretching or shrinking, and the operation stops when it is found to be continuous with the previous training video image frame. At this time, the obtained region is the marked training ROI. At this time, the frame length of the marked training ROI is the marked length, and the frame width is the marked width. For example, if it is determined that there is no obvious jump in the pixels of the two, the frames can be determined to be continuous.
[0102] Step S3: Compare the marked training ROI with the training video image frames to obtain the marked frame ratio.
[0103] In this step, by comparing the marked training ROI with the training video image frames, the marked frame ratio can be obtained. When the marked training ROI is a rectangle, the marked frame ratio includes the ratio of the marked length of the ROI to the frame length of the original video image frame and the ratio of the marked width of the ROI to the frame width of the original video image frame.
[0104] Step S4: Save the training radar ranging data and the marked frame ratio as a training set.
[0105] Save the training radar ranging data and the marked frame ratio as a training set. Subsequently, input the training radar ranging data and its corresponding marked frame ratio into the trained shallow neural network model to complete the training of the shallow neural network model.
[0106] Among them, the implementation method of performing fitting processing on the fitting model to obtain the processed fitting model may include:
[0107] Take each marked frame ratio as a known quantity and input it into the target fitting equation with the target display position as the unknown quantity, and iteratively solve the target display position, where the target fitting equation includes polynomial fitting coefficients;
[0108] When the deviation between the target display position and the marked display position of the marked training ROI on the training video image frame is less than the preset threshold, determine that the corresponding polynomial fitting coefficient value is the target fitting coefficient value, and use the target fitting equation determined by the target fitting coefficient value as the fitting model. The preset threshold can be, for example, 0.05 but is not limited to this.
[0109] In the above embodiments, the marked display position is the true value during the fitting process, the target display position is the predicted value during the fitting process, and the deviation can be the mean square error (MSE) or root mean square error (RMSE) between the target display position and the marked display position. The target fitting equation can be a linear equation or a non-linear equation. In addition, it should be noted that when the deviation between the target display position and the marked display position is less than the preset threshold, it can be considered that the deviation between the target display position and the marked display position is the smallest at this time.
[0110] For example, assuming that the marked screen ratio, i.e., the ratio data (k), and the marked display position, i.e., the position data (x, y), are linearly correlated, the target fitting equation can be exemplified as follows but is not limited thereto:
[0111] x = a + bk + ck^2 + dk^3 + ek^4 + fk^5 + gk^6
[0112] y = h + ik + jk^2 + lk^3 + mk^4 + nk^5 + pk^6
[0113] Among them, a, b, c, d, e, f, g in the first polynomial equation and h, i, j, l, m, n, p in the second polynomial equation are polynomial fitting coefficients.
[0114] The so-called fitting is to connect a series of points on a plane with a smooth curve. Since there are countless possibilities for this curve, there are various fitting methods. The fitting curve can generally be represented by a function, and there are different fitting names according to the different functions. Commonly used fitting methods include, for example, the least squares curve fitting method, etc. In the MATLAB (a mathematical software) tool, the polyfit function can also be used to fit polynomials. The polyfit function is based on the least squares method. In the above polynomial equations, each formula can have a solution as long as there are seven or more groups of non-linear data, that is, by using the polyfit function in MATLAB for fitting processing, the values of the polynomial fitting coefficients corresponding to the closest position data (x, y) to the true value (marked display position) can be obtained. Then, the values of the polynomial fitting coefficients corresponding at this time can be used as the target fitting coefficient values, and the target fitting equation determined by the target fitting coefficient values is used as the fitting model.
[0115] It should be noted that the specific process of using the polyfit function to fit polynomials in MATLAB can be implemented according to the existing technology, and this application does not limit it here. In addition, the above polynomial equations are exemplified by seven-term equations but are not limited thereto.
[0116] S303. Input the ranging data of the vehicle-mounted radar into a pre-trained shallow neural network model, and output the aspect ratio of the region of interest (ROI) to the original video image frame.
[0117] The shallow neural network model can be a basic neural network model that only includes an input layer, a hidden layer, and an output layer. Each layer uses the sigmoid (S-shaped function) as the activation function. Since the number of layers of the shallow neural network model is relatively small, compared with the deep neural network, the operation process of the shallow neural network model is much simpler. Therefore, the calculation speed is relatively fast, which can better meet the requirements of high real-time performance.
[0118] This step can input the ranging data of the vehicle-mounted radar into a pre-trained shallow neural network model for operation and processing, and output the aspect ratio of the region of interest (ROI) to the original video image frame. For the description of this step, please refer to the description in step S201, which will not be repeated here.
[0119] S304. Input the aspect ratio of the ROI to the original video image frame into a fitting model, and output the display position information of the ROI on the original video image frame.
[0120] In the embodiment of the present application, the role of the fitting model is to output the display position information of the ROI on the original video image frame when the aspect ratio of the ROI to the original video image frame is input.
[0121] In this step, input the aspect ratio of the ROI to the original video image frame into the fitting model for fitting operation, and the display position information of the ROI on the original video image frame can be output.
[0122] S305. According to the aspect ratio of the ROI to the original video image frame and the display position information of the ROI on the original video image frame, crop the target ROI from the original video image frame.
[0123] Since the size of the original video image frame is known or fixed, when the aspect ratio of the ROI to the original video image frame and the display position information of the ROI on the original video image frame are determined, it is very convenient to crop the target ROR, that is, the final ROI, from the original video image frame.
[0124] Among them, the processes of the above steps S303 to S305 can be referred to simultaneously Figure 5 as shown Figure 5 which is a schematic flowchart of applying the model in the video image stitching method shown in the embodiment of the present application.
[0125] S306. Stitch the target ROIs obtained by cropping into a target video image frame.
[0126] The target ROI obtained from the above process may be the original video image frames from the same camera or different cameras (for example, the front camera and the left or right camera). When splicing these target ROIs, one of the cameras can be used as a reference, such as the original video image frame captured by the front-facing camera of the vehicle or the target ROI of the video image frame with better quality.
[0127] Among them, splicing the target ROIs obtained by each cropping into a target video image frame can be achieved by using existing splicing processing methods. For example, splicing can be performed according to reference factors such as the pose or shooting time corresponding to the video image frame. This application does not limit this here.
[0128] It can be seen from this embodiment that this application uses the ranging data of the vehicle-mounted radar as the data input volume. The data volume of the vehicle-mounted radar ranging data itself is small, and the calculation amount input into the preset model is also relatively small. The vehicle-mounted radar ranging data can also reflect the different depths of field of the video image. In addition, this application also uses the operation of the shallow neural network model. Since the number of layers of the shallow neural network model itself is small, compared with the deep neural network, the operation process of the shallow neural network model is much simpler and the operation speed is fast. Therefore, the technical solution of this application can meet the high real-time requirements of video splicing, and can also make the display effect of the target video image spliced according to the target ROI better, thus improving the splicing effect of the video image.
[0129] Corresponding to the foregoing embodiment of the application function implementation method, this application also provides a video image splicing device, an electronic device, and corresponding embodiments.
[0130] Figure 6 It is a schematic structural diagram of the video image splicing device shown in the embodiment of this application.
[0131] See Figure 6 , a video image splicing device 60 provided by this application includes: a first output module 601, a second output module 602, a target area module 603, and a splicing module 604.
[0132] The first output module 601 is used to obtain the screen ratio of the region of interest ROI to the original video image frame according to the ranging data of the vehicle-mounted radar and the first preset model. Among them, the first preset model can be a shallow neural network model, and the shallow neural network model is trained in the following way: the shallow neural network model is trained using a training set to obtain a pre-trained shallow neural network model, where the training set includes labeled screen ratios and training radar ranging data, and the labeled screen ratio is the screen ratio of the labeled training ROI to the training video image frame.
[0133] The second output module 602 is configured to obtain the display position information of the ROI on the original video image frame according to the screen ratio and the second preset model. The second preset model may be a fitting model. The fitting model can be obtained in the following way: Fit the labeled display position of the ROI for training on the training video image frame and the labeled screen ratio to obtain the fitting model.
[0134] The target area module 603 is configured to crop the target ROI from the original video image frame according to the screen ratio obtained by the first output module 601 and the display position information obtained by the second output module 602. Since the screen size of the original video image frame is known or fixed, when the screen ratio of the ROI to the original video image frame and the display position information of the ROI on the original video image frame are determined, the target area module 603 can easily crop the target ROR, that is, the final ROI, from the original video image frame.
[0135] The splicing module 604 is configured to splice the target ROIs cropped by the target area module 603 into a target video image. Using the target ROI for splicing can make the picture of the spliced video image more continuous and the splicing effect better.
[0136] It can be seen from this embodiment that the video image splicing device provided by the present application uses the vehicle-mounted radar ranging data as the data input volume. The amount of data of the vehicle-mounted radar ranging data itself is small, and the calculation amount input to the preset model is also relatively small. In addition, the vehicle-mounted radar ranging data can also reflect the different depths of field of the video image. Therefore, the target ROI obtained according to the screen ratio of the region of interest ROI to the original video image frame and the display position information of the ROI on the original video image frame can make the subsequent spliced picture more continuous and make the display effect of the target video image spliced according to the target ROI better, thereby improving the splicing effect of the video image.
[0137] Figure 7 It is another structural schematic diagram of the video image splicing device shown in the embodiment of the present application.
[0138] See Figure 7 , a video image splicing device 60 provided by the present application includes: a first output module 601, a second output module 602, a target area module 603, a splicing module 604, a model training module 605, and a data collection module 606.
[0139] The functions of the first output module 601, the second output module 602, the target area module 603, and the splicing module 604 can be seen in Figure 6 the description in
[0140] Further, the first output module 601 can input the vehicle-mounted radar ranging data into a pre-trained shallow neural network model and output the aspect ratio of the region of interest (ROI) to the original video image frame.
[0141] The second output module 602 can input the aspect ratio into a fitting model and output the display position information of the ROI on the original video image frame.
[0142] The model training module 605 is used to train the shallow neural network model with a training set to obtain a pre-trained shallow neural network model, where the training set includes labeled aspect ratios and training radar ranging data, and the labeled aspect ratio is the aspect ratio of the labeled training ROI to the training video image frame. Among them, the training set can be obtained in the following way: select a set number of data from the collected radar ranging data as the training radar ranging data; according to the principle of continuous frames, label the ROI that is continuous with the previous frame of the training video image frame from the training video image frames to obtain the labeled training ROI; compare the labeled training ROI with the training video image frame to obtain the labeled aspect ratio; save the training radar ranging data and the labeled aspect ratio as the training set.
[0143] The model training module 605 can also perform fitting processing on the labeled display position and the labeled aspect ratio of the labeled training ROI on the training video image frame to obtain a fitting model. For example, input each labeled aspect ratio as a known quantity into a target fitting equation with the target display position as the unknown quantity, and iteratively solve the target display position, where the target fitting equation includes polynomial fitting coefficients; when the deviation between the target display position and the labeled display position of the labeled training ROI on the training video image frame is less than a preset threshold, determine the corresponding polynomial fitting coefficient value as the target fitting coefficient value, and use the target fitting equation determined by the target fitting coefficient value as the fitting model.
[0144] The data collection module 606 is used to collect the vehicle-mounted radar ranging data and the video image frames of the camera and perform preprocessing. In this application, when an object enters the effective range of the ultrasonic radar around the driving vehicle, all the current video image frame data and the vehicle-mounted radar ranging data of the camera can be recorded every set time, such as 5 s, and saved. To facilitate the training of the shallow neural network model for processing and further reduce the computational complexity, the data collection module 606 can also preprocess the vehicle-mounted radar ranging data. The preprocessing includes, for example, data cleaning (such as removing dirty point data, that is, ranging data that obviously does not meet the requirements) and data normalization, etc.; the normalized vehicle-mounted radar ranging data is more intuitive and convenient for researchers.
[0145] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0146] Figure 8 It is a schematic structural diagram of an electronic device shown in an embodiment of the present application.
[0147] See Figure 8 , the electronic device 800 includes a memory 810 and a processor 820.
[0148] The processor 820 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0149] The memory 810 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. Among them, the ROM may store static data or instructions required by the processor 820 or other modules of the computer. The permanent storage device may be a readable and writable storage device. The permanent storage device may be a non-volatile storage device that does not lose the stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device employs a mass storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In some other embodiments, the permanent storage device may be a removable storage device (such as a floppy disk, optical drive). The system memory may be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. The system memory may store some or all of the instructions and data required by the processor during operation. In addition, the memory 810 may include any combination of computer-readable storage media, including various types of semiconductor storage chips (such as DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks may also be used. In some embodiments, the memory 810 may include a removable storage device that is readable and / or writable, such as a compact disc (CD), read-only digital versatile disc (such as DVD-ROM, dual-layer DVD-ROM), read-only Blu-ray disc, super density disc, flash memory card (such as SD card, min SD card, Micro-SD card, etc.), magnetic floppy disk, etc. The computer-readable storage medium does not include carrier waves and instantaneous electronic signals transmitted wirelessly or by wire.
[0150] Executable code is stored on the memory 810, and when the executable code is processed by the processor 820, it may cause the processor 820 to execute some or all of the methods described above.
[0151] In addition, the method according to the present application may also be implemented as a computer program or computer program product, which includes computer program code instructions for executing some or all of the steps in the above method of the present application.
[0152] Alternatively, the present application may also be implemented as a computer-readable storage medium (or non-transitory machine-readable storage medium or machine-readable storage medium), on which executable code (or computer program or computer instruction code) is stored. When the executable code (or computer program or computer instruction code) is executed by a processor of an electronic device (or a server, etc.), it causes the processor to execute some or all of the steps of the above method according to the present application.
[0153] The embodiments of the present application have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A method for video image stitching, characterized in that, it includes: obtaining the aspect ratio of the region of interest (ROI) to the original video image frame according to the vehicle-mounted radar ranging data and a pre-trained shallow neural network model; obtaining the display position information of the ROI on the original video image frame according to the aspect ratio and a fitting model; cropping the target ROI from the original video image frame according to the aspect ratio and the display position information; stitching the cropped target ROIs into a target video image; wherein, the shallow neural network model is trained in the following manner: the shallow neural network model is trained with a training set to obtain the pre-trained shallow neural network model, the training set includes labeled aspect ratios and training radar ranging data, and the labeled aspect ratio is the aspect ratio of the labeled training ROI to the training video image frame; the fitting model is obtained in the following manner: performing a fitting process on the labeled display position of the labeled training ROI on the training video image frame and the labeled aspect ratio to obtain the fitting model; wherein, the training set is obtained in the following manner: selecting a set number of data from the collected radar ranging data as the training radar ranging data; stretching or shrinking a region in the training video image frame, and if the region is continuous with the previous training video image frame, stop the operation and use the obtained region as the labeled training ROI; when the labeled training ROI is a rectangle, the labeled aspect ratio includes the ratio of the labeled length of the labeled training ROI to the frame length of the original video image frame, and the ratio of the labeled width of the labeled training ROI to the frame width of the original video image frame; saving the training radar ranging data and the labeled aspect ratio as the training set.
2. The method according to claim 1, characterized in that: the obtaining the aspect ratio of the region of interest (ROI) to the original video image frame according to the vehicle-mounted radar ranging data and a pre-trained shallow neural network model includes: inputting the vehicle-mounted radar ranging data into the pre-trained shallow neural network model and outputting the aspect ratio of the region of interest (ROI) to the original video image frame; the obtaining the display position information of the ROI on the original video image frame according to the aspect ratio and a fitting model includes: inputting the aspect ratio into the fitting model and outputting the display position information of the ROI on the original video image frame.
3. The method according to claim 1, characterized in that, the performing a fitting process on the labeled display position of the labeled training ROI on the training video image frame and the labeled aspect ratio to obtain the fitting model includes: inputting each of the labeled aspect ratios as a known quantity into a target fitting equation with the target display position as the unknown quantity, and iteratively solving for the target display position, where the target fitting equation includes polynomial fitting coefficients; When the deviation between the target display position and the labeled display position of the labeled training ROI on the training video image frame is less than a preset threshold, determine that the corresponding polynomial fitting coefficient value is the target fitting coefficient value, and use the target fitting equation determined by the target fitting coefficient value as the fitting model.
4. A video image stitching device, characterized in that, it includes: A first output module, configured to obtain the aspect ratio of the region of interest ROI to the original video image frame according to the vehicle-mounted radar ranging data and a pre-trained shallow neural network model; A second output module, configured to obtain the display position information of the ROI on the original video image frame according to the aspect ratio and the fitting model; A target region module, configured to crop the target ROI from the original video image frame according to the aspect ratio obtained by the first output module and the display position information obtained by the second output module; A stitching module, configured to stitch the target ROIs cropped by the target region module into a target video image; wherein, the device further includes: A model training module, configured to train the shallow neural network model with a training set to obtain the pre-trained shallow neural network model, the training set includes labeled aspect ratios and training radar ranging data, the labeled aspect ratio is the aspect ratio of the labeled training ROI to the training video image frame; and, perform fitting processing on the labeled display position of the labeled training ROI on the training video image frame and the labeled aspect ratio to obtain the fitting model; wherein, the training set is obtained in the following manner: select a set number of data from the collected radar ranging data as the training radar ranging data; stretch or shrink a region in the training video image frame, if the region is continuous with the previous frame of the training video image frame, stop the operation, and use the obtained region as the labeled training ROI; when the labeled training ROI is a rectangle, the labeled aspect ratio includes the ratio of the labeled length of the labeled training ROI to the frame length of the original video image frame, and the ratio of the labeled width of the labeled training ROI to the frame width of the original video image frame; save the training radar ranging data and the labeled aspect ratio as the training set.
5. The device according to claim 4, characterized in that: The first output module inputs the vehicle-mounted radar ranging data into the pre-trained shallow neural network model and outputs the aspect ratio of the region of interest ROI to the original video image frame; The second output module inputs the aspect ratio into the fitting model and outputs the display position information of the ROI on the original video image frame.
6. An electronic device, characterized in that, it includes: A processor; and A memory, on which executable code is stored, and when the executable code is executed by the processor, the processor is caused to execute the method according to any one of claims 1 to 3.
7. A computer-readable storage medium having executable code stored thereon, which, when executed by a processor of an electronic device, causes the processor to execute the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Accurate correction processing method for image with rectangular frame
CN113763279A
Providing an image for display
US20100231609A1