Video data display method and device, equipment, medium and product

By parsing and rendering the video files generated by the vehicle surround view camera, a combined display screen containing images and vehicle driving information is generated, which solves the problem of difficulty in understanding spatial relationships when playing back surround view images and improves the interactivity of video data display.

CN121940562APending Publication Date: 2026-04-28AUTOCHIPS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
AUTOCHIPS
Filing Date
2025-12-08
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

The surround view images generated by the vehicle's surround view camera are difficult to understand intuitively when played back, resulting in poor interactive capabilities.

Method used

By parsing the video file to be displayed, the parsed image group and vehicle driving information are obtained, and then rendered to generate the target video file. The target video file contains a combined display of the parsed image group and vehicle driving information.

Benefits of technology

It improves the interactivity of video data display, enhances the information dimension of the target display screen, and enables users to more intuitively understand the spatial relationship around the vehicle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121940562A_ABST
    Figure CN121940562A_ABST
Patent Text Reader

Abstract

The invention discloses a video data display method and device, equipment, a medium and a product, the video data display method is applied to a vehicle, and the video data display method comprises the following steps: in response to a received video display instruction, obtaining a to-be-displayed video file corresponding to the video display instruction; analyzing the to-be-displayed video file to obtain a plurality of analyzed image groups and vehicle driving information matched with the analyzed image groups; the plurality of parsed image groups and the vehicle driving information matched with the parsed image groups are rendered to obtain a target video file, the target video file comprises target display pictures corresponding to the parsed image groups, and the target display pictures represent combinations of the parsed images in the parsed image groups and the corresponding vehicle driving information; and performing display processing on each target display picture in the target video file. According to the scheme, the interaction capability of video data display can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a video data display method, apparatus, device, medium, and product. Background Technology

[0002] The vehicle's surround-view camera can generate and record the vehicle's surround-view images. When playing back, the original video stream is often displayed as a single image or in a grid pattern. It is difficult for users to intuitively understand the spatial relationships and overall environment at the time of the incident, resulting in poor interactivity of the recorded surround-view images.

[0003] Therefore, there is an urgent need for an effective method for displaying video data. Summary of the Invention

[0004] This application provides at least one video data display method, apparatus, device, medium, and product that can improve the interactivity of video data display.

[0005] This application provides a video data display method applied to vehicles. The video data display method includes: in response to receiving a video display instruction, acquiring a video file to be displayed corresponding to the video display instruction, the video file to be displayed including a plurality of image groups to be parsed and vehicle driving information matched by each image group to be parsed, the image groups to be parsed including at least one frame of image to be parsed at the same acquisition time; parsing the video file to be displayed to obtain a plurality of parsed image groups and vehicle driving information matched by each parsed image group; rendering the plurality of parsed image groups and vehicle driving information matched by each parsed image group to obtain a target video file, the target video file including a target display screen corresponding to each parsed image group, the target display screen representing the combination of each parsed image in the parsed image group and the corresponding vehicle driving information; and displaying each target display screen in the target video file.

[0006] This application provides a video data display device, comprising: an acquisition module, a parsing module, a rendering module, and a display module; the acquisition module is used to acquire a video file to be displayed corresponding to a video display instruction received from the instruction, the video file to be displayed including a plurality of image groups to be parsed and vehicle driving information matched by each image group to be parsed, the image groups to be parsed including at least one frame of image to be parsed at the same acquisition time; the parsing module is used to parse the video file to be displayed to obtain a plurality of parsed image groups and vehicle driving information matched by each parsed image group; the rendering module is used to render the plurality of parsed image groups and vehicle driving information matched by each parsed image group to obtain a target video file, the target video file including a target display screen corresponding to each parsed image group, the target display screen representing the combination of each parsed image and the corresponding vehicle driving information in the parsed image group; the display module is used to display each target display screen in the target video file.

[0007] This application provides an electronic device, including a memory and a processor, wherein the processor is used to execute program instructions stored in the memory to implement the above-described video data display method.

[0008] This application provides a computer-readable storage medium storing program instructions thereon, which, when executed by a processor, implement the above-described video data display method.

[0009] This application provides a computer program product that, when executed by a processor, is used to implement the above-described video data display method.

[0010] The above scheme, in response to receiving a video display command, parses the acquired video file to be displayed to obtain several parsed image groups and vehicle driving information matching each parsed image group. It then renders the parsed image groups and their matching vehicle driving information to obtain a target video file. The target video file includes target display frames corresponding to each parsed image group, representing a combination of the parsed images and their corresponding vehicle driving information within each parsed image group. Displaying each target display frame in the target video file allows for the combination of the parsed images and their corresponding vehicle driving information within the target display frames through rendering, thereby increasing the information dimensionality of the target display frames and enhancing the interactivity when displaying video data.

[0011] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description

[0012] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.

[0013] Figure 1 This is a flowchart illustrating an embodiment of the video data display method of this application; Figure 2 yes Figure 1 A schematic diagram of the sub-process of step S13; Figure 3 yes Figure 2 A schematic diagram of the sub-process of step S23; Figure 4a This is a schematic diagram of the effect of the target display screen in one embodiment of the video data display method of this application. Figure 1 ; Figure 4b This is a schematic diagram of the effect of the target display screen in one embodiment of the video data display method of this application. Figure 2 ; Figure 5 This is a schematic diagram of the structure of an embodiment of the video data display device of this application; Figure 6 This is a schematic diagram of the structure of an embodiment of the electronic device of this application; Figure 7 This is a schematic diagram of the structure of an embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0014] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0015] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.

[0016] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this document means two or more. Moreover, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of objects. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0017] In this application, the entity executing the video data display method described herein can be a video data display device. For example, the video data display device can be a vehicle. The video data display device can be located in a terminal device, server, or other processing device. The terminal device can be a carrier device, mobile robot, user equipment (UE), user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, vehicle-mounted device, wearable device, etc. In some possible implementations, the video data display method can be implemented by a processor calling computer-readable instructions stored in memory.

[0018] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the video data display method of this application. Figure 1 As shown, the video data display method is applied to a vehicle. The video data display method provided in this embodiment may include the following steps: Step S11: In response to receiving a video display instruction, obtain the video file to be displayed corresponding to the video display instruction.

[0019] Video display commands are used to instruct the display of the corresponding video file. These commands can be generated in response to a user clicking a button on the video file to be displayed, or in response to the arrival of a fixed playback frequency or the triggering of a specific display event. For example, a fixed playback frequency could be the playback of video files acquired over a fixed period of time at a preset hourly interval, or the playback of video files stored in a preset database at a preset frequency. A specific display event could be the display of the corresponding video file when the video file acquired by the image acquisition device involves an accident or dispute (such as a vehicle collision or scrape) or an everyday accident (such as an abnormal road condition).

[0020] The video file to be displayed can be a file stored in a preset database, or it can be a video file captured in real time by at least one image acquisition device and bound to vehicle driving information.

[0021] The video file to be displayed includes several groups of images to be parsed and vehicle driving information matching each group. Each group of images to be parsed includes at least one frame of image to be parsed at the same acquisition time. All images in each group are acquired at the same acquisition time. Each image in each group is an image acquired at the same acquisition time by image acquisition devices located at different parts of the vehicle. For example, each image in each group might be an image acquired by a surround-view camera (four-channel camera, six-channel camera, etc.) in the vehicle.

[0022] Vehicle driving information refers to the driving information collected by vehicle sensors at a fixed sampling frequency. For example, vehicle driving information includes at least one of the following: motion state information and spatiotemporal state information. Motion state information represents state information related to vehicle motion, such as vehicle speed, steering angle, etc. Spatiotemporal state information represents state information related to the vehicle's location, such as the vehicle's latitude and longitude, and the vehicle's structured address information, etc.

[0023] Specifically, the vehicle driving information matched by the image group to be parsed is obtained based on the correspondence between the acquisition timestamps of the image group to be parsed and the acquisition timestamps of the vehicle driving information.

[0024] Specifically, the above-mentioned acquisition of the video file to be displayed corresponding to the video display instruction includes: a preset database containing several files to be displayed within a certain acquisition time period; retrieving the file to be displayed within the acquisition time period corresponding to the video display instruction from the preset database; or, acquiring several frames of images within the acquisition time period corresponding to the video display instruction in real time, and using each acquired image as an image to be parsed; receiving several vehicle driving information data within the acquisition time period corresponding to the video display instruction in real time; dividing the several images to be parsed into at least one group of images to be parsed according to the acquisition time of each image to be parsed; and binding the image groups to be parsed at the same acquisition time with the vehicle driving information to obtain the vehicle driving information matched by each image group to be parsed.

[0025] Step S12: The video file to be displayed is parsed to obtain several parsed image groups and vehicle driving information matched by each parsed image group.

[0026] The parsed image group corresponds to the image group to be parsed, representing the image group to be parsed after parsing processing. Each parsed image group includes parsed images corresponding to each image to be parsed. The vehicle driving information matched by the parsed image group corresponds to the vehicle driving information matched by the image group to be parsed, representing the vehicle driving information matched by the parsed image group after parsing processing. The parsing processing can involve parsing each image group to be parsed and the vehicle driving information matched by each image group in the video file to be displayed.

[0027] Step S13: Render the several parsed image groups and the vehicle driving information matched by each parsed image group to obtain the target video file.

[0028] The target video file includes target display frames corresponding to each group of parsed images. Each target display frame represents a combination of the parsed images in the parsed image group and their corresponding vehicle driving information. The timestamp of each target display frame corresponds to the acquisition timestamp of the parsed image group and / or the acquisition timestamp of the corresponding vehicle driving information, or the indices of each target display frame are arranged in chronological order according to the acquisition timestamps of the parsed image group and / or the corresponding vehicle driving information. Rendering processing is used to draw each parsed image and its corresponding vehicle driving information from the parsed image group at the same moment onto a preset display frame to obtain the target display frame corresponding to that moment.

[0029] In some application scenarios, step S13 above may involve performing the following steps for each parsed image group: filling the text filling area in a preset display screen with the vehicle driving information corresponding to the parsed image group to obtain a text filling result, including: the vehicle driving information includes motion state information and / or spatiotemporal state information, and filling at least a portion of the vehicle driving information sequentially into the text filling area in the preset display screen to obtain a text filling result. Filling each parsed image in the parsed image group into the image filling area in the preset display screen to obtain an image filling result includes: arranging each parsed image sequentially according to the orientation of the image acquisition device that acquired each parsed image to obtain an arranged parsed image, and filling the arranged parsed image into the image filling area in the preset display screen to obtain an image filling result; or, performing two-dimensional modeling processing and / or three-dimensional modeling processing on each parsed image to obtain an image modeling result, and filling the image modeling result into the image filling area in the preset display screen to obtain an image filling result. Determining the target display screen corresponding to the parsed image group based on the text fill result and the image fill result includes: superimposing the text fill result and the image fill result to obtain the target display screen, for example, superimposing the text fill result to a preset position in the previous layer of the image fill result, the preset position can be the upper right corner, the lower right corner, etc.; or, splicing the text fill result and the image fill result to obtain the target display screen, for example, arranging the text fill result in the surrounding area of ​​the image fill result, the surrounding area can be the right side or the top side.

[0030] Step S14: Display each target image in the target video file.

[0031] Specifically, step S14 above may involve sending the target video file to the vehicle's display module so that the vehicle's display module can process the display of each target screen in the target video file. The display module is used to display each target screen.

[0032] The display screen in the vehicle plays the images of each target in the target video file in chronological order of their timestamps.

[0033] The above scheme, in response to receiving a video display command, parses the acquired video file to be displayed to obtain several parsed image groups and vehicle driving information matching each parsed image group. It then renders the parsed image groups and their matching vehicle driving information to obtain a target video file. The target video file includes target display frames corresponding to each parsed image group, representing a combination of the parsed images and their corresponding vehicle driving information within each parsed image group. Displaying each target display frame in the target video file allows for the combination of the parsed images and their corresponding vehicle driving information within the target display frames through rendering, thereby increasing the information dimensionality of the target display frames and enhancing the interactivity when displaying video data.

[0034] Please see Figure 2 , Figure 2 yes Figure 1 A schematic diagram of the sub-process of step S13.

[0035] In some embodiments, the vehicle driving information matched by the parsed image group includes the motion state information matched by the parsed image group. Step S13 may include the following steps: Step S21: Obtain the initial vehicle-mounted model for vehicle matching. Step S22: Load the motion state information matched by the parsed image group into the initial vehicle-mounted model to obtain the target vehicle-mounted model corresponding to the parsed image group. Step S23: Determine the target display screen corresponding to each parsed image group based on each parsed image group and the target vehicle-mounted model corresponding to each parsed image group.

[0036] It is understood that step S13 above can be performed sequentially for each parsed image group from step S21 to step S23.

[0037] The initial vehicle model represents the virtual model / animation model to which the vehicle belongs. Specifically, step S21 above may be: obtaining the vehicle type to which the vehicle belongs, and using the preset vehicle model that matches the vehicle type from a number of preset vehicle models as the initial vehicle model for matching the vehicle.

[0038] The initial vehicle-mounted model contains parameters to be configured, which are used to simulate the vehicle's driving state. Step S22 can be used to configure the parameters to be configured in the initial vehicle-mounted model using the motion state information matched by the parsed image group to obtain the target vehicle-mounted model. The target vehicle-mounted model can perform motion simulation of the vehicle based on the motion state information matched by the parsed image group. Step S22 can include the following steps: the motion state information matched by the parsed image group includes the vehicle speed matched by the parsed image group; the parameters to be configured in the initial vehicle-mounted model include wheel configuration parameters; the wheel configuration parameters in the initial vehicle-mounted model are configured using the vehicle speed matched by the parsed image group to obtain the target vehicle-mounted model. And / or, the motion state information matched by the parsed image group includes the vehicle steering angle matched by the parsed image group; the parameters to be configured in the initial vehicle-mounted model include orientation configuration parameters; the orientation configuration parameters in the initial vehicle-mounted model are configured using the steering angle matched by the parsed image group to obtain the target vehicle-mounted model. It can be understood that the vehicle speed is used to configure the rotation state of each wheel in the initial vehicle-mounted model. The vehicle steering angle is used to configure the orientation of the initial vehicle-mounted model. It is understandable that the simulation results of the target vehicle model configured using the motion state information will differ depending on the motion state information.

[0039] In some application scenarios, step S23 above can be performed for each parsed image group as follows: Obtain the image filling result corresponding to each parsed image in the parsed image group. Determine the target display screen corresponding to the parsed image group based on the target vehicle model and the image filling result, including: superimposing the target vehicle model and the image filling result to obtain the target display screen, for example, superimposing the target vehicle model to a preset position in the previous layer of the image filling result, the preset position could be the upper right corner, lower right corner, etc.; or, stitching the target vehicle model and the image filling result together to obtain the target display screen, for example, arranging the target vehicle model in the surrounding area of ​​the image filling result, the surrounding area could be the right side, the top side. For example, arranging the target vehicle model in the middle area of ​​the image filling result, that is, placing the target vehicle model in the center position of the image filling result. Alternatively, filling the text filling area in the preset display screen with other information from the driving state information corresponding to the parsed image group, excluding motion state information, to obtain the text filling result. Determining the target display screen corresponding to the parsed image group based on the target vehicle model, text fill result, and image fill result includes: overlaying the target vehicle model, text fill result, and image fill result to obtain the target display screen, for example, overlaying the target vehicle model and text fill result to a preset position in the layer above the image fill result, the preset position could be the upper right corner, lower right corner, etc.; or, stitching the target vehicle model, text fill result, and image fill result together to obtain the target display screen. For example, arranging the text fill result in the surrounding area of ​​the image fill result, the surrounding area could be the right side or the top side. For example, arranging the target vehicle model in the middle area of ​​the image fill result, that is, placing the target vehicle model in the center position of the image fill result.

[0040] It can be argued that by combining the target display screen determined by the parsed image group and the target vehicle model corresponding to the parsed image group with the motion state information of the vehicle at the acquisition time of each parsed image in the parsed image group, the realism of the target display screen can be improved.

[0041] Please see Figure 3 , Figure 3 yes Figure 2 A schematic diagram of the sub-process of step S23.

[0042] In some embodiments, the parsed image group includes at least one parsed image frame acquired at the same time. Step S23 above may include the following steps: Step S31: Perform preset image processing on each of the parsed images in the parsed image group to obtain the target image corresponding to the parsed image group.

[0043] Preset image processing can optimize each parsed image to obtain a target image of higher quality compared to the original parsed image. Each group of parsed images corresponds to one frame of target image. The target image represents each parsed image after preset image processing.

[0044] In some application scenarios, step S31 above may involve performing image denoising and / or image correction processing on each of the parsed images in the parsed image group to obtain the image to be fused corresponding to each parsed image; and performing fusion processing on the image to be fused corresponding to each parsed image to obtain the target image corresponding to the parsed image group.

[0045] In some embodiments, step S31 may include the following steps: performing brightness adjustment processing on each of the parsed images in the parsed image group to obtain an adjusted image corresponding to each parsed image; and performing fusion processing on the adjusted images corresponding to each parsed image to obtain a target image corresponding to the parsed image group.

[0046] The adjusted image represents the resolved image after brightness adjustment processing. Brightness adjustment processing is used to ensure brightness balance between the resolved images from adjacent image acquisition devices.

[0047] In some application scenarios, the above steps of performing brightness adjustment processing on each of the parsed images in the parsed image group to obtain the adjusted image corresponding to each parsed image include: obtaining the average pixel value of each parsed image; for each parsed image, taking the product between the average pixel value of all regions in the parsed image or the average pixel value of the edge regions in the parsed image and the pixel value at each position in the parsed image as the updated pixel value at each position in the parsed image; and determining the adjusted image corresponding to the parsed image based on the updated pixel value at each position in the parsed image.

[0048] For example, the brightness adjustment processing performed on the parsed image in this application can be the execution of a brightness equalization algorithm. The brightness equalization algorithm can be local brightness equalization (e.g., brightness equalization of the edge region of the parsed image) or global brightness equalization. Taking global brightness equalization as an example, before implementing the fusion processing of each parsed image, the average brightness (or histogram) of the parsed image captured by each camera is calculated first, and then adjusted using a global gain factor. The specific brightness adjustment processing performed on the parsed image can refer to the following formula (1): Formula (1); Where αi represents the average pixel value of all regions in the parsed image of frame i, or the average pixel value of the edge regions in the parsed image of frame i. Ii(x) represents the pixel value at the x-th position of all regions or edge regions in the parsed image of frame i. I'i(x) represents the updated pixel value at the x-th position of all regions or edge regions in the parsed image of frame i. It represents multiplication.

[0049] In some application scenarios, the above steps of fusing the adjusted images corresponding to each parsed image to obtain the target image corresponding to the parsed image group specifically include performing the following steps for each parsed image group: directly using the stitching result of the adjusted images corresponding to each parsed image in the parsed image group as the target image corresponding to the parsed image group.

[0050] In some embodiments, the step of fusing the adjusted images corresponding to each parsed image to obtain the target image corresponding to the parsed image group may include the following steps: stitching the adjusted images corresponding to each parsed image to obtain an initial image corresponding to the parsed image group; responding to the existence of at least one overlapping region in the initial image corresponding to the parsed image group, smoothing each overlapping region in the initial image to obtain a smoothed image corresponding to each overlapping region; and determining the target image corresponding to the parsed image group based on the other regions in the initial image besides the overlapping regions and the smoothed images corresponding to each overlapping region.

[0051] The initial image represents the stitching result of the adjusted images corresponding to each resolved image in each group of resolved images. The initial image can be an image with at least one overlapping region.

[0052] In some application scenarios, the above-mentioned step of smoothing each overlapping region in the initial image to obtain the smoothed image corresponding to each overlapping region includes: dividing the adjusted images corresponding to the parsed images of adjacent image acquisition devices into a group of adjacent image groups. Each group of adjacent image groups includes the adjusted images of two frames of parsed images. For each overlapping region in the initial image, the following steps are performed: the adjacent image group to which the overlapping region belongs is taken as the target adjacent image group. Wherein, the overlapping region is the image region adjacent to the field of view of each adjusted image of the target adjacent image group. For example, each adjusted image of the adjacent image group is a first adjusted image associated with the front-view camera and a second adjusted image associated with the right-view camera in the surround-view camera; the overlapping region in the first adjusted image and the overlapping region in the second adjusted image are the overlapping regions in the initial image; the overlapping region in the first adjusted image is the right image region in the first adjusted image, and the overlapping region in the second adjusted image is the left image region in the second adjusted image. Determining the updated pixel values ​​of each position in the overlapping region of the initial image based on the pixel values ​​of each adjusted image in the target adjacent image group includes: obtaining the target pixel values ​​of the overlapping region or the entire image region in any adjusted image of the adjacent image group; using the product of the target pixel values ​​corresponding to each adjusted image and the weights of each position in each adjusted image as the updated pixel values ​​of each position; and determining the smoothed image corresponding to the overlapping region based on the updated pixel values ​​of each position. Alternatively, for any pixel position in the overlapping region, obtaining the initial pixel values ​​of each adjusted image in the target adjacent image group belonging to that pixel position; performing a weighted summation of the initial pixel values ​​of that pixel position to obtain the updated pixel value of that pixel position, wherein the weights of the weighted summation are determined based on the position of the pixel position in each adjusted image; and determining the smoothed image corresponding to the overlapping region based on the updated pixel values ​​of each pixel position in the overlapping region.

[0053] For example, during the fusion process of each parsed image, there is an overlapping area between the adjusted images corresponding to the parsed images of adjacent cameras. If smoothing is not performed, obvious "seams" may appear in the overlapping area. The core idea of ​​the fusion technology is to achieve a smooth transition between the adjusted images through weighted fusion. Linear weighted fusion is performed on the two adjacent adjusted images to which the overlapping area belongs to obtain the smoothed image corresponding to the overlapping area. Specifically, the method for determining the smoothed image corresponding to the overlapping area can refer to the following formula (2): Formula (2); Here, I1(x) and I2(x) can represent the pixel values ​​of the adjusted images from two adjacent cameras, respectively. w(x) represents the weight, which varies with the position of the pixel in the overlapping region; for example, it is close to 1 when the pixel is on the left side of the overlapping region and close to 0 when it is on the right side. I(x) represents the updated pixel value at each position in the smoothed image corresponding to the overlapping region. It represents multiplication.

[0054] In some application scenarios, the step of determining the target image corresponding to the parsed image group based on the regions other than the overlapping regions in the initial image and the smoothed images corresponding to the overlapping regions includes: replacing the images corresponding to the overlapping regions in the initial image with the smoothed images corresponding to the overlapping regions to obtain the target image corresponding to the initial image. Alternatively, the images of the regions other than the overlapping regions in the initial image and the smoothed images corresponding to the overlapping regions are stitched together to obtain the target image corresponding to the initial image.

[0055] In some embodiments, the step of stitching together the adjusted images corresponding to each resolved image to obtain the initial image corresponding to the resolved image group may include the following steps: performing two-dimensional perspective transformation on the adjusted images corresponding to each resolved image to obtain a two-dimensional transformed image corresponding to each resolved image; and / or performing three-dimensional perspective transformation on the adjusted images corresponding to each resolved image to obtain a three-dimensional transformed image corresponding to each resolved image; and performing image stitching on the two-dimensional transformed images corresponding to each resolved image and / or the three-dimensional transformed images corresponding to each resolved image to obtain the initial image corresponding to the resolved image group.

[0056] The image after 2D transformation represents the resolved image after 2D perspective transformation. 2D perspective transformation is used to map the adjusted image corresponding to the resolved image to the 2D transformed image according to a preset 2D perspective transformation relationship. The image after 3D transformation represents the resolved image after 3D perspective transformation. 3D perspective transformation is used to map the adjusted image corresponding to the resolved image to the 3D transformed image according to a preset 3D perspective transformation relationship.

[0057] In some application scenarios, two-dimensional perspective transformation is performed on the adjusted images corresponding to each parsed image to obtain the two-dimensional transformed images corresponding to each parsed image; image stitching is then performed on the two-dimensional transformed images corresponding to each parsed image to obtain the initial image corresponding to the parsed image group.

[0058] In other application scenarios, the adjusted images corresponding to each parsed image are subjected to 3D perspective transformation to obtain the 3D transformed images corresponding to each parsed image; the 3D transformed images corresponding to each parsed image are then stitched together to obtain the initial image corresponding to the parsed image group.

[0059] In other application scenarios, the initial images corresponding to the parsed image group include a 2D stitched image obtained through 2D modeling and a 3D stitched image obtained through 3D modeling. The adjusted images corresponding to each parsed image undergo 2D perspective transformation to obtain the 2D transformed images corresponding to each parsed image. The 2D transformed images corresponding to each parsed image are then stitched together to obtain a 2D stitched image. Furthermore, the adjusted images corresponding to each parsed image undergo 3D perspective transformation to obtain the 3D transformed images corresponding to each parsed image. Finally, the 3D transformed images corresponding to each parsed image are stitched together to obtain a 3D stitched image.

[0060] Step S32: Combine the target image corresponding to the parsed image group with the target vehicle model corresponding to the parsed image group to obtain the initial display screen corresponding to the parsed image group.

[0061] The initial display screen represents the combined result of the target image corresponding to the parsed image group and the target vehicle model corresponding to the parsed image group.

[0062] In some application scenarios, step S32 above can be performed for each parsed image group as follows: Obtain the target image corresponding to each parsed image in the parsed image group. Determine the target display screen corresponding to the parsed image group based on the target vehicle model and the target image, including: superimposing the target vehicle model and the target image to obtain the target display screen, for example, superimposing the target vehicle model to a preset position in the previous layer of the target image, the preset position could be the upper right corner, lower right corner, etc.; or, stitching the target vehicle model and the target image to obtain the target display screen, for example, arranging the target vehicle model in the surrounding area of ​​the target image, the surrounding area could be the right side, the top side. For example, arranging the target vehicle model in the middle area of ​​the target image, that is, placing the target vehicle model in the center of the target image.

[0063] Step S33: Determine the target display screen corresponding to the parsed image group based on the initial display screen corresponding to the parsed image group.

[0064] In some application scenarios, the initial display screen corresponding to the parsed image group is directly used as the target display screen corresponding to the parsed image group.

[0065] In some embodiments, the vehicle driving information matched by the parsed image group also includes the spatiotemporal state information matched by the parsed image group. Step S33 may include the following steps: filling the spatiotemporal state information matched by the parsed image group into a preset display area to obtain the text filling result corresponding to the parsed image group; and overlaying the text filling result corresponding to the parsed image group with the initial display screen corresponding to the parsed image group to obtain the target display screen corresponding to the parsed image group.

[0066] The preset display area represents a preset position in the top layer of the initial display screen. The preset position can be the upper right corner, the lower right corner, etc., or the preset display area represents a display area that is independent of the drawing layer, such as the area around the drawing layer (located to the right or upper layer of the display area outside the drawing layer).

[0067] In some application scenarios, the above-mentioned process of overlaying the text fill result corresponding to the parsed image group with the initial display screen corresponding to the parsed image group to obtain the target display screen corresponding to the parsed image group includes: determining the target display screen corresponding to the parsed image group based on the text fill result and the initial display screen, including: overlaying the text fill result and the initial display screen as layers to obtain the target display screen, for example, overlaying the text fill result to a preset position in the previous layer of the initial display screen, the preset position can be the upper right corner, the lower right corner, etc.; or, splicing the text fill result and the initial display screen to obtain the target display screen, for example, arranging the text fill result in the surrounding area of ​​the initial display screen, the surrounding area can be the right side or the top side.

[0068] For example, Figure 4a The "H" in the image represents the target vehicle model. The target display screen includes both two-dimensional and three-dimensional displays. Figure 4b In this context, K1 represents a two-dimensional display image, and K2 represents a three-dimensional display image. The target image corresponding to the parsed image group includes both two-dimensional and three-dimensional target images.

[0069] In some application scenarios, in response to the existence of at least one overlapping region in a 2D stitched image, smoothing is performed on each overlapping region to obtain a smoothed image corresponding to each overlapping region. Based on the other regions in the 2D stitched image besides the overlapping regions and the smoothed images corresponding to each overlapping region, the 2D target image corresponding to the parsed image group is determined. The 2D target image corresponding to the parsed image group is combined with the target vehicle model corresponding to the parsed image group to obtain the initial 2D display screen corresponding to the parsed image group. Based on the initial 2D display screen corresponding to the parsed image group, the 2D target display screen corresponding to the parsed image group is determined. For example, the 2D target display screen can be a panoramic 2D view, where, as...Figure 4b The 2D target display in K1 on the left is a vehicle-centric "bird's-eye view" panoramic view, highlighting the spatial relationships around the vehicle and the positions of obstacles in its environment at any given moment of image acquisition. The panoramic 2D view is achieved based on surround-view calibration and stitching algorithms. The system uses the vehicle's center as the origin, geometrically corrects and coordinates the images captured by the vehicle's four cameras, and then performs pre-defined fusion processing on overlapping areas using a stitching algorithm to ultimately generate a complete panoramic bird's-eye view image (i.e., the 2D target display). This pre-defined fusion processing includes weighted averaging of overlapping areas, brightness equalization, and edge smoothing to ensure that the stitched 2D target display has no obvious boundaries.

[0070] In other application scenarios, in response to the presence of at least one overlapping region in the 3D stitched image, smoothing is performed on each overlapping region to obtain a smoothed image corresponding to each overlapping region. Based on the other regions in the 3D stitched image besides the overlapping regions and the smoothed images corresponding to each overlapping region, the 3D target image corresponding to the parsed image group is determined. The 3D target image corresponding to the parsed image group is combined with the target vehicle model corresponding to the parsed image group to obtain the initial 3D display screen corresponding to the parsed image group. Based on the initial 3D display screen corresponding to the parsed image group, the 3D target display screen corresponding to the parsed image group is determined. For example, the 3D target display screen can be a panoramic 3D view, such as... Figure 4b The 3D target display in K2 on the right maps multi-camera textures onto a curved surface (such as a bowl-shaped model) in a virtual 3D space, generating a panoramic view with a sense of space from a virtual viewpoint. The target vehicle model in the 3D target display can be overlaid with the 3D vehicle model and body color within the 3D scene. The panoramic 3D view is not a real 3D scene, but is achieved by constructing a bowl-shaped 3D curved surface model. With the vehicle's center as the coordinate origin, for example, images from the front, rear, left, and right cameras of the vehicle are mapped onto the bowl wall and bottom of the bowl model, respectively. The overlapping areas use texture blending and brightness equalization algorithms. The video data display device uses OpenGL for rendering, calculates the surface vertices and texture coordinates through calibration parameters, and sets the virtual viewpoint to generate the final view. In terms of effect, the 2D view has a larger bowl bottom area, emphasizing a top-down perspective; the 3D view has a smaller bowl bottom, emphasizing a three-dimensional immersive experience.

[0071] In some application scenarios, the OpenGL rendering process can involve applying the parsed images from four vehicle cameras as textures onto a 3D mesh. The GPU's shader program calculates the blending weights in overlapping areas, and during rendering, the blending weight w(x) of each pixel in formula (2) above is dynamically calculated to achieve a smooth transition. Brightness gain compensation is applied to each pixel in the rendering pipeline to maintain consistent overall brightness.

[0072] During the reconstruction of the vehicle's driving scene, the video display device allows users to flexibly switch between various view modes and camera angles according to their needs. Specifically, the video display device responds to the user's view switching operation by triggering a view switching command. Whether you want to view panoramic information of the vehicle's surrounding environment (such as...) Figure 4b Both panoramic 2D and panoramic 3D views exist, or the focus is on a specific view (such as only viewing...). Figure 4a The system allows users to freely manipulate the target video file during playback (either displaying a single target image or a detailed image in a specific direction) to achieve precise observation of key areas and complete reconstruction of events.

[0073] In some application scenarios, vehicles acquire images at a fixed frequency and store the acquired images in a preset database. Specifically, the video display device includes a video encoding / decoding module, a recording component / recording service (mediacorder), and real-time acquisition of image frames (YUV format) by various cameras on the vehicle. The recording service is initialized, and the image frames acquired by each camera are distributed to the recording service through the corresponding transmission channel of each camera. For example, image frames are written to the header with "0x01" as the first data type, and image frames with the same timestamp are written after the header "0x01". The recording service receives vehicle driving information sent by the vehicle sensors, such as vehicle speed, vehicle steering angle, and current gear mode (e.g., P for parking, R for reversing, N for neutral, D for driving). For example, vehicle driving information is written to the header with "0x02" as the second data type, and the specific vehicle driving information is written after the header "0x02". The recording service creates two threads: thread 1: encorderThread, and thread 2: writeFileThread. The `encorderThread` is responsible for encoding frames into a compressed stream by the video codec module (MediaCodec or hardware H.264 encoder). The `writeFileThread` is responsible for writing the encoded data to local storage / a preset database (MP4 or TS format). In some application scenarios, when previewing stops, both `encorderThread` and `writeFileThread` are stopped to release video codec module resources and stop encoding. For example, a recording service records video files in three-minute increments and stores them in a preset database.

[0074] For example, when a user triggers the playback function (e.g., clicks on reversing recording), the video file for a specified time period is read through the video display device. Step S12 above is executed to decode the recorded (H.264 → YUV) video to obtain a parsed file. The parsed file includes each parsed image group and the corresponding vehicle driving information. The parsed file is played back on the display screen corresponding to the video display device, which provides a visual UI for user configuration options (e.g., configuration options related to viewpoint switching). The decoded frames are sent to the image processing module and displayed on the UI in a simulated form to obtain the target image corresponding to each parsed image group. The vehicle driving information corresponding to each parsed image group is also sent to the algorithm for processing, allowing the algorithm to perform rendering configuration on the vehicle's rotation speed, orientation, and headlight angles to obtain the target vehicle model and text filling results.

[0075] In some application scenarios, the various cameras on the vehicle generate raw image frames. These image frames are enqueued and placed into the buffer queue of the image processing module. The image processing module then passes the received frames to its internal algorithm processing module (AVM Algo) for processing (such as image optimization) to obtain the target images corresponding to each parsed image group. After processing, the target images corresponding to each parsed image group are dequeued and returned to the image processing module, which then sends each target image to the Preview Controller. The Preview Controller calls the target method (fillSurfaceBuffer method) to notify the vehicle's Hardware Abstraction Layer (HAL layer) to fill a surface buffer, which will be used for subsequent display. After executing the target method, the result is returned to the AVM Server, triggering a target callback (the notifyAlgoEvent callback event) to inform the algorithm layer that the task is complete. The AVM Server is a partition service process in the video data display module / video data display system. It acts as an intermediary layer connecting the upper-layer application and the lower-layer hardware driver, responsible for managing the buffer allocation and scheduling of the video stream and communicating with the HAL layer. The AVM Server calls `dequeueBuffer` to request a free display buffer from the Surface Flinger service. The Surface Flinger service allocates a handle to the buffer and returns it. The AVM Server uses this handle to store the processed image data, i.e., the target images corresponding to each group of images to be parsed. The AVM Server calls `queueBuffer` to submit the processed target images to the Surface Flinger service. Upon receiving the images, the Surface Flinger service composites them onto the screen to obtain the target display image and displays it. In other application scenarios, after completing an operation at the HAL layer, it will notify the upstream (such as the AVM Server or the preview controller) via the feedback message `notifyFillBufferDone`, indicating that the buffer is full and available for use.

[0076] For example, the image processing module described above processes each parsed image at each timestamp. When the image processing module is in the working state (STREAM_RUNNING state), it uses a data receiving thread (streamThread) to repeatedly call the parsed image streamFrame at each timestamp to produce panoramic data, i.e., the target image. Specifically, it retrieves a buffer of the surrounding environment from the camera stream, preparing to receive new frames. This is executed by the data receiving thread (streamThread) controlled by the image processing module. The raw camera data is copied or mapped to an EGL buffer (GPU-accessible memory) to obtain an empty output buffer for storing the processed results. The EGL buffer uses glTexImage2D or EGLImageKHR + glEGLImageTargetTexture2DOES technology for subsequent GPU-accelerated processing. The EGL buffer is bound to a specific texture or buffer (the intermediate Foo buffer) for use by subsequent algorithms. Internal GPU-related components perform image stitching and image processing algorithms (such as panorama stitching, HDR, stitching, geometric correction, and brightness blending) to obtain the target image corresponding to each parsed image group. The algorithm processing results are rendered onto the bound Foo buffer. The binding to the intermediate Foo buffer is then released, freeing up resources. The image processing module calls the EGL API to create a synchronization fence to ensure that the GPU completes its processing before proceeding to the next step. This prevents the CPU from prematurely submitting the buffer, which could cause image corruption. The processed buffer is placed in the output queue, and a semaphore notifies downstream modules that it can be retrieved and displayed on the screen.

[0077] In some application scenarios, binding an EGL buffer to a specific texture or buffer specifically includes binding an empty EGL output buffer to the current OpenGL FBO (Frame Buffer Object), so that subsequent image processing operations will be drawn directly to this output buffer. An FBO (Frame Buffer Object) is a logical object in OpenGL ES used to specify rendering attachments. The FBO itself does not directly store pixels; instead, it carries pixels by attaching color / depth / stencil attachments (such as Texture or Renderbuffer). In the GPU rendering pipeline, developers bind FBOs to direct rendering output to off-screen buffers.

[0078] The above scheme, in response to receiving a video display command, parses the acquired video file to be displayed to obtain several parsed image groups and vehicle driving information matching each parsed image group. It then renders the parsed image groups and their matching vehicle driving information to obtain a target video file. The target video file includes target display frames corresponding to each parsed image group, representing a combination of the parsed images and their corresponding vehicle driving information within each parsed image group. Displaying each target display frame in the target video file allows for the combination of the parsed images and their corresponding vehicle driving information within the target display frames through rendering, thereby increasing the information dimensionality of the target display frames and enhancing the interactivity when displaying video data.

[0079] Please see Figure 5 , Figure 5 This is a schematic diagram of an embodiment of the video data display device of this application. The video data display device 50 includes an acquisition module 51, a parsing module 52, a rendering module 53, and a display module 54. The acquisition module 51 is used to acquire a video file to be displayed corresponding to a video display command in response to receiving the video display command. The video file to be displayed includes a plurality of image groups to be parsed and vehicle driving information matched by each image group to be parsed. Each image group to be parsed includes at least one frame of image to be parsed at the same acquisition time. The parsing module 52 is used to parse the video file to be displayed to obtain a plurality of parsed image groups and vehicle driving information matched by each parsed image group. The rendering module 53 is used to render the plurality of parsed image groups and vehicle driving information matched by each parsed image group to obtain a target video file. The target video file includes a target display screen corresponding to each parsed image group. The target display screen represents the combination of each parsed image in the parsed image group and the corresponding vehicle driving information. The display module 54 is used to display each target display screen in the target video file.

[0080] In some embodiments, the vehicle driving information matched by the parsed image group includes the motion state information matched by the parsed image group; the rendering module 53 is used to perform rendering processing on several parsed image groups and the vehicle driving information matched by each parsed image group, including: obtaining the initial vehicle model matched by the vehicle; loading the motion state information matched by the parsed image group into the initial vehicle model to obtain the target vehicle model corresponding to the parsed image group; and determining the target display screen corresponding to each parsed image group based on each parsed image group and the target vehicle model corresponding to each parsed image group.

[0081] In some embodiments, the parsed image group includes at least one parsed image frame at the same acquisition time; the rendering module 53 is used to determine the target display screen corresponding to each parsed image group based on each parsed image group and the target vehicle model corresponding to each parsed image group, including: performing preset image processing on each parsed image in the parsed image group to obtain the target image corresponding to the parsed image group; combining the target image corresponding to the parsed image group with the target vehicle model corresponding to the parsed image group to obtain the initial display screen corresponding to the parsed image group; and determining the target display screen corresponding to the parsed image group based on the initial display screen corresponding to the parsed image group.

[0082] In some embodiments, the rendering module 53 is used to perform preset image processing on each of the parsed images in the parsed image group to obtain a target image corresponding to the parsed image group, including: performing brightness adjustment processing on each of the parsed images in the parsed image group to obtain an adjusted image corresponding to each parsed image; and performing fusion processing on the adjusted images corresponding to each parsed image to obtain a target image corresponding to the parsed image group.

[0083] In some embodiments, the rendering module 53 is configured to perform a fusion process on the adjusted images corresponding to each parsed image to obtain a target image corresponding to the parsed image group, including: performing a stitching process on the adjusted images corresponding to each parsed image to obtain an initial image corresponding to the parsed image group; responding to the existence of at least one overlapping region in the initial image corresponding to the parsed image group, performing a smoothing process on each overlapping region in the initial image to obtain a smoothed image corresponding to each overlapping region; and determining the target image corresponding to the parsed image group based on other regions in the initial image besides each overlapping region and the smoothed images corresponding to each overlapping region.

[0084] In some embodiments, the rendering module 53 is used to perform stitching processing on the adjusted images corresponding to each parsed image to obtain an initial image corresponding to the parsed image group, including: performing two-dimensional perspective transformation processing on the adjusted images corresponding to each parsed image to obtain two-dimensional transformed images corresponding to each parsed image; and / or performing three-dimensional perspective transformation processing on the adjusted images corresponding to each parsed image to obtain three-dimensional transformed images corresponding to each parsed image; and performing image stitching processing on the two-dimensional transformed images corresponding to each parsed image and / or the three-dimensional transformed images corresponding to each parsed image to obtain an initial image corresponding to the parsed image group.

[0085] In some embodiments, the vehicle driving information matched by the parsed image group also includes the spatiotemporal state information matched by the parsed image group; the rendering module 53 is used to determine the target display screen corresponding to the parsed image group based on the initial display screen corresponding to the parsed image group, including: filling the spatiotemporal state information matched by the parsed image group into a preset display area to obtain the text filling result corresponding to the parsed image group; and overlaying the text filling result corresponding to the parsed image group with the initial display screen corresponding to the parsed image group to obtain the target display screen corresponding to the parsed image group.

[0086] The above scheme, in response to receiving a video display command, parses the acquired video file to be displayed to obtain several parsed image groups and vehicle driving information matching each parsed image group. It then renders the parsed image groups and their matching vehicle driving information to obtain a target video file. The target video file includes target display frames corresponding to each parsed image group, representing a combination of the parsed images and their corresponding vehicle driving information within each parsed image group. Displaying each target display frame in the target video file allows for the combination of the parsed images and their corresponding vehicle driving information within the target display frames through rendering, thereby increasing the information dimensionality of the target display frames and enhancing the interactivity when displaying video data.

[0087] The functions of each module can be found in the video data display method embodiment, and will not be repeated here. For example, the video data display device can be a vehicle.

[0088] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of an embodiment of the electronic device of this application. The electronic device 60 includes a memory 61 and a processor 62. The processor 62 is used to execute program instructions stored in the memory 61 to implement the steps in any of the above-described video data display method embodiments. In a specific implementation scenario, the electronic device 60 may include, but is not limited to, electrical appliances, charging devices, microcomputers, and servers. In addition, the electronic device 60 may also include carrier devices such as laptops and tablets, which are not limited here.

[0089] Specifically, processor 62 controls itself and memory 61 to implement the steps in any of the video data display method embodiments described above. Processor 62 can also be referred to as a CPU (Central Processing Unit). Processor 62 may be an integrated circuit chip with signal processing capabilities. Processor 62 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 62 can be implemented using integrated circuit chips.

[0090] The above scheme, in response to receiving a video display command, parses the acquired video file to be displayed to obtain several parsed image groups and vehicle driving information matching each parsed image group. It then renders the parsed image groups and their matching vehicle driving information to obtain a target video file. The target video file includes target display frames corresponding to each parsed image group, representing a combination of the parsed images and their corresponding vehicle driving information within each parsed image group. Displaying each target display frame in the target video file allows for the combination of the parsed images and their corresponding vehicle driving information within the target display frames through rendering, thereby increasing the information dimensionality of the target display frames and enhancing the interactivity when displaying video data.

[0091] Please see Figure 7 , Figure 7 This is a schematic diagram of a computer-readable storage medium according to an embodiment of the present application. The computer-readable storage medium 70 stores program instructions 701 thereon, which, when executed by a processor, implement the steps in any of the above-described video data display method embodiments.

[0092] The above scheme, in response to receiving a video display command, parses the acquired video file to be displayed to obtain several parsed image groups and vehicle driving information matching each parsed image group. It then renders the parsed image groups and their matching vehicle driving information to obtain a target video file. The target video file includes target display frames corresponding to each parsed image group, representing a combination of the parsed images and their corresponding vehicle driving information within each parsed image group. Displaying each target display frame in the target video file allows for the combination of the parsed images and their corresponding vehicle driving information within the target display frames through rendering, thereby increasing the information dimensionality of the target display frames and enhancing the interactivity when displaying video data.

[0093] In some embodiments, this application also provides a computer program product, which, when executed by a processor, is used to implement the above-described video data display method.

[0094] Understandably, a computer program product can be a computer program product contained on a tangible computer-readable medium, which includes program code for performing the video data display method described above. In some embodiments, the computer program product can be downloaded and installed from a network, and can also be copied, transferred, and installed between different computer hardware. Its wireless transmission method can include the Internet, Bluetooth, WIFI, etc., and its wired transmission method can include USB, Lightning, Type-C, etc.

[0095] The above scheme, in response to receiving a video display command, parses the acquired video file to be displayed to obtain several parsed image groups and vehicle driving information matching each parsed image group. It then renders the parsed image groups and their matching vehicle driving information to obtain a target video file. The target video file includes target display frames corresponding to each parsed image group, representing a combination of the parsed images and their corresponding vehicle driving information within each parsed image group. Displaying each target display frame in the target video file allows for the combination of the parsed images and their corresponding vehicle driving information within the target display frames through rendering, thereby increasing the information dimensionality of the target display frames and enhancing the interactivity when displaying video data.

[0096] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0097] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0098] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, units or components may be combined or integrated into another system, or some features may be ignored or not executed. In another image location, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0099] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A video data display method, characterized in that, The method is applied to a vehicle, and the method includes: In response to receiving a video display instruction, the system obtains the video file to be displayed corresponding to the video display instruction. The video file to be displayed includes several groups of images to be parsed and vehicle driving information matched by each group of images to be parsed. Each group of images to be parsed includes at least one frame of image to be parsed at the same acquisition time. The video file to be displayed is parsed to obtain several parsed image groups and vehicle driving information matched by each parsed image group; The plurality of parsed image groups and the vehicle driving information matched by each parsed image group are rendered to obtain a target video file. The target video file includes a target display screen corresponding to each parsed image group. The target display screen represents the combination of each parsed image in the parsed image group and the corresponding vehicle driving information. The display of each target screen in the target video file is processed.

2. The method according to claim 1, characterized in that, The vehicle driving information matched by the parsed image group includes the motion state information matched by the parsed image group; The step of rendering the plurality of parsed image groups and the vehicle driving information matched by each parsed image group to obtain the target video file includes: Obtain the initial vehicle model matched to the vehicle; The motion state information matched by the parsed image group is loaded into the initial vehicle model to obtain the target vehicle model corresponding to the parsed image group; Based on each parsed image group and the corresponding target vehicle model, determine the target display screen corresponding to each parsed image group.

3. The method according to claim 2, characterized in that, The parsed image group includes at least one parsed image frame acquired at the same time. The step of determining the target display screen corresponding to each parsed image group based on each parsed image group and the target vehicle model corresponding to each parsed image group includes: The target image corresponding to the parsed image group is obtained by performing preset image processing on each parsed image in the parsed image group. The target image corresponding to the parsed image group is combined with the target vehicle model corresponding to the parsed image group to obtain the initial display screen corresponding to the parsed image group; Based on the initial display screen corresponding to the parsed image group, determine the target display screen corresponding to the parsed image group.

4. The method according to claim 3, characterized in that, The step of performing preset image processing on each parsed image in the parsed image group to obtain the target image corresponding to the parsed image group includes: Brightness adjustment processing is performed on each of the parsed images in the parsed image group to obtain the adjusted image corresponding to each parsed image; The adjusted images corresponding to each parsed image are fused to obtain the target image corresponding to the parsed image group.

5. The method according to claim 4, characterized in that, The step of fusing the adjusted images corresponding to each parsed image to obtain the target image corresponding to the parsed image group includes: The adjusted images corresponding to each parsed image are stitched together to obtain the initial image corresponding to the parsed image group; In response to the presence of at least one overlapping region in the initial image corresponding to the parsed image group, each overlapping region in the initial image is smoothed to obtain a smoothed image corresponding to each overlapping region. Based on the regions other than the overlapping regions in the initial image and the smoothed images corresponding to the overlapping regions, the target image corresponding to the parsed image group is determined.

6. The method according to claim 4, characterized in that, The step of stitching together the adjusted images corresponding to each parsed image to obtain the initial image corresponding to the parsed image group includes: Perform two-dimensional perspective transformation on the adjusted images corresponding to each resolved image to obtain the two-dimensional transformed images corresponding to each resolved image; and / or, Perform 3D perspective transformation on the adjusted images corresponding to each parsed image to obtain the 3D transformed images corresponding to each parsed image. Image stitching is performed on the two-dimensional transformed images and / or the three-dimensional transformed images corresponding to each of the parsed images to obtain the initial image corresponding to the parsed image group.

7. The method according to claim 3, characterized in that, The vehicle driving information matched by the parsed image group also includes the spatiotemporal state information matched by the parsed image group; The step of determining the target display screen corresponding to the parsed image group based on the initial display screen corresponding to the parsed image group includes: The spatiotemporal state information matched by the parsed image group is filled into a preset display area to obtain the text filling result corresponding to the parsed image group; The text filling result corresponding to the parsed image group is superimposed on the initial display screen corresponding to the parsed image group to obtain the target display screen corresponding to the parsed image group.

8. A video data display device, characterized in that, include: The acquisition module is used to acquire the video file to be displayed corresponding to the video display instruction in response to receiving the video display instruction. The video file to be displayed includes a plurality of image groups to be parsed and vehicle driving information matched by each image group to be parsed. The image group to be parsed includes at least one frame of image to be parsed at the same acquisition time. The parsing module is used to parse the video file to be displayed to obtain several parsed image groups and vehicle driving information matched by each parsed image group; The rendering module is used to render the plurality of parsed image groups and the vehicle driving information matched by each parsed image group to obtain a target video file. The target video file includes a target display screen corresponding to each parsed image group. The target display screen represents the combination of each parsed image in the parsed image group and the corresponding vehicle driving information. The display module is used to display the target display frames in the target video file.

9. An electronic device, characterized in that, include: A memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to perform the method as claimed in any one of claims 1-7.

10. A computer-readable storage medium / computer program product, characterized in that, The computer-readable storage medium stores program instructions that, when executed by a processor, are used to implement the method as described in any one of claims 1-7; The computer program product includes a computer program that, when executed by a processor, implements the method as described in any one of claims 1-7.