Video encoding method, apparatus, device, storage medium and autonomous vehicle

By using predetermined storage space address information to encode video data in autonomous vehicles, the problem of video acquisition difficulties caused by limited resources is solved, enabling the encoding and storage of multiple video data streams on conventional vehicles and reducing data acquisition costs.

CN116320466BActive Publication Date: 2026-04-21BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2023-03-16
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Conventional autonomous vehicles have limited resources and cannot effectively perform video acquisition tasks, leading to a reliance on data acquisition vehicles and increasing data acquisition costs.

Method used

By acquiring video data and encoding it according to the address information corresponding to the image frame data in the predetermined storage space, the system trades space and time for encoding performance, enabling video data caching and later encoding, which is suitable for autonomous driving systems.

Benefits of technology

Encoding and storing multi-channel video data on conventional autonomous vehicles reduces the difficulty and cost of data collection and increases flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116320466B_ABST
    Figure CN116320466B_ABST
Patent Text Reader

Abstract

This disclosure provides a video encoding method, apparatus, device, storage medium, and autonomous vehicle, relating to the field of artificial intelligence, and particularly to the field of autonomous driving. The specific implementation scheme includes: acquiring video data, which comprises multiple image frame data; for each image frame data, in response to detecting that the image frame data meets predetermined encoding conditions, encoding the image frame data according to address information corresponding to the image frame data in a predetermined storage space to obtain multiple encoded image data; wherein the address information corresponding to the image frame data indicates the address information of a target region in the predetermined storage space, and the target region stores a target image identical to the image frame data; and determining encoded video data of the target video data based on the multiple encoded image data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and more particularly to the field of autonomous driving. More specifically, this disclosure provides a video encoding method, apparatus, electronic device, storage medium, computer program product, and autonomous vehicle. Background Technology

[0002] The vehicles may include conventional autonomous driving vehicles for autonomous driving and data acquisition vehicles dedicated to data acquisition tasks. The data acquisition vehicles are equipped with the same types and number of sensors as the conventional autonomous driving vehicles, and the data acquisition vehicles have strong video encoding capabilities to compress and store data streams from multiple channels in real time. The data collected by the data acquisition vehicles using the sensors can be used for offline model training and validation.

[0003] Among various sensor data, video data has a large volume, and conventional autonomous vehicles have limited resources, which cannot provide enough resources for video encoding tasks. This makes it impossible for conventional autonomous vehicles to perform video acquisition tasks, thus requiring the reliance on data acquisition vehicles to perform video acquisition tasks, increasing the cost of data acquisition. Summary of the Invention

[0004] This disclosure provides a video coding method, apparatus, electronic device, storage medium, computer program product, and autonomous vehicle.

[0005] According to one aspect of this disclosure, a video encoding method is provided, comprising: acquiring video data, the video data including a plurality of image frame data; for each image frame data in the plurality of image frame data, in response to detecting that the image frame data satisfies a predetermined encoding condition, encoding the image frame data according to address information corresponding to the image frame data in a predetermined storage space to obtain a plurality of encoded image data; wherein the address information corresponding to the image frame data indicates address information of a target region in the predetermined storage space, the target region storing a target image identical to the image frame data; and determining encoded video data of the target video data based on the plurality of encoded image data.

[0006] According to another aspect of this disclosure, a video encoding apparatus is provided, comprising: an acquisition module, an encoding module, and a first determination module. The acquisition module is used to acquire video data, which includes multiple image frame data. The encoding module is used, for each image frame data in the multiple image frame data, in response to detecting that the image frame data satisfies predetermined encoding conditions, to encode the image frame data according to address information corresponding to the image frame data in a predetermined storage space, thereby obtaining multiple encoded image data; wherein the address information corresponding to the image frame data indicates the address information of a target region in the predetermined storage space, and the target region stores a target image identical to the image frame data. The first determination module is used to determine the encoded video data of the target video data based on the multiple encoded image data.

[0007] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the methods provided in this disclosure.

[0008] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the methods provided in this disclosure.

[0009] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods provided in this disclosure.

[0010] According to another aspect of this disclosure, an autonomous vehicle is provided, including: a camera and the aforementioned electronic equipment.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0013] Figure 1 This is a schematic diagram illustrating an application scenario of the video encoding method and apparatus according to embodiments of this disclosure;

[0014] Figure 2 This is a schematic flowchart of a video encoding method according to an embodiment of the present disclosure;

[0015] Figure 3 This is a schematic diagram illustrating the principle of a video encoding method according to an embodiment of the present disclosure;

[0016] Figure 4 This is a schematic structural block diagram of a video encoding apparatus according to embodiments of the present disclosure; and

[0017] Figure 5 This is a structural block diagram of an electronic device used to implement the video encoding method of the embodiments of this disclosure. Detailed Implementation

[0018] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0019] Figure 1 This is a schematic diagram illustrating an application scenario of the video encoding method and apparatus according to embodiments of this disclosure. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.

[0020] like Figure 1 As shown, the system architecture 100 according to this embodiment may include sensors 101, 102, and 103, a network 120, a server 130, and a roadside unit (RSU) 140. The network 120 serves as a medium for providing communication links between the sensors 101, 102, and 103 and the server 130. The network 120 may include various connection types, such as wired and / or wireless communication links, etc.

[0021] Sensors 101, 102, and 103 can interact with server 130 via network 120 to receive or send messages, etc.

[0022] Sensors 101, 102, and 103 can be functional components integrated on vehicle 110, such as cameras. Sensors 101, 102, and 103 can be used to acquire images of the surroundings of vehicle 110, such as environmental images and images of perceived objects (e.g., pedestrians, vehicles, obstacles). However, sensors can also be infrared sensors, ultrasonic sensors, millimeter-wave radar, lidar, inertial measurement units, etc.

[0023] Vehicle 110 can communicate with roadside unit 140, receive information from roadside unit 140, or send information to roadside unit 140.

[0024] Server 130 can be set at a remote location that can establish communication with the vehicle terminal. It can be implemented as a distributed server cluster consisting of multiple servers or as a single server.

[0025] Server 130 can be a server that provides various services. Applications such as map applications and data processing applications can be installed on server 130.

[0026] It should be noted that the video encoding method provided in this embodiment can generally be executed by the vehicle 110 or the server 130. Accordingly, the video encoding device provided in this embodiment can also be installed in the vehicle 110 or the server 130.

[0027] Understandable. Figure 1 The number of sensors, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of sensors, networks, and servers can be included.

[0028] Figure 2 This is a schematic flowchart of a video encoding method according to an embodiment of the present disclosure.

[0029] like Figure 2 As shown, the video encoding method 200 may include operations S210 to S230.

[0030] In operation S210, video data is acquired, which includes multiple image frame data.

[0031] For example, cameras can be installed on autonomous vehicles or other devices. After the cameras collect video data, they can send the collected video data to an electronic device that performs the video encoding method.

[0032] In operation S220, for each image frame data in the multiple image frame data, in response to detecting that the image frame data meets the predetermined encoding conditions, the image frame data is encoded according to the address information corresponding to the image frame data in the predetermined storage space to obtain multiple encoded image data.

[0033] For example, the address information corresponding to the image frame data indicates the address information of the target area in the predetermined storage space, and the target area stores the same target image as the image frame data.

[0034] For example, the predetermined encoding condition includes: the previous image frame data in the video data has been encoded. Using this predetermined encoding condition, multiple image frames in the video data can be processed sequentially according to the image acquisition order.

[0035] For example, predetermined encoding conditions can ensure that the resources required to encode image frame data are less than the remaining resources.

[0036] For example, the reserved storage space can be some space pre-allocated in memory or some space on the disk.

[0037] For example, during the encoding process of a certain image frame data, the target image can be obtained from the address information corresponding to the image frame data in the predetermined storage space, and then image encoding processing can be performed to obtain the encoded image of the certain image frame data.

[0038] In operation S230, the encoded video data of the target video data is determined based on multiple encoded image data.

[0039] For example, multiple images captured by the same camera can be encoded to obtain multiple encoded images. Combining these encoded images in the order they were captured yields encoded video data.

[0040] This disclosure provides a video encoding method applicable to autonomous driving systems. This method trades space (memory buffering) and time (non-real-time encoding) for encoding performance. Through video data caching and later encoding mechanisms, it enables autonomous driving systems with limited video encoding performance to achieve roadside data acquisition. This allows the video acquisition task of the camera to be performed on a conventional autonomous driving vehicle, achieving the effect of encoding and storing multiple video data on the autonomous driving system, while reducing the difficulty and cost of data acquisition and increasing flexibility.

[0041] According to another embodiment of this disclosure, before encoding image frame data, the video encoding method may further include a preprocessing stage, which may include processes such as allocating space, acquiring images, and storing images.

[0042] The following explains the space allocation process. For example, based on predetermined parameters, some space can be pre-allocated in memory as a predetermined storage space. This predetermined storage space can be used to store video data before and after encoding. Pre-allocation references can include the size of a single image, frame rate, number of video channels, and maximum video duration. For instance, the size of the predetermined storage space can be determined by multiplying the size of a single image, frame rate, number of video channels, and maximum video duration. This method allows for the reasonable allocation of the predetermined storage space in memory, avoiding insufficient predetermined storage space from affecting video capture quality, and also preventing excessive unused space during video capture, thus avoiding waste of storage resources.

[0043] The image acquisition process is described below. For example, a vehicle can be equipped with n cameras to acquire images, where n is an integer greater than or equal to 1. For instance, when a user discovers that the first scene around the vehicle (e.g., a traffic light ahead and many bicycles nearby) is the scene required for training the model, they can initiate a start operation, triggering the n cameras to acquire images of the surrounding environment. When the user deems it unnecessary to continue acquiring images of this scene, they can initiate a stop operation, triggering the n cameras to stop acquiring images of the surrounding environment. The real-time video data acquired by the cameras can be cached in the pre-allocated storage space mentioned above.

[0044] It should be noted that multiple video data sets can contain the same image data. For example, if one video data set is collected between 9:00 and 9:10, and another video data set is collected between 9:05 and 9:20, it can be seen that both video data sets include image frame data collected between 9:05 and 9:10.

[0045] Using the above methods, users can collect videos according to their actual needs, improving the flexibility of data collection and avoiding the problem of large, unprocessable data volumes caused by continuous collection.

[0046] The image storage process is described below. For example, a buffer queuing management module can be used to store images captured by the camera in a predetermined storage space and record the address information of each image in the predetermined storage space. It is understood that the camera can capture images in real time; this embodiment uses the processing of one image as an example for explanation.

[0047] An image captured by the camera and before being stored in a designated storage space is called a "to-be-stored image." The camera captures images in real time. After the encoding and queuing management module detects or receives a to-be-stored image, it can determine the address information of the image. For example, an available area for storing the image can be determined within the designated storage space, and then the image can be stored in that available area, with the address information of the available area recorded as the address information of the stored image. This method accurately records the address information of each newly acquired image to be stored. The following is a detailed explanation of how to determine the address information of the stored image.

[0048] First, you can determine the available area in the reserved storage space for storing the image to be stored.

[0049] In one example, the predetermined storage space can be pre-divided into multiple regions, and regions with remaining space larger than the size of the image to be stored can be identified as available regions.

[0050] In another example, available areas can be determined from the unoccupied areas within the predetermined storage space based on the status information of multiple regions. Identifying unoccupied areas as available areas prevents newly acquired images from overwriting images needed for subsequent encoding, thus avoiding disruption to the normal operation of the video encoding task.

[0051] For example, a predetermined storage space can be pre-divided into multiple regions, each of which can be used to store one or more images. The status of a region can include an unoccupied state and an occupied state. An unoccupied state indicates that no images are stored in the region, while an occupied state indicates that images are stored in the region. It is possible to detect whether there are any unoccupied regions within the predetermined storage space. During detection, unoccupied regions can be found based on the status information of each region, or multiple regions can be traversed to determine whether each region is unoccupied.

[0052] If there is no unoccupied area in the predetermined storage area, the images captured by the camera will stop entering the predetermined storage area. At this time, no new images will be added to the video data in the predetermined storage area. However, the encoding unit can continue to encode the images already stored in the predetermined storage area.

[0053] If there are unoccupied areas in the reserved storage area, then one of these areas can be selected as the available area.

[0054] After determining the available area, the status information of the available area can be updated to the occupied state, thereby preventing other images acquired later from occupying the available area if images are stored in that area.

[0055] After determining the available area, the image to be stored can be stored in the available area. For example, the buffer queuing management module can send the address information of the available area to the video subsystem, which can be a V4L2 (Video for Linux 2) system. Then, the video subsystem stores the image to be stored in the available area according to the address information. After the storage is completed, the video subsystem can notify the buffer queuing management module.

[0056] After determining the available area, the address information of the available area can be recorded as the address information for storing the images to be stored. For example, the buffer queuing management module can maintain management information that records the acquisition order of multiple images in each video data and the address information of each image in a predetermined storage area.

[0057] In one example, a data table structure can be used to record the aforementioned management information.

[0058] In another example, a linked list structure can be used to record the aforementioned management information. The linked list includes multiple nodes. Each time a new image is added to be stored and an available area is determined for that image, the head node of the linked list is used as the previous node, and a new current node corresponding to the image to be stored is added at the head node. Nodes in the linked list can include data fields and pointer fields. The pointer in the pointer field points from the previous node to the next node; this pointer can be called the NEXT pointer. The data field can record the address information of the available area. It is understandable that when adding a new current node, the tail node of the linked list can also be used as the previous node, and the new current node is added at the tail node.

[0059] This embodiment uses a linked list structure to record address information. Images captured by the same camera can correspond to a linked list. In the linked list, the next node can be determined based on the NEXT pointer of the previous node, and the order of each node in the linked list structure is consistent with the image acquisition order. Therefore, the image order can be accurately recorded based on the linked list structure. Furthermore, the address information corresponding to the image is recorded in the data field of the linked list structure, which can also accurately establish the correspondence between address information and images.

[0060] In some embodiments, the management information may further include an identifier set corresponding to the image, which includes identifiers for video data containing image frame data identical to the image. By recording the identifier set, it is possible to accurately determine whether the image's state can be updated to an idle state, thereby ensuring the normal execution of the video encoding task.

[0061] In the process of recording identifier sets, for example, when using a data table structure, the identifier set can be treated as a field in the data table. Alternatively, when using a linked list structure, the identifier set can be recorded in the data field of the current node. Recording the identifier set corresponding to the image in the data field of the linked list structure also allows for an accurate establishment of the correspondence between identifier sets and images.

[0062] The following explains how to determine the identifier set. For example, each time a user initiates a video data acquisition task, they can assign an identifier to the newly added video data and add that identifier to an array. Next, querying this array reveals the video data currently being acquired. Then, for each newly acquired image, the identifier from the array can be assigned to the data field of the current node in the linked list corresponding to that image, thus recording the identifier set in the data field of the linked list node. It's important to understand that the video data currently being acquired refers to video data that has not yet finished acquiring, not video data that has not been encoded.

[0063] In some embodiments, the management information may also include address information corresponding to the starting frame in each video data segment, whether each video data segment has been acquired, and whether each video data segment has been encoded.

[0064] In some embodiments, the space occupancy rate can be determined based on the amount of free space, the amount of occupied space, and the total reserved storage space, and the space occupancy rate can be displayed to prompt the user whether the reserved storage space is sufficient and whether video recording can continue.

[0065] The above provides a detailed explanation of the processes involved in the preprocessing stage, including space allocation, image acquisition, and image storage. Next, we will explain the encoding process following the preprocessing stage.

[0066] In this embodiment, during the encoding of image frame data, encoding tasks corresponding one-to-one with multiple video data can be generated. Multiple encoding tasks can be processed simultaneously or in batches. The encoding process is described below using the processing of one encoding task as an example.

[0067] In this embodiment, the image frame data can be encoded according to the address information corresponding to the image frame data in the predetermined storage space to obtain encoded image data.

[0068] The encoding dequeue management module can obtain the head or tail node of the linked list structure corresponding to each video data from the encoding enqueue management module. Then, it sends the address information in the data field of that node to the encoding unit. The encoding unit can then locate the target image in a predetermined storage space based on this address information, obtain the target image stored in the target area indicated by the address information, and load the target image. Next, the loaded target image can be encoded, and the encoded data can be used as the encoded image data for the image frame data.

[0069] In one example, if the video data includes other image data preceding the image frame data, these other image data can be the preceding image of the image frame data. For instance, the target image can be encoded based on a target image in a predetermined storage space that is identical to the image frame data, and also based on images in the predetermined storage space that are identical to other image data. For example, the target image can be encoded based on the differences between the target image and the preceding image. This encoding method can effectively reduce redundant information in the image and improve encoding efficiency. Alternatively, the target image can be encoded not based on images in the predetermined storage space that are identical to other image data, but only based on the target image in the predetermined storage space that is identical to the image frame data.

[0070] In another example, if the image frame data is the starting frame in the video data, since there are no other images before the image frame data in the video data, the target image can be encoded based on the target image. This encoding method can accurately obtain the encoded image of the image frame data.

[0071] According to another embodiment of this disclosure, after obtaining the encoded image for a certain image frame data, the encoding dequeue management module can also update the state of a target region in a predetermined storage area. As mentioned above, the target region is a region that stores a target image identical to the image frame data. The method of updating the state of the target region is described below.

[0072] For example, a target image corresponds to an identifier set, which includes identifiers for video data containing image frame data identical to that of the target image. For instance, if video data v1, video data v2, and video data v3 all contain image frame data identical to that of the target image p1, then the identifier set corresponding to the target image p1 may include v1, v2, and v3.

[0073] After obtaining the encoded image of the same image frame data as the target image p in the video data v1, v1 can be removed from the identifier set. It can be seen that the processed identifier set obtained after the removal includes v2 and v3.

[0074] Next, the status of the target area can be determined based on the processed identifier set.

[0075] For example, the processed identifier set obtained after the above deletion includes v2 and v3, that is, the processed identifier set is a non-empty set. The processed identifier set indicates that video data v2 and video data v3 still need to be encoded using target image p1. If target image p1 is deleted after video data v1 is encoded using target image p1, it will affect the encoding process of video data v2 and video data v3. Therefore, target image p1 stored in the target area can be left undeleted, and the state of the target area can be determined as occupied to avoid newly acquired images from covering target image p1.

[0076] For example, if the identifier set after the above processing is empty, it means that the remaining temporarily unencoded image frame data does not need to be encoded using the target image p1. In this case, the target image p1 stored in the target area can be deleted, and the status of the target area can be updated to unoccupied. It can be understood that after the target area is in an unoccupied state, newly acquired images can be stored in the target area.

[0077] As can be seen in this embodiment, when a target image is used during the encoding of video data, the identifier of that video data is removed from the identifier set corresponding to the target image. If the identifier set contains identifiers for other video data, the target area storing the target image remains occupied, and the target image is retained to prevent it from being deleted or covered by other images, thus affecting the encoding of other video data. If there are no other identifiers in the identifier set, the target image can be deleted, allowing the target area to store newly acquired images.

[0078] After encoding a certain image, the address information of the next image can be obtained from the linked list maintained by the encoding queuing management module, and then the next image can be encoded, until the encoding of the video data is completed.

[0079] Furthermore, after obtaining the encoded images, data can be written to disk. For example, the encoded images can be stored in a predetermined storage space, and the address information of each encoded image in the predetermined storage space can be recorded. Then, based on the address information, the encoded images can be transferred from the predetermined storage space to the disk.

[0080] Figure 3 This is a schematic diagram illustrating the video encoding method according to an embodiment of the present disclosure.

[0081] like Figure 3 As shown, in this embodiment, the vehicle can be equipped with n cameras 310, such as the first camera 311, the second camera 312, ..., the nth camera 313, where n is an integer greater than or equal to 1.

[0082] Multiple video data sets can be collected according to actual needs. For example, video data 1 was collected from 9:00 to 9:10, and video data 2 was collected from 9:05 to 9:20. The same video data set can include n sub-videos collected by n cameras 310.

[0083] In the actual encoding process, video streams captured by n cameras can be encoded simultaneously for the same video data.

[0084] For video data 1, the first camera 311 can capture video stream cam1, which includes multiple images, such as the first frame x11, the second frame y11, and the third frame z11. Cam2 to CamN are similar to Cam1. The second camera 312 captures video stream cam2, which includes multiple images, such as the first frame x12, the second frame y12, and the third frame z12. The nth camera 313 captures video stream camn, which includes multiple images, such as the first frame x1n, the second frame y1n, and the third frame z1n. It can be seen that the first frames x11, x12, and x1n are images from different perspectives captured by different cameras at the same time.

[0085] The relevant information of multiple images can be added to the queue according to the image acquisition order. It should be noted that the data acquired by the camera (e.g., the first frame x11, the second frame y11, and the third frame z11) is stored in a predetermined storage space. The buffer queuing management module 320 does not store each image acquired by the camera, but rather manages each image acquired by the camera, such as managing the address information of each image acquired by each camera in the predetermined storage area 330, and the acquisition order of multiple images acquired by the same camera. The buffer queuing management module 320 can use a linked list structure, a data table structure, or other structures to maintain the above management information. The linked list structure can be referred to above, and will not be repeated in this embodiment.

[0086] Images captured by the same camera at different times can correspond to an encoding task. During the execution of the encoding task, the encoding unit can obtain images from the predetermined storage space based on the management information maintained by the buffer queuing management module 320, perform encoding processing, and then output the encoded image to the predetermined storage area 330.

[0087] For example, if video data 1 includes multiple images captured by n cameras, then after encoding, the images x11, y11, and z11 captured by the first camera 311 are processed into encoded images x11', y11', and z11', respectively. Similarly, the images x12, y12, and z12 captured by the second camera 312 are processed into encoded images x12', y12', and z12', respectively, and the images x1n, y1n, and z1n captured by the nth camera 313 are processed into encoded images x1n', y1n', and z1n', respectively.

[0088] Similar to video data 1, video data 2 is acquired by the first camera 311, which captures a video stream cam1. This video stream cam1 includes multiple image frame data, such as the first frame x21, the second frame y21, and the third frame z21. These image frame data are processed into encoded images x21', y21', and z21', respectively. Cam2 to CamN are similar to Cam1. The second camera 312 captures a video stream cam2, which includes multiple image frame data, such as the first frame x22, the second frame y22, and the third frame z22. These images are processed into encoded images x22', y22', and z22', respectively. The nth camera 313 captures a video stream camn, which includes multiple image frame data, such as the first frame x2n, the second frame y2n, and the third frame z2n. These images are processed into encoded images x2n', y2n', and z2n', respectively.

[0089] The encoding dequeue management module 340 can also record the address information of each encoded image, such as the address information of multiple encoded images mentioned above. Furthermore, the encoding dequeue management module 340 can transfer encoded images from the predetermined storage area 330 to the disk 350, thereby completing video acquisition.

[0090] Figure 4 This is a schematic structural block diagram of a video encoding apparatus according to an embodiment of the present disclosure.

[0091] like Figure 4 As shown, the video encoding device 400 may include an acquisition module 410, an encoding module 420, and a first determination module 430.

[0092] The acquisition module 410 is used to acquire video data, which includes multiple image frame data.

[0093] The encoding module 420 is used to encode each image frame data in a plurality of image frame data in response to detecting that the image frame data meets a predetermined encoding condition, according to the address information corresponding to the image frame data in a predetermined storage space, to obtain a plurality of encoded image data; wherein, the address information corresponding to the image frame data indicates the address information of a target region in the predetermined storage space, and the target region stores the same target image as the image frame data.

[0094] The first determining module 430 is used to determine the encoded video data of the target video data based on multiple encoded image data.

[0095] According to another embodiment of this disclosure, the encoding module includes a loading submodule and an encoding submodule. The loading submodule is used to load a target image from a predetermined storage space according to address information corresponding to image frame data. The encoding submodule is used to encode the loaded target image and use the encoded data as encoded image data of the image frame data.

[0096] According to another embodiment of this disclosure, the encoding submodule includes a first encoding unit and a second encoding unit. The first encoding unit is configured to encode a target image based on a target image in response to detecting that the image frame data is the starting frame in the video data. The second encoding unit is configured to encode the target image based on the target image and images identical to other image data in a predetermined storage space in response to detecting that the video data includes other image data preceding the image frame data.

[0097] According to another embodiment of this disclosure, the predetermined encoding conditions include: the previous image frame data in the video data has been encoded.

[0098] According to another embodiment of this disclosure, multiple video data each include image frame data identical to the target image, the target image corresponds to an identifier set, and the identifier set corresponding to the target image includes the identifiers of each of the multiple video data. The apparatus further includes a deletion module and a second determination module. The deletion module is used to delete the identifiers of the target video data from the identifier set corresponding to the target image after encoding the image frame data, obtaining a processed identifier set. The second determination module is used to determine the state of the target region based on the processed identifier set. The target video data is the video data among the multiple video data that has been encoded with image frame data identical to the target image.

[0099] According to another embodiment of this disclosure, the second determining module includes an updating submodule and a determining submodule. The updating submodule is used to update the state of the target region to an unoccupied state in response to detecting that the processed identifier set is an empty set. The determining submodule is used to determine that the state of the target region is an occupied state in response to detecting that the processed identifier set is a non-empty set.

[0100] According to another embodiment of this disclosure, the above-described apparatus further includes: an allocation module, configured to allocate space in memory as a predetermined storage space based on the size of a single image, frame rate, number of video channels, and maximum video duration.

[0101] According to another embodiment of this disclosure, the apparatus further includes a third determining module, a storage module, and a first recording module. The third determining module is configured to determine an available area in a predetermined storage space in response to receiving an image to be stored. The storage module is configured to store the image to be stored in the available area. The first recording module is configured to record the address information of the available area as address information for storing the image to be stored.

[0102] According to another embodiment of this disclosure, the first recording module includes a adding submodule and a first recording submodule. The adding submodule is used to take the head node or tail node of the linked list as the previous node and add a current node corresponding to the image to be stored at the previous node; wherein, the nodes in the linked list include a data field. The first recording submodule is used to record the address information of the available area in the data field of the current node.

[0103] According to another embodiment of this disclosure, the apparatus further includes: a second recording module for recording an identifier set corresponding to an image to be stored; wherein the identifier set corresponding to the image includes: identifiers of video data containing image frame data identical to the image.

[0104] According to another embodiment of this disclosure, the image to be stored corresponds to the current node in a linked list, and the nodes in the linked list include a data field. The second recording module includes: a second recording submodule, used to record the identifier of video data containing image frame data identical to the image in the data field of the current node.

[0105] According to another embodiment of this disclosure, the third determining module includes: determining an available area from the unoccupied areas in the predetermined storage space based on the status information of each of the multiple areas in the predetermined storage space.

[0106] According to another embodiment of this disclosure, the above-mentioned apparatus further includes: an update module, configured to update the status of the available area to an occupied state after obtaining the available area.

[0107] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0108] In the technical solution disclosed herein, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.

[0109] According to embodiments of this disclosure, this disclosure also provides an electronic device, including at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the video encoding method described above.

[0110] According to embodiments of this disclosure, this disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the video encoding method described above.

[0111] According to embodiments of this disclosure, this disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described video encoding method.

[0112] According to embodiments of this disclosure, this disclosure also provides a vehicle, which may be an autonomous driving vehicle or a non-autonomous driving vehicle, and the vehicle may include a camera and the aforementioned electronic equipment.

[0113] Figure 5 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0114] like Figure 5 As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 502 or a computer program loaded from storage unit 508 into random access memory (RAM) 503. RAM 503 may also store various programs and data required for the operation of device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.

[0115] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0116] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as video encoding methods. For example, in some embodiments, the video encoding method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the video encoding method described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform video encoding methods by any other suitable means (e.g., by means of firmware).

[0117] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0118] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0119] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0120] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0121] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0122] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.

[0123] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0124] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A video encoding method, comprising: Acquire video data, which includes multiple image frame data; For each of the plurality of image frame data, in response to detecting that the image frame data meets a predetermined encoding condition, the target image stored in the target area indicated by the address information in the predetermined storage space is encoded according to the address information corresponding to the image frame data, to obtain a plurality of encoded image data; wherein, the address information corresponding to the image frame data indicates the address of the target area in the predetermined storage space, and the target area stores the target image that is the same as the image frame data; as well as Based on the plurality of encoded image data, the encoded video data of the video data is determined.

2. The method according to claim 1, wherein, The step of encoding the target image stored in the target area indicated by the address information according to the address information in the predetermined storage space to obtain multiple encoded image data includes: Based on the address information corresponding to the image frame data, the target image is loaded from the predetermined storage space; and The loaded target image is encoded, and the encoded data is used as the encoded image data of the image frame data.

3. The method according to claim 2, wherein, Encoding the target image includes: In response to detecting that the image frame data is the start frame in the video data, the target image is encoded based on the target image; and In response to detecting that the video data includes other image data preceding the image frame data, the target image is encoded based on the target image and images in the predetermined storage space that are identical to the other image data.

4. The method according to claim 1, wherein, The predetermined encoding conditions include: the previous image frame data in the video data has been encoded.

5. The method according to claim 1, wherein, The multiple video data sets each include image frame data identical to the target image, and the target image corresponds to an identifier set, the identifier set corresponding to the target image including the identifiers of each of the multiple video data sets; the method further includes: after encoding the image frame data... From the identifier set corresponding to the target image, the identifiers of the target video data are deleted to obtain the processed identifier set; and The state of the target region is determined based on the processed identifier set; The target video data refers to video data from the plurality of video data that has been encoded with image frame data identical to the target image.

6. The method according to claim 5, wherein, Determining the state of the target region based on the processed identifier set includes: In response to detecting that the processed identifier set is empty, the state of the target region is updated to an unoccupied state; and In response to detecting that the processed identifier set is not empty, the state of the target region is determined to be occupied.

7. The method according to claim 1, further comprising: Based on the size of a single image, frame rate, number of video channels, and maximum video duration, space is allocated in memory as the predetermined storage space.

8. The method according to any one of claims 1 to 7, further comprising: In response to receiving an image to be stored, an available area is determined in the predetermined storage space; Store the image to be stored in the available area; as well as The address information of the available area is recorded as the address information for storing the image to be stored.

9. The method according to claim 8, wherein, The step of recording the address information of the available area as the address information for storing the image to be stored includes: The head or tail node of the linked list is used as the preceding node, and a new current node corresponding to the image to be stored is added at the preceding node; wherein, the nodes in the linked list include data fields; and The address information of the available area is recorded in the data field of the current node.

10. The method according to claim 8, further comprising: Record the identifier set corresponding to the image to be stored; The identifier set corresponding to the image includes: identifiers for video data containing the same image frame data as the image.

11. The method according to claim 10, wherein, The image to be stored corresponds to the current node in the linked list, and the nodes in the linked list include a data field; The identifier set corresponding to the record and the image to be stored includes: The identifier of the video data containing the same image frame data as the image is recorded in the data field of the current node.

12. The method according to claim 8, wherein, Determining the available area in the predetermined storage space includes: Based on the status information of each of the multiple regions in the predetermined storage space, the available region is determined from the regions in the predetermined storage space that are in an unoccupied state.

13. The method of claim 8, further comprising: After obtaining the available area Update the status of the available area to "occupied".

14. A video encoding apparatus, comprising: The acquisition module is used to acquire video data, which includes multiple image frame data; An encoding module is configured to, for each of the plurality of image frame data, in response to detecting that the image frame data meets a predetermined encoding condition, encode a target image stored in a target area indicated by the address information in a predetermined storage space according to the address information corresponding to the image frame data, to obtain a plurality of encoded image data; wherein, the address information corresponding to the image frame data indicates the address of the target area in the predetermined storage space, and the target area stores the target image that is the same as the image frame data; as well as The first determining module is used to determine the encoded video data of the video data based on the plurality of encoded image data.

15. The apparatus according to claim 14, wherein, The encoding module includes: A loading submodule is configured to load the target image from the predetermined storage space according to the address information corresponding to the image frame data; and The encoding submodule is used to encode the loaded target image and use the encoded data as the encoded image data of the image frame data.

16. The apparatus according to claim 15, wherein, The encoding submodule includes: A first encoding unit is configured to encode the target image based on the target image in response to detecting that the image frame data is the start frame in the video data; and The second encoding unit is configured to encode the target image based on the target image and an image in the predetermined storage space that is identical to the other image data, in response to detecting that the video data includes other image data preceding the image frame data.

17. The apparatus according to claim 14, wherein, The predetermined encoding conditions include: the previous image frame data in the video data has been encoded.

18. The apparatus according to claim 14, wherein, The multiple video data sets each include image frame data identical to the target image, the target image corresponds to an identifier set, and the identifier set corresponding to the target image includes the identifiers of each of the multiple video data sets; the device further includes: The deletion module is configured to, after encoding the image frame data, delete the identifiers of the target video data from the identifier set corresponding to the target image, thereby obtaining a processed identifier set; and The second determining module is used to determine the state of the target region based on the processed identifier set; The target video data refers to video data from the plurality of video data that has been encoded with image frame data identical to the target image.

19. The apparatus according to claim 18, wherein, The second determining module includes: The update submodule is configured to update the state of the target region to an unoccupied state in response to detecting that the processed identifier set is empty; and The determination submodule is used to determine the state of the target region as occupied in response to detecting that the processed identifier set is a non-empty set.

20. The apparatus of claim 14, further comprising: The allocation module is used to allocate space in memory as the predetermined storage space based on the size of a single image, frame rate, number of video channels, and maximum video duration.

21. The apparatus according to any one of claims 14 to 20, further comprising: The third determining module is used to determine an available area in the predetermined storage space in response to receiving an image to be stored; A storage module is used to store the image to be stored in the available area; as well as The first recording module is used to record the address information of the available area as the address information for storing the image to be stored.

22. The apparatus according to claim 21, wherein, The first recording module includes: A new submodule is added to take the head node or tail node of the linked list as the previous node and add a current node corresponding to the image to be stored at the previous node; wherein, the nodes in the linked list include a data field; and The first recording submodule is used to record the address information of the available area in the data field of the current node.

23. The apparatus of claim 21, further comprising: The second recording module is used to record the identifier set corresponding to the image to be stored; The identifier set corresponding to the image includes: identifiers for video data containing the same image frame data as the image.

24. The apparatus according to claim 23, wherein, The image to be stored corresponds to the current node in the linked list, and the nodes in the linked list include a data field; The second recording module includes: The second recording submodule is used to record the identifier of video data containing the same image frame data as the image in the data field of the current node.

25. The apparatus according to claim 21, wherein, The third determining module includes: Based on the status information of each of the multiple regions in the predetermined storage space, the available region is determined from the regions in the predetermined storage space that are in an unoccupied state.

26. The apparatus of claim 21, further comprising: The update module is used to update the status of the available area to occupied after obtaining the available area.

27. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 13.

28. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 13.

29. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 13.

30. An autonomous vehicle, comprising: camera; as well as The electronic device according to claim 27.

Citation Information

Patent Citations

  • Video encoding method and device, and mobile platform

    CN111316643A