System and method for accelerated video-based training of machine learning models

By analyzing the bitstream of the video sequence and determining the timestamp and location of the image frame, the processor can quickly decode the video data and accelerate the training of the machine learning model, solving the problems of time-consuming and low resource utilization efficiency in the prior art, and realizing an efficient training process.

CN120153378APending Publication Date: 2025-06-13TESLA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380076697.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-09-29
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the prior art, the process of using video data to train machine learning models is very time-consuming and is less efficient in utilizing computing and storage resources, especially when predicting or perceiving the surroundings of autonomous vehicles, the training process is extremely time-consuming.

Method used

By receiving the bitstream of the video sequence and the image frame indicating the video sequence, the processor parses the bitstream to determine the timestamp, location and type of the image frame, and then determines the decoded segment, realizing the fast decoding of the video data and effective resource utilization, significantly accelerating the training and verification process of the machine learning model.

Benefits of technology

Accelerated training of machine learning models and more efficient use of computing and storage resources are realized, which significantly reduces training time and improves training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120153378A_ABST
    Figure CN120153378A_ABST
Patent Text Reader

Abstract

Systems and methods may include a computing system that receives a bitstream of a video sequence and one or more indications indicating one or more image frames of the video sequence, determines timestamps, locations, and types of image frames of the video sequence by parsing the bitstream, and transmits the timestamps, locations, and types of the image frames to the computing system. And determining one or more segments of the bitstream for decoding to extract one or more image frames using the one or more indications and the timestamps of the image frames of the video sequence, the location within the bitstream, and the type. For an image frame of the one or more image frames, a corresponding segment represents a corresponding reference chain of image frames in the video sequence. The computing system may decode one or more segments of the bitstream and train an ML model using one or more image frames.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related patent applications

[0002] This application claims priority to U.S. Provisional Application No. 63 / 377,954, filed Sep. 30, 2022, and U.S. Provisional Application No. 63 / 378,012, filed Sep. 30, 2022, the entire contents of each of which are hereby incorporated by reference for all purposes. Technical Field

[0003] The present disclosure generally relates to video training of machine learning (ML) models. Specifically, the present disclosure relates to systems and methods for accelerating the training of ML models using video data. Background Art

[0004] Due to the rapid development of computer technology, autonomous navigation technologies for autonomous vehicles and robots (collectively referred to as self-bodies) have become ubiquitous. These advancements allow for safer and more reliable autonomous navigation of self-bodies. Self-bodies typically need to navigate in complex and dynamic environments and terrains, which can include vehicles, traffic, pedestrians, cyclists, and various other static or dynamic obstacles. Understanding the surrounding environment of a self-body is necessary for informed and capable decision-making to avoid collisions. Summary of the Invention

[0005] The systems, devices, and methods described herein provide for accelerated training of machine learning (ML) models. Specifically, for ML models trained using video data, the systems, devices, and methods described herein achieve rapid decoding of video data and efficient use of computing and memory resources. For ML models or artificial intelligence (AI) models, such as occupancy networks, that are used to predict or perceive the surrounding environment of a self-body, the training of such models is extremely time-consuming. The systems, devices, and methods described herein significantly accelerate the training and / or validation of such models.

[0006] In one embodiment, a method may include: receiving, by a processor, a bitstream of a video sequence and one or more indications indicating one or more image frames of the video sequence for training a machine learning (ML) model; determining, by the processor, timestamps, positions, and types of image frames of the video sequence by parsing the bitstream; determining, by the processor, one or more segments of the bitstream for decoding to extract one or more image frames using the one or more indications and the timestamps, positions, and types of image frames of the video sequence, such that for an image frame among the one or more image frames, the corresponding segment represents a corresponding reference chain of the image frame of the video sequence; decoding, by the processor, the one or more segments of the bitstream; and training, by the processor, the ML model using the one or more image frames.

[0007] The method may further include: allocating a memory region by a processor within a memory of the processor; storing a bitstream by the processor in the allocated memory region; and storing one or more image frames by the processor in the allocated memory region after decoding one or more segments.

[0008] For each of the one or more image frames, a corresponding reference chain of the image frame may start at an intra-frame (I-frame) of the video sequence and end at the image frame among the one or more image frames.

[0009] Determining a timestamp, a position in the bitstream, and a type of an image frame of a video sequence may include generating one or more data structures that store: (i) for each image frame of the video sequence, a corresponding timestamp and a corresponding offset indicating a corresponding position of compressed data of the image frame in the bitstream, and (ii) for each image frame of a particular type of image frames in the video sequence, a corresponding indication of the particular type.

[0010] The processor may be a graphics processing unit (GPU). In some implementations, one or more segments of the bitstream may be decoded by a hardware decoder integrated in the processor.

[0011] One or more indications may include one or more time values, and determining one or more segments of the bitstream includes, for each of the one or more time values, may include: determining a first timestamp among timestamps of image frames of the video sequence that is closest to the time value, such that the first timestamp corresponds to a first image frame in the video sequence; determining a second timestamp of an I-frame of the video sequence, the second timestamp being determined as the I-frame timestamp closest to the first timestamp that is less than or equal to the first timestamp; using the second timestamp to determine a start position of the I-frame in the bitstream among positions of the image frames; using the first timestamp to determine an end position of the first image frame in the bitstream among positions of the image frames; and determining a segment of the bitstream that extends between the start position of the I-frame and the end position of the first image frame.

[0012] The bitstream can be the first bitstream of the first compressed video sequence captured by the first camera, and the method can further include: receiving, by the processor, a second bitstream of a second video sequence captured by a second camera; determining, by the processor, the timestamps, positions, and types of the image frames of the second video sequence by parsing the second bitstream; determining, by the processor, one or more segments of the second bitstream for decoding to extract one or more second image frames of the second video sequence using one or more indications and the timestamps, positions, and types of the image frames of the second video sequence, such that for an image frame among the one or more second image frames, the corresponding segment of the second bitstream represents the corresponding reference chain of the image frame of the second video sequence; decoding, by the processor, the one or more segments of the second bitstream; and training, by the processor, an ML model using the one or more second image frames. The first camera and the second camera may not be synchronized with each other.

[0013] At least two segments of the one or more segments of the first bitstream and the one or more segments of the second bitstream can be decoded in parallel by at least two hardware decoders integrated in the processor.

[0014] In another embodiment, a computing device can include a memory and processing circuitry. The processing circuitry can be configured to receive a bitstream of a video sequence and one or more indications indicating one or more image frames of the video sequence for training a machine learning (ML) model; determine the timestamps, positions, and types of the image frames of the video sequence by parsing the bitstream; determine one or more segments of the bitstream for decoding to extract one or more image frames using the one or more indications and the timestamps, positions, and types of the image frames of the video sequence, such that for an image frame among the one or more image frames, the corresponding segment represents the corresponding reference chain of the image frame of the video sequence; decode the one or more segments of the bitstream; and train the ML model using the one or more image frames.

[0015] The processing circuitry can further be configured to allocate a memory region within the memory; store the bitstream in the allocated memory region; and after decoding the one or more segments, store the one or more image frames in the allocated memory region.

[0016] For each image frame among the one or more image frames, the corresponding reference chain of the image frame can start at an intra-frame (I-frame) of the video sequence and end at the image frame among the one or more image frames.

[0017] When determining the timestamps, positions, and types of the image frames of a video sequence, the processing circuitry may be configured to generate one or more data structures. The one or more data structures may store: (i) for each image frame of the video sequence, a corresponding timestamp, a corresponding offset indicating the corresponding position of the compressed data of the image frame in the bitstream, and (ii) for each image frame of a particular type of image frames in the video sequence, a corresponding indication of that particular type.

[0018] The computing device may be a graphics processing unit (GPU). One or more segments of the bitstream may be decoded by a hardware decoder integrated in the GPU.

[0019] One or more indications may include one or more time values, and when determining one or more segments of the bitstream, the processing circuitry may be configured, for each time value of the one or more time values: determine a first timestamp among the timestamps of the image frames of the video sequence that is closest to the time value, the first timestamp corresponding to a first image frame in the video sequence; determine a second timestamp of an I-frame of the video sequence, the second timestamp being determined as the I-frame timestamp closest to the first timestamp that is less than or equal to the first timestamp; use the second timestamp to determine the start position of the I-frame in the bitstream among the positions of the image frames; use the first timestamp to determine the end position of the first image frame in the bitstream among the positions of the image frames; and determine the segment of the bitstream that extends between the start position of the I-frame and the end position of the first image frame.

[0020] The bitstream may be a first bitstream of a first compressed video sequence captured by a first camera, and the processing circuitry may also be configured to receive a second bitstream of a second video sequence captured by a second camera; determine the timestamps, positions, and types of the image frames of the second video sequence by parsing the second bitstream; use the one or more indications and the timestamps, positions, and types of the image frames of the second video sequence to determine one or more segments of the second bitstream for decoding to extract one or more second image frames of the second video sequence, for which the corresponding segments of the second bitstream represent the corresponding reference chains of the image frames of the second video sequence; decode the one or more segments of the second bitstream; and use the one or more second image frames to train an ML model. The first camera and the second camera may be out of sync with each other.

[0021] At least two segments of the one or more segments of the first bitstream and the one or more segments of the second bitstream may be decoded in parallel by at least two hardware decoders integrated in the computing device.

[0022] In yet another embodiment, a non-transitory computer-readable medium may store computer code instructions thereon. The computer code instructions, when executed by a processor, may cause the processor to: receive a bitstream of a video sequence and one or more indications indicative of one or more image frames of the video sequence for training a machine learning (ML) model; determine, by parsing the bitstream, timestamps, positions, and types of the image frames of the video sequence; use the one or more indications and the timestamps, positions, and types of the image frames of the video sequence to determine one or more segments of the bitstream for decoding to extract one or more image frames, for which the corresponding segments represent corresponding reference chains of the image frames of the video sequence; decode the one or more segments of the bitstream; and use the one or more image frames to train the ML model. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Non-limiting embodiments of the present disclosure are described by way of examples in connection with the accompanying drawings, which are schematic and not intended to be drawn to scale. Unless indicated as representing the background art, the drawings represent aspects of the present disclosure.

[0024] Figure 1A Illustrates components of an AI-enabled visual data analysis system according to an embodiment.

[0025] Figure 1B Illustrates various sensors associated with a vehicle according to an embodiment.

[0026] Figure 1C Illustrates components of a vehicle according to an embodiment.

[0027] Figure 2 Illustrates a block diagram of a video training system according to an embodiment.

[0028] Figure 3 Illustrates a flowchart of a method for accelerated training of a machine learning (ML) model using video data according to an embodiment.

[0029] Figure 4 Illustrates a diagram depicting a set of selected image frames in a video sequence and corresponding image frames to be decoded according to an embodiment.

[0030] Figure 5 Illustrates according to an embodiment corresponding to Figure 4 a diagram of a bitstream of a video sequence, compressed data corresponding to the selected image frames, and compressed data corresponding to the frames to be decoded. DETAILED DESCRIPTION

[0031] Reference will now be made to the illustrative embodiments depicted in the accompanying drawings, and specific language will be used to describe them. However, it should be understood that no limitation of the scope of the claims or the disclosure is thereby intended. Alterations and further modifications of the features of the invention shown herein, and additional applications of the principles of the subject matter shown herein, which would occur to one skilled in the relevant art and having possession of the present disclosure, are to be considered within the scope of the subject matter disclosed herein. Other embodiments may be utilized and / or other changes may be made without departing from the spirit or scope of the present disclosure. The illustrative embodiments described in the detailed description are not meant to limit the subject matter presented.

[0032] Training an ML model with video data is typically very time-consuming. Additionally, processing (e.g., decoding) a video sequence consumes significant processing power and memory. For an ML model or AI model to be trained using a large amount of video data, training (or validating) the model can be extremely time-consuming and highly demanding in terms of processing, memory, and bandwidth resources. For example, training an ML model or AI model to predict or perceive the surrounding environment of an ego involves using millions or even billions of video frames for training data. Video frames are typically stored as compressed video data. Decoding and processing such a large amount of video data to train the (multiple) ML models can take thousands of hours. The systems, devices, and methods described herein provide accelerated training and more efficient use of computing and memory resources. Specifically, the systems, devices, and methods described herein enable accelerated training and / or validation of occupancy networks configured to predict or perceive the three-dimensional surrounding environment of an ego.

[0033] Figure 1A is a non-limiting example of a component of a system in which the methods and systems discussed herein can be implemented. For example, an analysis server can train an AI model and use the trained AI model to generate an occupancy dataset and / or map for one or more egos. Figure 1A Illustrates the components of an AI-enabled visual data analysis system 100. The system 100 can include an analysis server 110a, a system database 110b, an administrator computing device 120, egos 140a to 140b (collectively referred to as (multiple) egos 140), ego computing devices 141a to 141c (collectively referred to as ego computing devices 141), and a server 160. The system 100 is not limited to the components described herein and can include additional or other components not shown for the sake of brevity, which will be considered within the scope of the embodiments described herein.

[0034] The above components can be connected via network 130. Examples of network 130 can include, but are not limited to, private or public LANs, WLANs, MANs, WANs, and the Internet. Network 130 can include wired and / or wireless communication according to one or more standards and / or via one or more transmission media.

[0035] Communication over network 130 can be performed according to various communication protocols such as Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), and IEEE communication protocols. In one example, network 130 can include wireless communication according to a Bluetooth specification set or another standard or proprietary wireless communication protocol. In another example, network 130 can also include communication over a cellular network, including, for example, GSM (Global System for Mobile Communications), CDMA (Code Division Multiple Access), or EDGE (Enhanced Data for Global Evolution) networks.

[0036] System 100 illustrates an example of a system architecture and components that can be used to train and execute one or more AI models, such as (multiple) AI models 110c. Specifically, as Figure 1A depicted and described herein, analysis server 110a can use the methods discussed herein to train (multiple) AI models 110c using data retrieved from the selfs 140 (e.g., by using data streams 172 and 174). When (multiple) AI models 110c have been trained, each self in selfs 140 can access and execute (multiple) trained AI models 110c. For example, vehicle 140a having a self computing device 141a can transmit its camera feed to (multiple) trained AI models 110c and can determine the occupancy status of its surrounding environment (e.g., data stream 174). Additionally, data ingested and / or predicted by (multiple) AI models 110c with respect to selfs 140 (at inference time) can also be used to improve (multiple) AI models 110c. Thus, system 100 depicts a continuous loop that can periodically improve the accuracy of (multiple) AI models 110c. Additionally, system 100 depicts a loop in which data received by selfs 140 can also be used in the training phase in addition to the inference phase.

[0037] The analysis server 110a can be configured to collect, process, and analyze navigation data (e.g., images captured during navigation) and various sensor data collected from the vehicle 140. The collected data can then be processed and prepared into a training dataset. The training dataset can then be used to train one or more AI models, such as the AI model 110c. The analysis server 110a can also be configured to collect visual data from the vehicle 140. Using the AI model 110c (trained using the methods and systems discussed herein), the analysis server 110a can generate a dataset and / or an occupancy map for the vehicle 140. The analysis server 110a can display the occupancy map on the vehicle 140 and / or transmit the occupancy map / dataset to the vehicle computing device 141, the administrator computing device 120, and / or the server 160.

[0038] In Figure 1A , the AI model 110c is illustrated as a component of the system database 110b, but the AI model 110c can be stored in a different or separate component, such as a cloud storage device or any other data repository accessible to the analysis server 110a.

[0039] The analysis server 110a can also be configured to display an electronic platform that illustrates various training attributes for training the AI model 110c. The electronic platform can be displayed on the administrator computing device 120 such that an analyst can monitor the training of the AI model 110c. An example of an electronic platform generated and hosted by the analysis server 110a can be a web-based application or website configured to display the training dataset collected from the vehicle 140 and / or the training status / metrics of the AI model 110c.

[0040] The analysis server 110a can be any computing device that includes a processor and a non-transitory machine-readable storage device capable of performing the various tasks and processes described herein. Non-limiting examples of such computing devices can include workstation computers, laptop computers, server computers, etc. Although the system 100 includes a single analysis server 110a, the system 100 can include any number of computing devices operating in a distributed computing environment, such as a cloud environment.

[0041] The vehicle body 140 may represent various electronic data sources that transfer data associated with its previous or current navigation session to the analysis server 110a. The vehicle body 140 may be any device configured for navigation, such as vehicle 140a and / or truck 140c. The vehicle body 140 is not limited to being a vehicle and may also include robotic devices. For example, the vehicle body 140 may include a robot 140b, which may represent a general-purpose, bipedal, autonomous humanoid robot capable of navigating various terrains. The robot 140b may be equipped with software for achieving balance, navigation, perception, or interaction with the physical world. The robot 140b may also include various cameras configured to transfer visual data to the analysis server 110a.

[0042] Although referred to herein as the "vehicle body", the vehicle body 140 may or may not be an autonomous device configured for autonomous navigation. For example, in some embodiments, the vehicle body 140 may be controlled by a human operator or a remote processor. The vehicle body 140 may include various sensors, such as Figure 1B the sensors depicted in. The sensors may be configured to collect data as the vehicle body 140 navigates various terrains (e.g., roads). The analysis server 110a may collect the data provided by the vehicle body 140. For example, the analysis server 110a may obtain navigation session and / or road / terrain data (e.g., an image of the vehicle body 140 navigating on a road) from various sensors, such that the collected data is ultimately used by the AI model 110c for training purposes.

[0043] As used herein, a navigation session corresponds to the journey of the vehicle body 140's travel route, regardless of whether the journey is autonomous or human-controlled. In some embodiments, the navigation session may be used for data collection and model training purposes. However, in some other embodiments, the vehicle body 140 may refer to a vehicle purchased by a consumer, and the purpose of the journey may be classified as daily use. When the vehicle body 140 moves from a non-moving position by more than a threshold distance (e.g., 0.1 mile, 100 feet) or at a rate exceeding a threshold (e.g., exceeding 0 mph, exceeding 1 mph, exceeding 5 mph), the navigation session may begin. When the vehicle body 140 returns to a non-moving position and / or is turned off (e.g., when the driver leaves the vehicle), the navigation session may end.

[0044] The self 140 can represent a collection of selves monitored by the analytics server 110a for training the AI model(s) 110c. For example, a driver of the vehicle 140a can authorize the analytics server 110a to monitor data associated with their respective vehicle. As a result, the analytics server 110a can use the various methods discussed herein to collect sensor / camera data and generate a training dataset to train the AI model(s) 110c accordingly. The analytics server 110a can then apply the trained AI model(s) 110c to analyze data associated with the self 140 and predict an occupancy map for the self 140. Additionally, additional / ongoing data associated with the self 140 can also be processed and added to the training dataset so that the analytics server 110a can recalibrate the AI model(s) 110c accordingly. Thus, the system 100 depicts a loop in which navigation data received from the self 140 can be used to train the AI model(s) 110c. The self 140 can include a processor that executes the trained AI model(s) 110c for navigation purposes. During navigation, the self 140 can collect additional data about its navigation session, and this additional data can be used to calibrate the AI model(s) 110c. That is, the self 140 represents a self that can be used to train, execute / use, and recalibrate the AI model(s) 110c. In a non-limiting example, the self 140 represents vehicles purchased by customers that can use the AI model(s) 110c for autonomous navigation while improving the AI model(s) 110c.

[0045] The self 140 can be equipped with various technologies that allow the self to collect data from its surrounding environment and (possibly) navigate autonomously. For example, the self 140 can be equipped with an inference chip to run autonomous driving software.

[0046] Various sensors for each self 140 can monitor the data collected associated with different navigation sessions and transmit this data to the analytics server 110a. Figures 1B to 1C A block diagram of sensors integrated within the self 140 according to an embodiment is illustrated. The number and location of each sensor discussed with respect to Figures 1B to 1C can depend on the type of self discussed in Figure 1A For example, the robot 140b can include different sensors than the vehicle 140a or the truck 140c. For example, the robot 140b may not include the airbag activation sensor 170q. Additionally, the sensors of the vehicle 140a and the truck 140c can be positioned differently than shown in Figure 1C

[0047] As discussed herein, the various sensors integrated within each ego body 140 can be configured to measure various data associated with each navigation session. The analytics server 110a can periodically collect the data monitored and collected by these sensors, where the data is processed according to the methods described herein and is used to train the AI model 110c and / or execute the AI model 110c to generate an occupancy map.

[0048] The ego body 140 can include a user interface 170a. The user interface 170a can refer to the user interface of an ego computing device (e.g., Figure 1A the ego computing device 141 in Figure 1B ). The user interface 170a can be implemented as a display screen, a head-up display, a touch screen, etc. integrated with or coupled to the interior of the vehicle. The user interface 170a can include input devices such as a touch screen, a knob, a button, a keyboard, a mouse, a gesture sensor, a steering wheel, etc. In various embodiments, the user interface 170a can be adapted to provide user input (e.g., as a kind of signal and / or sensor information) to other devices or sensors of the ego body 140 (e.g.,

[0049] the sensors shown

[0050] such as the controller 170c). The user interface 170a can also be implemented with one or more logic devices that can be adapted to execute instructions, such as software instructions, to implement any of the various processes and / or methods described herein. For example, the user interface 170a can be adapted to form a communication link, transmit and / or receive communications (e.g., sensor signals, control signals, sensor information, user input, and / or other information) or execute various other processes and / or methods. In another example, the driver can use the user interface 170a to control the temperature of the ego body 140 or activate its features (e.g., the autonomous driving or steering system 170o). Thus, the user interface 170a can, in combination with other sensors described herein, monitor and collect driving session data. The user interface 170a can also be configured to display various data generated / predicted by the analytics server 110a and / or the AI model 110c.

[0050] The orientation sensor 170b can be implemented as a compass, a float, an accelerometer, and / or one or more of other digital or analog devices capable of measuring the orientation of the ego body 140 (e.g., the magnitude and direction of roll, pitch, and / or yaw relative to one or more reference orientations such as gravity and / or magnetic north). The orientation sensor 170b can be adapted to provide heading measurements to the ego body 140. In other embodiments, the orientation sensor 170b can be adapted to provide roll, pitch, and / or yaw rates to the ego body 140 using a time series of orientation measurements. The orientation sensor 170b can be positioned and / or adapted to make orientation measurements relative to a particular coordinate system of the ego body 140.

[0051] The controller 170c can be implemented as any suitable logic device (e.g., a processing device, a microcontroller, a processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a memory storage device, a memory reader, or other device or combination of devices) that is adapted to execute, store, and / or receive appropriate instructions, such as software instructions that implement control loops for controlling various operations of the self 140. Such software instructions can also implement methods for processing sensor signals, determining sensor information, providing user feedback (e.g., via the user interface 170a), querying the device for operating parameters, selecting operating parameters for the device, or performing any of the various operations described herein.

[0052] The communication module 170e can be implemented as any wired and / or wireless interface configured to transmit sensor data, configuration data, parameters, and / or other data and / or signals to Figure 1A any of the features shown (e.g., the analytics server 110a). As described herein, in some embodiments, the communication module 170e can be implemented in a distributed manner such that portions of the communication module 170e are implemented within Figure 1B one or more of the elements and sensors shown. In some embodiments, the communication module 170e can delay the transmission of sensor data. For example, when the self 140 does not have a network connection, the communication module 170e can store the sensor data in a temporary data storage device and transmit the sensor data when the self 140 is identified as having an appropriate network connection.

[0053] The speed sensor 170d can be implemented as an electronic pitot tube, a metering gear or wheel, a water speed sensor, a wind speed sensor, a wind rate sensor (e.g., direction and magnitude), and / or other devices capable of measuring or determining the linear speed of the self 140 (e.g., in the surrounding medium and / or aligned with the longitudinal axis of the self 140) and providing such measurement as a sensor signal that can be transmitted to various devices.

[0054] The gyroscope / accelerometer 170f can be implemented as one or more electronic sextants, semiconductor devices, integrated chips, accelerometer sensors, or other systems or devices capable of measuring the angular velocity / acceleration and / or linear acceleration of the self 140 (e.g., direction and magnitude) and providing such measurement as a sensor signal that can be transmitted to various devices, such as the analytics server 110a. The gyroscope / accelerometer 170f can be positioned and / or adapted to make such measurements relative to a particular coordinate system of the self 140. In various embodiments, the gyroscope / accelerometer 170f can be associated with Figure 1BThe other elements depicted are implemented in a common housing and / or module to ensure a common reference system or known transformations between reference systems.

[0055] The Global Navigation Satellite System (GNSS) 170h may be implemented as a Global Positioning Satellite receiver and / or another device capable of determining the absolute and / or relative position of the vehicle body 140 based on wireless signals received, for example, from space and / or ground sources and capable of providing measurements such as sensor signals (which may be transmitted to various devices). In some embodiments, the GNSS 170h may be adapted to determine the rate, speed, and / or yaw rate of the vehicle body 140 (e.g., using a time series of position measurements), such as the yaw component of the absolute rate and / or angular rate of the vehicle body 140.

[0056] The temperature sensor 170i may be implemented as a thermistor, an electrical sensor, an electrical thermometer, and / or other devices capable of measuring the temperature associated with the vehicle body 140 and providing such measurement as a sensor signal. The temperature sensor 170i may be configured to measure the ambient temperature associated with the vehicle body 140, such as the cockpit or dashboard temperature, for example, which may be used to estimate the temperature of one or more elements of the vehicle body 140.

[0057] The humidity sensor 170j may be implemented as a relative humidity sensor, an electrical sensor, an electrical relative humidity sensor, and / or another device capable of measuring the relative humidity associated with the vehicle body 140 and providing such measurement as a sensor signal.

[0058] The steering sensor 170g may be adapted to physically adjust the heading of the vehicle body 140 according to one or more control signals provided by a logic device (such as the controller 170c) and / or user input. The steering sensor 170g may include one or more actuators and control surfaces of the vehicle body 140 (e.g., a rudder or other type of steering or trim mechanism) and may be adapted to physically adjust the control surface to various positive and / or negative steering angles / positions. The steering sensor 170g may also be adapted to sense the current steering angle / position of such a steering mechanism and provide such measurement.

[0059] The propulsion system 170k may be implemented as a propeller, a turbine, or other thrust-based propulsion system, a mechanical wheeled and / or tracked propulsion system, a wind / sail-based propulsion system, and / or other types of propulsion systems that may be used to power the vehicle body 140. The propulsion system 170k may also monitor the direction of the motive force and / or thrust of the vehicle body 140 relative to the reference coordinate system of the vehicle body 140. In some embodiments, the propulsion system 170k may be coupled to the sensor 170g and / or integrated with the steering sensor 170g.

[0060] The passenger restraint sensor 170l can monitor seat belt detection and locking / unlocking components, as well as other passenger restraint subsystems. The passenger restraint sensor 170l can include various environmental and / or status sensors, actuators, and / or other devices that facilitate the operation of safety mechanisms associated with the operation of the vehicle 140. For example, the passenger restraint sensor 170l can be configured to receive motion and / or status data from Figure 1B the other sensors depicted. The passenger restraint sensor 170l can determine whether a safety measure (e.g., seat belt) is being used.

[0061] As Figure 1C depicted, the camera 170m can refer to one or more cameras integrated within the vehicle 140, and can include multiple cameras integrated (or retrofitted) into the vehicle 140. The camera 170m can be an internal or external-facing camera of the vehicle 140. For example, as Figure 1C depicted, the vehicle 140 can include one or more internal-facing cameras that can monitor and collect footage of the passengers of the vehicle 140. The vehicle 140 can include eight external-facing cameras. For example, the vehicle 140 can include a front camera 170m-1, front side cameras 170m-2, 170m-3, rear side cameras 170m-4 on each front fender, cameras 170m-5 on each side (e.g., integrated within the B-pillar), and a rear camera 170m-6.

[0062] Referring to Figure 1B , the radar 170n and ultrasonic sensors 170p can be configured to monitor the distance of the vehicle 140 to other objects, such as other vehicles or immovable objects (e.g., trees or garage doors). The vehicle 140 can also include an autonomous driving or steering system 170o configured to self-navigate the vehicle 140 using data collected via various sensors (e.g., radar 170n, speed sensors 170d, and / or ultrasonic sensors 170p).

[0063] Thus, the autonomous driving or steering system 170o can analyze various data collected by one or more of the sensors described herein to identify driving data. For example, the autonomous driving or steering system 170o can calculate the risk of a forward collision based on the speed of the vehicle 140 and its distance to another vehicle on the road. The autonomous driving or steering system 170o can also determine whether the driver is touching the steering wheel. The autonomous driving or steering system 170o can transmit the analyzed data to various features discussed herein, such as an analysis server.

[0064] The airbag activation sensor 170q can predict or detect a collision and cause the activation or deployment of one or more airbags. The airbag activation sensor 170q can transmit data regarding the airbag deployment, including data associated with the event that caused the deployment.

[0065] Referring again to Figure 1A , the administrator computing device 120 can represent a computing device operated by a system administrator. The administrator computing device 120 can be configured to display data retrieved or generated by the analytics server 110a (e.g., various analytics metrics and risk scores), where the system administrator can monitor the various models utilized by the analytics server 110a, review feedback, and / or facilitate the training of the (multiple) AI models 110c maintained by the analytics server 110a.

[0066] (Multiple) agents 140 can be any device configured to navigate various routes, such as a vehicle 140a or a robot 140b. As discussed with respect to Figures 1B to 1C , the agent 140 can include various telemetry sensors. The agent 140 can also include an agent computing device 141. Specifically, each agent can have its own agent computing device 141. For example, a truck 140c can have an agent computing device 141c. For the sake of brevity, the agent computing devices are collectively referred to as the (multiple) agent computing devices 141. The agent computing device 141 can control the content presentation on the infotainment system of the agent 140, process commands associated with the infotainment system, aggregate sensor data, manage the communication of data to electronic data sources, receive updates, and / or transmit messages. In one configuration, the agent computing device 141 communicates with an electronic control unit. In another configuration, the agent computing device 141 is an electronic control unit. The agent computing device 141 can include a processor and a non-transitory machine-readable storage medium capable of performing the various tasks and processes described herein. For example, the (multiple) AI models 110c described herein can be stored and executed (or directly accessed) by the agent computing device 141. Non-limiting examples of the agent computing device 141 can include vehicle multimedia and / or display systems.

[0067] In one example of how video data can be used to accelerate the training of one or more AI models 110c and / or other ML models, the analytics server 110a can include multiple graphics processing units (GPUs), which are configured to train one or more AI models 110c in parallel. For example, the analytics server 110a can include a supercomputer. Each GPU can receive video data (e.g., one or more bitstreams) and an indication of video frames (or image frames) to be extracted from the video data and used to train one or more AI models 110c. The GPU can decode only the portion of the bitstream(s) required to decode the selected image frames and use the decoded form of the selected image frames to train one or more AI models 110c.

[0068] In some implementations, each GPU can be configured or designed to perform the training steps independently without using external resources. Specifically, the GPU can receive video data, decode the relevant portion or segments to extract the selected image frames, extract features from the selected image frames, and use the extracted features to train one or more AI models 110c without using any external memory or processing resources. In other words, all processing and data handling from the point of receiving the video data to the training of one or more AI models 110c can be performed internally and independently within the GPU.

[0069] Each GPU can include one or more hardware video decoders integrated therein to accelerate video decoding. The GPU can include multiple video decoders to parallelize video decoding. The parallelization can be achieved in various ways, such as by video segment, by bitstream, or by training session.

[0070] In some implementations, the GPU can have sufficient internal memory (e.g., cache memory) to store the data required to perform the training steps. The GPU can allocate a memory region within the memory to store data associated with the training steps and use the allocated memory region throughout the training steps for data storage.

[0071] Figure 2 A block diagram of a computer environment 200 for training an ML model according to an embodiment is illustrated. The computer environment 200 can include a training system 202 for training an ML model and a data storage system 204 for storing training data and / or validation data. The training system 202 can include multiple training nodes (or processing nodes) 206. Each training node 206 can include a corresponding data loader (or data loading device) 208 and a corresponding graphics processing unit (GPU) 210. Each GPU 210 can include a memory (e.g., cache memory) 212, processing circuitry 214, and one or more video decoders 216.

[0072] The data storage system 204 can include, or can be, a distributed storage system. For example, the data storage system 204 can have an infrastructure that can split data across multiple physical servers, such as supercomputers. The data storage system 204 can include one or more storage clusters of storage units, with mechanisms and infrastructure for parallel and accelerated access to data from multiple nodes or storage units of the (multiple) storage clusters. For example, the data storage system 204 can include sufficient data links and bandwidth to deliver data to the training nodes 206 in parallel or simultaneously.

[0073] The data storage system 204 can include sufficient memory capacity to store, for example, millions or even billions of video frames in a compressed form. For example, the data storage system 204 can have a memory capacity to store data of multiple petabytes (e.g., 10, 20, or 30 petabytes). The data storage system 204 can allow thousands of video sequences to move into and / or out of the data storage system 204 at any time instance. The relatively large size and bandwidth of the storage system 204 allow for parallel training of one or more ML models, as discussed below.

[0074] The training system 202 can be implemented as one or more physical servers, such as server 110a. For example, the training system 202 can be implemented as one or more supercomputers. Each supercomputer can include thousands of processing or training nodes 206. The training nodes 206 can be configured or designed to support parallel training of one or more ML models, such as the (multiple) AI models 110c. Each training node 206 can be communicatively coupled to the storage system 204 to access training data and / or validation data stored therein.

[0075] Each training node 206 can include a corresponding data loader 208 and a corresponding GPU 210 communicatively coupled to each other. The data loader 208 can be (or can include) a processor or a central processing unit (CPU) for handling data requests or data transfers between the corresponding GPU 210 (e.g., the GPU 210 in the same training node 206) and the data storage system 204. For example, the GPU 210 can request data by Figure 1COne or more video sequences captured by one or more of the described cameras 170m. For example, the front or forward cameras 170m-1, 170m-2, and 170m-3, the rear side camera 170m-4, the side camera 170m-5, and the rear camera 170m-6 can simultaneously capture video sequences and send the video sequences to the data storage system 204 and store the video sequences in the data storage system 204 for storage. In some implementations, the data storage system 204 can store the video sequences simultaneously captured by multiple cameras of the vehicle body 140 (such as the cameras 170m) as a set of video sequences or a combination of video sequences, and these video sequences can be delivered together to the training node 206. For example, the data storage system 204 can maintain additional data that indicates which video sequences are simultaneously captured by the cameras 170m, or represents the same scene from different camera angles. The data storage system 204 can maintain (e.g., for each stored video sequence) data indicating the vehicle body identifier, camera identifier, and time instance associated with the video sequence.

[0076] The training node 206 can use the video data captured by the cameras 170m of the vehicle body 140 and stored in the data storage system 204 to simultaneously (e.g., in parallel) train one or more ML models. In some implementations, the video data can be captured by the cameras 170m of multiple vehicle bodies 140. In the training node 206, the corresponding data loader 208 can request from the data storage system 204 the video data of one or more video sequences simultaneously captured by one or more cameras 170m of the vehicle body 140 during a time interval, and send the received video data to the corresponding GPU 210 for use in performing training steps (or validation steps) when training the ML model. The video data can be in a compressed form. For example, the video sequence can be encoded by an encoder integrated in or implemented in the vehicle body 140. Each data loader 208 can have sufficient processing power and bandwidth to deliver the video data to the corresponding GPU 210 in a manner that keeps the GPU 210 busy. In other words, the GPU 210 can be configured or designed to, for example, in terms of processing power and bandwidth, request the video data of a set of compressed video sequences from the data storage system 204, and deliver the video to the GPU 210 within a duration less than or equal to the average time consumed by the GPU 210 to process a set of video sequences.

[0077] Each GPU 210 may include a corresponding internal memory 212, such as a cache memory, to store executable instructions for performing the processes described herein, received video data of one or more video sequences, decoded video frames, features extracted from the decoded video frames, data of parameters or trained ML models, and / or other data used for training the ML models. The memory 212 may be large enough to store all data required to perform a single training step. As used herein, a training step may include receiving and decoding video data of one or more video sequences (e.g., a set of video sequences simultaneously captured by one or more cameras 170m of the vehicle 140), extracting features from the decoded video data, and using the extracted features to update the parameters of the ML model being trained or validated.

[0078] Each GPU 210 may include processing circuitry 214 to perform the processes or methods described herein. The processing circuitry 214 may include one or more microprocessors, multi-core processors, digital signal processors (DSPs), one or more logic circuits, or combinations thereof. The processing circuitry 214 may execute computer code instructions (e.g., stored in the memory 212) to perform the processes or methods described herein.

[0079] The GPU 210 may include one or more video decoders 216 for decoding video data received from the data storage system 204. The one or more video decoders 216 may include (a) hardware video decoder(s) integrated in the GPU 110 to accelerate video decoding. The one or more video decoders 216 may be part of the processing circuitry 214, or may include separate (a) electronic circuit(s) communicatively coupled to the processing circuitry 214.

[0080] Each GPU 210 may be configured or designed to process or perform a training step without using any external resources. The GPU 210 may include sufficient memory capacity and processing power to perform the training step. The processes performed by the training node 208 or the corresponding GPU 210 will be described in further detail below in connection with Figure 3 FIGs. 6.

[0081] Now refer to Figure 3, According to an embodiment, a flowchart of method 300 for accelerating the training of a machine learning (ML) model using video data. Briefly, method 300 may include receiving a bitstream of a video sequence and one or more indications indicating one or more image frames of the video sequence for training a machine learning (ML) model (step 302), and determining timestamps, positions within the bitstream, and types of the image frames of the video sequence by parsing the bitstream (step 304). Method 300 may include using the one or more indications and the timestamps, positions within the bitstream, and types of the image frames of the video sequence to determine one or more segments of the bitstream for decoding to extract one or more image frames (step 306). For an image frame among the one or more image frames, the corresponding segment may represent the corresponding reference chain of the image frame of the video sequence. Method 300 may include decoding the one or more segments of the bitstream (step 308), and training the ML model using the one or more image frames (step 310).

[0082] Method 300 may be fully implemented, executed, or carried out by GPU 210. GPU 210 may execute steps 302 to 310 without using any external memory or processing resources. Having method 300 fully executed by a single GPU 210 results in accelerated training of the (multiple) ML models. Specifically, by fully executing method 300 within a single GPU 210, the processing time can be reduced by avoiding data exchange between GPU 210 and any external resources. For example, decoding video data outside of GPU 210 or storing the decoded video data may introduce latency associated with the exchange of compressed and / or decoded video data between GPU 210 and any external resources.

[0083] Method 300 may include the GPU 210 receiving a bitstream of a video sequence and one or more indications indicating one or more image frames of the video sequence for training a machine learning (ML) model (step 302). In some implementations, the GPU 210 or the processing circuitry 214 may allocate a memory region within the memory 212 for performing the training steps. For example, before receiving the bitstream and the one or more indications, the processing circuitry 214 may allocate a memory region to store data for the next training step, such as compressed video data, decoded video, or image frames received from the storage system 204, features extracted from the video frames, and / or other data. In some implementations, the processing circuitry 214 may allocate a memory region at the start of each training step, or may allocate a memory region at the start of a training session, where the allocated memory region may be used for successive training steps. In some implementations, the processing circuitry 214 may overwrite segments of data stored in the allocated memory region (or the memory 212) that are no longer needed to efficiently utilize the memory 212.

[0084] The GPU 210 may receive one or more bitstreams of one or more video sequences from the storage system 204 via the loader 208. For example, the GPU 210 may receive multiple bitstreams of multiple compressed video sequences (e.g., simultaneously captured by the cameras 170m of the ego body 140). The compressed video sequences may be encoded at the ego body 140 and may not be synchronized. For example, each camera 170m may have a separate timeline (according to which image frames are captured) and a separate encoder for encoding the captured image frames. The image capture time instances for different cameras 170m may not be time-aligned. Additionally, the encoders associated with different cameras 170m may not be synchronized. Thus, when image frames captured by different cameras 170m at the same moment (or substantially at the same moment, considering any differences in the timelines for capturing image frames by different cameras 170m) are encoded into different bitstreams by separate encoders, the captured image frames may have different timestamps.

[0085] The GPU 210 can receive one or more indications that indicate image frames selected or to be selected from one or more video sequences for training an ML model. The one or more indicators can be specified by a user of the training system 202 and received by the GPU 210 as input (e.g., via the data loader 208). The one or more indications can include one or more time values. Each time value can indicate a separate image frame in each received bitstream. For example, if eight bitstreams of eight video sequences captured by eight different cameras 170m of the ego body 140 are received by the GPU 210, each time value will indicate eight image frames, e.g., the image frames in each video sequence. The time value can indicate the timestamp of the selected or to-be-selected image frame, but not necessarily be exactly equal to these timestamps. The GPU 210 can store the received (multiple) bitstreams and the one or more indications in an allocated memory region.

[0086] Method 300 can include the GPU 210 determining the timestamp, position, and type of the image frames of the video sequence by parsing the bitstream (step 304). Each video bitstream can include multiple headers distributed across the bitstream. The video bitstream can include a separate header for each compressed image frame in the bitstream. Each image frame header can be immediately before the compressed data of the image frame and can include data indicating information about the image frame and the corresponding compressed data. The header can include the timestamp of the image frame, the type of the image frame, the size of the compressed image frame data, the position of the compressed image frame data in the bitstream, and / or other data.

[0087] The type of the image frame can include an intra-frame (I-frame) type or a predicted frame (P-frame) type. The P-frame is encoded using data from another previously encoded image frame. To decode a P-frame, the decoder will decode any other image frame that the P-frame depends on before decoding the P-frame. The I-frame is also referred to as a reference frame and can be decoded independently of any other image frame. For an image frame, the corresponding timestamp represents the time of presentation of the image frame, e.g., relative to the first image frame in the video sequence. The position of the compressed image frame data in the bitstream can be an offset value, e.g., in bytes, which indicates where the compressed image frame data starts in the bitstream.

[0088] The processing circuitry 214 can parse each received bitstream to identify or determine the headers of each compressed image frame in the bitstream. The processing circuitry 214 can read the headers in each bitstream to determine the timestamp, the location, and the type of each image frame in each bitstream. Parsing the (multiple) bitstreams is much easier in terms of processing power and processing time compared to decoding a single bitstream or a portion thereof. The timestamps, the locations within the bitstreams, and the types of the image frames determined by parsing the (multiple) bitstreams enable a significant reduction in the amount of video data to be decoded to extract the (multiple) selected image frames, and thus significantly accelerate the training process.

[0089] When determining the timestamps, the locations in the bitstreams, and the types of the image frames of the (multiple) video sequences, the GPU 210 or the processing circuitry 214 can generate one or more data structures to store the timestamps, the locations in the bitstreams, and the types of the image frames of the (multiple) video sequences. The one or more data structures can include tables, data files, and / or some other types of data structures. For example, the processing circuitry 214 can generate a separate table for each bitstream, similar to Table 1 below. The leftmost first column of Table 1 can include the timestamps of all the image frames in the video sequence, e.g., in ascending order, the second column can include the location (or offset) of the compressed frame data of each image frame, and the rightmost column can include the type of each image frame. Each row of Table 1 corresponds to a separate image frame in the video sequence.

[0090]

[0091]

[0092] Table 1

[0093] In some implementations, the one or more data structures can include a first data structure for I-frames and a second data structure for all the frames in the video sequence. Each of the first and second data structures can be a table or a data file. An example of the first data structure can be Table 2 below. Each row of Table 2 represents a separate I-frame in the video sequence. The first column (e.g., the leftmost) can include the timestamp of the I-frame (e.g., in ascending order), and the second column can include the corresponding location or offset (e.g., in bytes). The data in Table 2 allows for a quick determination of the location of the compressed data for any I-frame in the bitstream.

[0094]

[0095] Table 2

[0096] An example of the first data structure can be Table 3 below. Each row of Table 3 represents a separate image frame in the video sequence. The first column (e.g., the leftmost) can include the timestamp of the image frame (e.g., in ascending order), and the second column can include the corresponding position or offset (e.g., in bytes). The data in Table 3 allows determination of the position of the compressed data for any image frame in the bitstream.

[0097]

[0098]

[0099] Table 3

[0100] The data of any table in the above table can be stored in a data file. Compared with Table 1, Table 2 and Table 3 may not include an indication of the image frame type. Instead, I-frames can be identified from Table 2 specific to I-frames. In some implementations, other data structures can be generated (e.g., instead of or in combination with any of Tables 1, 2, and / or 3). Once generated, the GPU 210 or the processing circuitry 214 can store one or more data structures in the memory 212. One or more data structures can indicate: (i) for each image frame of the video sequence, the corresponding timestamp and corresponding offset indicating the corresponding position of the compressed data of the image frame in the bitstream, and (ii) the specific type of image frame in the video sequence.

[0101] Method 300 can include the GPU 210 using one or more indications and the timestamp, position, and type of the image frames of the video sequence to determine one or more segments of the bitstream for decoding to extract one or more image frames (step 306). One or more indications can include one or more time values for use in identifying or determining selected image frames for use in training the ML model. Each time value can indicate one or more corresponding image frames or (multiple) corresponding timestamps in one or more received bitstreams. If multiple bitstreams corresponding to multiple cameras 170m are received, each indicator or time value can indicate a separate image frame or corresponding timestamp in each of the received bitstreams.

[0102] For each of one or more time values, the GPU 210 or the processing circuitry 214 may determine, for each received bitstream, a corresponding timestamp that is closest to the time value. For example, the processing circuitry 214 may use Table 3 (or Table 1) for a given bitstream to determine the timestamp of that bitstream that is closest to the time value. The processing circuitry 214 may determine, for each bitstream, a separate timestamp that is closest to the time value by using a corresponding data structure (e.g., Table 3 or Table 1). Each determined timestamp indicates a corresponding image frame in the corresponding bitstream indicated or selected via the time value. For example, if eight bitstreams corresponding to eight cameras 170m are received, the processing circuitry 214 may determine, for each time value (or indicator), eight corresponding timestamps that indicate eight selected image frames, with one selected image frame in each bitstream. If the GPU 210 receives three indicators and eight bitstreams corresponding to eight cameras 170m, the processing circuitry 214 may determine a total of 24 timestamps that correspond to 24 selected image frames, where three timestamps corresponding to three selected image frames are determined for each bitstream.

[0103] The processing circuitry 214 may determine, for each time value (or each indicator), a separate second timestamp for each received bitstream. For each bitstream, the second timestamp indicates the corresponding I-frame of the bitstream. For each time value, the processing circuitry 214 may determine the second timestamp for a given bitstream as the I-frame timestamp of that bitstream that is closest to the time value (or the timestamp of the selected frame of the bitstream indicated by the time value), and the I-frame timestamp is less than or equal to the time value (or less than or equal to the timestamp of the selected frame of the bitstream indicated by the time value). For example, the processing circuitry 214 may use Table 2 (or Table 1) of the bitstream to determine the I-frame timestamp of the bitstream that is closest to a given time value (or indicator) and is less than or equal to the time value (or indicator).

[0104] For a given time value (or indicator) and a given bitstream, if the timestamp of the selected image frame and the corresponding I-frame timestamp are equal, it means that the selected image frame (or the image frame indicated by the time value) is an I-frame. However, if the two timestamps are different, the processing circuitry 214 may determine that the selected image frame (or the image frame indicated by the time value) is not an I-frame.

[0105] For each time value (or indicator) and each received bitstream, the processing circuitry 214 can use the I-frame timestamp and one or more data structures to determine the position or offset (e.g., starting position) of the I-frame in the bitstream. For example, the processing circuitry 214 can use Table 2 (or Table 1) to determine the position or offset corresponding to the determined I-frame timestamp. For example, if the determined I-frame timestamp is T 3,I , the corresponding offset is O 3,I .

[0106] The processing circuitry 214 can, for each time value (or indicator) and each received bitstream, use the timestamp of the corresponding selected image frame to determine the end position of the corresponding selected image frame. The processing circuitry 214 can determine the end position of the selected image frame in the bitstream as the starting position of the next (or subsequent) image frame in the bitstream. The processing circuitry 214 can determine the end position of the selected image frame in the bitstream as the starting position of the selected image frame plus the size of the compressed data of the bitstream. The processing circuitry 214 can determine the size of each image frame in the bitstream by parsing the bitstream or frame header and recording the size in one or more data structures.

[0107] The processing circuitry 214 can, for each time value (or indicator) and each received bitstream, determine the corresponding segment of the bitstream that extends between the starting position of the corresponding I-frame and the end position of the corresponding selected image frame. The determined segment represents the minimum amount of compressed video data that is to be decoded to decode the selected image frame. If two or more segments corresponding to different time values (or different indicators) but from the same bitstream overlap, the processing circuitry 214 can consider the longest segment for decoding and ignore the shorter segment(s).

[0108] Note that while the indicator has been described above as a time value, other implementations are possible. For example, the indicator received by the GPU 210 can be an index of an image frame.

[0109] Now refer to Figure 4, according to an embodiment, a diagram depicting a set of selected image frames in video sequence 400 and corresponding image frames to be decoded is shown. Using three different indicators or time values, processing circuitry 214 can determine (e.g., as discussed above) image frames 402, 406, and 412 as selected image frames or image frames indicated to be selected by the indicator or time value. Image frame 402 is an I-frame, while image frames 406 and 412 are P-frames. With respect to selected image frame 402, processing circuitry 214 determines that only image frame 402 is to be decoded, as it is an I-frame and does not depend on any other image frames. For selected image frame 406 (which is a P-frame), processing circuitry 214 determines that the closest previous I-frame is image frame 404, and both frames 404 and 406 are to be decoded in order to obtain the selected image frame 406 in decoded form. For selected image frame 412 (which is a P-frame), processing circuitry 214 determines that the closest previous I-frame is image frame 408, and frames 408, 410, and 412 are to be decoded in order to obtain the selected image frame 412 in decoded form.

[0110] Figure 5 A diagram illustrating, according to an embodiment, Figure 4 a bitstream 500 corresponding to video sequence 400, compressed data corresponding to selected image frames, and compressed data corresponding to frames to be decoded. Compressed data segment 502 represents the compressed data of image frame 402. Compressed data segment 504 represents the compressed data of image frames 404 and 406. Compressed data segment 506 represents the compressed data of image frames 408, 410, and 412. Processing circuitry 214 can feed only compressed data segments 502, 504, and 506 to (a) video decoder(s) 216 for decoding, rather than decoding the entire bitstream 500.

[0111] For each selected image frame (or each image frame indicated for selection), the corresponding compressed data segment to be decoded can be considered to represent the corresponding reference chain of the image frame. The reference chain starts with the closest I-frame before the selected frame and ends with the selected image frame. The reference chain represents a chain of interdependent image frames with dependencies (frame references) that start at the selected image frame and go backward to the first encountered I-frame. For example, in the reference chain formed by image frames 408, 410, and 412, image frame 412 references image frame 410, and the latter references image frame 410, which is an I-frame.

[0112] Referring again to Figure 3, Method 300 may include the GPU 210 decoding one or more segments of the bitstream (step 308). The processing circuitry 214 may provide or feed the determined compressed data segments to the video decoder(s) 216 for decoding. By decoding only the compressed data required to decode the selected image frames, the GPU 210 significantly reduces the processing time and processing power consumed in decoding the selected image frames (e.g., compared to decoding the entire bitstream). The processing circuitry 214 may store the decoded video data in the memory 212. For efficient use of the memory resources in the GPU 210, the video processing circuitry 214 may overwrite the decoded video data that is no longer needed. For example, when decoding the compressed data segment 506 and once the image frame 410 is decoded, the processing circuitry 214 may determine that the data of the decoded image frame 410 is no longer needed and delete the decoded image frame 410 to free up memory space. Once the compressed data segments corresponding to the selected image frames are decoded, only the decoded data for the selected image frames may be retained in the memory 212 or the allocated memory region, while other decoded image frames (non-selected image frames) may be deleted to free up memory space.

[0113] As discussed above, the GPU 210 may include multiple video decoders that can operate in parallel. The GPU 210 may perform parallel video decoding in various ways. For example, the processing circuitry 214 may assign different compressed data segments (regardless of the corresponding bitstream) to different video decoders 216 to keep all the video decoders 216 continuously busy and speed up the video decoding of the segments. In some implementations, the processing circuitry 214 may assign different bitstreams (or their compressed data segments) to different video decoders 216. In some implementations, the GPU 210 may receive video bitstreams for multiple sessions simultaneously (e.g., bitstreams captured by different autosomes 140 or bitstreams captured at different time intervals). The processing circuitry 214 may assign different video decoders 216 to decode the video data of different sessions.

[0114] Method 300 may include the GPU 210 using one or more image frames to train an ML model (step 310). Once the selected image frames are decoded, the processing circuitry 214 may extract one or more features from each selected image frame and feed the extracted features to a training module configured to train the ML model. In response, the training module may modify or update one or more parameters of the ML model.

[0115] Although method 300 was described above as being performed or implemented by GPU 210, in general, method 300 can be performed or implemented by any computing system that includes a memory and one or more processors. Additionally, another type of processor can be used in place of GPU 210.

[0116] The various illustrative logical blocks, modules, circuits, and algorithmic steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure or the claims.

[0117] Embodiments implemented in computer software can be implemented in software, firmware, middleware, microcode, hardware description language, or any combination thereof. Code segments or machine-executable instructions can represent a procedure, function, subprogram, program, routine, subroutine, module, software package, class, or any combination of instructions, data structures, or program statements. A code segment can be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. can be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0118] The actual software code or specialized control hardware used to implement these systems and methods does not limit the claimed features or the present disclosure. Accordingly, the operations and behavior of the systems and methods have been described without reference to specific software code, it being understood that software and control hardware can be designed to implement the systems and methods based on the description herein.

[0119] When implemented in software, functions can be stored as one or more instructions or code on a non-transitory computer-readable or processor-readable storage medium. The steps of the methods or algorithms disclosed herein can be embodied in processor-executable software modules, which can reside on a computer-readable or processor-readable storage medium. Non-transitory computer-readable or processor-readable media include both computer storage media and tangible storage media, and tangible storage media facilitate the transfer of a computer program from one place to another. The non-transitory processor-readable storage medium can be any available medium accessible by a computer. By way of example and not limitation, such non-transitory processor-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other tangible storage medium that can be used to store the desired program code in the form of instructions or data structures and can be accessed by a computer or a processor. As used herein, disk and optical disk include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), Blu-ray disc, and floppy disk, where "disk" generally reproduces data magnetically, while "optical disk" reproduces data optically with a laser. Combinations of the above should also be included within the scope of computer-readable media. Additionally, the operations of the methods or algorithms can reside as code and / or instructions in one or any combination or set on a non-transitory processor-readable medium and / or computer-readable medium, which can be incorporated into a computer program product.

[0120] The foregoing description of the disclosed embodiments is provided to enable a person skilled in the art to make or use the embodiments and their variations described herein. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the principles defined herein can be applied to other embodiments without departing from the spirit or scope of the subject matter disclosed herein. Therefore, this disclosure is not intended to be limited to the embodiments shown herein, but should be accorded the widest scope consistent with the following claims and the principles and novel features disclosed herein.

[0121] Although various aspects and embodiments have been disclosed, other aspects and embodiments may also be contemplated. The various aspects and embodiments disclosed are for illustrative purposes and are not intended to be limiting, and the true scope and spirit are indicated by the following claims.

Claims

1. A method, comprising: receiving, by a processor, a bitstream of a video sequence and one or more indications indicative of one or more image frames of the video sequence for training a machine learning model; determining, by the processor, timestamps, positions within the bitstream, and types of the image frames of the video sequence by parsing the bitstream; determining, by the processor, one or more segments of the bitstream for decoding to extract the one or more image frames using the one or more indications and the timestamps, positions within the bitstream, and types of the image frames of the video sequence, wherein for an image frame among the one or more image frames, a corresponding segment represents a corresponding reference chain of the image frame of the video sequence; decoding, by the processor, the one or more segments of the bitstream; and training, by the processor, the machine learning model using the one or more image frames.

2. The method according to claim 1, further comprising: allocating, by the processor, a memory region within a memory of the processor; storing, by the processor, the bitstream in the allocated memory region; and storing, by the processor, the one or more image frames in the allocated memory region after decoding the one or more segments.

3. The method according to claim 1, wherein for each image frame among the one or more image frames, a corresponding reference chain of the image frame starts at an intra-frame (I-frame) of the video sequence and ends at the image frame among the one or more image frames.

4. The method according to claim 1, wherein determining the timestamps, positions in the bitstream, and types of the image frames of the video sequence includes generating one or more data structures indicative of: for each image frame of the video sequence, a corresponding timestamp and a corresponding offset representing a corresponding position of compressed data of the image frame in the bitstream; and specific types of image frames in the video sequence.

5. The method according to claim 1, wherein the processor is a graphics processing unit.

6. The method according to claim 1, wherein the one or more segments of the bitstream are decoded by a hardware decoder integrated in the processor.

7. The method according to claim 1, wherein the one or more indications include one or more time values, and for each time value among the one or more time values, determining the one or more segments of the bitstream comprising: determining a first timestamp among the timestamps of the image frames of the video sequence that is closest to the time value, the first timestamp corresponding to a first image frame in the video sequence; determining a second timestamp of an I-frame of the video sequence, the second timestamp being determined as the I-frame timestamp closest to the first timestamp that is less than or equal to the first timestamp; using the second timestamp to determine a start position of the I-frame in the bitstream among the positions of the image frame; Use the first timestamp to determine the end position of the first image frame in the bitstream among the positions of the image frames; and Determine the segment of the bitstream extending between the start position of the I-frame and the end position of the first image frame.

8. The method according to claim 1, wherein the bitstream is the first bitstream of a first compressed video sequence captured by a first camera, and the method further comprises: Receiving, by the processor, a second bitstream of a second video sequence captured by a second camera; Determining, by the processor, the timestamps, positions within the second bitstream, and types of the image frames of the second video sequence by parsing the second bitstream; Determining, by the processor, one or more segments of the second bitstream for decoding to extract one or more second image frames of the second video sequence using the one or more indications and the timestamps, positions within the second bitstream, and types of the image frames of the second video sequence, wherein for an image frame among the one or more second image frames, the corresponding segment of the second bitstream represents the corresponding reference chain of the image frame of the second video sequence; Decoding, by the processor, the one or more segments of the second bitstream; and Training, by the processor, the ML model using the one or more second image frames.

9. The method according to claim 8, wherein the first camera and the second camera are not synchronized with each other.

10. The method according to claim 9, wherein at least two segments among the one or more segments of the first bitstream and the one or more segments of the second bitstream are decoded in parallel by at least two hardware decoders integrated in the processor.

11. A computing device, comprising: A memory; and Processing circuitry configured to: Receive a bitstream of a video sequence and one or more indications indicating one or more image frames of the video sequence for training a machine learning model; Determine the timestamps, positions within the bitstream, and types of the image frames of the video sequence by parsing the bitstream; Use the one or more indications and the timestamps, positions within the bitstream, and types of the image frames of the video sequence to determine one or more segments of the bitstream for decoding to extract the one or more image frames, wherein for an image frame among the one or more image frames, the corresponding segment represents the corresponding reference chain of the image frame of the video sequence; Decode the one or more segments of the bitstream; and Train the machine learning model using the one or more image frames.

12. The computing device according to claim 11, wherein the processing circuitry is further configured to: Allocate a memory area within the memory; Store the bitstream in the allocated memory area; and After decoding the one or more segments, store the one or more image frames in the allocated memory area.

13. The computing device according to claim 11, wherein for each of the one or more image frames, the corresponding reference chain of the image frame starts at an intra-frame (I-frame) of the video sequence and ends at the image frame among the one or more image frames.

14. The computing device according to claim 11, wherein when determining the timestamp, position, and type of the image frame of the video sequence in the bitstream, the processing circuitry is configured to generate one or more data structures that indicate: for each image frame of the video sequence, a corresponding timestamp and a corresponding offset indicating the corresponding position of the compressed data of the image frame in the bitstream; and image frames of a specific type in the video sequence.

15. The computing device according to claim 11, wherein the computing device is a graphics processing unit (GPU).

16. The computing device according to claim 15, wherein the one or more segments of the bitstream are decoded by a hardware decoder integrated in the GPU.

17. The computing device according to claim 11, wherein the one or more indications include one or more time values, and when determining the one or more segments of the bitstream, the processing circuitry is configured, for each of the one or more time values: determine a first timestamp among the timestamps of the image frames of the video sequence that is closest to the time value, the first timestamp corresponding to a first image frame in the video sequence; determine a second timestamp of an I-frame of the video sequence, the second timestamp being determined as the I-frame timestamp closest to the first timestamp that is less than or equal to the first timestamp; use the second timestamp to determine the start position of the I-frame in the bitstream among the positions of the image frames; use the first timestamp to determine the end position of the first image frame in the bitstream among the positions of the image frames; and determine the segment of the bitstream that extends between the start position of the I-frame and the end position of the first image frame.

18. The computing device according to claim 11, wherein the bitstream is a first bitstream of a first compressed video sequence captured by a first camera, and the processing circuitry is further configured to: receive a second bitstream of a second video sequence captured by a second camera; determine to parse the second bitstream, the timestamp of the image frames of the second video sequence, the position, and the type within the second bitstream; use the one or more indications and the timestamp of the image frames of the second video sequence, the position within the second bitstream, and the type to determine one or more segments of the second bitstream for decoding to extract one or more second image frames of the second video sequence, and for an image frame among the one or more second image frames, the corresponding segment of the second bitstream represents the corresponding reference chain of the image frame of the second video sequence; Decode the one or more segments of the second bitstream; And Use the one or more second image frames to train the machine learning model.

19. The computing device according to claim 18, wherein the first camera and the second camera are not synchronized with each other.

20. The computing device according to claim 19, wherein at least two segments of the one or more segments of the first bitstream and the one or more segments of the second bitstream are decoded in parallel by at least two hardware decoders integrated in the computing device.

21. A non-transitory computer-readable medium having computer code instructions stored thereon, the computer code instructions, when executed by a processor, cause the processor to: Receive a bitstream of a video sequence and one or more indications indicating one or more image frames of the video sequence for training a machine learning model; Determine the timestamps, positions within the bitstream, and types of the image frames of the video sequence by parsing the bitstream; Use the one or more indications and the timestamps, positions within the bitstream, and types of the image frames of the video sequence to determine one or more segments of the bitstream for decoding to extract the one or more image frames, and for an image frame among the one or more image frames, the corresponding segment represents the corresponding reference chain of the image frame of the video sequence; Decode the one or more segments of the bitstream; And Use the one or more image frames to train the machine learning model.