Traffic analysis method and device

By using convolutional neural networks and transformation models to analyze traffic videos on edge devices, the limitations of edge device computing power are solved, enabling real-time and efficient traffic analysis, reducing latency and improving accuracy.

CN121884290APending Publication Date: 2026-04-17LEOTEK CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LEOTEK CORP
Filing Date
2025-09-28
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing technologies, traffic analysis is usually performed on backend devices due to the limited computing power of edge devices, which leads to latency issues.

Method used

Convolutional neural networks and transformation models are deployed on edge devices. A first machine learning model analyzes video clips and embeds traffic descriptions, while a second machine learning model is used for further analysis. An appropriate model is selected in conjunction with the traffic signal stage to optimize the utilization of computing resources.

Benefits of technology

It enables real-time and rapid traffic analysis on edge devices, reducing latency and improving the accuracy and efficiency of traffic condition assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884290A_ABST
    Figure CN121884290A_ABST
Patent Text Reader

Abstract

The invention provides a traffic analysis method and device. The traffic analysis device may include a camera, a processor, and a memory. The camera can be used for extracting a video associated with a road scene. The memory may be coupled to the processor. The processor may be used to analyze a first segment of the movie using a first machine learning model to generate a first traffic description, embed the first traffic description in a first movie encoding parameter synchronized with the first segment of the movie, and using a second machine learning model to analyze a second segment of the movie according to the first movie coding parameter to generate a second traffic description.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a wireless communication technology, and more particularly to a traffic analysis based on a machine learning model or an artificial intelligence model. Background Technology

[0002] Unless otherwise stated herein, the methods described in this section do not constitute prior art with respect to the claims of this invention, nor are they to be recognized as prior art.

[0003] In current technologies, artificial intelligence (AI) is widely used in various applications. For example, a machine learning model or AI analysis model can be applied to traffic analysis. However, due to the limitations of edge device computing power, most traffic analysis operations performed using machine learning or AI analysis models are typically executed on a backend device (e.g., a remote server or a cloud server). In other words, the edge device needs to transmit video data to the backend device before the backend device can use the machine learning or AI analysis model to perform traffic analysis based on the video. Therefore, delays in traffic analysis may occur.

[0004] Therefore, how to perform traffic analysis more instantly and quickly on edge devices is a topic worth discussing. Summary of the Invention

[0005] The following overview is illustrative only and is not intended to be limiting in any way. That is, it is provided to introduce the concepts, key points, benefits, and advantages of the novel and non-obvious techniques described herein. Selected embodiments will be further described in the detailed description below. Therefore, the following summary is not intended to identify the essential features of the claimed subject matter, nor is it intended to define the scope of the claimed subject matter.

[0006] One object of the present invention is to provide schemes, concepts, designs, systems, methods, and apparatus related to traffic analysis in mobile communications. It is believed that by implementing one or more of the proposed schemes described herein, the aforementioned problems can be avoided or mitigated.

[0007] One embodiment of the present invention provides a traffic analysis apparatus. The traffic analysis apparatus may include a camera, a processor, and a memory. The camera can be used to extract video associated with a road scene. The memory may be coupled to the processor. The processor can be used to analyze a first segment of the video using a first machine learning model to generate a first traffic description, embed the first traffic description into first video encoding parameters synchronized with the first video segment, and use a second machine learning model to analyze a second segment of the video based on the first video encoding parameters to generate a second traffic description.

[0008] In some embodiments, the first machine learning model may include a convolutional neural network (CNN) model, and the second machine learning model may include a transformer model.

[0009] In some embodiments, the processor can be further configured to generate a first traffic description during a first time period, and to generate a second traffic description of the video from the first segment to the second segment during a second time period. The second time period may be longer than the first time period.

[0010] In some embodiments, the processor can also be used to obtain a first position of an object in a first segment of a video using a first machine learning model, encode the first position into a first traffic description, obtain the trajectory of the object from the first segment to the second segment using a second machine learning model, and encode the trajectory into a second traffic description.

[0011] In some embodiments, the processor can also be used to determine a first category of objects in the first segment using a first machine learning model, encode the first category into a first traffic description, and generate a description of objects from the first segment to the second segment based on the first category and the trajectory using a second machine learning model.

[0012] In some embodiments, the processor can further be used to determine the computing resources of the traffic analysis device, and based on the computing resources, determine whether to generate a first category and a first location of the object, and when the computing resources are insufficient, prioritize generating the first location.

[0013] In some embodiments, the second segment of the video may continue from the first segment of the video, and the processor may further be used to embed a second traffic description into a second video encoding parameter synchronized with the second segment of the video.

[0014] In some embodiments, the processor can be further configured to obtain a second location and a second classification of an object in the second segment using a first machine learning model, and to analyze the second segment, the first video encoding parameters, the second location, and the second classification using a second machine learning model to generate a description associated with the object.

[0015] In some embodiments, the video may include a traffic signal, and the processor may further be used to determine the traffic signal stage of the traffic signal using a first machine learning model.

[0016] In some embodiments, the processor is further configured to obtain the traffic signal phase of the traffic signal from a traffic controller connected to the traffic signal.

[0017] In some embodiments, the processor can be further configured to switch from one machine learning model to another when a server indicates that the accuracy of the first machine learning model is lower than that of another machine learning model.

[0018] In some embodiments, the processor can further be configured to determine, based on instructions from the server, whether to update the first machine learning model and the second machine learning model. The server may calculate the computational resources of the traffic analysis device to generate instructions.

[0019] In some embodiments, the first segment and the second segment may each include a first group of images and a second group of images. The first video encoding parameters and the second video encoding parameters may each include a first supplemental enhancement information (SEI) following the first group of images and a second supplemental enhancement information following the second group of images.

[0020] In some embodiments, when the traffic signal is green, the processor can further use a traffic flow calculation model to calculate the number of vehicles on the road in a segment preceding the first segment of the video. In some embodiments, when the traffic signal is yellow, the processor can further use a fast vehicle trajectory model to generate the trajectories of vehicles in a segment of the video. In some embodiments, when the traffic signal is red, the processor can further use a queue length calculation model to generate the total length of vehicles waiting at the red light on the road in a segment of the video.

[0021] In some embodiments, the processor can also be used to generate a distributed traffic description based on multiple videos from different cameras using a first machine learning model, embed the distributed traffic description into the video encoding parameters of each video, and determine whether to switch to another machine learning model based on the distributed traffic description.

[0022] An embodiment of the present invention provides a traffic analysis method. The traffic analysis method can be applied to a traffic analysis device. The traffic analysis method may include the following steps: The traffic analysis device may extract video associated with a road scene. Then, the traffic analysis device may use a first machine learning model to analyze a first segment of the video to generate a first traffic description. Next, the traffic analysis device may embed the first traffic description into a first video encoding parameter synchronized with the first video segment. Then, the traffic analysis device may use a second machine learning model, based on the first video encoding parameter, to analyze a second segment of the video to generate a second traffic description.

[0023] Other additional features and advantages of the present invention can be obtained by those skilled in the art through modifications and refinements based on the traffic analysis method and apparatus disclosed in the embodiments of this invention, without departing from the spirit and scope of the present invention. Attached Figure Description

[0024] Figure 1 This is a block diagram of a traffic analysis system according to an embodiment of the present invention.

[0025] Figure 2 This is a partial diagram of a traffic analysis device according to an embodiment of the present invention.

[0026] Figure 3 This is a schematic diagram of a video clip format according to an embodiment of the present invention.

[0027] Figure 4 This is a flowchart of a traffic analysis procedure according to an embodiment of the present invention.

[0028] Figure 5A and Figure 5B This is a flowchart of a machine learning model switching method for traffic signal phases based on a traffic signal, according to an embodiment of the present invention.

[0029] Figure 6 This is a schematic diagram of a traffic analysis at an intersection according to an embodiment of the present invention.

[0030] Figure 7 This is a flowchart of a traffic analysis method according to an embodiment of the present invention.

[0031] The attached figures are labeled as follows:

[0032] 100: Traffic Analysis System

[0033] 110, 200: Traffic Analysis Device

[0034] 120: Server

[0035] 210: Wireless transceiver

[0036] 211: Baseband Processing Device

[0037] 212: Radio Frequency Devices

[0038] 213: Antenna

[0039] 220: Processor

[0040] 230: Storage device

[0041] 240: Camera

[0042] 300: Fragment Format

[0043] 400, 500, 700: Flowchart

[0044] S410~S7470, S501~S514, S710~S740: Steps Detailed Implementation

[0045] This section describes preferred embodiments of the invention and is intended to illustrate the spirit of the invention rather than to limit its scope of protection. The scope of protection of the invention shall be determined by the appended claims.

[0046] Figure 1 This is a block diagram of a traffic analysis system 100 according to an embodiment of the present invention. Figure 1 As shown, the traffic analysis system 100 may include a traffic analysis device 110 and a server (or a backend device) 120. It should be noted that... Figure 1 The block diagrams shown are merely for illustrating embodiments of the present invention, and the present invention is not intended to be construed as such. Figure 1 Limited to.

[0047] According to an embodiment of the present invention, the traffic analysis device 110 can be located at an edge device in a network architecture. The traffic analysis device 110 can communicate with a server 120 via a wireless communication technology. The traffic analysis device 110 can provide traffic analysis results to the server 120 via wireless communication technology. The traffic analysis device 110 can be installed at an intersection.

[0048] Figure 2 This is a block diagram of a traffic analysis device 200 according to an embodiment of the present invention. The traffic analysis device 200 can be applied to... Figure 1 The traffic analysis device 110 is shown. (For example...) Figure 2 As shown, the traffic analysis device 200 may include a wireless transceiver 210, a processor 220, a storage device 230, and at least a camera 240.

[0049] The wireless transceiver 210 can be used to perform wireless transmission and reception of the traffic analysis device 200.

[0050] Specifically, the wireless transceiver 210 may include a baseband processing device 211, a radio frequency (RF) device 212, and an antenna 213, wherein the antenna 213 may include an antenna array.

[0051] The baseband processing device 211 can be used to perform baseband signal processing, such as analog-to-digital conversion (ADC) / digital-to-analog conversion (DAC), gain adjustment, modulation / demodulation, encoding / decoding, etc. The baseband processing device 211 may include multiple hardware components, such as a baseband processor, to perform baseband signal processing.

[0052] Radio frequency (RF) device 212 can receive RF wireless signals via antenna 213, convert the received RF wireless signals into baseband signals to be processed by baseband processing device 211, or receive baseband signals from baseband processing device 211 and convert the received baseband signals into RF wireless signals to be subsequently transmitted via antenna 213. RF device 212 may include multiple hardware components to perform wireless frequency conversion. For example, RF device 212 may include a power amplifier, a mixer, an ADC converter / DAC converter, etc.

[0053] According to one embodiment of the present invention, the radio frequency device 212 and the baseband processing device 211 can be considered as a whole as a wireless module capable of communicating with a wireless network to provide wireless communication services based on a radio access technology (RAT). Note that in some embodiments of the present invention, the traffic analysis device 200 can be further expanded to include multiple antennas and / or multiple wireless modules; however, the present invention does not necessarily imply... Figure 2 The above is the limit.

[0054] Processor 220 may be a general-purpose processor, a central processing unit (CPU), a microcontroller unit (MCU), an application processor, a digital signal processor (DSP), a graphics processing unit (GPU), a holographic processing unit (HPU), a neural processing unit (NPU), or other similar devices. Processor 220 may include various circuits to provide underlying functions such as data processing and computation, controlling communication between wireless transceiver 210 and network node 110, storing data (e.g., program code) to and retrieving data from storage device 230, and controlling one or more cameras 240 to extract a video (or multiple videos) associated with the road scene.

[0055] Specifically, the processor 220 can coordinate the operations of the aforementioned transceiver 210, storage device 230, and camera 240 to execute the method of the present invention.

[0056] As will be understood by those skilled in the art, the circuitry of processor 220 may include transistors configured to control the operation of the circuitry according to the functions and operations described herein. As will be further understood, the specific structure or interconnection of the transistors may be determined by a compiler, such as a Register Transfer Language (RTL) compiler. An RTL compiler can be operated by the processor based on scripts very similar to assembly language code to compile the scripts into a form suitable for the layout or fabrication of the final circuitry. Indeed, RTL is renowned for its role and use in facilitating the design process of electronic and digital systems.

[0057] Storage device 230 may be a non-transient machine-readable storage medium, including memory, such as flash memory or non-volatile random access memory (NVRAM), or magnetic storage devices, such as hard disks or magnetic tapes, or optical discs, or any combination thereof, for storing data, instructions and / or application software program code, communication protocols and / or methods of the present invention.

[0058] Camera (or multiple cameras) 240 can be used to extract a video (or multiple videos) of the scene associated with the road for use in traffic analysis.

[0059] It should be understood that, Figure 2 The components described in the embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. For example, a traffic analysis device may include more components, such as another camera. Alternatively, a traffic analysis device may include fewer components.

[0060] According to one embodiment of the present invention, a traffic analysis device 110 can extract a video (or a video footage) associated with a scene of a road. Then, the traffic analysis device 110 can use a first machine learning model to analyze a first segment of the video to generate a first traffic description. Next, the traffic analysis device 110 can embed the first traffic description into a first video encoding parameter synchronized with the first video segment. Furthermore, the traffic analysis device 110 can use a second machine learning model to analyze the first video segment based on the first video encoding parameters to generate a second traffic description. In this embodiment, the second segment may continue from the first segment. The traffic analysis device 110 can further embed the second traffic description into a second video encoding parameter synchronized with the second video segment. According to one embodiment of the present invention, each traffic description may have a...

[0061] JSON (JavaScript Object Notation) file format.

[0062] According to one embodiment of the present invention, the traffic analysis device 110 can perform an image compression technique (e.g., H.264, meaning the video can be an H.264 video or an H.264 stream, but the present invention is not limited thereto) on the video. Furthermore, according to one embodiment of the present invention, the video encoding parameters (e.g., first video encoding parameters) used to encode the video (e.g., H.264 video) may include supplemental enhancement information (SEI). The SEI may include traffic description, and the SEI may be embedded in the H.264 video. According to an embodiment of the present invention, the traffic analysis device 110 can analyze the video encoding parameters (e.g., SEI) to obtain the identification result of a machine learning model or artificial intelligence (AI) model for a segment of the video (i.e., the traffic description embedded in the video encoding parameters), and based on the identification result, determine whether to switch the machine learning model to analyze the next segment of the video.

[0063] According to one embodiment of the present invention, each segment of a video may include a group of pictures (GOP) and video encoding parameters (e.g., a SEI) following the GOP. For example, the first segment and the second segment may respectively include a first GOP and a second GOP, and the first video encoding parameters and the second video encoding parameters may respectively include a first SEI following the first GOP and a second SEI following the second GOP.

[0064] Figure 3This is a schematic diagram of a video clip format 300 according to an embodiment of the present invention. Figure 3 As shown, each segment of the video may include a Group of Pictures (GOP) and a Sequence Instance (SEI) following the GOP. Each GOP may include an I-frame and multiple P-frames (e.g., 1 I-frame and 49 P-frames). Each segment also includes a sequence parameterset (SPS) and a picture parameter set. Each SEI may include different information about the segment. Furthermore, each SEI may be associated with a different machine learning model. In addition, segment format 300 also includes an advanced video coding (AVC) sequence header used for decoding the video.

[0065] According to one embodiment of the present invention, the first machine learning model may include a convolutional neural network (CNN) model (e.g., YOLOv8, normalized object coordinate space (NOCS), ResNet, or DenseNet, but the invention is not limited thereto), and the second machine learning model may include a transformer model (e.g., LLaMa, LLaVA, or other video understanding models, but the invention is not limited thereto). For example, in another embodiment, the first machine learning model may include a transformer model, and the second machine learning model may include a CNN model. In another embodiment, the first and second machine learning models may each include more than one model. That is, the traffic analysis device 110 can use multiple machine learning models to analyze a segment of a video.

[0066] Furthermore, according to one embodiment of the present invention, based on the traffic light phase, each of the first machine learning model and the second machine learning model may include a traffic counting model, a fast wheel trajectory model, and a fleet length calculation model, but the present invention is not limited thereto.

[0067] Traffic flow calculation models may include a CNN model combined with a transformation model, a fine-tuned YOLOv5 or YOLOv8 object detection model for vehicle counting, and a model with a focal loss function for dense traffic scenes.

[0068] The invention may include RetinaNet models, a vision transformer (ViT) model trained on aerial traffic datasets, or a graph neural network (GNN) model that models vehicle interactions to improve computational accuracy, but is not limited thereto.

[0069] Fast vehicle trajectory models may include a CNN model, a 3D CNN model for motion pattern analysis, an optical flow-based model using FlowNet2, a recurrent neural network (RNN) model or a long short-term memory (LSTM) model for transient trajectory prediction, or a transformer-based motion prediction model trained from wheel movement sequences, but the present invention is not limited thereto.

[0070] Fleet length calculation models may include a CNN model combined with a transformation model, a 3D CNN model for spatiotemporal feature extraction, a hybrid model combining CNN and LSTM for sequential vehicle detection, a finely tuned ViT model for vehicle segmentation, a multi-task learning model that jointly detects vehicle counts and lengths using shared convolutional backbones and attention mechanisms, or a depth estimation model that uses stereo vision or monocular depth prediction to infer vehicle space and fleet length, but the present invention is not limited thereto.

[0071] For example, a traffic flow calculation model can be associated with the green light phase to calculate traffic flow on the road; a fast vehicle trajectory model can be associated with the yellow light phase to identify the trajectory of objects in a video clip; and a queue length calculation model can be associated with the red light phase to calculate the queue length (i.e., the length of vehicles waiting at a red light on the road) in a video clip. In other words, the traffic analysis device 110 can determine the type of machine learning model (e.g., a first machine learning model or a second machine learning model) based on the traffic signal phase in the video clip.

[0072] According to one embodiment of the present invention, a traffic signal may be included in a video clip. Therefore, the traffic analysis device 110 can determine the traffic signal stage of the traffic signal in the video clip via a machine learning model.

[0073] According to one embodiment of the present invention, the traffic analysis device 110 can obtain the traffic signal phase of the traffic signal from a traffic controller connected to the traffic signal.

[0074] According to one embodiment of the present invention, the traffic analysis device 110 can generate a first traffic description based on a first segment of a video during a first time period (e.g., 1 millisecond (ms), but this invention is not limited thereto) using a first machine learning model. Furthermore, the traffic analysis device 110 can generate a second traffic description of the video from the first segment to the second segment during a second time period (e.g., 30 milliseconds, but this invention is not limited thereto) based on a second segment of the video and first video encoding parameters using a second machine learning model. In one embodiment, the second time period may be longer than the first time period, but this invention is not limited thereto. The lengths of the first and second time periods can be determined based on the machine learning model employed.

[0075] According to one embodiment of the present invention, each traffic description (e.g., a first traffic description and a second traffic description) may include at least one of the following: an indicator of the type of scene associated with a segment of a video, location information of an object (or multiple objects) in a segment of the video (e.g., coordinate information in a bounding box), trajectory information of an object (or multiple objects) in a segment of the video (e.g., the trajectory of a car), classification information (or type information) of an object (or multiple objects) in a segment of the video, a description of the scene associated with a segment of the video, and time information associated with a segment of the video (e.g., a timestamp), but the present invention is not limited thereto.

[0076] For example, in some embodiments, the traffic analysis device 110 can obtain a first location of an object (e.g., a car) in a first segment of a video using a first machine learning model, and then the traffic analysis device 110 can encode the first location into a first traffic description. Furthermore, the traffic analysis device 110 can obtain the trajectory of the object from the first segment to a second segment using a second machine learning model, and then the traffic analysis device 110 can encode the trajectory of the object into a second traffic description.

[0077] Furthermore, according to some embodiments of the present invention, the traffic analysis device 110 can determine a first category of objects in the first segment using a first machine learning model, and then the traffic analysis device 110 can encode the first category into a first traffic description. Additionally, the traffic analysis device 110 can generate a description of objects moving from the first segment to the second segment based on the first category and trajectory using a second machine learning model.

[0078] In another example, according to some embodiments of the present invention, the traffic analysis device 110 may obtain a second location and a second classification of an object in a second segment through a machine learning model (e.g., a first machine learning model). Furthermore, the traffic analysis device 110 may analyze the second segment, first video encoding parameters, second location, and second classification through another machine learning model (e.g., a second machine learning model) to generate a description associated with the object.

[0079] According to one embodiment of the present invention, the traffic analysis device 110 can determine its computational resources (or computational capability or computational power), for example, the computational resources of its processor. Then, based on its computational resources, the traffic analysis device 110 can determine whether to generate or calculate the classification and position of each object in the video segment. In one embodiment, when computational resources are insufficient, the traffic analysis device 110 can prioritize generating the position of each object without first generating its classification, but this is not a limitation of the present invention. Specifically, if the traffic analysis device 110's computational resources are insufficient, it can determine which operation needs to be executed first based on the current scene of the video segment. For example, when the traffic analysis device 110 determines from the video segment that the traffic light is green and its current computational resources are insufficient, it can prioritize generating the position of each object in the video segment to calculate the traffic flow on the road.

[0080] Based on computing resources, determine whether to generate the first category and first position of the object, and when computing resources are insufficient, prioritize generating the first position.

[0081] Figure 4 This is a flowchart 400 of a traffic analysis program according to an embodiment of the present invention. The traffic analysis program can be applied to a traffic analysis device 110. Figure 4 As shown, in step S410, the traffic analysis device 110 can extract a video associated with a road and a scene.

[0082] In step S420, the traffic analysis device 110 may analyze a segment of the video via a machine learning model (e.g., a CNN model, and / or a transformation model, but the present invention is not limited thereto) to generate a traffic description associated with the segment.

[0083] In step S430, the traffic analysis device 110 may embed the traffic description into the video encoding parameters (e.g., SEI) associated with this segment.

[0084] In step S440, the traffic analysis device 110 can encode video encoding parameters (e.g., SEI) into the video (e.g., H.264 video or H.264 stream).

[0085] In step S450, the traffic analysis device 110 can synchronize information of the video encoding parameters (e.g., the SEI describing the traffic signal phase). This synchronization is to ensure consistency of contextual information (e.g., traffic signal status in the video data), rather than to achieve precise time synchronization. By embedding or associating traffic signal phase metadata with the video, the downstream device or model can more accurately understand traffic conditions when relevant signals change.

[0086] In step S460, the traffic analysis device 110 can determine whether it is the next segment of the video based on the information of the video encoding parameters (e.g., SEI) and switch the machine learning model.

[0087] In step S470, the traffic analysis device 110 may store segments of traffic descriptions with embedded video encoding parameters (e.g., SEI) for use in subsequent video analysis.

[0088] Figure 5A and Figure 5B This is a flowchart 500 of a machine learning model switching method based on traffic signal phases according to an embodiment of the present invention. The machine learning model switching method can be applied to a traffic analysis device 110. Figure 5A Figure 5B As shown, in step S501, the traffic analysis device 110 can extract a first segment of a video associated with a road and a scene.

[0089] In step S502, the traffic analysis device 110 can analyze a first segment of the video using a CNN model to generate a first traffic description associated with the first segment, and determine whether the traffic signal phase is a yellow light phase based on the first segment or information from a traffic controller. That is, in this embodiment, when the traffic signal is a yellow light phase, the traffic analysis device 110 can use a CNN model to perform object detection and object classification within a short time period (e.g., 1 millisecond) to quickly generate the first traffic description.

[0090] In step S503, the traffic analysis device 110 may embed the first traffic description into the first video encoding parameter (e.g., SEI) associated with the first segment.

[0091] In step S504, the traffic analysis device 110 can encode the first video encoding parameter (e.g., SEI) into the video (e.g., H.264 video or H.264 stream).

[0092] In step S505, the traffic analysis device 110 can synchronously describe the information in the first video encoding parameters (e.g., SEI) of the traffic signal phase.

[0093] In step S506, the traffic analysis device 110 may determine, based on information from the first video encoding parameters (e.g., SEI), whether to switch the machine learning model for the next segment of the video (e.g., the second segment). For example, based on the information from the first video encoding parameters (e.g., SEI), the traffic analysis device 110 may determine that in the next segment, the traffic signal phase may change from a yellow light to a red light. Therefore, the traffic analysis device 110 may decide to switch to another machine learning model suitable for analyzing the next segment, but the present invention is not limited thereto.

[0094] In step S507, the traffic analysis device 110 may store a first segment having a first traffic description with embedded first video encoding parameters (e.g., SEI).

[0095] In step S508, the traffic analysis device 110 can extract a second segment of the video.

[0096] In step S509, the traffic analysis device 110 can analyze the second segment of the video using a CNN model and a transformation model to generate a second traffic description associated with the second segment, and determine whether the traffic signal phase is a red light phase based on the second segment or information from the traffic controller. Specifically, in this embodiment, when the traffic signal is a red light phase, the traffic analysis device 110 can use a CNN model to perform object detection and object classification during a first time period (e.g., 1 millisecond), and use a transformation model to analyze the trajectory of each object from the first segment to the second segment. In one example, because the yellow light phase is short (e.g., 5 seconds), the traffic analysis device 110 needs to quickly obtain the trajectory information of each object in the first segment to generate the first traffic description. Then, when the traffic signal phase is a red light phase with a longer duration (e.g., 1 minute), the traffic analysis device 110 can have sufficient time to analyze the trajectory of each object from the first segment to the second segment based on the second traffic description associated with the second segment and pre-stored information obtained during the yellow light phase.

[0097] In step S510, the traffic analysis device 110 may embed the second traffic description into the second video encoding parameters (e.g., SEI) associated with the next segment (i.e., the second segment).

[0098] In step S511, the traffic analysis device 110 can encode the second video encoding parameter (e.g., SEI) into the video (e.g., H.264 video or H.264 stream).

[0099] In step S512, the traffic analysis device 110 can simultaneously describe the information in the second video encoding parameters (e.g., SEI) of the traffic signal phase.

[0100] In step S513, the traffic analysis device 110 can determine whether to switch the machine learning model for the next segment of the video based on the information of the second video encoding parameters (e.g., SEI).

[0101] In step S514, the traffic analysis device 110 may store a second segment having a second traffic description with embedded second video encoding parameters (e.g., SEI).

[0102] According to one embodiment of the present invention, when server 120 indicates that the accuracy of the currently used machine learning model is lower than that of another machine learning model, traffic analysis device 110 may switch the currently used machine learning model to the other machine learning model. Specifically, server 120 may obtain a video segment with a traffic description having embedded video encoding parameters (e.g., SEI) and analyze the traffic description to determine whether the accuracy of the currently used machine learning model is sufficient (e.g., whether the accuracy of the currently used machine learning model is lower than a threshold value). When the accuracy of the machine learning model currently used by traffic analysis device 110 is insufficient, server 120 may instruct traffic analysis device 110 to use another machine learning model with higher accuracy (e.g., a machine learning model with accuracy higher than a threshold value) to process the video segment.

[0103] According to one embodiment of the present invention, based on the instructions of the server 120, the traffic analysis device 110 may determine whether to update the first machine learning model and / or the second machine learning model. The server 120 may calculate the computing resources (or computing power) of the traffic analysis device 110 to generate instructions.

[0104] According to one embodiment of the present invention, when the traffic analysis device 110 includes more than one camera, the traffic analysis device 110 can generate a distributed traffic depiction based on video footage from different cameras using a machine learning model (e.g., a first machine learning model). Then, the traffic analysis device 110 can embed the distributed traffic depiction into a video encoding parameter (e.g., SEI) of each video to synchronize each video. Furthermore, the traffic analysis device 110 can determine whether to switch to another machine learning model based on the distributed traffic depiction.

[0105] According to one embodiment of the present invention, different traffic analysis devices 110 can be configured at each corner (or bend) of an intersection. A server (or backend device) 120 can receive video containing traffic descriptions from each traffic analysis device 110. The server 120 then analyzes the video encoding parameters (e.g., SEI) of each video to synchronize videos from different traffic analysis devices 110. Furthermore, the server 120 can analyze the video encoding parameters (e.g., SEI) of each video to determine whether to update the machine learning model of each traffic analysis device 110. In another embodiment, each traffic analysis device 110 can also transmit the video stream or extracted traffic features, along with the traffic description, to other traffic analysis devices 110 configured at the intersection. According to one embodiment of the present invention, traffic analysis devices 110 configured at the same intersection can also jointly generate a unified traffic description. Cooperation between traffic analysis devices 110 can be implemented via a structured data exchange protocol (e.g., protocol buffer, Protobuf). Structured data exchange protocols enable efficient, low-latency, and platform-independent communication between devices.

[0106] Figure 6 This is a schematic diagram of a traffic analysis at an intersection according to an embodiment of the present invention. Figure 6 As shown, different traffic analysis devices 110 are respectively configured at the four corners (or turns) of the intersection. A server (or backend device) 120 can receive video clips with traffic descriptions from each traffic analysis device 110. Then, the server 120 can analyze the video encoding parameters (e.g., SEI) of each video clip to synchronize videos from different traffic analysis devices 110 and determine whether to update the machine learning model of each traffic analysis device 110.

[0107] Figure 7 This is a flowchart 700 of a traffic analysis method according to an embodiment of the present invention. Figure 7 The traffic analysis method shown is applicable to traffic analysis device 110. For example... Figure 7 As shown, in step S710, the traffic analysis device 110 can extract a video associated with a road and a scene.

[0108] In step S720, the traffic analysis device 110 may use a first machine learning model to analyze a first segment of the video to generate a first traffic description.

[0109] In step S730, the traffic analysis device 110 may embed the first traffic description into a first video encoding parameter that is synchronized with the first segment of the video.

[0110] In step S740, the traffic analysis device 110 may use a second machine learning model to analyze a second segment of the video based on the first video encoding parameters to generate a second traffic description.

[0111] According to one embodiment of the present invention, in the traffic analysis method, the first machine learning model may include a convolutional neural network (CNN) model, and the second machine learning model may include a transformation model.

[0112] According to an embodiment of the present invention, in a traffic analysis method, a traffic analysis device 110 can generate a first traffic description during a first time period, and generate a second traffic description of a video clip from the first clip to the second clip during a second time period. The second time period may be longer than the first time period.

[0113] According to an embodiment of the present invention, in a traffic analysis method, a traffic analysis device 110 can obtain a first position of an object in a first segment of a video through a first machine learning model, encode the first position into a first traffic description, obtain the trajectory of the object from the first segment to the second segment through a second machine learning model, and encode the trajectory into a second traffic description.

[0114] According to an embodiment of the present invention, in a traffic analysis method, a traffic analysis device 110 can determine a first category of an object in a first segment through a first machine learning model, encode the first category into a first traffic description, and generate a description of the object from the first segment to the second segment based on the first category and the trajectory through a second machine learning model.

[0115] According to an embodiment of the present invention, in a traffic analysis method, a traffic analysis device 110 can determine the computing resources of the traffic analysis device, and based on the computing resources, determine whether to generate a first category and a first location of an object, and when the computing resources are insufficient, prioritize generating the first location.

[0116] According to one embodiment of the present invention, in a traffic analysis method, a second segment of a video can be continued from the first segment of the video, and the processor is further used to embed a second traffic description into a second video encoding parameter synchronized with the second segment of the video.

[0117] According to an embodiment of the present invention, in a traffic analysis method, a traffic analysis device 110 can obtain a second location and a second classification of an object in a second segment through a first machine learning model, and analyze the second segment, the first video encoding parameters, the second location, and the second classification through a second machine learning model to generate a description associated with the object.

[0118] According to one embodiment of the present invention, in the traffic analysis method, the video may include a traffic signal, and the traffic analysis device 110 may determine the traffic signal stage of the traffic signal through a first machine learning model.

[0119] According to one embodiment of the present invention, in a traffic analysis method, a traffic analysis device 110 can obtain the traffic signal phase of a traffic signal from a traffic controller connected to the traffic signal.

[0120] According to one embodiment of the present invention, in a traffic analysis method, the traffic analysis device 110 can switch the first machine learning model to another machine learning model when a server indicates that the accuracy of the first machine learning model is lower than the accuracy of another machine learning model.

[0121] According to one embodiment of the present invention, in a traffic analysis method, a traffic analysis device 110 can determine whether to update a first machine learning model and a second machine learning model based on instructions from a server. The server can calculate the computing resources of the traffic analysis device to generate instructions.

[0122] According to one embodiment of the present invention, in the traffic analysis method, the first segment and the second segment may each include a first GOP and a second GOP. The first video encoding parameters and the second video encoding parameters may each include a first SEI following the first group of images and a second SEI following the second group of images.

[0123] According to one embodiment of the present invention, in a traffic analysis method, a traffic analysis device 110 may use a traffic flow calculation model to calculate the number of vehicles on the road in a segment preceding the first segment of a video. When the traffic signal is in the yellow light phase, the traffic analysis device 110 may use a rapid vehicle trajectory model to generate the trajectories of vehicles in a segment of the video. When the traffic signal is in the red light phase, the traffic analysis device 110 may use a queue length calculation model to generate the total length of vehicles waiting at the red light on the road in a segment of the video.

[0124] According to an embodiment of the present invention, in a traffic analysis method, a traffic analysis device 110 can generate a distributed traffic description based on multiple videos from different cameras using a first machine learning model, embed the distributed traffic description into the video encoding parameters of each video, and determine whether to switch to another machine learning model based on the distributed traffic description.

[0125] According to the traffic analysis method provided by embodiments of the present invention, the analysis results of machine learning (i.e., traffic descriptions) can be embedded or compressed into the video encoding parameters (e.g., SEI) of a video. Therefore, the traffic analysis device and / or server can determine whether to switch the machine learning model based on the video encoding parameters (e.g., SEI) of the video. Furthermore, according to the traffic analysis method provided by embodiments of the present invention, the traffic analysis device and / or server can determine road conditions more instantly.

[0126] The steps of the methods and algorithms disclosed in this specification can be directly applied to hardware and software modules, or a combination of both, by executing a processor. A software module (including execution instructions and related data) and other data can be stored in a data memory, such as random access memory (RAM), flash memory, read-only memory (ROM), erasable programmable read-only memory (EPROM), electronically erasable programmable read-only memory (EEPROM), registers, hard disks, portable optical discs, optical disc read-only memory (CD-ROM), DVDs, or any other computer-readable storage media format in the art. A storage media can be coupled to a machine device, for example, a computer / processor (referred to as a processor in this specification for convenience), which can read information (such as program code) and write information to the storage media. A storage media can integrate a processor. An application-specific integrated circuit (ASIC) includes a processor and a storage media. A user equipment includes an application-specific integrated circuit. In other words, the processor and storage media are included in the user equipment in a manner that does not directly connect to the user equipment. Furthermore, in some embodiments, any product suitable for computer programs includes readable storage media, wherein the readable storage media includes program code associated with one or more of the disclosed embodiments. In some embodiments, the product of the computer program may include packaging material.

[0127] Furthermore, those skilled in the art should understand that, generally, the terms used herein, and especially in the appended claims (e.g., the body of the appended claims), are intended to be “open-ended” terms; for example, the term “including” should be interpreted as “including but not limited to,” the term “having” should be interpreted as “having at least,” and the term “include” should be interpreted as “including but not limited to,” etc. Those skilled in the art should also understand that if a specific number of introductory claims are desired, this intention will be explicitly stated in the claims, and without such enumeration, this intention does not exist. For example, to aid understanding, the appended claims may contain the use of introductory phrases “at least one” and “one or more” to introduce the enumeration of claims. However, the use of such phrases should not be considered as implying that a list of claims introduced by the indefinite article "a(a)" or "an" will limit any particular claim containing such an introductory list to containing only one implementation of such a list, even when the same claim includes the introductory phrase "one or more" or "at least one" and indefinite articles such as "a(a)" or "an," for example, "a(a)" or "an" should be interpreted as meaning "at least one" or "one or more"; it remains true for the use of definite articles for introductory list of claims. Furthermore, even when a specific number of introductory list of claims is explicitly stated, those skilled in the art should recognize that such a list should be interpreted as meaning at least the stated number; for example, a bare list of "two lists" without other modifiers means at least two lists, or two or more lists. Furthermore, in instances using conventions such as "at least one of A, B, and C," this syntactic structure generally aims to ensure that a person skilled in the art understands the meaning of this convention. For example, "a system having at least one of A, B, and C" should include, but is not limited to, systems having a single A, a single B, a single C, A and B together, A and C together, B and C together, and / or A, B, and C together. In instances using conventions such as "at least one of A, B, or C," this syntactic structure generally aims to ensure that a person skilled in the art understands the meaning of this convention. For example, "a system having at least one of A, B, or C" should include, but is not limited to, systems having a single A, a single B, a single C, A and B together, A and C together, B and C together, and / or A, B, and C together.Those skilled in the art should also understand that, in practice, any transition words and / or phrases presenting two or more alternative terms (whether in the specification, claims, or drawings) should be understood to imply the possibility of including one, any, or both of these terms. For example, the phrase "A or B" should be understood to include the possibility of including "A" or "B" or "A and B".

[0128] It should be noted that, although not explicitly specified, one or more steps of the methods described herein may include storage, display, and / or output steps as needed for a particular application. In other words, any data, records, fields, and / or intermediate results discussed in the methods may be stored, displayed, and / or output to another device as needed for a particular application. While the foregoing descriptions are directed to embodiments of the invention, other and more embodiments of the invention may be designed without departing from its basic scope. The various embodiments or portions thereof given herein may be combined to create more embodiments. The foregoing descriptions of embodiments of the invention present the best mode for carrying out the invention. The foregoing descriptions of embodiments of the invention are for illustrative purposes only and should not be construed as limiting the invention. The scope of protection of the invention is determined by the appended claims.

[0129] The preceding paragraphs use multiple levels of description. Clearly, the teachings herein can be implemented in various ways, and any particular architecture or functionality disclosed in the examples is merely a representative case. Based on the teachings herein, those skilled in the art will understand that the various levels disclosed herein can be implemented independently or that two or more levels can be implemented in combination.

[0130] Although this disclosure has been provided above with reference to embodiments, it is not intended to limit this disclosure. Any person skilled in the art may make some modifications and refinements without departing from the spirit and scope of this disclosure. Therefore, the scope of protection of this invention shall be determined by the appended claims.

Claims

1. A traffic analysis device, comprising: A camera used to capture a video of a scene associated with a road; One processor; as well as A memory, coupled to the aforementioned processor, The aforementioned processor is used for: A first segment of the aforementioned video is analyzed using a first machine learning model to generate a first traffic description; Embed the first traffic description into a first video encoding parameter that is synchronized with the first segment of the aforementioned video; and Using a second machine learning model, a second segment of the aforementioned video is analyzed based on the first video encoding parameters to generate a second traffic description.

2. The traffic analysis device as claimed in claim 1, wherein the first machine learning model may include a convolutional neural network, and the second machine learning model may include a transformation model.

3. The traffic analysis apparatus of claim 2, wherein the processor is further configured to: During a first time period, the aforementioned first traffic description is generated; and During a second time period, the aforementioned second traffic description is generated from the aforementioned first segment to the aforementioned second segment in the aforementioned video. The second time period mentioned above is longer than the first time period mentioned above.

4. The traffic analysis apparatus of claim 2, wherein the processor is further configured to: Using the first machine learning model described above, a first position of an object in the first segment of the aforementioned video is obtained; The first location is encoded into the first traffic description. Using the second machine learning model described above, a trajectory of the object from the first segment to the second segment is obtained; and The above trajectory is encoded into the second traffic description mentioned above.

5. The traffic analysis apparatus of claim 4, wherein the processor is further configured to: Using the first machine learning model described above, a first category is determined for the object in the first segment described above; The first classification code is assigned to the first traffic description; and Using the second machine learning model described above, a description of the object from the first segment to the second segment is generated based on the first classification and the trajectory described above.

6. The traffic analysis apparatus of claim 5, wherein the processor is further configured to: Determine the computing resources of the aforementioned traffic analysis device; Based on the aforementioned computing resources, determine whether the aforementioned first category and the aforementioned first position of the object have been generated; and When the aforementioned computing resources are insufficient, the first position will be generated first.

7. The traffic analysis apparatus of claim 1, wherein the second segment of the video can be continued from the first segment of the video, and the processor is further configured to embed the second traffic description into a second video encoding parameter synchronized with the second segment of the video.

8. The traffic analysis apparatus of claim 1, wherein the processor is further configured to: The first machine learning model described above is used to obtain a second position and a second classification of an object in the second segment described above; and The second machine learning model is used to analyze the second segment, the first video encoding parameters, the second position, and the second classification to generate a description associated with the object.

9. The traffic analysis apparatus of claim 1, wherein the video may include a traffic signal, and the processor is further configured to determine a traffic signal stage of the traffic signal using the first machine learning model.

10. The traffic analysis apparatus of claim 1, wherein the processor is further configured to: A traffic signal stage that obtains the aforementioned traffic signal from a traffic controller connected to a traffic signal.

11. The traffic analysis apparatus of claim 1, wherein the processor is further configured to: When a server indicates that the accuracy of the first machine learning model is lower than the accuracy of another machine learning model, the first machine learning model is switched to the other machine learning model.

12. The traffic analysis apparatus of claim 1, wherein the processor is further configured to: Based on an instruction from a server, determine whether to update the first machine learning model and the second machine learning model mentioned above. The server calculates a computing resource of the traffic analysis device to generate the instruction.

13. The traffic analysis apparatus of claim 1, wherein the first segment and the second segment respectively comprise a first group of images and a second group of images, and wherein the first video encoding parameter and the second video encoding parameter respectively comprise a first supplementary enhancement information following the first group of images and a second supplementary enhancement information following the second group of images.

14. The traffic analysis apparatus of claim 1, wherein the processor is further configured to: When a traffic signal is in a green light phase, a traffic flow calculation model is used to calculate the number of vehicles on the road in the segment preceding the first segment of the aforementioned video. When the traffic signal is in the yellow light phase, a fast vehicle trajectory model is used to generate the vehicle trajectories in a segment of the aforementioned video; and When the traffic signal is red, a convoy length calculation model is used to generate the total length of vehicles waiting at the red light on the road in the above video clip.

15. The traffic analysis apparatus of claim 1, wherein the processor is further configured to: Using the first machine learning model described above, a decentralized traffic description is generated based on multiple videos from different cameras. Embed the aforementioned dispersed traffic descriptions into a video encoding parameter for each of the aforementioned videos; and Based on the above description of dispersed traffic, determine whether to switch to another machine learning model.

16. A traffic analysis method applicable to a traffic analysis device, comprising: Using a camera in the aforementioned traffic analysis device, a video of a scene associated with a road is extracted; A processor of the aforementioned traffic analysis device uses a first machine learning model to analyze a first segment of the aforementioned video to generate a first traffic description. The processor embeds the first traffic description into a first video encoding parameter synchronized with the first segment of the video; and Using a second machine learning model, the processor analyzes a second segment of the video based on the first video encoding parameters to generate a second traffic description.

17. The traffic analysis method of claim 16, wherein the first machine learning model may include a convolutional neural network, and the second machine learning model may include a transformation model.

18. The traffic analysis method as described in claim 17, further comprising: The processor described above generates the first traffic description during a first time period. as well as The processor generates the second traffic description of the video from the first segment to the second segment during a second time period. The second time period mentioned above is longer than the first time period mentioned above.

19. The traffic analysis method as described in claim 17, further comprising: The processor obtains a first position of an object in the first segment of the video using the first machine learning model. The processor encodes the first location into the first traffic description. The processor, through the second machine learning model, obtains a trajectory of the object from the first segment to the second segment; and The aforementioned processor encodes the aforementioned trajectory into the aforementioned second traffic description.

20. The traffic analysis method as described in claim 19, further comprising: The processor uses the first machine learning model to determine a first category of the object in the first segment. The processor encodes the first category into the first traffic description; and The processor, through the second machine learning model, generates a description of the object from the first segment to the second segment based on the first classification and the trajectory.

21. The traffic analysis method as described in claim 20, further comprising: The processor determines the computing resources of the traffic analysis device. The processor, based on the aforementioned computing resources, determines whether the aforementioned first category and the aforementioned first position of the object have been generated; and When the aforementioned computing resources are insufficient, the aforementioned processor will be used to generate the aforementioned first position first.

22. The traffic analysis method of claim 16, wherein the second segment of the video can be followed by the first segment of the video, and the processor is further configured to embed the second traffic description into a second video encoding parameter synchronized with the second segment of the video.

23. The traffic analysis method as described in claim 16, further comprising: The processor described above obtains a second position and a second classification of an object in the second segment using the first machine learning model described above. as well as The processor analyzes the second segment, the first video encoding parameters, the second position, and the second classification using the second machine learning model to generate a description associated with the object.

24. The traffic analysis method of claim 16, wherein the video may include a traffic signal, and the processor is further configured to determine a traffic signal stage of the traffic signal using the first machine learning model.

25. The traffic analysis method as described in claim 16, further comprising: The processor obtains a traffic signal phase from a traffic controller connected to a traffic signal.

26. The traffic analysis method as described in claim 16, further comprising: When a server indicates that the accuracy of the first machine learning model is lower than the accuracy of another machine learning model, the processor switches the first machine learning model to the other machine learning model.

27. The traffic analysis method as described in claim 16, further comprising: Based on an instruction from a server, the processor determines whether to update the first machine learning model and the second machine learning model. The server calculates a computing resource of the traffic analysis device to generate the instruction.

28. The traffic analysis method of claim 16, wherein the first segment and the second segment respectively comprise a first group of images and a second group of images, and wherein the first video encoding parameter and the second video encoding parameter respectively comprise a first supplementary enhancement information following the first group of images and a second supplementary enhancement information following the second group of images.

29. The traffic analysis method as described in claim 16, further comprising: When a traffic signal is in a green light phase, the processor uses a traffic flow calculation model to calculate the number of vehicles on the road in a segment preceding the first segment of the video. When the traffic signal is in the yellow light phase, the processor uses a fast vehicle trajectory model to generate the trajectory of the vehicle in a segment of the video. as well as When the traffic signal is red, the processor uses a convoy length calculation model to generate the total length of vehicles waiting at the red light on the road in the aforementioned video clip.

30. The traffic analysis method as described in claim 16, further comprising: The processor described above uses the first machine learning model to generate a distributed traffic description based on multiple videos from different cameras. The processor described above embeds the dispersed traffic description into a video encoding parameter of each of the aforementioned videos; and Based on the aforementioned description of dispersed traffic, the processor determines whether to switch to another machine learning model.