Methods, systems and storage medium for traffic condition recognition

By leveraging temporal correlation between traffic scenario and participant features, the method and system improve the accuracy and adaptability of traffic condition recognition, addressing the limitations of existing methods in handling complex real-world scenarios.

WO2026020986A1PCT designated stage Publication Date: 2026-01-29ZHEJIANG DAHUA TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/098091
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-23
Filing Date
2025-05-29
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing traffic condition recognition methods face challenges due to the complexity and diversity of real-world scenarios, leading to limited applicability and low accuracy in image feature extraction and logical determination.

Method used

A method and system that leverage temporal correlation between global features of traffic scenarios and local features of traffic participants, using a traffic condition recognition model to adaptively learn rules for different scenarios, improving accuracy by generating and utilizing first and second temporal features.

Benefits of technology

Enhances the accuracy of traffic condition recognition by comprehensively considering local and global features, demonstrating strong adaptability and improving the precision of traffic condition recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025098091_29012026_PF_FP_ABST
    Figure CN2025098091_29012026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is a method and system for traffic condition recognition and a storage medium. The method comprises: obtaining a plurality of images of a target region acquired at consecutive time points in a time period; obtaining a first initial feature of a target traffic participant and a second initial feature of a traffic scenario in each of at least a portion of the plurality of images by processing the plurality of images; generating, based on the first initial feature, a first temporal feature of the target traffic participant; generating, based on the second initial feature, a second temporal feature of the traffic scenario; and determining, based on the first temporal feature and the second temporal feature, a traffic condition recognition result of the target region.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS, SYSTEMS AND STORAGE MEDIUM FOR TRAFFIC CONDITION RECOGNITIONCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to Chinese Application No. 202410988200.6, filed on July 23, 2024, the entire contents of which are incorporated herein by reference.TECHNICAL FIELD

[0002] The present disclosure relates to the field of intelligent transportation technology, and in particular to a method, a system, and storage medium for traffic condition recognition.BACKGROUND

[0003] With the rapid development of intelligent transportation systems, the demand for accurate traffic condition recognition is also increasing. In this context, traffic condition recognition is typically achieved through image feature extraction or logical determination. However, the complexity and diversity of real-world traffic scenarios pose significant challenges to these methods, resulting in limited applicability and low accuracy in traffic condition recognition.

[0004] Accordingly, the present disclosure provides a method and system for traffic condition recognition to solve or enhance the above problems.SUMMARY

[0005] One or more embodiments of the present disclosure provide a method for traffic condition recognition. The method may comprise: obtaining a plurality of images of a target region acquired at consecutive time points in a time period; obtaining a first initial feature of a target traffic participant and a second initial feature of a traffic scenario in each of at least a portion of the plurality of images by processing the plurality of images; generating, based on the first initial feature, a first temporal feature of the target traffic participant; generating, based on the second initial feature, a second temporal feature of the traffic scenario; and determining, based on the first temporal feature and the second temporal feature, a traffic condition recognition result of the target region.

[0006] One or more embodiments of the present disclosure provide a system for traffic condition recognition. The system may comprise a storage device and at least one processor. The storage device may be configured to store executable instructions. The at least one processor may include: an image acquisition module configured to obtain a plurality of images of a target region acquired at consecutive time points in a time period; a first feature acquisition module configured to obtain a first initial feature of a target traffic participant and a second initial feature of a traffic scenario in each of at least a portion of the plurality of images by processing the plurality of images; a second feature acquisition module configured to generate, based on the first initial feature, a first temporal feature of the target traffic participant; a third feature acquisition module configured to generate, based on the second initial feature, a second temporal feature of the traffic scenario; and an identification module configured to determine, based on the first temporal feature and the second temporal feature, a traffic condition recognition result of the target region.

[0007] One or more embodiments of the present disclosure provide a non-transitory computer-readable storage medium, comprising executable instructions that, when read by at least one processor, may direct the at least one processor to implement a method for traffic condition recognition. The method may include: obtaining a plurality of images of a target region acquired at consecutive time points in a time period; obtaining a first initial feature of a target traffic participant and a second initial feature of a traffic scenario in each of at least a portion of the plurality of images by processing the plurality of images; generating, based on the first initial feature, a first temporal feature of the target traffic participant; generating, based on the second initial feature, a second temporal feature of the traffic scenario; and determining, based on the first temporal feature and the second temporal feature, a traffic condition recognition result of the target region.BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The present disclosure will be further illustrated by way of exemplary embodiments, which will be described in detail by means of the accompanying drawings. These embodiments are not limiting, and in these embodiments, the same numbering indicates the same structure, wherein:

[0009] FIG. 1 is a schematic diagram illustrating an application scenario of a system for traffic condition recognition according to some embodiments of the present disclosure;

[0010] FIG. 2 is a flowchart illustrating an exemplary process for traffic condition recognition according to some embodiments of the present disclosure;

[0011] FIG. 3 is a flowchart illustrating an exemplary process of determining a target traffic participant according to some embodiments of the present disclosure;

[0012] FIG. 4 is a flowchart illustrating an exemplary process of determining a first initial feature and a second initial feature according to some embodiments of the present disclosure;

[0013] FIG. 5 is a flowchart illustrating an exemplary process of determining a local temporal feature according to some embodiments of the present disclosure;

[0014] FIG. 6 is a schematic diagram illustrating an exemplary process of determining a local temporal feature according to some embodiments of the present disclosure;

[0015] FIG. 7 is a flowchart illustrating an exemplary process of determining a traffic condition recognition result of a target region according to some embodiments of the present disclosure;

[0016] FIG. 8 is a schematic diagram illustrating an exemplary traffic condition recognition model according to some embodiments of the present disclosure;

[0017] FIG. 9 is a schematic structural diagram illustrating an attention model according to some embodiments of the present disclosure;

[0018] FIG. 10 is a flowchart illustrating an exemplary process for traffic condition recognition according to some embodiments of the present disclosure;

[0019] FIG. 11 is a flowchart illustrating an exemplary process of determining a first initial feature and a second initial feature according to some embodiments of the present disclosure;

[0020] FIG. 12 is a flowchart illustrating an exemplary process of determining a first temporal feature and a second temporal feature according to some embodiments of the present disclosure;

[0021] FIG. 13 is a flowchart illustrating an exemplary process of determining a traffic condition recognition result according to some embodiments of the present disclosure;

[0022] FIG. 14 is a block diagram illustrating an apparatus for traffic condition recognition according to some embodiments of the present disclosure;

[0023] FIG. 15 is a schematic diagram illustrating a system for traffic condition recognition according to some embodiments of the present disclosure; and

[0024] FIG. 16 is a schematic diagram illustrating a non-transitory computer-readable storage medium according to some embodiments of the present disclosure.DETAILED DESCRIPTION

[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required to be used in the description of the embodiments are briefly described below. Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present disclosure, and it is possible for a person of ordinary skill in the art to apply the present disclosure to other similar scenarios in accordance with these drawings without creative labor. Unless obviously obtained from the context or the context illustrates otherwise, the same numeral in the drawings refers to the same structure or operation.

[0026] It should be understood that the terms “system, ” “device, ” “unit” and / or “module” used herein are a way to distinguish between different components, elements, parts, sections, or assemblies at different levels. However, the terms may be replaced by other expressions if other words accomplish the same purpose.

[0027] As shown in the present disclosure and in the claims, unless the context clearly suggests an exception, the words “one, ” “a, ” “an, ” “one kind, ” and / or “the” do not refer specifically to the singular, but may also include the plural. Generally, the terms “including” and “comprising” suggest only the inclusion of clearly identified steps and elements, however, the steps and elements that do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0028] Flowcharts are used in the present disclosure to illustrate the operations performed by a system according to embodiments of the present disclosure, and the related descriptions are provided to aid in a better understanding of the magnetic resonance imaging method and / or system. It should be appreciated that the preceding or following operations are not necessarily performed in an exact sequence. Instead, steps can be processed in reverse order or simultaneously. Also, it is possible to add other operations to these processes or to remove a step or steps from these processes.

[0029] In the related technologies, methods for traffic condition recognition often relay on image feature extraction or logical determination is often used. However, a logical determination module tends to be complex and exhibit limited applicability. While a general video identification method may be utilized, it does not specifically account for the unique characteristics of road traffic scenarios and fails to effectively harness the dynamic changes of traffic participants, leading to low accuracy.

[0030] According to the method for traffic condition recognition disclosed in some embodiments of the present disclosure, traffic condition recognition is achieved through global encoding that comprehensively leverages a temporal correlation between global features of the traffic scenarios and local features of the traffic participants, to allow for the adaptive extraction of temporal transformation between the traffic scenarios and the traffic participants during a specific traffic incident, thereby effectively improving the accuracy of traffic condition recognition.

[0031] Specifically, in the process of traffic condition recognition, a local feature, a position feature, and a category feature of a target traffic participant and a global feature of a traffic scenario in a plurality of images are comprehensively considered. By utilizing a traffic condition recognition model, the model can adaptively learn the rules or logics pertinent to different traffic scenarios, or different traffic incidents, or the like, thereby demonstrating strong adaptability and further improving the accuracy of traffic condition recognition.

[0032] FIG. 1 is a schematic diagram illustrating an application scenario of a system for traffic condition recognition according to some embodiments of the present disclosure.

[0033] In some embodiments, as shown in FIG. 1, a system 100 for traffic condition recognition (also referred to as a traffic condition recognition system 100) may include an image acquisition device 110, a network 120, a terminal 130, a processor 140, a storage device 150, etc.

[0034] The image acquisition device 110 may include various devices with image acquisition and data transmission functions. For example, the image acquisition device 110 may include a visible light imaging device, an infrared imaging device, a thermal imaging device, a 3D imaging device, or the like.

[0035] In some embodiments, the image acquisition device 110 may be configured to obtain a plurality of images of a target region. For example, the image acquisition device 110 may obtain a plurality of images representing the target region including a target traffic participant and a traffic scenario. In some embodiments, the image acquisition device 110 may obtain a video of the target region. The plurality of images of the target region may be a plurality of image frames extracted from the video.

[0036] The network 120 may include any suitable network that facilitates information and / or data exchange of the system for traffic condition recognition. In some embodiments, one or more components (e.g., the image acquisition device 110, the network 120, the terminal 130, the processor 140, the storage device 150, etc. ) of the system for traffic condition recognition may perform information and / or data communication with one or more other components of the system via the network 120. For example, the image acquisition device 110 may send the plurality of images (or the video) to the processor 140 via the network 120.

[0037] The terminal 130 may be configured to provide a functional component related to user interaction and realize a user interaction function (such as providing or displaying information and data to a user) . The user refers to management personnel of the system for traffic condition recognition, processing personnel of a traffic incident, a traffic participant, etc. In some embodiments, the terminal 130 may display a traffic condition recognition result to the user. Merely by way of example, the terminal 130 may include a mobile device, a tablet computer, a laptop computer, a desktop computer, or other devices with input and / or output functions, or any combination thereof.

[0038] The processor 140 may be configured to process information and / or data related to the system for traffic condition recognition to implement one or more functions described in the present disclosure. In some embodiments, the processor 140 may obtain the plurality of images of the target region acquired at consecutive time points in a time period; obtain a first initial feature of a target traffic participant and a second initial feature of a traffic scenario in each of at least a portion of the plurality of images by processing the plurality of images; generate, based on the first initial feature, a first temporal feature of the target traffic participant; generate, based on the second initial feature, a second temporal feature of the traffic scenario; determine, based on the first temporal feature and the second temporal feature, a traffic condition recognition result of the target region. More description may be found in the present disclosure below (e.g., FIGs. 2-7 and the related descriptions thereof) .

[0039] In some embodiments, the processor 140 may include a central processing unit (CPU) , a digital signal processor (DSP) , a system on chip (SoC) , a microcontroller unit (MCU) , a computer, a user console, or the like, or any combination thereof. In some embodiments, the processor 140 may include a single server or a server group. The server group may be centralized or distributed. In some embodiments, the processor 140 may be local or remote. In some embodiments, the processor 140 may be implemented on a cloud platform. Merely by way of example, the cloud platform may include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an internal cloud, a multi -layer cloud, or the like, or any combination thereof. In some embodiments, the processor 140 may be integrated or included in one or more other components of the system 100 for traffic condition recognition.

[0040] The storage device 150 may be configured to store data, instructions, and / or any other information. In some embodiments, the storage device 150 may store data obtained from the image acquisition device 110 and  / or the processor 140. For example, the storage device 150 may store the plurality of images of the target region acquired at the consecutive time points in the time period. As another example, the storage device 150 may store the traffic condition recognition result of the target region, etc. In some embodiments, the storage device 150 may include a large-capacity memory, a removable memory, a volatile read-write memory, a read-only memory (ROM) , or the like, or any combination thereof. In some embodiments, the storage device 150 may be executed on the cloud platform. In some embodiments, the storage device 150 may be connected to the network 120 to communicate with one or more other components (e.g., the image acquisition device 110, the processor 140, etc. ) of the system for traffic condition recognition.

[0041] It should be noted that the system 100 for traffic condition recognition is provided only for illustrative purposes and is not intended to limit the scope of the present disclosure. For those having ordinary skills in the art, various changes and modifications can be made according to the description of the present disclosure. For example, the system 100 for traffic condition recognition may further include a database, an information source, etc. As another example, the system 100 for traffic condition recognition may be implemented on other devices to achieve similar or different functions. However, these changes and modifications will not deviate from the scope of the present disclosure.

[0042] FIG. 2 is a flowchart illustrating an exemplary process for traffic condition recognition according to some embodiments of the present disclosure. In some embodiments, a process 200 may be performed by a processing device (e.g., the processor 140) . As shown in FIG. 2, the process 200 may include the following operations.

[0043] In 210, a plurality of images of a target region acquired at consecutive time points in a time period may be obtained.

[0044] The target region refers to a region where traffic monitoring for managing or analyzing vehicle movement and pedestrian activity is required. The target region may be a region within a range of traffic monitoring. For example, the target region may be a road intersection, a community access, a parking lot access, etc., which is not limited in the embodiments of the present disclosure.

[0045] In some embodiments, the processor 140 may obtain the plurality of images of the target region through the image acquisition device 110 (e.g., a still camera, a video camera, a surveillance camera, etc. ) . For example, the processor 140 may obtain the plurality of images of the target region by performing image acquisition on the target region using the image acquisition device 110. The plurality of images are consecutive images. As another example, the plurality of images obtained through the image acquisition device 110 may be stored in the storage device 150 or a database, and the processor may obtain the plurality of images of the target region by reading from the storage device or the database.

[0046] In some embodiments, the processor 140 may obtain a target video of the target region; and obtain the plurality of images from the target video according to a preset sliding window.

[0047] The target video refers to video data acquired by the image acquisition device 110 whose field of view includes the target region. The target video may be configured for traffic condition recognition. In some embodiments, the target video may be real-time or non-real-time.

[0048] The preset sliding window refers to a sequence data processing technique for image extraction on an image sequence of the target video. The preset sliding window may be configured to extract image fragments by gradually moving a fixed-size window on the image sequence of the target video.

[0049] The preset sliding window may be configured as a sliding window for traffic condition recognition, and the processor 140 may perform traffic condition recognition on the plurality of images obtained by each preset sliding window. For example, each preset sliding window includes T images, and a spacing between different preset sliding windows (i.e., a count of images between first images of two adjacent preset sliding windows) is L images. If the target video includes 10 images, T is 5 images, and L is 2 images, the processor 140 uses 1st to 5th images as T images of a first preset sliding window, and 3rd to 7th images as T images of a second preset sliding window, and so on, to obtain the plurality of images of each of a plurality of preset sliding windows.

[0050] In some embodiments, only one preset sliding window may be provided, and the plurality of images may be obtained by moving the preset sliding window in a certain step size, such as 2 images, 3 images, etc., and performing image extraction in the preset sliding window after each movement.

[0051] In some embodiments, the processor 140 may obtain the plurality of consecutive images by segmenting the target video of the target region, and perform traffic condition recognition on the plurality of images corresponding to each segment of the target video.

[0052] It is understood that the processor 140 may determine the plurality of images of the target region in other ways, which is not limited in the embodiments of the present disclosure. For ease of description, the embodiments of the present disclosure take the plurality of images as T images as an example for description.

[0053] In 220, a first initial feature of each of one or more target traffic participants and a second initial feature of a traffic scenario in each of at least a portion of the plurality of images may be obtained by processing the plurality of images.

[0054] The one or more target traffic participants satisfy a preset tracking condition. In some embodiments, the preset tracking condition may include appearing in a preset count of consecutive images. For example, each of the one or more target traffic participants appears in each of the preset count of consecutive images. In some embodiments, the count (second count) of the one or more traffic participants appearing in the consecutive images may be greater than a count threshold. The preset count of consecutive images may be less than or equal to T images. The count threshold may be greater than or equal to 2. For example, T may be 5. The second count may be an integer greater than 1, which is not limited in the embodiments of the present disclosure.

[0055] The traffic participant refers to a subject in a motion state or is about to enter the motion state in the target region. In some embodiments, the traffic participant may include a subject that has a direct or indirect relationship with the traffic incident, such as a user (e.g. a driver) of a vehicle, a movable subject (e.g., a vehicle, a pedestrian, etc. ) , etc. For example, the traffic participant may mainly include movable subjects, users of various vehicles, and subjects that use part of the public road traffic resources corresponding to the needs. For example, the traffic participant may include a pedestrian, a motor vehicle, a non-motor vehicle, etc. More descriptions regarding determining the target traffic participant may be found in FIG. 3 and the related descriptions thereof.

[0056] The traffic scenario refers to a specific environment or situation for traffic condition recognition. In some embodiments, the traffic scenario may include the target region. For example, the traffic scenario may include the traffic participant and the target region where the traffic participant is located, which is not limited in the embodiments of the present disclosure.

[0057] The first initial feature refers to feature information extracted from one of the plurality of images that can describe attributes of the target traffic participant represented in the image. In some embodiments, the first initial feature may include a position feature, a category feature, and / or a local feature of the target traffic participant. More descriptions regarding the position feature, the category feature, and the local feature may be found in FIG. 3 and the related descriptions thereof.

[0058] In some embodiments, a plurality of first initial features may be provided, and each of the plurality of first initial features may correspond to and be extracted from one of the plurality of images.

[0059] The second initial feature refers to a global feature of the traffic scenario. The global feature may be configured to describe the overall attributes of the entire traffic scenario. In some embodiments, the global feature may include information such as a shape, a color distribution, a texture pattern, or the like, of one of the plurality of images corresponding to the traffic scenario. In some embodiments, a plurality of second initial features may be provided, and each of the plurality of second initial features may correspond to and be extracted from one of the plurality of images.

[0060] In some embodiments, the processor 140 may obtain the first initial feature of the target traffic participant and the second initial feature of the traffic scenario in each of the at least a portion of the plurality of images in the following manner. For each of the at least a portion of the plurality of images, the position feature and the category feature of the target traffic participant may be extracted from the image by performing target detection on the image. The local feature of the target traffic participant may be obtained by performing a downsampling operation with a first multiple on the image during the target detection. The second initial feature of the traffic scenario may be obtained by performing a downsampling operation with a second multiple on the image during the target detection. The second initial feature of the traffic scenario is the global feature of the traffic scenario, the second multiple is greater than the first multiple, and the first initial feature includes the position feature, the category feature, and the local feature of the target traffic participant. More descriptions regarding the first initial feature and the second initial feature may be found in FIG. 4 and the related descriptions thereof.

[0061] It should be noted that each of the at least a portion of the plurality of images corresponds to one first initial feature and one second initial feature. Unless otherwise specified, the first initial feature, the second initial feature described in the present disclosure, etc., refer to features corresponding to each of the at least a portion of the plurality of images.

[0062] In 230, a first temporal feature of the target traffic participant may be generated based on first initial features obtained based on the at least a portion of the plurality of images.

[0063] The first temporal feature refers to a feature that reflects a change in the first initial features in a time dimension, such as a feature that the first initial features changes over time. In some embodiments, the first temporal feature may include a temporal sequence between the first initial features reflecting a change of the first initial features at different times corresponding to the at least a portion of the plurality of images. For example, the temporal sequence between the first initial features may include multiple elements and each of the multiple elements represents a change between first initial features acquired from two adjacent images among the at least a portion of the plurality of images. Each of the at least a portion of the plurality of images may be acquired at a time point and have a timestamp, such that the temporal sequence may represent multiple changes of the first initial features at the different time points when the at least a portion of the plurality of images are acquired. In some embodiments, the first temporal feature may include a position temporal feature, a category temporal feature, and / or a local temporal feature.

[0064] The position temporal feature refers to a feature that reflects a change in the position feature of the target traffic participant in the at least a portion of the plurality of images in the time dimension. In some embodiments, the position temporal feature may be configured to characterize a position of the target traffic participant in each of the at least a portion of the plurality of images. For example, the position temporal feature reflects a change in the position of the target traffic participant in the at least a portion of the plurality of images over time.

[0065] The category temporal feature refers to a feature that reflects a change in the category feature of the target traffic participant in the at least a portion of the plurality of images in the time dimension. In some embodiments, the category temporal feature may be configured to characterize a category of the target traffic participants in the at least a portion of the plurality of images. For example, in the at least a portion of the plurality of images, the category and a count of the target traffic participants may change over time (e.g., a vehicle or a pedestrian leaves the target region, resulting in a decrease in the count and the category of the target traffic participant) .

[0066] The local temporal features refer to a feature that reflects a change in the local feature of the target traffic participant in the at least a portion of the plurality of images the time dimension. In some embodiments, the local temporal feature may be configured to characterize a change in the image information of the at least a portion of the plurality of images, such as a texture, a color, a shape, or the like, of the image.

[0067] In some embodiments, the processor 140 may obtain the first temporal feature of the target traffic participant by synthesizing the first initial features of the at least a portion of the plurality of images (e.g. the T images) . For example, the processor 140 may respectively synthesize the position features, the category features, and the local features of the target traffic participant in the at least a portion of the plurality of images to obtain the corresponding position temporal feature, the category temporal feature, and the local temporal feature. The synthesizing may be splicing of features. For example, the processor 140 may obtain the first temporal feature of the target traffic participant by splicing the first initial feature of each target traffic participant in each of the at least a portion of the plurality of images (e.g., the T images) .

[0068] In some embodiments, the processor 140 may obtain the position temporal feature of the target traffic participant by splicing the position feature of the target traffic participant in each of the at least a portion of the plurality of images, and obtain the category temporal feature of the target traffic participant by splicing the category feature of the target traffic participant in each of the at least a portion of the plurality of images.

[0069] The splicing may include directly adding the position features of the plurality of images. Similarly, the splicing may include directly adding the category features of the plurality of images. Adding may include directly combining features together. In some embodiments, the processor 140 may extract, based on a position range of the position temporal feature of the target traffic participant, an intermediate feature corresponding to the position range from the local feature of the target traffic participant and determine, based on the intermediate feature, the local temporal feature. As described above, the position feature may be represented by the coordinates and the size of a detection frame representing the position of the target traffic participant. The position range of the target traffic participant in each of the at least a portion of the plurality of images may be determined based on the position feature, and the local temporal feature of the target traffic participant may be determined based on the intermediate feature corresponding to the position range.

[0070] For example, the processor 140 may obtain the local temporal feature by extracting, based on the position range (e.g., coordinates corresponding to the position temporal feature F1) of a position temporal feature F1 of the target traffic participant, the candidate feature corresponding to the position range from the local feature of the target traffic participant. As a further example, the processor 140 may extract, based on the coordinates corresponding to the position temporal feature F1 in each of the at least a portion of the plurality of images, the intermediate feature of the position range corresponding to the position temporal feature F1 of the target traffic participant from the local feature of each of the at least a portion of the plurality of images. Due to different sizes or difference sizes of corresponding detection frame sizes of different target traffic participants (i.e., M*T intermediate features of different sizes are obtained) , after global pooling is performed on the intermediate features, the local temporal feature may be obtained, which is denoted as F3 (N 1, M, T) , where N1 denotes a dimension of the local feature.

[0071] In some embodiments, the processor 140 may obtain a first position temporal feature by performing dimensional transformation on the position temporal feature of the target traffic participant; determine, based on a position range of the target traffic participant in the first position temporal feature, a feature extraction range in the local feature of the target traffic participant; and extract, based on the feature extraction range, the intermediate feature corresponding to the position range of the position temporal feature of the target traffic participant from the local feature of the target traffic participant. More descriptions may be found in FIG. 5 and FIG. 6 and the related descriptions thereof.

[0072] In 240, a second temporal feature of the traffic scenario may be generated based on the second initial features obtained based on the at least a portion of the plurality of images.

[0073] The second temporal feature refers to a feature that reflects a change in the second initial feature of the traffic scenario in the plurality of images in the time dimension. For example, the second temporal feature represents a change in the global feature of the traffic scenario in the at least a portion of the plurality of images in the time dimension.

[0074] In some embodiments, the processor 140 may obtain the second temporal features of the traffic scenario by synthesizing the second initial features of the plurality of images (e.g., the T images) . In some embodiments, the second temporal feature may include a temporal sequence between the second initial features reflecting a change of the second initial features at different times corresponding to the at least a portion of the plurality of images. For example, the temporal sequence between the second initial features may include multiple elements and each of the multiple elements represents a change between second initial features acquired from two adjacent images among the at least a portion of the plurality of images. Each of the at least a portion of the plurality of images may be acquired at a time point and have a timestamp, such that the temporal sequence may represent multiple changes of the second initial features at the different time points when the at least a portion of the plurality of images are acquired.

[0075] For example, the processor 140 may obtain the second temporal feature of the traffic scenario by splicing the second initial feature of the traffic scenario in each of the at least a portion of the plurality of images (e.g., the T images) . For example, the processor 140 may obtain the second temporal feature F4 (N3, T) of the traffic scenario by splicing the global feature (i.e., the second initial feature) of the traffic scenario in each of the at least a portion of the plurality of images (e.g. the T images) , wherein, N3 denotes a dimension of the second temporal feature after splicing.

[0076] In 250, a traffic condition recognition result of the target region may be determined based on the first temporal feature and the second temporal feature.

[0077] The traffic condition recognition result refers to a conclusion obtained after performing traffic condition recognition for the target region. The traffic condition recognition result may include a traffic incident type, a traffic incident location, a traffic incident occurrence time, etc.

[0078] The traffic condition type may include a traffic accident, traffic congestion, a vehicle breakdown, a pedestrian crossing the road, a vehicle violation (e.g., running red lights and driving in the wrong direction) , etc.

[0079] The traffic incident position refers to position information where the traffic incident occurs. For example, the incident position may include a road name, an intersection number, longitude and latitude coordinates, or a relative position description (e.g., the middle of a road section, being close to a landmark building, etc. ) .

[0080] In some embodiments, the processor 140 may input the first temporal feature and the second temporal feature into a traffic condition recognition model, the traffic condition recognition model may generate and output the traffic condition recognition result of the target region by performing traffic condition recognition based on a temporal correlation between the first temporal feature and the second temporal feature. More descriptions regarding the temporal correlation between the first temporal feature and the second temporal feature may be found in the present disclosure below.

[0081] When traffic condition recognition is performed based on the first temporal feature and the second temporal feature, since the count of target traffic participants appearing in different preset sliding windows T may be different, the processor may perform filling processing on the first temporal feature corresponding to each of the preset sliding windows. Specifically, the first temporal feature may be filled to make the filled first temporal feature be of the preset target count. The filling processing may be performed with a value of 0. The preset target count M2 is a maximum target count of detected and tracked traffic participant. For example, the preset target count M2 is the maximum count of traffic participants detected and tracked in the T images. Alternatively, the preset target count M2 is a maximum count of traffic participants detected and tracked in the target video, which is not limited in the embodiments of the present disclosure. For each sliding window, the processor 140 may perform the same operations, so that each sliding window (i.e., the T images) may obtain a categorization result to determine whether there is a traffic accident in the T images.

[0082] For example, the processor 140 may perform the filling processing on the first temporal feature using the preset target count M2. Specifically, the processor 140 may perform the filling processing on the position temporal feature F1 (4, M, T) , the category temporal feature F2 (1, M, T) , and the local temporal feature F3 (N1, M, T) of the traffic participant, respectively, to obtain features F’1 (4, M2, T) , F’2 (1, M 2, T) , and F’3 (N1, M2, T) after the filling processing. For example, taking F1 as an example, the original feature of F1 is denoted as (4, M, T) , and filling to (4, M2, T) means to fill M2-M, zeros, which actually means to fill each image (T images in total) with (M2-M) features, where each feature has a dimensionality of 4.

[0083] Furthermore, the processor 140 may obtain the traffic condition recognition result of the target region by inputting the second temporal feature of the traffic scenario and the position temporal feature, the category temporal feature, and the local temporal feature after the filling processing into the traffic condition recognition model to perform traffic condition recognition.

[0084] In some embodiments, the processor 140 may perform a temporal transformation on the first temporal feature and the second temporal feature to obtain a transformation feature, generate an incident identification feature based on the transformation feature obtained after the temporal transformation, and determine the traffic condition recognition results of the target region based on the incident identification feature. More descriptions may be found in FIG. 7 and the related descriptions thereof.

[0085] In some embodiments, the processor 140 may provide a corresponding traffic management strategy based on the traffic condition recognition result.

[0086] The traffic management strategy refers to a treatment measure / method formulated for a traffic incident. In some embodiments, the traffic management strategy may include one or more of a traffic prompt, an accident alarm, or a violation record.

[0087] In some embodiments, different traffic condition recognition results may correspond to different traffic management strategies. For example, if the traffic condition recognition result is a traffic accident, the processor 140 may provide an accident alarm strategy; as another example, if the traffic condition recognition result is a vehicle violation, the processor 140 may provide a violation record strategy; as another example, if the traffic condition recognition result is a pedestrian running a red light, the processor 140 may provide a traffic prompt strategy (e.g., prompting the pedestrian through a voice broadcast or other ways) .

[0088] In some embodiments of the present disclosure, a plurality of images acquired in a target region are acquired and processed to obtain a first initial feature of traffic participants and a second initial feature of the traffic scenario in each of the plurality of images. The first initial features and the second initial features of the plurality of images are respectively integrated to obtain a first temporal feature of traffic participants and a second temporal feature of the traffic scenario. Thus, the temporal features of traffic participants and traffic scenarios in the plurality of images can be obtained. Then, the temporal correlation between the first temporal feature and the second temporal feature is comprehensively utilized to perform traffic condition recognition and obtain a traffic condition recognition result for the target region. The traffic condition recognition result comprehensively considers the features of traffic participants and traffic scenarios in the plurality of images and can adapt to different application scenarios for traffic condition recognition. The result has strong adaptability and can improve the accuracy of traffic condition recognition.

[0089] FIG. 3 is a flowchart illustrating an exemplary process of determining a target traffic participant according to some embodiments of the present disclosure. In some embodiments, a process 300 may be performed by a processing device (e.g., the processor 140) . As shown in FIG. 3, the process 300 may include the following operations.

[0090] In 310, a detection result of one or more traffic participants may be obtained by processing a plurality of images using a target detection model.

[0091] The target detection model is a machine learning model configured to detect the one or more traffic participants in the plurality of images. For example, the target detection model may include any one of a neural network model (NN) , a deep neural network model (DNN) , a convolutional neural network model (CNN) , a Region-CNN (R-CNN) , a YOLO model, or the like, or any combination thereof.

[0092] In some embodiments, the processor 140 may input the plurality of images into the target detection model, and the target detection model may output the detection result of the one or more traffic participants. The detection result of the one or more traffic participants may include a position feature and a category feature of each of the one or more traffic participants in each of at least a portion of the plurality of images.

[0093] The position feature refers to spatial position information of a participant (e.g. the traffic participant and / or the target traffic participant) in each of the at least a portion of the plurality of images. In some embodiments, the target detection model may identify the position of one of the one or more traffic participants in one of the at least a portion of the plurality of images through a detection frame. For example, the position feature represents position coordinates of the detection frame where the traffic participant is located, such as a size of the detection frame, coordinate points (e.g., center point coordinates) , any corner coordinates (e.g., upper left corner coordinates) of the detection frame, etc., which is not limited in the embodiments of present disclosure. For example, the position feature may be expressed as (x, y, w, h) , where x and y respectively represent the coordinate points of the detection frame enclosing the traffic participant, and w and h respectively represent a width and a height of the detection frame enclosing the traffic participant.

[0094] The category feature refers to information indicating the category of a participant (e.g., the traffic participant and / or the target traffic participant) . For example, the category may be denoted as a pedestrian, a motor vehicle, a non-motor vehicle, etc.

[0095] In some embodiments, the target detection model may be obtained by model training through a first training set. The first training set may include a plurality of first samples with first labels. Each of the first samples may include one or more sample traffic images. The sample traffic images may be historical traffic images. A first label of a first sample may include position features and category features of one or more traffic participants in the one or more sample traffic images.

[0096] In 320, one or more target traffic participants may be determined based on the detection result of the one or more traffic participants.

[0097] Since the detection result of the one or more traffic participants includes the position and the category of each of the one or more traffic participants detected from at least a portion of the plurality of images, the target traffic participant may be determined from the one or more traffic participants by screening the one or more traffic participants based on the position and the category.

[0098] In some embodiments, the processor 140 may obtain a tracking result of each of the one or more traffic participants by performing target tracking based on the plurality of images based on the position feature and the category feature of each of the one or more traffic participants; and determine, based on the tracking result, the one or more target traffic participants. The target traffic participants including one of the one or more traffic participants that present in each of the at least a portion of the plurality of images and the at least a portion of the plurality of images includes consecutive images among the plurality of images. In other words, a target traffic participant is represented in each of the consecutive images among the plurality of images.

[0099] Target tracking is a technique applied to the field of computer vision and artificial intelligence (AI) , which aims to track position and motion trajectories of a target subject in different images in real time by detecting and analyzing the target subject in a video or image sequence. In this embodiment, target tracking is configured to track position and motion trajectories of the one or more traffic participants in the plurality of images. The target traffic participant refers to a traffic participant that satisfies a condition. More descriptions regarding the condition may be found in FIG. 2 and the related descriptions thereof (e.g., the operation 220) .

[0100] In some embodiments, the processor 140 may perform target tracking on the plurality of images based on the position feature and the category feature of each of the one or more traffic participants. In some embodiments, the processor 140 may perform target tracking based on the position feature and the category feature of each of the one or more traffic participants and a local feature of each of the one or more traffic participants. The tacking result of each of the one or more traffic participants may be obtained by performing target tracking on the plurality of images.

[0101] The local feature refers to image information corresponding to a region correlated with traffic participants (e.g., the one or more traffic participants and / or the target traffic participant) in the plurality of images. For example, the local feature may include information such as a texture, a color, a shape, etc. of an image of a region where the traffic participants are located in the plurality of images. More descriptions regarding determining the local features may be found in the related descriptions below.

[0102] In some embodiments, the processor 140 may perform target tracking on the plurality of images based on a target tracking network. The target tracking network may include a Mean-Shift algorithm, a Kernelized Correlation Filters (KCF) algorithm, a Simple Online and Realtime Tracking (SORT) algorithm, etc., which is not limited in the embodiments of the present disclosure.

[0103] The tracking result includes tracking and recording information such as position and motion trajectories of the traffic participants in the plurality of images. The tracking result may include the motion trajectory of each of the traffic participants in the plurality of images and the positions of each of the traffic participants in each of the images, and the category of each of the traffic participants. In some embodiments, in the process of performing target tracking on the plurality of images, denoising, auxiliary tracking, or the like, may be performed selectively based on the local feature to improve the accuracy of the tracking result. Being selectively means that the local feature may be used or not. For example, denoising, auxiliary tracking, or the like, may be performed based on the tracking of the position features of the traffic participants in each of images, such as loss of a tracking target, a matching degree or similarity of the tracking target is lower than a preset matching value, and the local features of the traffic participants may also be input into a tracking algorithm such as Kalman filtering for target tracking to obtain the tracking results of the plurality of images.

[0104] In some embodiments, the processor 140 may determine the target traffic participant based on the tracking result. For example, the processor 140 may determine one of the one or more traffic participants that present in each of the at least a portion of the plurality of images as the target traffic participant.

[0105] FIG. 4 is a flowchart illustrating an exemplary process of determining a first initial feature and a second initial feature according to some embodiments of the present disclosure. In some embodiments, a process 400 may be performed by a processing device (e.g., the processor 140) . As shown in FIG. 4, the process 400 may include the following operations.

[0106] In 410, a position feature and a category feature of a target traffic participant may be obtained by performing target detection on each of a plurality of images.

[0107] In some embodiments, the processor 140 may input each of the plurality of images into a target detection model. For example, the processor 140 may input each of the plurality of images into the target detection model. The target detection model may output the position feature and the category feature of the target traffic participant in each of the plurality of images.

[0108] In 420, a local feature of the target traffic participant may be obtained by performing a downsampling operation with a first multiple during the target detection.

[0109] The target detection refers to a process of the target detection model processing an input image. For example, the target detection model may perform operations such as feature extraction, convolution, or the like, on the input image.

[0110] The downsampling operation with the first multiple refers to perform a downsampling operation on a first intermediate layer feature extracted by the target detection model during the target detection based on the first multiple.

[0111] For example, during the target detection, the processor 140 may store an output feature of an intermediate layer of the target detection model as the first intermediate layer feature of the target traffic participant. The first intermediate layer feature may be a feature output by a first layer of the target detection model (also referred to as the target detection network) . For example, the first layer may be a layer other than an input layer or an output layer of the target detection model. For example, the first layer may be a layer before the last layer in the target detection model. As another example, the first layer may be a layer next to the input layer of the target detection model, etc. The intermediate layer feature is not limited in the embodiments of the present disclosure.

[0112] In some embodiments, the processor 140 may extract a shallow feature. For example, a result of the downsampling operation on the first intermediate layer feature based on the first multiple may be used as the shallow feature. As another example, convolution is performed on the result of the downsampling operation on the first intermediate layer feature based on the first multiple to obtain a result of convolution, and the result of the convolution may be designated as the shallow feature the local feature of the target traffic participant may be obtained based on the shallow feature. The shallow feature represents a feature extracted by a shallow network of the target detection model. The shallow feature is closer to the input layer and contains more pixel point information and fine granularity information (e.g., color, texture, edge, or angular information) . The shallow feature has a relatively small receptive field and can capture more details.

[0113] In some embodiments, the processor 140 may obtain the result of the downsampling operation with the first multiple based on the first intermediate layer feature output by the target detection network. The result of the downsampling operation with the first multiple may be obtained by performing the downsampling operation with the first multiple (i.e., the first multiple) P1 on the first intermediate layer feature. Furthermore, the processor 140 may obtain the shallow feature of the target traffic participant based on the first intermediate layer feature. The shallow feature may be configured to extract the local feature of the target traffic participant, so as to obtain the local feature of the target traffic participant. For example, assuming that a height and a width of each of the images are H and W, respectively, and a count of channels of the shallow feature is N1, the shallow feature may be denoted as where the first multiple P1 may be determined based on specific application scenarios and requirements, which is not limited in the embodiments of the present disclosure.

[0114] In some embodiments, the first intermediate layer feature with the downsampling operation with the first multiple of 4 may be directly used as the shallow feature of the target traffic participant, and the shallow feature may be used as local feature of the target traffic participant. Furthermore, the processor 140 may use the position feature, the category feature, and the local feature of the target traffic participant as the first initial feature of the target traffic participant.

[0115] In 430, a second initial feature of a traffic scenario may be obtained by performing a downsampling operation with a second multiple on each of the at least a portion of the plurality of images during the target detection.

[0116] The downsampling operation with the second multiple refers to performing the downsampling operation on a second intermediate layer feature extracted by the target detection model during the target detection based on the second multiple, the second multiple being greater than the first multiple.

[0117] In some embodiments, the processor 140 may obtain a result of the downsampling operation with the second multiple based on a second intermediate layer feature output by the target detection network. The result of the downsampling operation with the second multiple may be obtained by performing the downsampling operation with the second multiple P2 on the second intermediate layer feature. For example, the processor 140 may obtain a deep feature of the traffic scenario by taking the second intermediate layer feature with a second multiple of 32, denote the deep feature as and use the deep feature as the second initial feature of the traffic scenario. The global feature of the traffic scenario may be obtained based on the second initial feature. The second multiple P2 may be determined based on specific application scenarios and requirements, which is not limited in the embodiments of the present disclosure.

[0118] In some embodiments, after the deep feature is obtained, the processor 140 may perform global pooling processing on the deep feature to obtain a feature, where N2 denotes a dimension, the feature generated by global pooling processing may be used as global feature of the traffic scenario, so as to obtain the second initial feature of the traffic scenario.

[0119] In some embodiments, during the target detection, the processor 140 may obtain the global feature of the traffic scenario by extracting the deep feature of the target detection model. The deep feature represents a feature extracted by a deep network of the target detection model. The deep feature is closer to the output layer and contains some coarse granularity information and more abstract semantic information. The deep feature has a relatively large receptive field and can obtain some information about the overall image and has stronger semantic information.

[0120] In some embodiments, the processor 140 may obtain the deep feature of the traffic scenario by performing the downsampling operation with the second multiple P2 on the second intermediate layer feature output by the target detection network. For example, assuming that a height and a width of each of the images are H and W, respectively, and a count of channels of the deep feature is N2, the deep feature may be denoted as where the second initial feature of the traffic scenario may be the global feature of the traffic scenario. More descriptions regarding the second initial feature and the global feature may be found in FIG. 2 and the related descriptions thereof.

[0121] In some embodiments of the present disclosure, by fully considering the position feature, the category feature, and local feature of the target traffic participant, and the global feature of the traffic scenario when the first initial feature and the second initial feature are determined, traffic condition recognition is performed based on more abundant features, thereby improving the accuracy of traffic condition recognition.

[0122] FIG. 5 is a flowchart illustrating an exemplary process of determining a local temporal feature according to some embodiments of the present disclosure. In some embodiments, a process 500 may be performed by a processing device (e.g., the processor 140) . As shown in FIG. 5, the process 500 may include the following operations.

[0123] In 510, a first position temporal feature may be obtained by performing dimensional transformation on a position temporal feature of a target traffic participant.

[0124] The first position temporal feature refers to a feature obtained after performing the dimensional transformation on the position temporal feature of the target traffic participant.

[0125] The dimensional transformation refers to changing a dimension of a feature. For example, the dimensional transformation may include a dimension increase or dimension reduction operation. For example, the processor 140 may obtain the first position temporal feature by performing the dimensional transformation on the position temporal feature of the target traffic participant through one or more of principal component analysis, linear discriminant analysis, factor analysis, kernel PCA, feature interaction construction, etc.

[0126] In some embodiments, the first position temporal feature has the same dimension as the local temporal feature.

[0127] In 520, a feature extraction range may be determined from a local feature of the target traffic participant based on a position range of the target traffic participant in the first position temporal feature.

[0128] The position range of the target traffic participant refers to a coordinate range of a detection frame enclosing the target traffic participant in each of at least a portion of the plurality of images.

[0129] The feature extraction range refers to a coordinate range used to extract the local feature of the target traffic participant in each of the at least a portion of the plurality of images.

[0130] In some embodiments, the processor 140 may use the position range of the target traffic participant in the first position temporal feature as the feature extraction range in the local feature of the target traffic participant. In some embodiments, the feature extraction range may be a coordinate range corresponding to the detection frame of the position temporal feature after the dimensional transformation.

[0131] FIG. 6 is a schematic diagram illustrating an exemplary process of determining a local temporal feature according to some embodiments of the present disclosure. In some embodiments, as shown in FIG. 6, the dimensions of the position temporal feature and the local feature may be different, and it is difficult to directly determine the coordinate range corresponding to the local feature of the target traffic participant based on the position range of the detection frame in the position temporal feature. In some embodiments of the present disclosure, the processor 140 may obtain the first position temporal feature with the same dimension as the local feature by performing dimensional transformation on the position temporal feature, and determine the feature extraction range of the local temporal feature based on the position range of the target traffic participant in the first position temporal feature, such that the local temporal feature extracted subsequently is more accurate.

[0132] In 530, a feature corresponding to the position range of the position temporal feature of the target traffic participant may be extracted from the local feature of the target traffic participant based on the feature extraction range and the feature corresponding to the position range from the local feature of the target traffic participant may be determined as a local temporal feature.

[0133] In some embodiments, as shown in FIG. 6, the processor 140 may extract the local temporal feature from the local feature based on the feature extraction range in the local feature. For example, the processor 140 may extract the feature of the position range corresponding to the first position temporal feature of the target traffic participant from the local feature of each of the at least a portion of the plurality of images based on coordinates corresponding to the feature extraction range in each of the at least a portion of the plurality of images, and obtain the local temporal feature by performing global pooling processing on the features of the target traffic participant from the local features of the at least a portion of the plurality of images.

[0134] FIG. 7 is a flowchart illustrating an exemplary process of determining a traffic condition recognition result of a target region according to some embodiments of the present disclosure. In some embodiments, a process 700 may be performed by a processing device (e.g., the processor 140) . As shown in FIG. 7, the process 700 may include the following operations.

[0135] In 710, a first transformation feature of a target traffic participant may be generated by performing temporal transformation on a first temporal feature.

[0136] The first transformation feature refers to a feature obtained after performing the temporal transformation on the first temporal feature.

[0137] The temporal transformation refers to a process of reconstructing or transforming a temporal feature to extract underlying regularities in a time dimension. In some embodiments, the temporal transformation may enhance the accuracy of subsequent traffic condition recognition.

[0138] In some embodiments, the temporal transformation may include performing global encoding on the first temporal feature. For example, the processor 140 may perform the global encoding on the first temporal feature to obtain a global encoding feature, and designate the global encoding feature of the target traffic participant as the first transformation feature.

[0139] In some embodiments, the first temporal feature may include a position temporal feature, a category temporal feature, and a local temporal feature. The processor 140 may obtain the first transformation feature of the target traffic participant by performing the global encoding on the position temporal feature, the category temporal feature, and the local temporal feature. In this way, the position temporal feature, the category temporal feature, and the local temporal feature of the target traffic participant may be unified.

[0140] In some embodiments, the processor 140 may perform channel transformation on the position temporal feature, the category temporal feature, and the local temporal feature, respectively to obtain a transformed position temporal feature, a transformed category temporal feature, and a transformed local temporal feature, and then obtain the first transformation feature of the target traffic participant by synthesizing the transformed position temporal feature, the transformed category temporal feature, and the transformed local temporal feature.

[0141] In some embodiments, the processor 140 may generate a first mapping feature corresponding to the position temporal feature, a second mapping feature corresponding to the category temporal feature, and a third mapping feature corresponding to the local temporal feature by performing channel mapping on the position temporal feature, the category temporal feature, and the local temporal feature to the same preset channel count; and generating the first transformation feature of the target traffic participant based on the first mapping feature, the second mapping feature, and the third mapping feature.

[0142] The performing channel mapping on the category temporal feature may include mapping the category temporal feature that is discrete to a continuous vector space.

[0143] The channel mapping refers to mapping features with different channel counts to features with the same preset channel count.

[0144] FIG. 8 is a schematic diagram illustrating an exemplary traffic condition recognition model according to some embodiments of the present disclosure.

[0145] A traffic condition recognition model 800 is a machine learning model configured for traffic condition recognition. In some embodiments, the traffic condition recognition model 800 may be various models, such as a neural network model or a deep neural network model. The processor 140 may input a first temporal feature (including a position temporal feature, a category temporal feature, and a local temporal feature) of the target traffic participant and a second temporal feature of a traffic scenario into the traffic condition recognition model 800, and the traffic condition recognition model 800 may output a traffic condition recognition result.

[0146] In some embodiments, the traffic condition recognition model 800 may be a model of a custom structure. For example, the model structure of the traffic condition recognition model 800 may include a custom structure such as a multilayer perceptron layer, an Embedding layer, an Add layer, a Transformer encoder, a self-attention layer, etc.

[0147] For example, as shown in FIG. 8, the traffic condition recognition model 800 may include a global encoding module 810, a first multilayer perceptron layer 820, a connection layer 830, a self-attention layer 840, and a classification layer 850. The global encoding module 810 may include a second multilayer perceptron layer 811, a third multilayer perceptron layer 812, an embedding layer 813, and an Add layer 814.

[0148] In some embodiments, as shown in FIG. 8, the processor 140 may obtain a first transformation feature of the target traffic participant by performing temporal transformation on a first temporal feature using the global encoding module 810. Specifically, the global encoding module 810 may generate a first mapping feature corresponding to the position temporal feature, a second mapping feature corresponding to the category temporal feature, and a third mapping feature corresponding to the local temporal feature by performing channel mapping on the position temporal feature, the category temporal feature, and the local temporal feature to the same preset channel count, and obtain the first transformation feature of the target traffic participant by adding the first mapping feature, the second mapping feature, and the third mapping feature.

[0149] In some embodiments, values of each of the position temporal feature and the local temporal feature may be continuous, and values of the category temporal feature may be discrete. In some embodiments, the processor 140 performing channel mapping on the category temporal feature may include mapping the discrete category temporal feature to a continuous vector space. For example, the processor 140 may obtain a mapping feature of a preset channel count by performing channel mapping on the category temporal feature using the Embedding layer. The Embedding layer is a network layer of a deep learning model and configured to map a discrete input (e.g., a word, a category, etc. ) to the continuous vector space.

[0150] In some embodiments, the processor 140 may input the position temporal feature F1 (4, M, T) into the third multilayer perceptron layer 812 for channel mapping, and the third multilayer perceptron layer 812 may output the first mapping feature.

[0151] In some embodiments, the processor 140 may input the category temporal feature F2 (1, M, T) into the embedding layer 813 for channel mapping, and the embedding layer 813 may output the second mapping feature.

[0152] In some embodiments, the processor 140 may input the local temporal feature F3 (N1, M, T) into the second multilayer perceptron layer 811 for channel mapping, and the second multilayer perceptron layer 811 may output the third mapping feature.

[0153] The processor 140 may generate the first mapping feature corresponding to the position temporal feature, the second mapping feature corresponding to the category temporal feature, and the third mapping feature corresponding to the local temporal feature by performing channel mapping on the position temporal feature, the category temporal feature, and the local temporal feature to the same preset channel count. The preset channel count may be determined based on channel counts of the position temporal feature, the category temporal feature, and the local temporal feature, respectively. For example, the preset channel count may be a maximum value, a sum value, or the like, among the preset channel counts of the position temporal feature, the category temporal feature, and the local temporal feature. As another example, the preset channel count may be predetermined based on a channel count of the local temporal feature. It is understood that the processor 140 may determine the same preset channel count in other ways, and the preset channel count, the mapping manner, etc., are not limited in the embodiments of the present disclosure. Each of the second multilayer perceptron layer 811 and the third multilayer perceptron layer 812 may be a multilayer perceptron (MLP) structure. The embedding layer 813 may be the structure of the Embedding layer, which is not limited in the embodiments of the present disclosure.

[0154] In some embodiments, the processor 140 may add the first mapping feature, the second mapping feature, and the third mapping feature by inputting the first mapping feature, the second mapping feature, and the third mapping feature into the add layer 814 (e.g., the Add layer) . The add layer 814 may output a first transformation feature of the target traffic participant F5 (N, M2, T) , where N denotes the preset channel count. In some embodiments, the first transformation feature of the target traffic participant may also be used as the global feature of the target traffic participant.

[0155] In the embodiments of the present disclosure, with the global encoding module 810 of the traffic condition recognition model 800, global encoding can be performed based on the position temporal feature, the category temporal feature, and the local temporal feature of the target traffic participant, thereby obtaining the global encoding feature of the target traffic participant, and effectively unifying all features of the target traffic participant.

[0156] In some embodiments, the processor 140 may obtain the first transformation feature of the target traffic participant by inputting the first mapping feature, the second mapping feature, and the third mapping feature into a first neural network.

[0157] In some embodiments, the first neural network may be obtained by model training through a second training set. The second training set may include a plurality of second samples with second labels. Each of the second samples may include one or more sample first mapping features, one or more sample second mapping features, and one or more sample third mapping features. The sample first mapping features, the sample second mapping features, and the sample third mapping features may be obtained through historical traffic images. The second label of a second sample may include a sample first transformation feature of the target traffic participant in the historical traffic images.

[0158] In some embodiments of the present disclosure, by obtaining the first transformation feature through the first neural network instead of simply directly adding the features, feature fusion can be achieved, thereby improving the accuracy of subsequent traffic condition recognition.

[0159] In some embodiments, the processor 140 may obtain the first transformation features of the target traffic participant by inputting the first temporal feature (including the position temporal feature, the category temporal feature, and the local temporal feature of the target traffic participant) into a second neural network.

[0160] In some embodiments, the second neural network may be obtained by model training through a third training set. The third training set may include a plurality of third samples with third labels. Each of the third samples may include a sample first temporal feature. The sample first temporal feature may include a sample position temporal feature, a sample category temporal feature, and a sample local temporal feature. The sample first temporal feature may be obtained through the historical traffic images. The third label of a third sample may include a sample first transformation feature of the target traffic participant in the historical traffic images.

[0161] In some embodiments of the present disclosure, the first transformation feature can be determined more quickly and accurately through the second neural network. In addition, the second neural network can automatically learn key features from the data, has strong generalization ability, and can obtain more accurate results even when the data may be biased.

[0162] In 720, a second transformation feature of a traffic scenario may be generated by performing temporal transformation on a second temporal feature.

[0163] The second transformation feature refers to a feature obtained after performing the temporal transformation on the second temporal feature.

[0164] In some embodiments, the processor 140 may obtain the second transformation feature of the traffic scenario by performing channel count transformation on the second temporal feature based on the preset channel count. The preset channel count may be the same as the channel count when the channel count transformation is performed on the first temporal feature.

[0165] In some embodiments, the processor 140 may obtain the second transformation feature of the traffic scenario by performing the temporal transformation on the second temporal feature of the traffic scenario using the first multilayer perceptron layer 820. For example, the processor 140 may obtain a second transformation feature F6 (N, T) of the traffic scenario by performing the channel count transformation on the second temporal feature F4 (N 3, T) using the first multilayer perceptron layer 820 based on a preset channel count N.

[0166] By transforming the second temporal feature into the second transformation feature having the same channel count as the first transformation feature, subsequent processing of the first transformation feature and the second transformation feature can be facilitated.

[0167] In 730, an incident identification feature may be generated by performing a feature encoding operation based on the first transformation feature and the second transformation feature.

[0168] The incident identification feature is feature data that is encoded for traffic condition recognition.

[0169] The incident identification feature may include a temporal correlation between the first temporal feature and the second temporal feature. The temporal correlation represents a dynamic change in the first temporal feature and the second temporal feature in a plurality of images over time, etc. For example, the temporal correlation represents a dynamic change in the target traffic participant and the traffic scenario over time, etc.

[0170] In some embodiments, the processor 140 may obtain the incident identification feature by performing a preprocessing operation such as normalization and / or temporal alignment on the first transformation feature and the second transformation feature to obtain a preprocessed first transformation feature and a preprocessed second transformation feature, then fusing the preprocessed first transformation feature and the preprocessed second transformation feature through manners such as vector splicing, attention weighted fusion, neural network, etc., to obtain a fused feature and finally performing a feature encoding operation on the fused feature through manners of time encoding and space encoding.

[0171] In some embodiments, the processor 140 may generate a third transformation feature by performing the dimensional transformation on the first transformation feature of the target traffic participant; generate a fourth transformation feature by performing the dimensional transformation on the second transformation feature of the traffic scenario; and generate the incident identification feature based on the third transformation feature and the fourth transformation feature.

[0172] The third transformation feature is a feature obtained by performing the dimensional transformation on the first transformation feature. In some embodiments, the third transformation feature may include a global encoding feature of the target traffic participant in the plurality of images. The dimensional transformation on a feature refers to changing the dimension of the feature. For example, the processor 140 may generate the third transformation feature F7 (M2, N*T) by performing the dimensional transformation on the first transformation feature F5 (N, M2, T) of the target traffic participant.

[0173] The fourth transformation feature is a feature obtained by performing the dimensional transformation on the second transformation feature. In some embodiments, the fourth transformation feature may include the global encoding feature of the traffic scenario in the plurality of images. For example, the processor 140 may generate the fourth transformation feature F8 (1, N*T) by performing the dimensional transformation on the second transformation feature F6 (N, T) of the target traffic participant.

[0174] In some embodiments, the processor 140 may generate the incident identification feature by splicing the third transformation feature and the fourth transformation feature using the connection layer 830 (e.g., a Concat layer) of the incident identification model. For example, the processor 140 may generate the incident identification feature F9 (1 + M2, N*T) by splicing the third transformation feature F7 (M2, N*T) and the fourth transformation feature F8 (1, N*T) using the connection layer 830 (e.g., the Concat layer) . In some embodiments of the present disclosure, the processor 140 may transform the first transformation feature and the second transformation feature into the third transformation feature and the fourth transformation feature, respectively, through a dimensional transformation operation, which can eliminate the dimensional difference between the first transformation feature and the second transformation feature, and facilitate the determination of the incident identification feature; in addition, the dimensional transformation operation can also remove redundant information and suppress noise, thereby improving the efficiency and accuracy of the subsequent determination of the traffic condition recognition result.

[0175] In 740, a traffic condition recognition result of a target region may be determined based on the incident identification feature.

[0176] In some embodiments, the processor 140 may determine the traffic condition recognition result of the target region based on a vector distance between the incident identification feature and each of historical incident identification features. For example, a plurality of historical incident identification features and corresponding traffic condition recognition results may be pre-stored in the storage device 150, and the processor 140 may determine the traffic condition recognition result based on the vector distance between the incident identification feature and each of the historical incident identification features. For example, the processor 140 may determine a traffic condition recognition result corresponding to a historical incident identification feature with the smallest vector distance to the incident identification feature or with vector distance to the incident identification feature satisfying a distance threshold as the traffic condition recognition result. The historical incident identification features may be determined based on historical traffic images. The vector distance may include a cosine distance, a Euclidean distance, a Manhattan distance, a Chebyshev distance, etc.

[0177] In some embodiments, the processor 140 may generate an attention feature by processing the incident identification feature using an attention model; and determine the traffic condition recognition result of the target region based on the attention feature.

[0178] The attention model is a model constructed based on an attention mechanism. For example, the attention model may be a Transformer Encoder. In some embodiments, the attention model may be a part of the traffic condition recognition model 800. For example, the attention model may be the self-attention layer 840 of the traffic condition recognition model 800.

[0179] In some embodiments, the attention model may be obtained by training a fourth training set. The fourth training set may include a plurality of fourth samples with fourth labels. Each of the fourth samples may include a sample incident identification feature. The sample incident identification features may be obtained through the historical traffic images. The fourth label of a fourth sample may include a sample attention feature in one of the historical traffic images.

[0180] The attention feature refers to a feature obtained after performing self-attention processing on the incident identification feature.

[0181] In some embodiments, the processor 140 may perform the self-attention processing on the incident identification feature using the attention model (e.g., the Transformer Encoder) , and generate the attention feature by performing self-attention interaction on the global feature of the traffic scenario and a global feature of each of M2 target traffic participants.

[0182] FIG. 9 is a schematic structural diagram illustrating an attention model according to some embodiments of the present disclosure. The attention model shown in FIG. 9 may also be the self-attention layer 840 of the traffic condition recognition model 800. Referring to FIG. 9, the incident identification feature F9 (1+ M2, N*T) may include a global feature of a traffic scenario and a global feature of each of M2 target traffic participants. The incident identification feature F9 (1+ M2, N*T) may be input into the attention model Transformer Encoder for processing. For example, taking M2 = 9 traffic participants as an example, position 0 represents the global feature (1, N*T) of the traffic scenario, and positions 1 to 9 represent the global features (M2, N*T) of the M2 target traffic participants, respectively.

[0183] By performing self-attention processing on the incident identification feature F9 (1 + M2, N*T) through the attention model Transformer Encoder, an attention feature F10 (1 + M2, N*T) may be generated. In some embodiments, the attention feature F10 represents the global feature of the traffic scenario after the self-attention interaction.

[0184] In some embodiments, the processor 140 may input the attention feature F10 (1 + M2, N*T) into a classification layer of the incident identification model.

[0185] In some embodiments, the processor 140 may obtain a traffic condition recognition result of a target region by performing traffic incident classification using the attention feature. For example, the processor 140 may obtain the traffic condition recognition result of the target region by performing traffic incident classification on the attention feature using a classification layer (e.g., an MLP classification head) of the incident identification model. More descriptions regarding the traffic condition recognition result may be found in FIG. 2 and the related descriptions thereof (e.g., the operation 250) .

[0186] In some embodiments, the processor 140 may extract a temporal correlation feature from the attention feature based on a second temporal feature; and obtain the traffic condition recognition result of the target region by performing traffic condition recognition using the temporal correlation feature.

[0187] The temporal correlation feature refers to a feature that represents a dynamic change in a first temporal feature and the second temporal feature in time in a plurality of images. For example, the temporal correlation feature may be a characteristic representation of the temporal correlation between the first temporal feature and the second temporal feature.

[0188] In some embodiments, the processor 140 may generate the attention feature by performing a self-attention operation (e.g., performing an attention operation through the attention model or the self-attention layer 840 of the traffic condition recognition model 800) on the incident identification feature. For example, the processor 140 may input the incident identification feature into the attention model for the self-attention operation, and the attention model may output the attention feature.

[0189] Furthermore, the processor 140 may determine a position of the second temporal feature in the incident identification feature, such as a position corresponding to “0*” in FIG. 9, and extract a feature corresponding to the same position from the attention feature, such as a position corresponding to “0: ” in the attention feature, and the extracted feature may be a temporal correlation feature F11 (0, : ) .

[0190] In some embodiments, the processor 140 may input the temporal correlation feature F11 (0, : ) into the classification layer 850 of the traffic condition recognition model 800, and the classification layer 850 may output the traffic condition recognition result.

[0191] In some embodiments of the present disclosure, some temporal correlation information in the global feature (i.e., the third transformation features) related to the target traffic participant in the incident identification features may be introduced into the temporal correlation feature F11 through the attention operation. When traffic condition recognition is performed, it only needs to perform traffic condition recognition based on the temporal correlation feature F11, which guarantees the identification accuracy, and improves the identification efficiency and speed.

[0192] In some embodiments, the traffic condition recognition model 800 may be obtained by training through a fifth training set. The fifth training set may include a plurality of fifth samples with fifth labels. Each of the fifth samples may include one or more sample first temporal features (including a sample position temporal feature, a sample category temporal feature, and a sample local temporal feature) and sample second temporal features. The sample first temporal features and the sample second temporal features may be obtained through historical traffic images. A fifth label of a fifth sample may include sample attention features in the historical traffic images. Each module / layer of the traffic condition recognition model 800 may be obtained through separate training or joint training.

[0193] It should be noted that the above descriptions of the processes 200, 300, 400, 500, and 700 are only for example and explanation, and do not limit the scope of application of the present disclosure. For those skilled in the art, various modifications and changes can be made to the processes 200, 300, 400, 500, and 700 under the guidance of the present disclosure. However, these modifications and changes are still within the scope of the present disclosure. Taking the process 400 as an example, the operation 430 may be performed first, and then the operation 410 and the operation 420; the operation 420 may also be performed first, and then the operation 410 and the operation 430, etc.

[0194] FIG. 10 is a flowchart illustrating an exemplary method for traffic condition recognition according to some embodiments of the present disclosure. As shown in FIG. 10, a process 1000 may include the following operations.

[0195] In 1010, a plurality of images of a target region acquired at consecutive time points in a time period may be obtained.

[0196] In 1020, a first initial feature of a target traffic participant and a second initial feature of a traffic scenario in each of at least a portion of the plurality of images may be obtained by processing the plurality of images.

[0197] FIG. 11 is a flowchart illustrating an exemplary process of determining a first initial feature and a second initial feature according to some embodiments of the present disclosure. In some embodiments, referring to FIG. 11, a process 1100 may further extend the operation 1020 of the above embodiment. Obtaining a first initial feature of a target traffic participant and a second initial feature of a traffic scenario in each of at least a portion of the plurality of images by processing the plurality of images may include the following operations.

[0198] In 1110, a position feature and a category feature of a target traffic participant may be obtained by performing target detection on each of a plurality of images.

[0199] In addition, a local feature of the target traffic participant and a global feature of the traffic scenario may be obtained respectively by extracting a shallow feature and a deep feature of each of the plurality of images, including the following operations.

[0200] In 1120, a local feature of the target traffic participant may be obtained by generating an intermediate layer feature of a downsampling operation with a first multiple based on the intermediate layer feature during target detection.

[0201] In 1130, a second initial feature of a traffic scenario may be obtained by generating an intermediate layer feature of a downsampling operation with a second multiple, the second initial feature of the traffic scenario being a global feature of the traffic scenario, and the second multiple being greater than the first multiple.

[0202] Referring to FIG. 10, after the operation 1020, the process may further include the following operations.

[0203] In 1030, a first temporal feature of the target traffic participant and a second temporal feature of the traffic scenario may be obtained by synthesizing the first initial feature and the second initial feature of each of the plurality of images, respectively.

[0204] FIG. 12 is a flowchart illustrating an exemplary process of determining a first temporal feature and a second temporal feature according to some embodiments of the present disclosure. In some embodiments, referring to FIG. 12, a process 1200 may further extend the operation 1030 of the above embodiment. Obtaining the first temporal feature of the target traffic participant and the second temporal feature of the traffic scenario by synthesizing the first initial feature and the second initial feature of the plurality of images, respectively, may include the following operations.

[0205] In 1210, a second temporal feature of a traffic scenario may be obtained by splicing second initial features obtained based on at least a portion of a plurality of images.

[0206] In 1220, a target traffic participant may be determined by performing target tracking on the plurality of images, the target traffic participant satisfying a preset tracking condition.

[0207] In 1230, a first temporal feature of the target traffic participant may be generated by synthesizing a first initial feature of the target traffic participant in each of the plurality of images.

[0208] It is understood that the embodiments of the present disclosure do not limit the order of the above operations 1210-1230. For example, the operation 1210 may be performed first, and then the operations 1220-1230. Alternatively, the operations 1220-1230 may be performed first, and then the operation 1210 may be performed.

[0209] Referring to FIG. 10, after the above the operation 1030, the process may further include the following operations.

[0210] In 1040, a traffic condition recognition result of the target region may be obtained by performing traffic condition recognition based on a temporal correlation between the first temporal feature and the second temporal feature.

[0211] FIG. 13 is a flowchart illustrating an exemplary process of determining a traffic condition recognition result according to some embodiments of the present disclosure. In some embodiments, referring to FIG. 13, a process 1300 may further extend the operation 1040 of the above embodiment. Obtaining the traffic condition recognition result of the target region by performing traffic condition recognition based on a temporal correlation between the first temporal feature and the second temporal feature may include the following operations.

[0212] In 1310, a first transformation feature of a target traffic participant may be generated by performing temporal transformation on a first temporal feature, and a second transformation feature of a traffic scenario may be generated by performing temporal transformation on a second temporal feature.

[0213] In 1320, an incident identification feature may be generated by splicing the first transformation feature and the second transformation feature.

[0214] In 1330, an attention feature may be generated by performing self-attention processing on the incident identification feature.

[0215] In 1340, a traffic condition recognition result of a target region may be determined by performing traffic incident classification using the attention feature.

[0216] More descriptions regarding the above processes 1000, 1100, 1200 and 1300, may be found in the related descriptions above (e.g., FIGs. 2-9) , which are not repeated here.

[0217] FIG. 14 is a block diagram illustrating a processor of a system for traffic condition recognition according to some embodiments of the present disclosure. As shown in FIG. 14, the processor 140 may include an image acquisition module 141, a first feature acquisition module 142, a second feature acquisition module 143, a third feature acquisition module 144, and an identification module 145.

[0218] The image acquisition module 141 may be configured to obtain a plurality of images of a target region acquired at consecutive time points in a time period.

[0219] In some embodiments, the image acquisition module 141 may be further configured to obtain a target video of the target region; and obtain a plurality of images from the target video based on a preset sliding window.

[0220] The first feature acquisition module 142 may be configured to obtain a first initial feature of a target traffic participant and a second initial feature of a traffic scenario in each of at least a portion of the plurality of images by processing the plurality of images.

[0221] In some embodiments, the first feature acquisition module 142 may be further configured to obtain a detection result of one or more traffic participants by processing the plurality of images using a target detection model; and determine, based on the detection result of the one or more traffic participants, the target traffic participant, the target traffic participant satisfying a preset tracking condition.

[0222] In some embodiments, the first feature acquisition module 142 may be further configured to obtain a tracking result of each of the one or more traffic participants by performing target tracking on the plurality of images based on a position feature and a category feature of each of the one or more traffic participants; determine, based on the tracking result, the target traffic participant, the target traffic participant including one of the one or more traffic participants that present in each of the at least a portion of the plurality of images and the at least a portion of the plurality of images includes consecutive images among the plurality of images. In other words, a target traffic participant is represented in each of the consecutive images among the plurality of images.

[0223] In some embodiments, the first feature acquisition module 142 may be further configured to obtain a position feature and a category feature of the target traffic participant by performing target detection on each of the plurality of images; obtain a local feature of the target traffic participant by performing a downsampling operation with a first multiple on each of the plurality of images during the target detection; and obtain the second initial feature of the traffic scenario by performing a downsampling operation with a second multiple on each of the plurality of images during the target detection, the second initial feature of the traffic scenario being a global feature of the traffic scenario, the second multiple being greater than the first multiple, and the first initial feature including the position feature, the category feature, and the local feature of the target traffic participant.

[0224] The second feature acquisition module 143 may be configured to generate, based on the first initial feature, a first temporal feature of the target traffic participant.

[0225] In some embodiments, the first initial feature may include the position feature, the category feature, and the local feature of the target traffic participant. In some embodiments, the second feature acquisition module 143 may be further configured to obtain a position temporal feature of the target traffic participant by splicing the position feature of the target traffic participant in each of the at least a portion of the plurality of images; obtain a category temporal feature of the target traffic participant by splicing the category feature of the target traffic participant in each of the at least a portion of the plurality of images; extract, based on a position range of the position temporal feature of the target traffic participant, a feature corresponding to the position range from the local feature of the target traffic participant, the feature corresponding to the position range from the local feature of the target traffic participant being determined as a local temporal feature, the first temporal feature including the position temporal feature, the category temporal feature, and the local temporal feature.

[0226] In some embodiments, the second feature acquisition module 143 may be further configured to obtain a first position temporal feature by performing dimensional transformation on the position temporal feature of the target traffic participant; determine, based on a position range of the target traffic participant in the first position temporal feature, a feature extraction range from the local feature of the target traffic participant; and extract, based on the feature extraction range, the feature corresponding to the position range of the position temporal feature of the target traffic participant from the local feature of the target traffic participant, the feature corresponding to the position range from the local feature of the target traffic participant being determined as the local temporal feature.

[0227] The third feature acquisition module 144 may be configured to generate, based on the second initial feature, a second temporal feature of the traffic scenario.

[0228] The identification module 145 may be configured to determine a traffic condition recognition result of a target region based on the first temporal feature and the second temporal feature.

[0229] In some embodiments, the identification module 145 may be further configured to generate a first transformation feature of the target traffic participant by performing temporal transformation on the first temporal feature; generate a second transformation feature of the traffic scenario by performing temporal transformation on the second temporal feature; generate an incident identification feature by performing a feature encoding operation on the first transformation feature and the second transformation feature, the incident identification feature including a temporal correlation between the first temporal feature and the second temporal feature; and determine the traffic condition recognition result of the target region based on the incident identification feature.

[0230] In some embodiments, the identification module 145 may be further configured to generate an attention feature by processing the incident identification feature using an attention model; and determine the traffic condition recognition result of the target region based on the attention feature.

[0231] In some embodiments, the identification module 145 may be further configured to extract, based on the second temporal feature, a temporal correlation feature from the attention feature; and obtain the traffic condition recognition result of the target region by performing traffic condition recognition based on the temporal correlation feature.

[0232] In some embodiments, the position temporal feature may be configured to characterize a position of the target traffic participant in each of the plurality of images; the category temporal feature may be configured to characterize a category of the target traffic participant; and the local temporal feature may be configured to characterize an image feature in each of the plurality of images.

[0233] In some embodiments, the identification module 145 may be further configured to generate a first mapping feature corresponding to the position temporal feature, a second mapping feature corresponding to the category temporal feature, and a third mapping feature corresponding to the local temporal feature by performing channel mapping on the position temporal feature, the category temporal feature, and the local temporal feature to the same preset channel count; performing channel mapping on the category temporal feature including mapping a discrete category temporal feature to a continuous vector space; and generating the first transformation feature of the target traffic participant based on the first mapping feature, the second mapping feature, and the third mapping feature.

[0234] In some embodiments, the identification module 145 may be further configured to generate the first transformation feature of the target traffic participant by adding the first mapping feature, the second mapping feature, and the third mapping feature.

[0235] In some embodiments, the identification module 145 may be further configured to generate the first transformation feature of the target traffic participant by inputting the first mapping feature, the second mapping feature, and the third mapping feature into a first neural network model.

[0236] In some embodiments, the identification module 145 may be further configured to generate the first transformation feature of the target traffic participant by inputting the first temporal feature into a second neural network model.

[0237] In some embodiments, the identification module 145 may be further configured to generate the second transformation feature of the traffic scenario by performing channel count transformation on the second temporal feature based on a preset channel count.

[0238] In some embodiments, the identification module 145 may be further configured to generate a third transformation feature by performing dimensional transformation on the first transformation feature of the target traffic participant; generate a fourth transformation feature by performing dimensional transformation on the second transformation feature of the traffic scenario; and generate the incident identification feature based on the third transformation feature and the fourth transformation feature; the third transformation feature including a global encoding feature of the target traffic participant in each of the at least a portion of the plurality of images, and the fourth transformation feature including a global encoding feature of the traffic scenario in each of the at least a portion of the plurality of images.

[0239] In some embodiments, the identification module 145 may be further configured to perform, based on the traffic condition recognition result, a corresponding traffic management strategy, the traffic management strategy including at least one of a traffic prompt, an accident alarm, or a violation record.

[0240] FIG. 15 is a schematic diagram illustrating a system for traffic condition recognition according to some embodiments of the present disclosure.

[0241] As shown in FIG15, a system 1500 for traffic condition recognition may include the storage device 150 and the processor 140. The storage device 150 and the processor 140 may be coupled with each other. The storage device 150 may be configured to store executable instructions. The processor 140 may be configured to read the executable instructions to perform the following operations: obtaining a plurality of images of a target region acquired at consecutive time points in a time period; obtaining a first initial feature of a target traffic participant and a second initial feature of a traffic scenario in each of at least a portion of the plurality of images by processing the plurality of images; generating, based on the first initial feature, a first temporal feature of the target traffic participant; generating, based on the second initial feature, a second temporal feature of the traffic scenario; and determining, based on the first temporal feature and the second temporal feature, a traffic condition recognition result of the target region.

[0242] FIG. 16 is a schematic diagram illustrating a non-transitory computer-readable storage medium according to some embodiments of the present disclosure. As shown in FIG. 16, a non-transitory computer-readable storage medium 1600 may include executable instructions 161 that, when read by at least one processor 140, direct the at least one processor 140 to implement a method for traffic condition recognition. The method may include: obtaining a plurality of images of a target region acquired at consecutive time points in a time period; obtaining a first initial feature of a target traffic participant and a second initial feature of a traffic scenario in each of at least a portion of the plurality of images by processing the plurality of images; generating, based on the first initial feature, a first temporal feature of the target traffic participant; generating, based on the second initial feature, a second temporal feature of the traffic scenario; and determining, based on the first temporal feature and the second temporal feature, a traffic condition recognition result of the target region.

[0243] In this embodiment, the non-temporary computer-readable storage medium 1600 may be a medium that stores the executable instructions 161, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM) , a random access memory (RAM) , a disk, an optical disk, or the like, or a server that stores the executable instructions 161. The server may send the executable instructions 161 to other devices for execution, or execute the executable instructions 161.

[0244] Having thus described the basic concepts, it may be rather apparent to those skilled in the art after reading this detailed disclosure that the foregoing detailed disclosure is intended to be presented by way of example only and is not limiting. Various alterations, improvements, and modifications may occur and are intended to those skilled in the art, though not expressly stated herein. These alterations, improvements, and modifications are intended to be suggested by this disclosure and are within the spirit and scope of the exemplary embodiments of this disclosure.

[0245] Similarly, it should be appreciated that in the foregoing description of embodiments of the present disclosure, various features are sometimes grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure aiding in the understanding of one or more of the various embodiments. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claimed subject matter requires more features than are expressly recited in each claim. Rather, claimed subject matter may lie in less than all features of a single foregoing disclosed embodiment.

[0246] For each patent, patent application, patent application publication, or other materials cited in the present disclosure, such as articles, books, specifications, publications, documents, or the like, the entire contents of which are hereby incorporated into the present disclosure as a reference. The application history documents that are inconsistent or conflict with the content of the present disclosure are excluded, and the documents that restrict the broadest scope of the claims of the present disclosure (currently or later attached to the present disclosure) are also excluded. It should be noted that if there is any inconsistency or conflict between the description, definition, and / or use of terms in the auxiliary materials of the present disclosure and the content of the present disclosure, the description, definition, and / or use of terms in the present disclosure is subject to the present disclosure.

[0247] Finally, it should be understood that the embodiments described in the present disclosure are only used to illustrate the principles of the embodiments of the present disclosure. Other variations may also fall within the scope of the present disclosure. Therefore, as an example and not a limitation, alternative configurations of the embodiments of the present disclosure may be regarded as consistent with the teaching of the present disclosure. Accordingly, the embodiments of the present disclosure are not limited to the embodiments introduced and described in the present disclosure explicitly.

Claims

1.A method for traffic incident identification, comprising:obtaining a plurality of images of a target region acquired at consecutive time points in a time period;obtaining a first initial feature of a target traffic participant and a second initial feature of a traffic scenario in each of at least a portion of the plurality of images by processing the plurality of images;generating, based on the first initial feature, a first temporal feature of the target traffic participant;generating, based on the second initial feature, a second temporal feature of the traffic scenario; anddetermining, based on the first temporal feature and the second temporal feature, a traffic condition recognition result of the target region.2.The method of claim 1, wherein the determining, based on the first temporal feature and the second temporal feature, a traffic condition recognition result of the target region includes:generating a first transformation feature of the target traffic participant by performing temporal transformation on the first temporal feature;generating a second transformation feature of the traffic scenario by performing temporal transformation on the second temporal feature;generating an incident identification feature by performing a feature encoding operation on the first transformation feature and the second transformation feature, the incident identification feature including a temporal correlation between the first temporal feature and the second temporal feature; anddetermining the traffic condition recognition result of the target region based on the incident identification feature.3.The method of claim 2, wherein the determining the traffic condition recognition result of the target region based on the incident identification feature includes:generating an attention feature by processing the incident identification feature using an attention model; anddetermining the traffic condition recognition result of the target region based on the attention feature.4.The method of claim 3, wherein the determining the traffic condition recognition result of the target region based on the attention feature includes:extracting, based on the second temporal feature, a temporal correlation feature from the attention feature; andobtaining the traffic condition recognition result of the target region by performing traffic condition recognition based on the temporal correlation feature.5.The method of claim 2, wherein the first temporal feature includes a position temporal feature, a category temporal feature, and a local temporal feature;the position temporal feature is configured to characterize a position of the target traffic participant in each of at least a portion of the plurality of images;the category temporal feature is configured to characterize a category of the target traffic participant;the local temporal feature is configured to characterize an image feature of each of the plurality of images.6.The method of claim 5, wherein the generating a first transformation feature of the target traffic participant by performing temporal transformation on the first temporal feature includes:generating a first mapping feature corresponding to the position temporal feature, a second mapping feature corresponding to the category temporal feature, and a third mapping feature corresponding to the local temporal feature by performing channel mapping on the position temporal feature, the category temporal feature, and the local temporal feature to the same preset channel count; wherein performing channel mapping on the category temporal feature includes mapping a discrete category temporal feature to a continuous vector space;generating the first transformation feature of the target traffic participant based on the first mapping feature, the second mapping feature, and the third mapping feature.7.The method of claim 6, wherein the generating the first transformation feature of the target traffic participant based on the first mapping feature, the second mapping feature, and the third mapping feature includes:generating the first transformation feature of the target traffic participant by adding the first mapping feature, the second mapping feature, and the third mapping feature.8.The method of claim 6, wherein the generating the first transformation feature of the target traffic participant based on the first mapping feature, the second mapping feature, and the third mapping feature includes:generating the first transformation feature of the target traffic participant by inputting the first mapping feature, the second mapping feature, and the third mapping feature into a first neural network model.9.The method of claim 2, wherein the generating a first transformation feature of the target traffic participant by performing temporal transformation on the first temporal feature includes:generating the first transformation feature of the target traffic participant by inputting the first temporal feature into a second neural network model.10.The method of any one of claims 2-9, wherein the generating a second transformation feature of the traffic scenario by performing temporal transformation on the second temporal feature includes:generating the second transformation feature of the traffic scenario by performing channel count transformation on the second temporal feature based on a preset channel count.11.The method of any one of claims 2-10, wherein the generating an incident identification feature by performing a feature encoding operation on the first transformation feature and the second transformation feature includes:generating a third transformation feature by performing dimensional transformation on the first transformation feature of the target traffic participant;generating a fourth transformation feature by performing dimensional transformation on the second transformation feature of the traffic scenario; andgenerating the incident identification feature based on the third transformation feature and the fourth transformation feature; whereinthe third transformation feature includes a global encoding feature of the target traffic participant in each of the at least a portion of the plurality of images, and the fourth transformation feature includes a global encoding feature of the traffic scenario in each of the at least a portion of the plurality of images.12.The method of any one of claims 1-11, further comprising:obtaining a detection result of one or more traffic participants by processing the plurality of images using a target detection model; anddetermining, based on the detection result of the one or more traffic participants, the target traffic participant, the target traffic participant satisfying a condition.13.The method of claim 12, wherein the detection result of the one or more traffic participants includes a position feature and a category feature of each of the one or more traffic participants; and the determining the target traffic participant based on the detection result of the one or more traffic participants includes:obtaining a tracking result of each of the one or more traffic participants by performing target tracking on the plurality of images based on the position feature and the category feature of each of the one or more traffic participants; anddetermining, based on the tracking result, the target traffic participant, the target traffic participant including one of the one or more traffic participants that present in each of the at least a portion of the plurality of images.14.The method of any one of claims 1 -13, wherein the first initial feature includes a position feature, a category feature, and a local feature of the target traffic participant; the generating, based on the first initial feature of each of the at least a portion of the plurality of images, a first temporal feature of the target traffic participant includes:obtaining a position temporal feature of the target traffic participant by splicing the position feature of the target traffic participant in each of the at least a portion of the plurality of images;obtaining a category temporal feature of the target traffic participant by splicing the category feature of the target traffic participant in each of the at least a portion of the plurality of images; andextracting, based on a position range of the position temporal feature of the target traffic participant, a feature corresponding to the position range from the local feature of the target traffic participant, the feature corresponding to the position range from the local feature of the target traffic participant being determined as the local temporal feature, the first temporal feature including the position temporal feature, the category temporal feature, and the local temporal feature.15.The method of claim 14, wherein the extracting, based on a position range of the position temporal feature of the target traffic participant, a feature corresponding to the position range from the local feature of the target traffic participant includes:obtaining a first position temporal feature by performing dimensional transformation on the position temporal feature of the target traffic participant;determining, based on a position range of the target traffic participant in the first position temporal feature, a feature extraction range from the local feature of the target traffic participant; andextracting, based on the feature extraction range, the feature corresponding to the position range of the position temporal feature of the target traffic participant from the local feature of the target traffic participant.16.The method of any one of claims 1-15, wherein the obtaining a first initial feature of a target traffic participant and a second initial feature of a traffic scenario in each of the at least a portion of the plurality of images by processing the plurality of images includes:obtaining a position feature and a category feature of the target traffic participant by performing target detection on each of the plurality of images;obtaining a local feature of the target traffic participant by performing a downsampling operation with a first multiple on each of the plurality of images during the target detection; andobtaining the second initial feature of the traffic scenario by performing a downsampling operation with a second multiple on each of the plurality of images during the target detection, wherein the second initial feature of the traffic scenario is a global feature of the traffic scenario, the second multiple is greater than the first multiple, and the first initial feature includes the position feature, the category feature, and the local feature of the target traffic participant.17.The method of any one of claims 1-16, wherein the obtaining a plurality of images of a target region includes:obtaining a target video of the target region; andobtaining, based on a preset sliding window, the plurality of images from the target video.18.The method of claim 1, further comprising:performing, based on the traffic condition recognition result, a corresponding traffic management strategy, the traffic management strategy including at least one of a traffic prompt, an accident alarm, or a violation record.19.A system for traffic condition recognition, comprising a storage device and at least one processor, wherein the storage device is configured to store executable instructions, and the at least one processor is configured to perform operations including:obtaining a plurality of images of a target region acquired at consecutive time points in a time period;obtaining a first initial feature of a target traffic participant and a second initial feature of a traffic scenario in each of at least a portion of the plurality of images by processing the plurality of images;generating, based on the first initial feature, a first temporal feature of the target traffic participant;generating, based on the second initial feature, a second temporal feature of the traffic scenario; anddetermining, based on the first temporal feature and the second temporal feature, a traffic condition recognition result of the target region.20.A non-transitory computer-readable storage medium, comprising executable instructions that, when read by at least one processor, direct the at least one processor to implement the method for traffic condition recognition, wherein the method includes:obtaining a plurality of images of a target region acquired at consecutive time points in a time period;obtaining a first initial feature of a target traffic participant and a second initial feature of a traffic scenario in each of at least a portion of the plurality of images by processing the plurality of images;generating, based on the first initial feature, a first temporal feature of the target traffic participant;generating, based on the second initial feature, a second temporal feature of the traffic scenario; anddetermining, based on the first temporal feature and the second temporal feature, a traffic condition recognition result of the target region.

Citation Information

Patent Citations

  • Accident detection method and device, electronic equipment and storage medium

    CN114677618A

  • Road intersection dynamic traffic scene intelligent generation method

    CN116503571A

  • Driving detection method, device and equipment based on automobile data recorder and storage medium

    CN117423093A

  • Small sample behavior recognition method and system based on event camera

    CN118314625A

  • Traffic event identification method, computer equipment and storage medium

    CN118521945A