Identification information generation method and device, electronic equipment and computer readable medium

By real-time processing and feature extraction of vehicle driving videos, the action hazard identification information is generated, and the problem of huge vehicle driving data is solved, resulting in low video viewing efficiency, and efficient action hazard identification and warning are achieved.

CN120107935APending Publication Date: 2025-06-06ADDX (BEIJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510168633.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the case of huge vehicle driving data, the efficiency of technicians to view videos is limited, resulting in more errors and leaks in vehicles with dangerous actions.

Method used

By obtaining the real-time video of the target vehicle, performing video preprocessing to generate a frame image sequence, inputting the frame image into the pre-trained object segmentation model to generate an object segmentation image sequence, extracting object action feature information and image time information, performing feature information clustering to generate action feature information clusters, performing action information checksum recognition, and generating action hazard identification information.

Benefits of technology

It realizes the precise and efficient generation of action hazard identification information for the target object, and provides dangerous driving warnings to avoid subsequent dangerous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107935A_ABST
    Figure CN120107935A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an identification information generation method and device, electronic equipment and a computer readable medium. A specific embodiment of the method comprises the following steps: acquiring a real-time shot video; performing video preprocessing on the real-time shot video to generate a frame image sequence; inputting each frame image into a pre-trained object segmentation model to generate an object segmentation image; generating object action feature information and image time information corresponding to each object segmented image; executing feature information clustering to generate an object action feature information cluster sequence; generating object action information corresponding to each object action feature information cluster; performing action information verification on the object action information sequence to generate verification information; in response to the representation passing verification, generating object action recognition information; and generating action danger identification information. According to the embodiment, the action danger identification information for the target object can be accurately and efficiently generated, dangerous driving warning is performed, and follow-up dangerous driving is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular to a method, device, electronic device, and computer-readable medium for generating identification information. Background Art

[0002] At present, in the field of vehicle driving, dangerous actions of vehicle drivers often lead to traffic accidents. For the detection of dangerous action information, the method usually adopted is: through the method of manual video viewing by relevant technicians, the behavior information of dangerous actions is screened out.

[0003] However, when using the above method to detect dangerous action information, the following technical problems often occur:

[0004] Given the large amount of vehicle driving data, the efficiency of video review by technicians is limited, which may result in a large number of errors and omissions of vehicles with dangerous driving actions.

[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the invention

[0006] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.

[0007] Some embodiments of the present disclosure propose identification information generation methods, devices, electronic devices, and computer-readable media to solve one or more of the technical problems mentioned in the above background technology section.

[0008] In a first aspect, some embodiments of the present disclosure provide a method for generating identification information, comprising: acquiring a real-time video of a target vehicle; performing video preprocessing on the real-time video to generate a frame image sequence for the target object; inputting each frame image in the frame image sequence into a pre-trained object segmentation model to generate an object segmentation image, thereby obtaining an object segmentation image sequence; generating object action feature information and image time information corresponding to each object segmentation image in the object segmentation image sequence to obtain an object action feature information sequence and an image time information sequence; performing feature information clustering on the object action feature information sequence and the image time information sequence to generate an object action feature information cluster sequence, wherein each object action feature information has a corresponding object action feature information. The image time information corresponding to each object action information in the feature information cluster is continuous, and the number of object action information corresponding to each object action feature information cluster is the target number; the object action information corresponding to each object action feature information cluster in the above object action feature information cluster sequence is generated to obtain the object action information sequence; the object action information sequence is verified to generate verification information; in response to determining that the verification information representation passes the verification, according to the image feature information sequence corresponding to the above object segmentation image sequence, using a pre-trained first object action recognition model, object action recognition information is generated; according to the above object action information sequence and the above object action recognition information, action hazard recognition information for the above target object is generated.

[0009] In a second aspect, some embodiments of the present disclosure provide an identification information generating device, comprising: an acquisition unit, configured to acquire a real-time captured video of a target vehicle; a video preprocessing unit, configured to perform video preprocessing on the real-time captured video to generate a frame image sequence for the target object; an input unit, configured to input each frame image in the frame image sequence into a pre-trained object segmentation model to generate an object segmentation image and obtain an object segmentation image sequence; a first generation unit, configured to generate object action feature information and image time information corresponding to each object segmentation image in the object segmentation image sequence to obtain an object action feature information sequence and an image time information sequence; an execution unit, configured to perform feature information clustering for the object action feature information sequence and the image time information sequence to generate an object action feature information cluster sequence, wherein each frame image in the frame image sequence is a pre-trained object segmentation model, and a pre-trained object segmentation model is used to generate an object segmentation image sequence; a first generation unit, configured to generate object action feature information and image time information corresponding to each object segmentation image in the object segmentation image sequence, and obtain an object action feature information sequence and an image time information sequence; and a execution unit, configured to perform feature information clustering for the object action feature information sequence and the image time information sequence to generate an object action feature information cluster sequence. The image time information corresponding to each object action information in each object action feature information cluster is continuous, and the number of object action information corresponding to each object action feature information cluster is the target number; the second generating unit is configured to generate the object action information corresponding to each object action feature information cluster in the above-mentioned object action feature information cluster sequence to obtain the object action information sequence; the verification unit is configured to perform action information verification on the above-mentioned object action information sequence to generate verification information; the third generating unit is configured to generate object action recognition information according to the image feature information sequence corresponding to the above-mentioned object segmentation image sequence in response to determining that the above-mentioned verification information representation has passed the verification, using the pre-trained first object action recognition model; the fourth generating unit is configured to generate action hazard recognition information for the above-mentioned target object according to the above-mentioned object action information sequence and the above-mentioned object action recognition information.

[0010] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.

[0011] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner in the first aspect is implemented.

[0012] The above-mentioned embodiments of the present disclosure have the following beneficial effects: through the identification information generation method of some embodiments of the present disclosure, the action hazard identification information for the target object can be accurately and efficiently generated to warn of dangerous driving and avoid the occurrence of subsequent dangerous driving. Specifically, the reason why the relevant action hazard identification information is not accurate and efficient is that on the basis of the relatively large amount of vehicle driving data, the efficiency of the technicians in video viewing is limited, which may lead to the omission of vehicles with action hazard driving. Based on this, the identification information generation method of some embodiments of the present disclosure first obtains a real-time video shot for the target vehicle. Here, the obtained real-time video is used as a data basis to extract the action behavior information corresponding to the target object. Then, the real-time video is preprocessed to generate a frame image sequence for the target object, which is processed into a frame image form to facilitate subsequent object segmentation processing. Then, each frame image in the frame image sequence is input into a pre-trained object segmentation model to accurately generate an object segmentation image to obtain an object segmentation image sequence. Then, the object action feature information and image time information corresponding to each object segmentation image in the object segmentation image sequence are generated to obtain an object action feature information sequence and an image time information sequence. Here, by determining the object action feature information and image time information corresponding to each object segmentation image, it is convenient to perform clustering processing on the image content semantics corresponding to the object segmentation image in the subsequent step, and combine the object feature information with similar image content semantics and continuous time, that is, perform feature information clustering on the above object action feature information sequence and the above image time information sequence to generate an object action feature information cluster sequence, wherein the image time information corresponding to each object action information in each object action feature information cluster is continuous, and the number of object action information corresponding to each object action feature information cluster is the target number. Here, by performing feature information clustering, the corresponding action of the target object in the time period corresponding to the real-time video can be effectively parsed, so as to generate object action information, that is, object action information sequence, in a more detailed manner. Secondly, the object action information sequence is verified to generate verification information to determine the accuracy of the object action information in the object action information sequence. Further, in response to determining that the verification information indicates that the verification is passed, according to the image feature information sequence corresponding to the object segmentation image sequence, the object action recognition information can be accurately generated using the pre-trained first object action recognition model. Finally, based on the object action information sequence and the object action recognition information, the action hazard recognition information for the target object is accurately generated. In summary, through the clustering of object action feature information and image time information, the object segmentation image sequence can be refined and decomposed into action details to generate an accurate object action information sequence.In addition, the first object action recognition model is used to consider the overall video to generate object action recognition information in the overall video state. Thus, based on the object action information sequence and the object action recognition information, action hazard recognition information for the target object can be accurately generated. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0014] Figure 1 is a flow chart of some embodiments of the identification information generation method according to the present disclosure;

[0015] Figure 2 is a schematic diagram of the structure of some embodiments of the identification information generating device according to the present disclosure;

[0016] Figure 3 It is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0018] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.

[0019] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0020] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0021] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0022] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0023] refer to Figure 1 , shows a process 100 of some embodiments of the identification information generation method according to the present disclosure. The identification information generation method comprises the following steps:

[0024] Step 101, obtaining a real-time video of a target vehicle.

[0025] In some embodiments, the execution subject of the above identification information generation method can obtain a real-time video of the target vehicle through a wired connection or a wireless connection. The target vehicle can be a vehicle to be subjected to object dangerous action detection. The real-time video can be a real-time in-vehicle video of the target vehicle.

[0026] Step 102: perform video preprocessing on the real-time captured video to generate a frame image sequence for the target object.

[0027] In some embodiments, the execution subject may perform video preprocessing on the real-time video to generate a frame image sequence for the target object. The video preprocessing may include but is not limited to at least one of the following: video decoding processing and video frame extraction processing. The frame image in the frame image sequence may be an image showing the target object. In practice, the target object may be a driver. That is, the target object may be a driver driving a target vehicle.

[0028] Step 103: input each frame image in the frame image sequence into a pre-trained object segmentation model to generate an object segmentation image, thereby obtaining an object segmentation image sequence.

[0029] In some embodiments, the execution subject may input each frame image in the frame image sequence into a pre-trained object segmentation model to generate an object segmentation image and obtain an object segmentation image sequence. The object segmentation model may be a neural network model that targets the target object and performs target segmentation. In practice, the object segmentation model may be a lightweight U-net model. The object segmentation model may be a model trained based on a conventional training method (e.g., a training method based on a gradient descent method).

[0030] Step 104 : generating object motion feature information and image time information corresponding to each object segmentation image in the above object segmentation image sequence, and obtaining an object motion feature information sequence and an image time information sequence.

[0031] In some embodiments, the execution subject may generate object motion feature information and image time information corresponding to each object segmentation image in the object segmentation image sequence, and obtain an object motion feature information sequence and an image time information sequence. The object motion feature information may be feature information representing the motion semantic features corresponding to the target object in the object segmentation image. The image time information may be frame time information of the object segmentation image in the real-time video.

[0032] As an example, the execution subject may input each object segmentation image in the object segmentation image sequence into an image action feature extraction model to generate object action feature information and obtain an object action feature information sequence. The image action feature extraction model may be an attention mechanism model based on a convolutional neural network model.

[0033] Step 105 , performing feature information clustering on the object action feature information sequence and the image time information sequence to generate an object action feature information cluster sequence.

[0034] In some embodiments, the execution subject may perform feature information clustering for the object action feature information sequence and the image time information sequence to generate an object action feature information cluster sequence. The image time information corresponding to each object action information in each object action feature information cluster is continuous. The number of object action information corresponding to each object action feature information cluster is the target number. The similarity between each object action feature information in the object action feature information cluster is high. The action information corresponding to each object action feature information in the object action feature information cluster is similar. In practice, the target number is a value of 1. That is, the object action feature information cluster is a feature information cluster for a certain action information. The action information may be the action posture of the target object.

[0035] In some optional implementations of some embodiments, the above-mentioned performing feature information clustering for the above-mentioned object action feature information sequence and the above-mentioned image time information sequence to generate an object action feature information cluster sequence may include the following steps:

[0036] In the first step, the k-means clustering algorithm is used to cluster the object action feature information in the object action feature information sequence to generate an initial object action feature information cluster set, wherein the initial object action feature information in the initial object action feature information cluster has similar action feature semantic content.

[0037] The second step is to determine the image time information corresponding to each cluster center feature information in the cluster center feature information set, and obtain the image time information set. The cluster center feature information set is the cluster center set corresponding to the initial object action feature information cluster set. There is a one-to-one correspondence between the initial object action feature information cluster in the initial object action feature information cluster set and the cluster center in the cluster center set. There is a feature correspondence between the cluster center feature information and the object segmentation image. There is a time correspondence between the object segmentation image and the image time information.

[0038] In the third step, for each initial object motion feature information cluster in the initial object motion feature information cluster set, the following second generation step is performed:

[0039] Sub-step 1: determining the cluster center feature information corresponding to the initial object action feature information cluster.

[0040] Sub-step 2: determining the image time information corresponding to the cluster center feature information as the first image time information.

[0041] Sub-step 3, removing the cluster center feature information from the initial object action feature information cluster to obtain a feature information set after removal.

[0042] Sub-step 4: determining the time difference between the image time information corresponding to each feature information in the feature information set after removal and the first image time information, to obtain a time difference set.

[0043] Sub-step 5: determining the feature information in the above removed feature information set whose corresponding time difference is greater than a predetermined value as pending feature information. The predetermined value may be a preset time value. For example, the predetermined value may be 0.5.

[0044] Step 4: for each undetermined feature information in the obtained undetermined feature information set, perform the following third generation step:

[0045] Sub-step 1: determine the image time information corresponding to the above-mentioned undetermined feature information as the second image time information.

[0046] Sub-step 2: select the first image time information with the shortest time interval with the second image time information and the smallest cosine distance between corresponding features from the obtained first image time information set as the target image time information. The cosine distance can represent the similarity between feature information.

[0047] Sub-step 3: generating attribution information indicating that the undetermined feature information belongs to the initial object action feature information cluster corresponding to the target image time information.

[0048] In the fifth step, based on the obtained attribution information set, the feature information of the pending feature information set in the initial object action feature information cluster set is classified to generate a classified object action feature information cluster set as the object action feature information cluster set.

[0049] Step 106: Generate object action information corresponding to each object action feature information cluster in the above object action feature information cluster sequence to obtain an object action information sequence.

[0050] In some embodiments, the execution subject may generate object action information corresponding to each object action feature information cluster in the object action feature information cluster sequence to obtain an object action information sequence. The object action information may be action information of an object action of a target object corresponding to the object action feature information cluster. Each object action feature information in the object action feature information cluster has unique corresponding object action information.

[0051] In some optional implementations of some embodiments, the generating of the object action information corresponding to each object action feature information cluster in the object action feature information cluster sequence may include the following steps:

[0052] In the first step, each object action feature information in the object action feature information cluster is time-sorted to generate an object action feature information subsequence.

[0053] As an example, the execution subject may perform time sorting on each object action feature information in the object action feature information cluster in the order of corresponding time from early to late, so as to generate an object action feature information subsequence.

[0054] In the second step, each object motion feature information in the object motion feature information subsequence is grouped to generate an object motion feature information group sequence, wherein the image time information corresponding to each object motion feature information in the object motion feature information group is continuous.

[0055] As an example, the execution subject divides each object motion feature information in the object motion feature information subsequence into 4 groups on average to generate an object motion feature information group sequence, wherein the image time information corresponding to each object motion feature information in the object motion feature information group is continuous.

[0056] In the third step, the action recognition information and action recognition probability information corresponding to each object action feature information group in the above-mentioned object action feature information group sequence are determined by using the pre-trained second object action recognition model, and an action recognition information sequence and an action recognition probability information sequence are obtained. Among them, the second object action recognition model can be a neural network model that generates object action recognition related information. Object action recognition related information may include: action recognition information and action recognition probability information. The action recognition information may be the action information of the identified target object. The action recognition probability information may be the precise probability value of the action information of the identified target object. The action recognition probability information may be a value between 0 and 1. The larger the value, the higher the corresponding recognition accuracy.

[0057] The fourth step is to generate the object action information according to the action recognition information sequence and the action recognition probability information sequence.

[0058] As an example, first, the execution subject may select action recognition information with corresponding action recognition probability information greater than 80% from the action recognition information sequence to obtain at least one action recognition information. Then, the action recognition information with the highest corresponding frequency is selected from the at least one action recognition information as the object action information.

[0059] Step 107: perform action information verification on the object action information sequence to generate verification information.

[0060] In some embodiments, the execution subject may perform action information verification on the object action information sequence to generate verification information, wherein the verification information includes: information indicating that the verification has been passed and information indicating that the verification has not been passed.

[0061] In some optional implementations of some embodiments, the above-mentioned performing action information verification on the above-mentioned object action information sequence to generate verification information may include the following steps:

[0062] The first step is to determine the cluster time corresponding to each object action feature information cluster in the object action feature information cluster sequence to obtain a cluster time sequence.

[0063] As an example, the execution subject may use the image time information corresponding to the cluster center in each object motion feature information cluster in the object motion feature information cluster sequence as the cluster time to obtain the cluster time sequence.

[0064] In the second step, for each object action feature information cluster in the above object action feature information cluster sequence, the following first generation step is performed:

[0065] The first sub-step is to determine the cluster time corresponding to the object action feature information cluster as the target cluster time.

[0066] The second sub-step is to determine the previous cluster time and the next cluster time corresponding to the target cluster time. The cluster position of the target cluster time corresponding to the object action feature information cluster is located after the cluster position of the previous cluster time corresponding to the object action feature information cluster. The cluster position of the target cluster time corresponding to the object action feature information cluster is located before the cluster position of the next cluster time corresponding to the object action feature information cluster.

[0067] The third sub-step is to determine the object action feature information cluster corresponding to the previous cluster time as the first object action feature information cluster, and to determine the object action feature information cluster corresponding to the next cluster time as the second object action feature information cluster.

[0068] The fourth sub-step is to select a first number of object action feature information from the first object action feature information cluster to obtain a first object action feature information sub-cluster. The first number may be a preset number. For example, the first number may be 10.

[0069] The fifth sub-step is to select a second number of object action feature information from the second object action feature information cluster to obtain a second object action feature information sub-cluster. The second number may be a preset number. For example, the second number may be 15.

[0070] The sixth sub-step is to fuse the feature information of the first object motion feature information sub-cluster, the second object motion feature information sub-cluster and the object motion feature information cluster to generate a fused feature information sequence.

[0071] The seventh sub-step is to generate candidate object action information based on the above fused feature information sequence.

[0072] As an example, the execution subject may generate candidate object action information based on the fused feature information sequence using a pre-trained first object action recognition model. In practice, the first object action recognition model may be a neural network model for recognizing object action information. The object action information may be information about the action posture performed by the object. The first object action recognition model may be trained based on a conventional model training method.

[0073] In an eighth sub-step, in response to determining that the candidate object action information is the same as the object action information corresponding to the object action feature information cluster, generating verification sub-information indicating that verification has been passed.

[0074] The third step is to generate the above verification information according to the obtained verification sub-information sequence.

[0075] As an example, first, in response to determining that all the check information in the syndrome information sequence is information indicating that the check has been passed, check information indicating that the check has been passed is generated.

[0076] Step 108 , in response to determining that the verification information representation passes the verification, object action recognition information is generated according to the image feature information sequence corresponding to the object segmentation image sequence using a pre-trained first object action recognition model.

[0077] In some embodiments, in response to determining that the verification information representation passes the verification, the execution subject may generate object action recognition information based on the image feature information sequence corresponding to the object segmentation image sequence using a pre-trained first object action recognition model. The first object action recognition model may be a neural network model for generating object action recognition information. In practice, the first object action recognition model may be a convolutional neural network model + a recurrent neural network model. The object action recognition information may be an action recognition result corresponding to the target object. Specifically, the action recognition result may be a driving posture action recognition result.

[0078] As an example, the execution entity may directly input the image feature information sequence into the first object action recognition model to generate object action recognition information.

[0079] In some optional implementations of some embodiments, the generating of the object action recognition information according to the image feature information sequence corresponding to the object segmentation image sequence using a pre-trained first object action recognition model may include the following steps:

[0080] In the first step, for the image feature information in the image feature information sequence, the following fourth generation step is performed:

[0081] Sub-step 1: determining the time step corresponding to the above image feature information as the target time step, wherein the target time step may be the input time step of the first object action recognition model under the temporal neural network.

[0082] Sub-step 2: determining the output image feature information corresponding to the previous time step, wherein the previous time step is the time step at the previous time of the target time step.

[0083] Sub-step 3: splicing the output image feature information with the image feature information to generate spliced ​​feature information.

[0084] Sub-step 4, inputting the above-mentioned spliced ​​feature information into the feature extraction unit corresponding to the above-mentioned target time step included in the above-mentioned first object action recognition model to generate output image feature information corresponding to the above-mentioned target time step. Among them, the feature extraction unit can be a convolutional neural network. The output image feature information can represent the summary information of the semantic feature information of the object action image at the current time step.

[0085] Sub-step 5, in response to determining that the target time step is a multiple of the target value, inputting the output image feature information into the action recognition feature information generation model corresponding to the target time step and included in the first object action recognition model to generate the action recognition feature information corresponding to the target time step. In practice, the target value can be a preset value. For example, the target value can be a value of 3.

[0086] Sub-step 6, in response to determining that the above-mentioned image feature information is the image feature information of the target position in the image feature information sequence, removing the action recognition feature information of the target position from the obtained action recognition feature information sequence to obtain a post-removal action recognition feature information sequence. The target position may be the position corresponding to the image feature information with the latest corresponding image time information in the image feature information sequence.

[0087] Sub-step 7, according to the time step sequence corresponding to the above-mentioned action recognition feature information sequence after removal, set the corresponding feature weight information sequence. There is a one-to-one correspondence between the feature weight information in the feature weight information sequence and the time step in the time step sequence. The feature weight information can represent the importance of the action recognition feature information at the corresponding time step.

[0088] As an example, the execution entity may use a time step and weight information association table to determine a feature weight information sequence corresponding to a time step sequence.

[0089] Sub-step 8, inputting the above-mentioned feature weight information sequence and the above-mentioned removed action recognition feature information sequence into the candidate action recognition information generation model based on the attention mechanism included in the above-mentioned first object action recognition model to generate the first candidate action recognition information.

[0090] Sub-step 9, inputting the action recognition feature information of the target position into the multi-layer serially connected fully connected layer included in the first object action recognition model to output second candidate action recognition information.

[0091] Sub-step 10, input each removed action recognition feature information in the above-mentioned removed action recognition feature information sequence into the corresponding fully connected layer included in the above-mentioned first object action recognition model to output the third candidate action recognition information to obtain the third candidate action recognition information sequence.

[0092] Sub-step 11, generating summary action recognition information for the third candidate action recognition information sequence.

[0093] Sub-step 12, generating object action recognition information according to the first candidate action recognition information, the second candidate action recognition information and the summary action recognition information.

[0094] In the second step, in response to determining that the image feature information is not the image feature information of the target position in the image feature information sequence, the next image feature information corresponding to the image feature information is used as the image feature information, and the fourth generation step is continued.

[0095] Step 109: Generate motion hazard identification information for the target object based on the object motion information sequence and the object motion identification information.

[0096] In some embodiments, the execution subject may generate the motion hazard identification information for the target object according to the object motion information sequence and the object motion identification information.

[0097] In some optional implementations of some embodiments, the generating of the motion hazard identification information for the target object according to the object motion information sequence and the object motion identification information may include the following steps:

[0098] The first step is to determine at least one summary object action information corresponding to the above object action information sequence.

[0099] As an example, first, the execution subject may perform action information deduplication processing on each object action information in the object action information sequence to generate a deduplicated object action information set. Then, the occurrence frequency of each object action information in the deduplicated object action information set in the object action information sequence is determined to obtain a frequency set. Finally, the deduplicated object action information set and the frequency set are correspondingly combined to generate summary object action information, and at least one summary object action information is obtained.

[0100] As another example, the execution subject may perform action information deduplication processing on each object action information in the object action information sequence to generate a deduplicated object action information set as at least one summary object action information.

[0101] In the second step, the at least one summary object action information and the object action recognition information are fused to generate an action information set.

[0102] As an example, the execution subject may perform deduplication fusion on at least one summary object action information and object action recognition information to generate an action information set.

[0103] In the third step, each action information in the above action information set and the set target scene information are input into the pre-trained action hazard identification information generation model to generate the above action hazard identification information. Among them, the target scene information can be vehicle driving scene information. The target scene information can be a scene identifier. For example, the scene identifier can be 001, corresponding to the vehicle driving scene. Among them, the action hazard identification information generation model can be a neural network model for generating action hazard identification information. The action hazard identification information can be identification information that characterizes the existence of danger in the target scene of the corresponding action information. In practice, the action hazard identification information can include: an action hazard identifier and an action hazard level identifier. In practice, the action hazard identification information generation model can be a conventional pre-set action hazard identification corresponding rule. That is, the action hazard identification information model characterizes the dangerous situation of the action information in the corresponding scene. That is, whether there is danger and the level of danger. The action hazard identification information generation model can be a summary based on the historical relevant expert experience, characterizing the dangerous situation of the action information in the corresponding scene.

[0104] In some optional implementations of some embodiments, after step 109, the steps further include:

[0105] In the first step, in response to determining that the verification information representation fails the verification, object action information in the object action information sequence that fails the action information verification is determined as target object action information, and at least one target object action information is obtained.

[0106] The second step is to determine at least one action object characteristic information cluster corresponding to the at least one target object action information, wherein the target object action information in the at least one target object action information and the action object characteristic information cluster in the at least one action object characteristic information cluster have a one-to-one correspondence.

[0107] Step 3: for each action object feature information cluster in the at least one action object feature information cluster, perform the following fifth generation step:

[0108] Sub-step 1: Determine the cluster semantic information corresponding to the action object feature information cluster.

[0109] As an example, first, the above-mentioned execution entity can input the action object feature information cluster into a pre-trained common semantic feature information generation model to generate action object common semantic feature information. Then, the action object common semantic feature information is determined as the above-mentioned cluster semantic information. Among them, the common semantic feature information represents the feature semantic information related to the action object between each action object feature information. Among them, the action object common semantic feature information generation model can be a neural network model that generates common semantic feature information. Specifically, the common semantic feature information generation model can be a multi-layer series of convolutional layers.

[0110] In practice, the common semantic feature information generation model can be obtained based on the training data set using a conventional label training method. Specifically, the action object information generation layer can be used to assist the model training of the common semantic feature information generation model. The action object information generation layer can be a network layer located after the common semantic feature information generation model and used to output object action information. In practice, the action object information generation layer can be a classification layer with multiple layers connected in series. For example, a fully connected layer.

[0111] Sub-step 2: Based on the cluster semantic information and the action object feature information cluster, a generative and adversarial neural network model is used to generate at least one candidate action feature information for the action object feature information cluster. In practice, the generative and adversarial neural network model can be a GAN model.

[0112] As an example, first, the above-mentioned execution subject can input cluster semantic information into the generative model in the generative and adversarial neural network model to generate at least one initial candidate action feature information. Among them, the vector dimension corresponding to the initial candidate action feature information is the same as the vector dimension corresponding to the action object feature information. Then, determine the cluster center corresponding to the action object feature information cluster as the target cluster center. Next, determine the cosine distance between the target cluster center and each initial candidate action feature information as the candidate cosine distance, and obtain at least one candidate cosine distance. Remove the initial candidate action feature information whose corresponding candidate cosine distance is greater than the preset distance from at least one initial candidate action feature information to obtain at least one candidate action feature information.

[0113] Sub-step 3: adding the at least one candidate action feature information to the action object feature information cluster to generate an added action object feature information cluster.

[0114] Sub-step 4: determining the cluster center feature information corresponding to the action object feature information cluster added as the first cluster center feature information.

[0115] As an example, the execution subject may re-confirm the cluster center feature information corresponding to the added action object feature information cluster based on the distance calculation method of the clustering algorithm as the first cluster center feature information.

[0116] Sub-step 5: determining the cluster semantic information corresponding to the action object feature information cluster added as the target cluster semantic information.

[0117] In the fourth step, in response to determining that the above-mentioned first cluster center feature information is the same as the cluster center feature information corresponding to the above-mentioned action object feature information cluster, and the target cluster semantic information is the same as the above-mentioned cluster semantic information, generate the object action information corresponding to the above-mentioned added action object feature information cluster as the object action information corresponding to the action object feature information cluster.

[0118] The above-mentioned embodiments of the present disclosure have the following beneficial effects: through the identification information generation method of some embodiments of the present disclosure, the action hazard identification information for the target object can be accurately and efficiently generated to warn of dangerous driving and avoid the occurrence of subsequent dangerous driving. Specifically, the reason why the relevant action hazard identification information is not accurate and efficient is that on the basis of the relatively large amount of vehicle driving data, the efficiency of the technicians in video viewing is limited, which may lead to the omission of vehicles with action hazard driving. Based on this, the identification information generation method of some embodiments of the present disclosure first obtains a real-time video shot for the target vehicle. Here, the obtained real-time video is used as a data basis to extract the action behavior information corresponding to the target object. Then, the real-time video is preprocessed to generate a frame image sequence for the target object, which is processed into a frame image form to facilitate subsequent object segmentation processing. Then, each frame image in the frame image sequence is input into a pre-trained object segmentation model to accurately generate an object segmentation image to obtain an object segmentation image sequence. Then, the object action feature information and image time information corresponding to each object segmentation image in the object segmentation image sequence are generated to obtain an object action feature information sequence and an image time information sequence. Here, by determining the object action feature information and image time information corresponding to each object segmentation image, it is convenient to perform clustering processing on the image content semantics corresponding to the object segmentation image in the subsequent step, and combine the object feature information with similar image content semantics and continuous time, that is, perform feature information clustering on the above object action feature information sequence and the above image time information sequence to generate an object action feature information cluster sequence, wherein the image time information corresponding to each object action information in each object action feature information cluster is continuous, and the number of object action information corresponding to each object action feature information cluster is the target number. Here, by performing feature information clustering, the corresponding action of the target object in the time period corresponding to the real-time video can be effectively parsed, so as to generate object action information, that is, object action information sequence, in a more detailed manner. Secondly, the object action information sequence is verified to generate verification information to determine the accuracy of the object action information in the object action information sequence. Further, in response to determining that the verification information indicates that the verification is passed, according to the image feature information sequence corresponding to the object segmentation image sequence, the object action recognition information can be accurately generated using the pre-trained first object action recognition model. Finally, based on the object action information sequence and the object action recognition information, the action hazard recognition information for the target object is accurately generated. In summary, through the clustering of object action feature information and image time information, the object segmentation image sequence can be refined and decomposed into action details to generate an accurate object action information sequence.In addition, the first object action recognition model is used to consider the overall video to generate object action recognition information in the overall video state. Thus, based on the object action information sequence and the object action recognition information, action hazard recognition information for the target object can be accurately generated.

[0119] Further references Figure 2 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an identification information generating device. These device embodiments are similar to Figure 1 Corresponding to the method embodiments shown, the identification information generating device can be specifically applied to various electronic devices.

[0120] like Figure 2 As shown, an identification information generating device 200 includes: an acquisition unit 201, a video preprocessing unit 202, an input unit 203, a first generating unit 204, an execution unit 205, a second generating unit 206, a verification unit 207, a third generating unit 208 and a fourth generating unit 209. Among them, the acquisition unit 201 is configured to acquire a real-time captured video of the target vehicle; the video preprocessing unit 202 is configured to perform video preprocessing on the above-mentioned real-time captured video to generate a frame image sequence for the target object; the input unit 203 is configured to input each frame image in the above-mentioned frame image sequence into a pre-trained object segmentation model to generate an object segmentation image and obtain an object segmentation image sequence; the first generation unit 204 is configured to generate object action feature information and image time information corresponding to each object segmentation image in the above-mentioned object segmentation image sequence to obtain an object action feature information sequence and an image time information sequence; the execution unit 205 is configured to perform feature information clustering for the above-mentioned object action feature information sequence and the above-mentioned image time information sequence to generate an object action feature information cluster sequence, wherein each object action feature information cluster has a plurality of image segments, each of which ... The image time information corresponding to each object action information is continuous, and the number of object action information corresponding to each object action feature information cluster is the target number; the second generation unit 206 is configured to generate the object action information corresponding to each object action feature information cluster in the above-mentioned object action feature information cluster sequence to obtain the object action information sequence; the verification unit 207 is configured to perform action information verification on the above-mentioned object action information sequence to generate verification information; the third generation unit 208 is configured to generate object action recognition information according to the image feature information sequence corresponding to the above-mentioned object segmentation image sequence in response to determining that the above-mentioned verification information representation has passed the verification, using the pre-trained first object action recognition model; the fourth generation unit 209 is configured to generate action hazard recognition information for the above-mentioned target object according to the above-mentioned object action information sequence and the above-mentioned object action recognition information.

[0121] It can be understood that the units recorded in the identification information generating device 200 are similar to the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the identification information generating device 200 and the units included therein, and will not be described in detail here.

[0122] Reference below Figure 3 , which shows a structural schematic diagram of an electronic device (eg, an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0123] like Figure 3 As shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0124] Typically, the following devices may be connected to the I / O interface 305: input devices 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 308 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 309. The communication devices 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 3 Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0125] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.

[0126] It should be noted that the computer-readable medium in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0127] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0128] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist independently without being assembled into the electronic device. The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: obtains a real-time video of the target vehicle; performs video preprocessing on the above-mentioned real-time video to generate a frame image sequence for the target object; inputs each frame image in the above-mentioned frame image sequence into a pre-trained object segmentation model to generate an object segmentation image and obtain an object segmentation image sequence; generates object action feature information and image time information corresponding to each object segmentation image in the above-mentioned object segmentation image sequence to obtain an object action feature information sequence and an image time information sequence; performs feature information clustering on the above-mentioned object action feature information sequence and the above-mentioned image time information sequence to generate an object action feature information cluster sequence. , wherein each image time information corresponding to each object action information in each object action feature information cluster is continuous, and the number of object action information corresponding to each object action feature information cluster is a target number; generating object action information corresponding to each object action feature information cluster in the above object action feature information cluster sequence to obtain an object action information sequence; performing action information verification on the above object action information sequence to generate verification information; in response to determining that the above verification information representation passes the verification, generating object action recognition information based on the image feature information sequence corresponding to the above object segmentation image sequence using a pre-trained first object action recognition model; generating action hazard recognition information for the above target object based on the above object action information sequence and the above object action recognition information.

[0129] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0130] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0131] The units described in some embodiments of the present disclosure may be implemented by software or by hardware. The described units may also be provided in a processor, for example, may be described as: a processor including an acquisition unit, a video preprocessing unit, an input unit, a first generation unit, an execution unit, a second generation unit, a verification unit, a third generation unit, and a fourth generation unit. The names of these units do not, in some cases, constitute limitations on the units themselves, for example, the acquisition unit may also be described as a "unit for acquiring real-time captured video of a target vehicle".

[0132] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0133] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) and the technical solutions formed.

Claims

1. A method for generating identification information, comprising: Obtain real-time video footage of the target vehicle; Performing video preprocessing on the real-time captured video to generate a frame image sequence for the target object; Inputting each frame image in the frame image sequence into a pre-trained object segmentation model to generate an object segmentation image, thereby obtaining an object segmentation image sequence; Generating object motion feature information and image time information corresponding to each object segmentation image in the object segmentation image sequence to obtain an object motion feature information sequence and an image time information sequence; Performing feature information clustering on the object action feature information sequence and the image time information sequence to generate an object action feature information cluster sequence, wherein each image time information corresponding to each object action information in each object action feature information cluster is continuous, and the number of object action information corresponding to each object action feature information cluster is a target number; Generating object action information corresponding to each object action feature information cluster in the object action feature information cluster sequence to obtain an object action information sequence; Performing action information verification on the object action information sequence to generate verification information; In response to determining that the verification information representation passes verification, generating object action recognition information using a pre-trained first object action recognition model according to an image feature information sequence corresponding to the object segmentation image sequence; Motion hazard identification information for the target object is generated according to the object motion information sequence and the object motion identification information.

2. The method according to claim 1, wherein: The generating the object action information corresponding to each object action feature information cluster in the object action feature information cluster sequence comprises: Time-sorting each object action feature information in the object action feature information cluster to generate an object action feature information subsequence; Grouping the object action feature information in the object action feature information subsequence into feature information groups to generate an object action feature information group sequence, wherein the image time information corresponding to the object action feature information in the object action feature information group is continuous; Using a pre-trained second object action recognition model, determine the action recognition information and action recognition probability information corresponding to each object action feature information group in the object action feature information group sequence, and obtain an action recognition information sequence and an action recognition probability information sequence; The object action information is generated according to the action recognition information sequence and the action recognition probability information sequence.

3. The method according to claim 1, wherein: The performing action information verification on the object action information sequence to generate verification information includes: Determine the cluster time corresponding to each object action feature information cluster in the object action feature information cluster sequence to obtain a cluster time sequence; For each object action feature information cluster in the object action feature information cluster sequence, the following first generation step is performed: Determine a cluster time corresponding to the object action feature information cluster as a target cluster time; Determine the previous cluster time and the next cluster time corresponding to the target cluster time; Determine the object action feature information cluster corresponding to the previous cluster time as the first object action feature information cluster, and determine the object action feature information cluster corresponding to the next cluster time as the second object action feature information cluster; Filtering a first number of object action feature information from the first object action feature information cluster to obtain a first object action feature information sub-cluster; Filtering a second number of object action feature information from the second object action feature information cluster to obtain a second object action feature information sub-cluster; fusing feature information of the first object motion feature information sub-cluster, the second object motion feature information sub-cluster, and the object motion feature information cluster to generate a fused feature information sequence; generating candidate object action information according to the fused feature information sequence; In response to determining that the candidate object action information is the same as the object action information corresponding to the object action feature information cluster, generating syndrome information indicating that the syndrome passes the verification; The verification information is generated according to the obtained syndrome information sequence.

4. The method according to claim 1, wherein: The performing of clustering the feature information of the object motion feature information sequence and the image time information sequence to generate an object motion feature information cluster sequence includes: Using a k-means clustering algorithm, clustering each object action feature information in the object action feature information sequence to generate an initial object action feature information cluster set; Determine the image time information corresponding to each cluster center feature information in the cluster center feature information set to obtain the image time information set, wherein the cluster center feature information set is the cluster center set corresponding to the initial object action feature information cluster set; For each initial object motion feature information cluster in the initial object motion feature information cluster set, the following second generation step is performed: Determine the cluster center feature information corresponding to the initial object action feature information cluster; Determine the image time information corresponding to the cluster center feature information as the first image time information; Removing the cluster center feature information from the initial object action feature information cluster to obtain a post-removal feature information set; Determine the time difference between the image time information corresponding to each feature information in the removed feature information set and the first image time information to obtain a time difference set; Determine the feature information in the removed feature information set whose corresponding time difference is greater than a predetermined value as pending feature information; For each undetermined feature information in the obtained undetermined feature information set, the following third generation step is performed: Determine the image time information corresponding to the undetermined feature information as the second image time information; Filter out the first image time information set obtained, the first image time information with the shortest time interval with the second image time information and the smallest cosine distance between corresponding features, as the target image time information; Generate belonging information indicating that the undetermined feature information belongs to the initial object action feature information cluster corresponding to the target image time information; According to the obtained attribution information set, the feature information of the pending feature information set in the initial object action feature information cluster set is classified to generate a classified object action feature information cluster set as the object action feature information cluster set.

5. The method according to claim 1, wherein: The generating object action recognition information according to the image feature information sequence corresponding to the object segmentation image sequence using a pre-trained first object action recognition model comprises: For the image feature information in the image feature information sequence, the following fourth generation step is performed: Determine a time step corresponding to the image feature information as a target time step; Determine output image feature information corresponding to a previous time step, wherein the previous time step is a time step previous to the target time step; Performing feature information splicing on the output image feature information and the image feature information to generate spliced ​​feature information; Inputting the splicing feature information into a feature extraction unit corresponding to the target time step included in the first object action recognition model to generate output image feature information corresponding to the target time step; In response to determining that the target time step is a multiple of the target value, inputting the output image feature information into an action recognition feature information generation model corresponding to the target time step and included in the first object action recognition model to generate action recognition feature information corresponding to the target time step; In response to determining that the image feature information is the image feature information of the target position in the image feature information sequence, removing the action recognition feature information of the target position from the obtained action recognition feature information sequence to obtain a post-removal action recognition feature information sequence; According to the time step sequence corresponding to the removed action recognition feature information sequence, setting a corresponding feature weight information sequence; Inputting the feature weight information sequence and the removed action recognition feature information sequence into a candidate action recognition information generation model based on an attention mechanism included in the first object action recognition model to generate first candidate action recognition information; Inputting the action recognition feature information of the target position into a multi-layer serially connected fully connected layer included in the first object action recognition model to output second candidate action recognition information; Inputting each removed action recognition feature information in the removed action recognition feature information sequence into a corresponding fully connected layer included in the first object action recognition model to output third candidate action recognition information to obtain a third candidate action recognition information sequence; generating summary action recognition information for the third candidate action recognition information sequence; generating object action recognition information according to the first candidate action recognition information, the second candidate action recognition information and the summary action recognition information; In response to determining that the image feature information is not the image feature information of the target position in the image feature information sequence, the next image feature information corresponding to the image feature information is used as the image feature information, and the fourth generating step is continued.

6. The method according to claim 1, wherein: The step of generating the action hazard identification information for the target object according to the object action information sequence and the object action identification information comprises: Determine at least one summary object action information corresponding to the object action information sequence; fusing the at least one summary object action information and the object action recognition information to generate an action information set; Each action information in the action information set and the set target scene information are input into a pre-trained action hazard identification information generation model to generate the action hazard identification information.

7. The method according to claim 1, wherein: The method further comprises: In response to determining that the verification information representation fails verification, determining object action information in the object action information sequence that fails the action information verification as target object action information, and obtaining at least one target object action information; Determine at least one action object feature information cluster corresponding to the at least one target object action information; For each action object feature information cluster in the at least one action object feature information cluster, the following fifth generation step is performed: Determining cluster semantic information corresponding to the action object feature information cluster; According to the cluster semantic information and the action object feature information cluster, using a generative and adversarial neural network model to generate at least one candidate action feature information for the action object feature information cluster; adding the at least one candidate action feature information to the action object feature information cluster to generate an added action object feature information cluster; Determine the cluster center feature information corresponding to the added action object feature information cluster as the first cluster center feature information; Determine the cluster semantic information corresponding to the added action object feature information cluster as the target cluster semantic information; In response to determining that the first cluster center feature information is the same as the cluster center feature information corresponding to the action object feature information cluster, and the target cluster semantic information is the same as the cluster semantic information, object action information corresponding to the added action object feature information cluster is generated as the object action information corresponding to the action object feature information cluster.

8. An identification information generating device, comprising: An acquisition unit is configured to acquire a real-time video of a target vehicle; A video preprocessing unit, configured to perform video preprocessing on the real-time captured video to generate a frame image sequence for a target object; An input unit is configured to input each frame image in the frame image sequence into a pre-trained object segmentation model to generate an object segmentation image and obtain an object segmentation image sequence; A first generating unit is configured to generate object motion feature information and image time information corresponding to each object segmentation image in the object segmentation image sequence, and obtain an object motion feature information sequence and an image time information sequence; an execution unit, configured to perform feature information clustering on the object action feature information sequence and the image time information sequence to generate an object action feature information cluster sequence, wherein each image time information corresponding to each object action information in each object action feature information cluster is continuous, and the number of object action information corresponding to each object action feature information cluster is a target number; A second generating unit is configured to generate object action information corresponding to each object action feature information cluster in the object action feature information cluster sequence to obtain an object action information sequence; a verification unit, configured to perform action information verification on the object action information sequence to generate verification information; a third generating unit configured to generate object action recognition information by using a pre-trained first object action recognition model according to an image feature information sequence corresponding to the object segmentation image sequence in response to determining that the verification information representation passes the verification; The fourth generating unit is configured to generate action hazard identification information for the target object according to the object action information sequence and the object action identification information.

9. An electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

10. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.