Video processing methods, apparatus, computer equipment and storage media

By decomposing and analyzing motion event information in logistics sorting videos, identifying target video segments, and using machine learning models to determine motion events, the problem of high computational resource consumption and poor performance in the recognition of violent sorting actions in logistics is solved, and efficient motion event recognition is achieved.

CN114758271BActive Publication Date: 2026-05-26JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
Filing Date
2022-03-24
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies consume a lot of computational resources and have poor recognition results in recognizing violent sorting actions in logistics.

Method used

By decomposing the video, multiple video segments are obtained, the action event information corresponding to the video segments is determined, the target video segment is identified and the object description information is extracted from it, and a machine learning model is used to determine whether the target action event has occurred.

Benefits of technology

It effectively reduces the computational resource consumption for target action event recognition and judgment, and improves the recognition effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114758271B_ABST
    Figure CN114758271B_ABST
Patent Text Reader

Abstract

This disclosure proposes a video processing method, apparatus, computer device, and storage medium. The method includes: decomposing a video to obtain multiple video segments; determining multiple action event information corresponding to each of the multiple video segments; determining a target video segment from the multiple video segments based on the action event information; identifying object description information from the target video segment; and determining whether a target action event has occurred in the scene described by the video based on the object description information. Because the video is first decomposed, and a target video segment is determined based on the decomposed video segments, and target action events in the scene are identified and judged around the target video segment, the method achieves rapid identification and judgment of target action events around the target video segment, effectively reducing the computational resources consumed in target action event identification and judgment, and improving the identification and judgment effect of target action events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer vision technology, and in particular to a video processing method, apparatus, computer equipment, and storage medium. Background Technology

[0002] In the scenario of recognizing violent sorting actions in logistics, violent sorting actions on packages are usually target actions such as kicking, throwing, or tossing packages. In practical applications, it is necessary to identify and judge these violent sorting target action events from logistics sorting videos to assist in the supervision of logistics operations.

[0003] In related technologies, it is common to use object detection and object tracking to determine action events in a scene. For example, frame-by-frame detection is used to detect and process action events in video.

[0004] This approach results in high computational resource consumption for recognition and judgment, and poor performance in recognizing and judging action events in the scene, leading to low efficiency. Summary of the Invention

[0005] This disclosure aims to at least partially address one of the technical problems in the related art.

[0006] Therefore, the purpose of this disclosure is to propose a video processing method, apparatus, computer device, and storage medium. Since the video is first decomposed, a target video segment is determined based on the multiple video segments obtained from the decomposition. The target action events in the scene are then identified and judged around the target video segment, thereby achieving rapid identification and judgment of target action events around the target video segment. This effectively reduces the computing resources consumed in the identification and judgment of target action events and improves the identification and judgment effect of target action events.

[0007] The video processing method proposed in the first aspect of this disclosure includes: decomposing a video to obtain multiple video segments; determining multiple action event information corresponding to the multiple video segments respectively, and determining a target video segment from the multiple video segments based on the action event information; identifying object description information from the target video segment; and determining whether a target action event has occurred in the scene described by the video based on the object description information.

[0008] The video processing method proposed in the first aspect of this disclosure decomposes a video to obtain multiple video segments, determines multiple action event information corresponding to each of the multiple video segments, identifies a target video segment from the multiple video segments based on the action event information, identifies object description information from the target video segment, and determines whether a target action event has occurred in the scene described by the video based on the object description information. Since the video is decomposed first, and the target video segment is determined based on the multiple video segments obtained from the decomposition, the target action event in the scene is identified and judged around the target video segment. This achieves rapid identification and judgment of the target action event around the target video segment, effectively reduces the computing resources consumed in the identification and judgment of the target action event, and improves the identification and judgment effect of the target action event.

[0009] The video processing apparatus according to the second aspect of this disclosure includes: a decomposition module for decomposing a video to obtain multiple video segments; a first determination module for determining multiple action event information corresponding to the multiple video segments respectively, and determining a target video segment from the multiple video segments based on the action event information; a first identification module for identifying object description information from the target video segment; and a second determination module for determining whether a target action event has occurred in the scene described by the video based on the object description information.

[0010] The video processing apparatus proposed in the second aspect of this disclosure decomposes a video to obtain multiple video segments, determines multiple action event information corresponding to each of the multiple video segments, identifies a target video segment from the multiple video segments based on the action event information, identifies object description information from the target video segment, and determines whether a target action event has occurred in the scene described by the video based on the object description information. Since the video is decomposed first, and the target video segment is determined based on the multiple video segments obtained from the decomposition, the target action event in the scene is identified and judged around the target video segment. This achieves rapid identification and judgment of the target action event around the target video segment, effectively reducing the computing resources consumed in the identification and judgment of the target action event, and improving the identification and judgment effect of the target action event.

[0011] A third aspect of this disclosure provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the video processing method as proposed in the first aspect of this disclosure.

[0012] The fourth aspect of this disclosure provides a non-transitory computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the video processing method as described in the first aspect of this disclosure.

[0013] A fifth aspect of this disclosure provides a computer program product that, when executed by an instruction processor, performs a video processing method as described in a first aspect of this disclosure.

[0014] Additional aspects and advantages of this disclosure will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this disclosure. Attached Figure Description

[0015] The above and / or additional aspects and advantages of this disclosure will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, in which:

[0016] Figure 1 This is a schematic flowchart of a video processing method according to an embodiment of the present disclosure;

[0017] Figure 2 This is a schematic flowchart of a video processing method according to another embodiment of this disclosure;

[0018] Figure 3 This is a schematic flowchart of a video processing method according to another embodiment of this disclosure;

[0019] Figure 4 This is a schematic diagram of the target action event determination process in an embodiment of this disclosure;

[0020] Figure 5 This is a schematic diagram of the structure of a video processing apparatus according to an embodiment of the present disclosure;

[0021] Figure 6 This is a schematic diagram of the structure of a video processing apparatus according to another embodiment of the present disclosure;

[0022] Figure 7 A block diagram of an exemplary computer device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation

[0023] Embodiments of this disclosure are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are used only to explain this disclosure, and should not be construed as limiting this disclosure. Rather, embodiments of this disclosure include all variations, modifications, and equivalents falling within the spirit and scope of the appended claims.

[0024] Figure 1 This is a schematic flowchart of a video processing method proposed in an embodiment of this disclosure.

[0025] It should be noted that the execution subject of the video processing method in this embodiment is a video processing device, which can be implemented by software and / or hardware. The device can be configured in a computer device, which may include, but is not limited to, a terminal, a server, etc.

[0026] like Figure 1 As shown, the video processing method includes:

[0027] S101: Decompose the video to obtain multiple video segments.

[0028] The embodiments disclosed herein can be applied to logistics sorting scenarios, such as e-commerce companies, logistics companies, or express delivery companies using sorting equipment or manual processing to sort logistics packages for subsequent logistics transportation or delivery. There are no limitations on this.

[0029] The video to be decomposed can be a real-time video of a logistics sorting scenario captured by a camera device, or a video transmitted from other electronic devices. The video processing device can be pre-configured with a video acquisition device to capture real-time video of a logistics sorting scenario and then decompose the video. Alternatively, the video processing device can be configured with a data transmission interface to receive logistics sorting scenario videos transmitted from other electronic devices and then decompose the video. There are no restrictions on this.

[0030] In this embodiment of the disclosure, when decomposing a video, the video can be segmented into non-overlapping fixed-length segments. The length of the video segment can be set as a configurable variable. This configurable variable can be configured as needed for videos in different scenarios. For example, the length of the video segment can be configured to 10 seconds or 20 seconds. Then, the video is segmented into fixed-length segments to obtain multiple non-overlapping video segments. There are no restrictions on this.

[0031] In other embodiments, before decomposing the video, some preprocessing operations can be performed on the acquired video. These preprocessing operations can be, for example, filtering the video to remove still images and images with poor quality, such as distorted images, flickering images, fast-forwarded images, and shaky images. The execution frequency of various preprocessing operations can be configured as needed using various filtering algorithms. For example, the Laplacian operator can be used for distorted image detection. Then, the preprocessed video can be segmented into multiple non-overlapping video segments.

[0032] S102: Determine multiple action event information corresponding to multiple video segments respectively, and determine the target video segment from the multiple video segments based on the action event information.

[0033] Among them, action events in video clips refer to events in which objects in video clips perform interactive movements and generate movement trajectories. The objects are people, objects, or other subjects in the video, and there are no restrictions on this.

[0034] For example, the action events in the video clip could be events involving the sorting of logistics packages. Depending on different operational requirements, these sorting events can be categorized as normal sorting events or abnormal sorting events. Normal sorting events could be, for example, sorting events that comply with logistics operation supervision regulations, while abnormal sorting events could be, for example, sorting events involving kicking, throwing, or tossing packages. There are no restrictions on these.

[0035] Among them, action event information can be used to describe the above-mentioned action events, that is, action event information is used to characterize the features of the corresponding action events.

[0036] For example, action event information can include the type of action event, the start time of the action event, and the score of the action event. Action event information can be used as reference information to help identify a target video segment from multiple video segments.

[0037] Among them, the video segment that is suitable for the needs of the video processing task is identified from multiple video segments. The target video segment can be, for example, a video segment that is determined to have a high probability of describing an action event from multiple video segments, or a video segment that is determined to carry continuous action event information from multiple video segments. There is no restriction on this.

[0038] In this embodiment of the disclosure, when determining multiple action event information corresponding to multiple video segments respectively, a Temporal Action Localization (TAL) filter can be used to process the multiple video segments to obtain the action event information of each video segment output by the TAL algorithm.

[0039] In other embodiments, Gaussian Temporal Awareness Networks (GTAN) can be used to perform motion localization processing on multiple video segments respectively. The spatiotemporal features of multiple video segments are extracted using a basic feature network, and Gaussian kernel convolutional layers are used to predict the action event category and the start and end time of the action event based on the spatiotemporal features, so as to determine the action event information of each video segment.

[0040] In this embodiment of the disclosure, after determining multiple action event information corresponding to multiple video segments respectively, a target video segment can be determined from the multiple video segments based on the action event information. The target video segment may be a video segment with a higher action event score in the corresponding action event information, or it may be a video segment that is determined to have a high probability of containing action events, or it may be a video segment that is determined to carry continuous action event information. There are no restrictions on this.

[0041] In this embodiment of the disclosure, when determining the target video segment from multiple video segments based on action event information, the multiple video segments can be filtered based on the action event score in the action event information. For example, a score threshold can be set for the action event score, and then the action event scores corresponding to multiple video segments can be compared and judged with reference to the score threshold to filter out video segments with action event scores less than the score threshold in order to obtain the target video segment. There are no restrictions on this.

[0042] S103: Identify object description information from the target video segment.

[0043] The object can be an element to be identified and judged in the video clip, such as a person or a package in the video clip. The object description information can be some spatiotemporal feature information related to the element to be identified and judged in the video clip. This object description information can be used to assist in the identification and judgment of target action events, and there are no restrictions on it.

[0044] For example, object description information can be used to locate the specific location of a target action event in a video frame. The object in the target video clip can be a person or an object. The object description information can be the category information of the person, object, etc. in the target video clip, the interaction relationship between the objects, and the motion state information of the objects. This motion state information can be, for example, motion trajectory information, speed information, acceleration information, and motion angle information, or any other possible descriptive information of the object can be used as the object description information, without any restrictions.

[0045] In this embodiment of the disclosure, when identifying object description information from a target video segment, a fine-grained video analysis algorithm can be used to process the target video segment. The images in the target video segment can be analyzed pixel by pixel to detect the target object. The target object in each frame of the image can be labeled to track the target object. Feature values ​​in each frame of the image can be extracted, and then multiple features can be fused to obtain depth feature values. The corresponding object description information can be determined based on the depth feature values ​​of the target video segment.

[0046] In other embodiments, a trajectory analysis processing model can be used to analyze and process the motion trajectory of objects in the target video to obtain motion state information of the target object as object description information. Alternatively, any other method that can analyze and generate classification results, location information, and motion state physical quantity information of people and objects can be used to identify object description information from the target video segment.

[0047] S104: Based on the object description information, determine whether a target action event has occurred in the scene described by the video.

[0048] Among them, the target action event refers to the interactive motion event between objects that need to be identified and judged from the video clip. The target action event can be adaptively set according to the regulatory requirements of the actual logistics sorting scenario, and there are no restrictions on it.

[0049] For example, the target action event can be an action event that sorts the package using abnormal sorting methods such as kicking, throwing, or tossing, or it can be any other action event that performs abnormal sorting of the package, without any restrictions.

[0050] In this embodiment, after identifying object description information from the target video segment, it can determine whether a target action event has occurred in the scene described by the video based on the object description information.

[0051] In this embodiment of the disclosure, when determining whether a target action event has occurred in the scene described by the video based on the object description information, the object description information can be input into a machine learning classifier model. The machine learning classifier model can then be used to judge the object physical quantity information in the object description information to obtain the judgment result of the object description information. This judgment result can be used to determine whether a target action event has occurred in the scene described by the video.

[0052] In other embodiments, other machine learning algorithms can also be used to determine whether a target action event has occurred in the scene described by the video. For example, Gradient Boosting Decision Tree (GBDT) or binary classification decision tree can be used to judge and process the object description information to determine whether a target action event has occurred in the scene described by the video. Alternatively, a neural network model can be used to judge and process the object description information. This neural network model can be, for example, a convolutional neural network model, a two-stream convolutional network model, or a dilated convolutional network model to determine whether a target action event has occurred in the scene described by the video. There are no restrictions on this.

[0053] In this embodiment, the video is decomposed to obtain multiple video segments. Multiple action event information corresponding to each video segment is determined. Based on the action event information, a target video segment is identified from the multiple video segments. Object description information is identified from the target video segment. Based on the object description information, it is determined whether a target action event has occurred in the scene described by the video. Since the video is decomposed first, and the target video segment is determined based on the multiple video segments obtained from the decomposition, the target action event in the scene is identified and judged around the target video segment. This achieves rapid identification and judgment of the target action event around the target video segment, effectively reducing the computing resources consumed by the identification and judgment of the target action event and improving the identification and judgment effect of the target action event.

[0054] Figure 2 This is a schematic flowchart of a video processing method proposed in another embodiment of this disclosure.

[0055] like Figure 2 As shown, the video processing method includes:

[0056] S201: Determine the total duration of the video.

[0057] The total duration refers to the duration of the acquired video before it is decomposed. This total duration can be the duration of a single acquired video segment or the total duration of a video composed of multiple acquired videos with a continuous time sequence.

[0058] In this disclosure, when determining the total duration of a video, the total duration of the video before it is decomposed can be obtained in the background using a time acquisition method, or the start and end times of multiple video segments can be obtained, and the durations of the multiple videos obtained can be accumulated in chronological order to obtain the total duration of the video stream, or the total duration of the video can be annotated when the video is acquired to determine the total duration of the video through time annotation. There are no restrictions on this.

[0059] S202: Divide the total duration into multiple duration segments.

[0060] In this embodiment of the disclosure, when the video is segmented to obtain multiple video segments, a duration segment variable for segmenting the total duration can be configured. The duration segment variable is set according to the duration requirement of the video segments based on the time positioning algorithm, and then the total duration is segmented to obtain multiple duration segments. There are no restrictions on this.

[0061] S203: The video is segmented according to multiple duration segments to obtain multiple video clips.

[0062] In this embodiment of the disclosure, after the total duration is divided into multiple duration segments as described above, the video can be segmented according to the multiple duration segments to obtain multiple video clips.

[0063] In other embodiments, multiple video segments obtained from the segmentation process can be input into a time-localization algorithm model for prediction. If the duration of the multiple video segments obtained from the segmentation process is greater than the input duration set by the model, the video can be further segmented and sampled. A sampling frequency can be set, such as 10 frames per second, and multiple video segments can be sampled according to the sampling frequency to obtain the processed video segments as the video segments after the video has been decomposed.

[0064] S204: Determine multiple action event information corresponding to multiple video segments respectively, and determine the target video segment from the multiple video segments based on the action event information.

[0065] For a detailed description of S204, please refer to the above embodiments, which will not be repeated here.

[0066] S205: Identify at least two video segments to be merged from multiple target video segments.

[0067] In this embodiment of the disclosure, after determining multiple target video segments from multiple video segments based on action event information, at least two video segments to be merged can be identified from the multiple target video segments, and then the identified video segments to be merged can be merged.

[0068] In this embodiment of the disclosure, when identifying video segments to be merged from multiple target video segments, a video processing model can be used to identify multiple target video segments from the start time to the end time of the video segments, so as to identify whether there is an action event in the target video segment that spans video segments at the end time. The action event can span two video segments or more than two video segments. Then, multiple video segments that completely contain the action event can be used as video segments to be merged.

[0069] Optionally, in some embodiments, the action event information includes: action event type, start and end time of the action event. When at least two video segments to be merged are identified from multiple target video segments, the interval value of the start and end time between adjacent target video segments can be determined. If the interval value is less than or equal to the interval threshold, and the action event type of the adjacent target video segments is the same, then the adjacent target video segments are taken as video segments to be merged. Thus, the video segments to be merged that can be merged can be identified according to the interval value of the start and end time between adjacent target video segments, and the video segments to be merged are merged, avoiding the destruction of the effective information carried in the original video data, ensuring the integrity of the original video data, thereby helping to improve the accuracy of determining whether a target action event has occurred in the video.

[0070] Among them, the action event type can be used to describe the classification of action events. For example, the action event type refers to the sorting operation of logistics packages in the video clip. This action event type can be the normal sorting operation of packages or the abnormal violent sorting operation. The abnormal sorting operation event type can be the sorting of packages by kicking, throwing or tossing, etc. There are no restrictions on this.

[0071] The start and end time interval threshold between adjacent target video segments can be configured as needed to determine whether there is a possibility of merging adjacent target video segments.

[0072] In this embodiment of the disclosure, when at least two video segments to be merged are identified from multiple target video segments, a time interval threshold between the start and end times of adjacent target video segments can be preset. This time interval threshold can be set to, for example, 0.5 seconds. Then, based on the interval threshold and the action event type of the adjacent target video segments, it can be determined whether there is a possibility of merging the adjacent target video segments. If the time interval between the start and end times of adjacent target video segments is less than or equal to the interval threshold, and the action event type of the adjacent target video segments is the same, then the two adjacent target video segments are regarded as video segments to be merged. Alternatively, multiple target video segments that meet the conditions can be regarded as videos to be merged.

[0073] In this embodiment of the disclosure, after identifying at least two video segments to be merged from multiple target video segments, the determined video segments to be merged can be merged, as detailed in subsequent embodiments.

[0074] S206: Merge at least two video segments to be merged to obtain a merged video segment.

[0075] Among them, the video segment obtained by merging at least two video segments to be merged can be referred to as a merged video segment.

[0076] In this embodiment of the disclosure, after identifying at least two video segments to be merged from multiple target video segments, a video processing model can be used to merge the two identified video segments to obtain a merged video segment. Alternatively, the multiple identified video segments to be merged can be merged in chronological order to obtain a merged video segment. The merged video segment and the target video segments can be processed to obtain object description information, which can be used to assist in the identification and determination of target action events.

[0077] S207: Identify object description information from the target video segment and the merged video segment.

[0078] After obtaining the merged video segments as described above, the embodiments of this disclosure can perform identification processing on the target video segments and the merged video segments to obtain object description information in the target video segments and the merged video segments.

[0079] In this embodiment of the disclosure, when identifying object description information from target video segments and merged video segments, a fine-grained video analysis algorithm can be used to identify and analyze the target video segments and merged video segments. The target video segments and merged video segments can be analyzed pixel by pixel to detect the target object. The target object in each frame of the image is labeled for target object tracking. The feature values ​​in each frame of the image are extracted separately, and then multiple features are fused to obtain depth feature values. The corresponding object description information is determined based on the depth feature values ​​of the target video segments and merged video segments.

[0080] In other embodiments, object motion trajectory data can be extracted from the target video segment and the merged video segment, and the motion trajectory data can be analyzed and processed using a trajectory analysis and processing model to obtain the classification results of people and objects in the target video segment and the merged video segment, as well as location information and motion state physical quantity information, etc., as object description information. Alternatively, any other possible method can be used to identify object description information from the target video segment and the merged video segment, without any limitation.

[0081] S208: Based on the object description information, determine whether a target action event has occurred in the scene described by the video.

[0082] For a detailed description of S208, please refer to the above embodiments, which will not be repeated here.

[0083] In this embodiment, the video is decomposed to obtain multiple video segments. Multiple action event information corresponding to each video segment is determined. Based on the action event information, a target video segment is identified from the multiple video segments. Object description information is identified from the target video segment. Based on the object description information, it is determined whether a target action event has occurred in the scene described by the video. Since the video is first decomposed, and the target video segment is determined based on the multiple video segments obtained from the decomposition, the target action event in the scene is identified and judged around the target video segment. This achieves rapid identification and judgment of the target action event around the target video segment, effectively reducing the computational resources consumed in the identification and judgment of the target action event, and improving the identification and judgment effect of the target action event. The length of the video segments obtained from the video segmentation process is adaptively configured, so that video segments of optimal length can be used for subsequent video processing logic, assisting in improving the computational efficiency of operations such as object description information recognition for video segments. Target video segments that can be merged are identified based on the start and end time interval between adjacent target video segments and merged, avoiding damage to the original video data and ensuring the integrity of the original video data. This helps to improve the accuracy of judging whether a target action event has occurred in the video.

[0084] Figure 3 This is a schematic flowchart of a video processing method proposed in another embodiment of this disclosure.

[0085] like Figure 3 As shown, the video processing method includes:

[0086] S301: Decompose the video to obtain multiple video segments.

[0087] For a detailed description of S301, please refer to the above embodiments, which will not be repeated here.

[0088] S302: Input multiple video clips into a pre-trained temporal motion localization model to obtain multiple action event information output by the temporal motion localization model. The action event information also includes: recognition score value.

[0089] In this process, the initial first artificial intelligence model is trained in advance using sample video clips and sample action event information of the sample video clips until the first artificial intelligence model converges. The trained first artificial intelligence model is then used as the time action localization model.

[0090] The artificial intelligence model used to train the time-action localization model can be called the first artificial intelligence model. The first artificial intelligence model can be a machine learning model or a neural network model, and it has the function of information extraction.

[0091] In this embodiment of the disclosure, sample video clips and sample action event information of the sample video clips can be obtained in advance. The sample video clips and sample action event information of the sample video clips can come from multiple public test sets. The obtained sample video clips and sample action event information of the sample video clips are used to train a first artificial intelligence model, and the model parameters are iteratively updated until the first artificial intelligence model converges. The converged first artificial intelligence model has a faster inference speed and better inference speed. The first artificial intelligence model trained to convergence is used as the time action localization model.

[0092] The temporal action localization model can be used to analyze and process video clips, and has the ability to extract action event information from video clips. The temporal action localization model can be, for example, an anchor-free saliency-based detector (AFSD), a Gaussian temporal awareness network (GTAN) model, or any other temporal localization algorithm model or a convolutional neural network model that can perform temporal action localization, etc., without any restrictions.

[0093] The action event information also includes an identification score, which is used to characterize the probability of an action event occurring in a video clip. Video clips with higher identification scores are more likely to have that action event occurring.

[0094] In this embodiment of the disclosure, after decomposing the video into multiple video segments, the multiple video segments can be input into a pre-trained temporal action localization model. The temporal action localization model is used to perform feature extraction, coarse prediction of action events, and fine prediction of action events on the video segments to predict the recognition score value of the action events in the video segments. The recognition score value can be used as the action event information in the corresponding video segments.

[0095] Optionally, in some embodiments, the time-motion localization model includes: a video feature extraction sub-model, a time regression sub-model, and a category classification sub-model connected to the video feature extraction sub-model. When multiple video segments are input into the pre-trained time-motion localization model to obtain multiple action event information output by the time-motion localization model, the multiple video segments can be input into the video feature extraction sub-model to obtain multiple video features output by the video feature extraction sub-model. The time regression sub-model is used to identify the start and end times of the action events for each of the multiple video features, obtaining multiple start and end times corresponding to each of the multiple action events. The category classification sub-model is used to identify the types of the action events for each of the multiple video features, obtaining multiple start and end times corresponding to each of the multiple action events. For multiple corresponding action event types, the identification of action events from the corresponding video features is scored based on the action event type and the start and end times of the action event, resulting in an identification score value. The action event type, start and end times of the action event, and identification score are collectively used as action event information. This allows for the use of a time-based motion localization model with superior processing performance to process segmented video segments and obtain action event identification scores as corresponding action event information. This action event information can be used to filter video segments, thereby improving the efficiency of filtering video segments to obtain target video segments. Furthermore, subsequent target action event judgment on target video segments can save computational resources and improve the efficiency of target action event judgment and processing.

[0096] The time-motion localization model includes: a video feature extraction sub-model, a time regression sub-model connected to the video feature extraction sub-model, and a category classification sub-model.

[0097] The video feature extraction sub-model is used to extract the spatiotemporal features of video segments and convert the spatiotemporal features of video segments into feature pyramids. This video feature extraction sub-model can be an Inflated 3D ConvNet (I3D) model.

[0098] Among them, the time regression sub-model connected to the video feature extraction sub-model can be used to perform preliminary estimation of the start and end times of action events in video segments, and the category classification sub-model connected to the video feature extraction sub-model can be used to perform preliminary estimation of the categories of action events in video segments. The time regression sub-model and the category classification sub-model connected to the video feature extraction sub-model can be trained using the spatiotemporal feature pyramid of video segments extracted by the video feature sub-model.

[0099] In this embodiment of the disclosure, when multiple video segments are input into a pre-trained temporal motion localization model to obtain multiple action event information output by the temporal motion localization model, the multiple video segments can be input into a video feature extraction sub-model. The video feature extraction sub-model processes the video segments to obtain spatiotemporal features corresponding to the multiple video segments, and converts the spatiotemporal features into a feature pyramid. Then, the feature pyramid information is used to train a temporal regression sub-model and a category classification sub-model. The temporal regression sub-model is used to identify the start and end times of action events for the multiple video features. The output of the temporal regression sub-model is then fine-tuned to obtain the processed output of the temporal regression sub-model as multiple start and end times corresponding to the multiple action events. The category classification sub-model is used to identify the types of action events for the multiple video features. The output of the category classification sub-model is then fine-tuned to obtain the processed output of the category classification sub-model as multiple action event types corresponding to the multiple action events.

[0100] In this embodiment of the disclosure, a machine learning model can be trained using supervised data, and then the machine learning model can be used to extract video features from multiple video segments. Based on the corresponding video features, action events in the video segments can be identified, and an identification score value for the action events can be given.

[0101] In this embodiment of the disclosure, after processing and obtaining the action event type, the start and end time of the action event, and the recognition score of the action event, the action event type, the start and end time of the action event, and the recognition score can be used together as action event information. Then, the action event information can be used to filter multiple video segments to obtain the target video segment, as detailed in the following embodiments.

[0102] S303: Select the video segment corresponding to the action event information whose recognition score value is greater than or equal to the scoring threshold as the target video segment.

[0103] The scoring threshold is a pre-set threshold for the identification score value, used to determine the scoring limit value that meets the requirements of the actual video processing and detection scenario. This scoring threshold is a numerical threshold that can be used to examine the video segments corresponding to action event information in order to select the target video segment from multiple video segments. Video segments with an identification score value greater than or equal to the scoring threshold can indicate that the target action event is more likely to occur in the video segment, while video segments with an identification score value less than the scoring threshold can indicate that the target action event is less likely to occur in the video segment.

[0104] In this embodiment of the disclosure, after inputting multiple video segments into a pre-trained temporal action localization model to obtain multiple action event information output by the temporal action localization model, a comparison and judgment can be triggered based on a scoring threshold to identify the recognition score values ​​in the action event information corresponding to the multiple video segments. Non-maximum suppression (NMS) can be used to filter the multiple video segments in the time dimension, and the video segments corresponding to the action event information whose recognition score values ​​are greater than or equal to the scoring threshold can be selected as the target video segments.

[0105] S304: Identify multiple target objects from the target video segment.

[0106] The target object refers to the object in which the target action event occurs in the video clip. Specifically, the target object can be people or objects in the video clip.

[0107] In this embodiment of the disclosure, after selecting the video segment corresponding to the action event information to which the identification score value is greater than or equal to the scoring threshold is selected as the target video segment, multiple target objects can be identified from the target video segment.

[0108] In this embodiment of the disclosure, when identifying multiple target objects from a target video segment, a target detection model can be used to perform target detection processing on the target video segment to identify multiple target objects in the video segment and to label the target objects.

[0109] S305: Determine multiple object categories and multiple location information corresponding to multiple target objects.

[0110] Among them, the object category corresponding to the target object can be the type of object in which the target action event occurs in the video clip. The object category corresponding to the target object can be specifically, for example, a person or an object in the video clip. Multiple location information can be the specific location information of the target object in the video frame when the target action event occurs.

[0111] In this embodiment of the disclosure, when determining the multiple object categories and multiple location information corresponding to multiple target objects, target detection and tracking technology can be used. The target video segment is input into the target detection and tracking model, and the target object is tracked, detected and identified using the target object's tag, so as to obtain the multiple object categories and multiple location information corresponding to the multiple target objects.

[0112] S306: Determine the interaction state information and relative motion information between different target objects.

[0113] Among them, the interaction state information between different objects can be the interaction state information between people and objects in the video clip. The interaction state information can be specific, such as the state in which people and objects in the video clip are in contact, or the state in which they are a certain spatial distance apart, or the state in which they are not in contact, etc. The relative motion information between different objects can be represented by the physical quantities of motion state between the objects. These physical quantities of motion state can be the motion state distance, motion speed, acceleration, and motion angle between people and objects, etc., without any restrictions.

[0114] In this embodiment of the disclosure, when determining the interaction state information and relative motion information between different target objects, the target detection model can be used to analyze the video segments in the video segment frame by frame and track the target objects in the video segment to obtain the interaction state information and relative motion information between different target objects. There are no limitations on this.

[0115] S307: Multiple object categories, multiple location information, interaction state information, and relative motion information are used together as object description information.

[0116] In this embodiment of the disclosure, after determining the multiple object categories and multiple location information corresponding to the multiple target objects, and determining the interaction state information and relative motion information between different target objects, the multiple object categories, multiple location information, interaction state information and relative motion information can be used together as object description information. The object description information can be used to assist in the identification and judgment of target action events.

[0117] S308: Input multiple object categories, multiple location information, interaction state information, and relative motion information into the pre-trained action event recognition model to obtain the recognition result of whether the target action event has occurred in the scene described by the instruction video, as output by the action event recognition model.

[0118] In this process, the initial second artificial intelligence model is trained in advance using sample object description information and the sample video to which the sample object description information belongs, until the second artificial intelligence model converges. The trained second artificial intelligence model is then used as the action event recognition model.

[0119] The artificial intelligence model used to train the action event recognition model can be called the second artificial intelligence model. The second artificial intelligence model can be a machine learning model or a neural network model. This second artificial intelligence model has the ability to recognize and process target action events.

[0120] In this embodiment of the disclosure, when determining whether a target action event has occurred in the scene described by the video based on the object description information, labeled training data in the scene described by the video can be obtained in advance. The labeled training data can be the pre-obtained sample object description information and the sample video to which the sample object description information belongs. Then, the sample object description information and the sample video to which the sample object description information belongs can be used to train the second artificial intelligence model, iteratively update the model parameters until the second artificial intelligence model converges, and use the trained second artificial intelligence model as the action event recognition model.

[0121] In this embodiment of the disclosure, after obtaining the trained action event recognition model, multiple object categories, multiple location information, interaction state information and relative motion information can be input into the pre-trained action event recognition model. The action event recognition model is used to process the multiple object categories, multiple location information, interaction state information and relative motion information to obtain the recognition result of the action event recognition model. The recognition result can indicate whether a target action event has occurred in the scene described by the video.

[0122] Optionally, in some embodiments, business scenario requirement information can be received, and based on the business scenario requirement information, annotation recognition results corresponding to the sample object description information can be generated. The annotation recognition results indicate whether the target action event has occurred in the scene described by the sample video. The annotation recognition results are used to determine the convergence time of the second artificial intelligence model, thereby generating annotation recognition results corresponding to the sample object description information in a targeted manner according to the business scenario requirement information. This can solve the problem that different business parties have different judgment standards and that some judgment standards are difficult to quantify. Determining the convergence time of the second artificial intelligence model based on the annotation recognition results can improve the accuracy of the action event recognition model in judging the target action event.

[0123] Among them, business scenarios refer to different scenarios that require judgment of target action events. For example, it can be different logistics business scenarios, security scenarios, or other scenarios that use video or continuous image frames to monitor whether defined events or actions have occurred. There are no restrictions on this.

[0124] Among them, business scenario requirement information can be the judgment rule information set for the occurrence of target action events in a business scenario.

[0125] For example, in a logistics sorting business scenario, the business scenario requirement information could be that an object is thrown 1.5 meters, which would indicate that a target action event has occurred in the business scenario. Alternatively, the business scenario requirement information can be configured adaptively for the needs of other business scenarios without any restrictions.

[0126] In this embodiment of the disclosure, when receiving business scenario requirement information, a data transmission interface can be configured on the video processing device to receive the business scenario requirement information. After receiving the business scenario requirement information, a labeling and recognition result corresponding to the sample object description information can be generated based on the business scenario requirement information.

[0127] In this embodiment of the disclosure, when generating the annotation and recognition results corresponding to the sample object description information based on the business scenario requirements, the business side reviewers can be introduced to interact with the machine learning algorithm. The business side can define and set the judgment criteria under the business scenario, and perform labeling processing on the sample object description information according to the defined judgment criteria to obtain the labeling results after labeling processing. The labeling results are used as tags, and combined with the analysis features of the sample video to which the extracted sample object description information belongs, the annotation and recognition results corresponding to the sample object description information are generated. The annotation and recognition results are used to determine the timing of the convergence of the second artificial intelligence model.

[0128] For example, such as Figure 4 As shown, Figure 4 This is a schematic diagram of the target action event determination process in an embodiment of this disclosure. When determining whether a target action event has occurred in the scene described by the video, the video can first be filtered, and then the processed video can be segmented to obtain multiple video segments. The multiple video segments are processed using a time positioning algorithm to obtain action event information corresponding to the multiple video segments. Based on the action event information, a target video segment is selected from the multiple video segments for fine-grained analysis to obtain object description information in the target video segment. The action event recognition and processing model is trained based on the business requirements information in the corresponding business scenario, and the target action event is determined in the corresponding business scenario to obtain the determination result.

[0129] In this embodiment, the video is decomposed to obtain multiple video segments. Multiple action event information corresponding to each video segment is determined. Based on the action event information, a target video segment is identified from the multiple video segments. Object description information is identified from the target video segment. Based on the object description information, it is determined whether a target action event has occurred in the scene described by the video. Since the video is first decomposed, and the target video segment is determined based on the decomposed video segments, the target action event in the scene is identified and judged around the target video segment. This achieves rapid identification and judgment of the target action event around the target video segment, effectively reducing the computational resources consumed in the identification and judgment of the target action event, and improving the identification and judgment effect of the target action event. A time-motion localization model with superior processing performance is used to process the segmented video segments to obtain action event recognition scores, etc., as corresponding action event information. This action event information can be used for video... The process of filtering video segments improves the efficiency of obtaining target video segments. Furthermore, subsequent target action event judgment on these segments saves computational resources and enhances the efficiency of target action event judgment. A pre-trained action event recognition model determines whether a target action event has occurred in the scene described in the video. Since this model can be trained adaptively using sample data from various application scenarios, it is applicable to a wide range of scenarios, effectively improving the accuracy of target action event judgment. Based on business scenario requirements, targeted annotation and recognition results corresponding to the sample object description information are generated, addressing the issues of differing judgment standards among different business parties and the difficulty in quantifying certain standards. The timing of the second artificial intelligence model's convergence is determined based on the annotation and recognition results, further improving the accuracy of the action event recognition model in judging target action events.

[0130] Figure 5 This is a schematic diagram of the structure of a video processing apparatus according to an embodiment of the present disclosure.

[0131] like Figure 5 As shown, the video processing apparatus 50 includes:

[0132] The decomposition module 501 is used to decompose the video into multiple video segments;

[0133] The first determining module 502 is used to determine multiple action event information corresponding to multiple video segments respectively, and to determine the target video segment from the multiple video segments based on the action event information.

[0134] The first recognition module 503 is used to identify object description information from the target video segment; and

[0135] The second determining module 504 is used to determine whether a target action event has occurred in the scene described by the video, based on the object description information.

[0136] In some embodiments of this disclosure, such as Figure 6 As shown, Figure 6 This is a schematic diagram of a video processing apparatus according to another embodiment of the present disclosure, wherein the number of target video segments is multiple, and it also includes:

[0137] The second identification module 505 is used to identify at least two video segments to be merged from multiple target video segments before identifying object description information from the target video segments;

[0138] The merging module 506 is used to merge at least two video segments to be merged to obtain a merged video segment;

[0139] The first identification module 503 is specifically used for:

[0140] Identify object description information from target video segments and merged video segments.

[0141] In some embodiments of this disclosure, the action event information includes: action event type and start and end times of the action event;

[0142] The second identification module 505 is specifically used for:

[0143] Determine the start and end time intervals between adjacent target video segments;

[0144] If the interval value is less than or equal to the interval threshold, and the action event types of adjacent target video segments are the same, then the adjacent target video segments are used as video segments to be merged.

[0145] In some embodiments of this disclosure, the decomposition module 501 is specifically used for:

[0146] Determine the total duration of the video;

[0147] The total duration is divided into multiple duration segments;

[0148] The video is segmented into multiple segments based on its duration, resulting in multiple video clips.

[0149] In some embodiments of this disclosure, the first determining module 502 is specifically used for:

[0150] Multiple video clips are input into a pre-trained temporal action localization model to obtain multiple action event information output by the temporal action localization model. The action event information also includes: recognition score value.

[0151] Select the video segments corresponding to the action event information whose recognition score is greater than or equal to the scoring threshold as the target video segments;

[0152] In this process, the initial first artificial intelligence model is trained in advance using sample video clips and sample action event information of the sample video clips until the first artificial intelligence model converges. The trained first artificial intelligence model is then used as the time action localization model.

[0153] In some embodiments of this disclosure, the time-motion localization model includes: a video feature extraction sub-model, a time regression sub-model, and a category classification sub-model connected to the video feature extraction sub-model.

[0154] The first determining module 502 is further used for:

[0155] Multiple video clips are input into a pre-trained temporal action localization model to obtain multiple action event information output by the model, including:

[0156] Multiple video segments are input into the video feature extraction sub-model to obtain multiple video features output by the video feature extraction sub-model;

[0157] A time regression sub-model is used to identify the start and end times of action events for multiple video features, resulting in multiple start and end times corresponding to each action event.

[0158] A category classification sub-model is used to identify the types of action events from multiple video features, thereby obtaining multiple action event types corresponding to each action event.

[0159] Based on the type of action event and the start and end time of the action event, the identification of action events from the corresponding video features is scored to obtain an identification score value;

[0160] Among them, the action event type, the start and end time of the action event, and the recognition score are collectively used as action event information.

[0161] In some embodiments of this disclosure, the first identification module 503 is specifically used for:

[0162] Identify multiple target objects from a target video clip;

[0163] Determine the multiple object categories and multiple location information corresponding to multiple target objects;

[0164] Determine the interaction state information and relative motion information between different target objects;

[0165] Multiple object categories, multiple location information, interaction state information, and relative motion information are used together as object description information.

[0166] In some embodiments of this disclosure, the second determining module 504 is specifically used for:

[0167] Multiple object categories, multiple location information, interaction state information, and relative motion information are input into a pre-trained action event recognition model to obtain the recognition result of whether a target action event has occurred in the scene described by the video, as indicated by the action event recognition model output.

[0168] In this process, the initial second artificial intelligence model is trained in advance using sample object description information and the sample video to which the sample object description information belongs, until the second artificial intelligence model converges. The trained second artificial intelligence model is then used as the action event recognition model.

[0169] In some embodiments of this disclosure, it also includes:

[0170] The receiving module 507 is used to receive business scenario requirement information;

[0171] The generation module 508 is used to generate annotation recognition results corresponding to the sample object description information based on the business scenario requirements information. The annotation recognition results indicate whether the target action event has occurred in the scene described by the sample video.

[0172] The annotation and recognition results were used to determine when the second artificial intelligence model converged.

[0173] With the above Figures 1 to 4 Corresponding to the video processing method provided in the embodiments, this disclosure also provides a video processing apparatus. Since the video processing apparatus provided in the embodiments of this disclosure is similar to the one described above... Figures 1 to 4 The video processing method provided in the embodiments corresponds to the video processing apparatus provided in the embodiments of this disclosure, and will not be described in detail in the embodiments of this disclosure.

[0174] In this embodiment, the video is decomposed to obtain multiple video segments, and multiple action event information corresponding to each video segment is determined. Based on the action event information, a target video segment is determined from the multiple video segments, object description information is identified from the target video segment, and based on the object description information, it is determined whether a target action event has occurred in the scene described by the video. Since the video is decomposed first, and the target video segment is determined based on the multiple video segments obtained from the decomposition, the target action event in the scene is identified and judged around the target video segment. This achieves rapid identification and judgment of the target action event around the target video segment, effectively reducing the computing resources consumed in the identification and judgment of the target action event, and improving the identification and judgment effect of the target action event.

[0175] To implement the above embodiments, this disclosure also proposes a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the video processing method proposed in the foregoing embodiments of this disclosure.

[0176] To implement the above embodiments, this disclosure also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the video processing method proposed in the foregoing embodiments of this disclosure.

[0177] To implement the above embodiments, this disclosure also proposes a computer program product that, when executed by an instruction processor, performs the video processing method as described in the foregoing embodiments of this disclosure.

[0178] Figure 7 A block diagram of an exemplary computer device suitable for implementing embodiments of the present disclosure is shown. Figure 7 The computer device 12 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.

[0179] like Figure 7 As shown, the computer device 12 is represented in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, system memory 28, and a bus 18 connecting different system components (including system memory 28 and processing unit 16).

[0180] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.

[0181] Computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by computer device 12, including volatile and non-volatile media, removable and non-removable media.

[0182] Memory 28 may include computer system readable media in the form of volatile memory, such as Random Access Memory (RAM) 30 and / or cache memory 32. Computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (…). Figure 7 Not shown; usually referred to as a "hard drive".

[0183] although Figure 7 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disc drive for reading and writing to a removable non-volatile optical disc (e.g., a compact disc read-only memory (CD-ROM), a digital video disc read-only memory (DVD-ROM), or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of this disclosure.

[0184] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of this disclosure.

[0185] Computer device 12 can also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), and with one or more devices that enable a user to interact with computer device 12, and / or with any device that enables computer device 12 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 22. Furthermore, computer device 12 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 20. As shown, network adapter 20 communicates with other modules of computer device 12 via bus 18. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with computer device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0186] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the video processing method mentioned in the foregoing embodiments.

[0187] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0188] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

[0189] It should be noted that in the description of this disclosure, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this disclosure, unless otherwise stated, "a plurality of" means two or more.

[0190] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of preferred embodiments of this disclosure includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the function involved, as will be understood by those skilled in the art to which embodiments of this disclosure pertain.

[0191] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0192] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0193] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0194] The storage media mentioned above can be read-only memory, disk, or optical disk, etc.

[0195] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0196] Although embodiments of the present disclosure have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present disclosure.

Claims

1. A video processing method, characterized in that, include: The video is decomposed into multiple video segments; Determine multiple action event information corresponding to the multiple video segments respectively, and determine the target video segment from the multiple video segments based on the action event information. The action events include events in which objects in the video segments perform interactive movements and generate motion trajectories. At least two video segments to be merged are identified from multiple target video segments. The multiple target video segments are identified from the start time to the end time to identify whether the action event in the target video segment spans multiple video segments at the end time. If the action event spans two or more video segments, then the multiple video segments that completely contain the action event are taken as the video segments to be merged. The at least two video segments to be merged are merged to obtain a merged video segment; Identifying object description information from the target video segment includes: performing pixel-by-pixel analysis on the target video segment and the merged video segment, labeling the target object in each frame image for target object tracking, extracting feature values ​​from each frame image separately, then fusing multiple features to obtain depth feature values, and determining corresponding object description information based on the depth feature values ​​of the target video segment and the merged video segment. And based on the object description information, determine whether a target action event has occurred in the scene described by the video.

2. The method as described in claim 1, characterized in that, The action event information includes: action event type and start and end times of the action event; The step of identifying at least two video segments to be merged from the plurality of target video segments includes: Determine the interval value of the start and end times between adjacent target video segments; If the interval value is less than or equal to the interval threshold, and the action event types of the adjacent target video segments are the same, then the adjacent target video segments are taken as the video segments to be merged.

3. The method as described in claim 1, characterized in that, The video is decomposed to obtain multiple video segments, including: Determine the total duration of the video; The total duration is divided into multiple duration segments; The video is segmented according to the multiple duration segments to obtain the multiple video segments.

4. The method as described in claim 2, characterized in that, The step of determining multiple action event information corresponding to the multiple video segments respectively, and determining the target video segment from the multiple video segments based on the action event information, includes: The multiple video segments are respectively input into a pre-trained temporal motion localization model to obtain multiple motion event information output by the temporal motion localization model. The motion event information also includes: recognition score value. Select the video segment corresponding to the action event information to which the identification score value is greater than or equal to the scoring threshold as the target video segment; Specifically, a first artificial intelligence model is trained in advance using sample video clips and sample action event information of the sample video clips until the first artificial intelligence model converges, and the trained first artificial intelligence model is used as the time action localization model.

5. The method as described in claim 4, characterized in that, The time-based motion localization model includes: a video feature extraction sub-model, a time regression sub-model, and a category classification sub-model connected to the video feature extraction sub-model. The step of inputting the multiple video segments into a pre-trained temporal motion localization model to obtain multiple action event information output by the temporal motion localization model includes: The multiple video segments are respectively input into the video feature extraction sub-model to obtain multiple video features output by the video feature extraction sub-model; The time regression sub-model is used to identify the start and end times of action events for the multiple video features, thereby obtaining multiple start and end times corresponding to the multiple action events. The category classification sub-model is used to identify the types of action events in the multiple video features respectively, thereby obtaining multiple action event types corresponding to the multiple action events; Based on the action event type and the start and end times of the action event, a score is given for the identification of the action event from the corresponding video features to obtain an identification score value; The action event type, the start and end times of the action event, and the recognition score are collectively used as the action event information.

6. The method as described in claim 1, characterized in that, The step of identifying object description information from the target video segment includes: Multiple target objects were identified from the target video segment; Determine the multiple object categories and multiple location information corresponding to the multiple target objects; Determine the interaction state information and relative motion information between different target objects; The multiple object categories, the multiple location information, the interaction state information, and the relative motion information are collectively used as the object description information.

7. The method as described in claim 6, characterized in that, The step of determining whether a target action event has occurred in the scene described by the video based on the object description information includes: The multiple object categories, the multiple location information, the interaction state information, and the relative motion information are respectively input into a pre-trained action event recognition model to obtain the recognition result output by the action event recognition model, which indicates whether the target action event has occurred in the scene described by the video. Specifically, a second artificial intelligence model is initially trained using sample object description information and the sample video to which the sample object description information belongs, until the second artificial intelligence model converges, and the trained second artificial intelligence model is used as the action event recognition model.

8. The method as described in claim 7, characterized in that, The method further includes: Receive business scenario requirement information; Based on the business scenario requirements information, an annotation and recognition result corresponding to the sample object description information is generated. The annotation and recognition result indicates whether the target action event has occurred in the scene described by the sample video. The annotation and recognition results are used to determine when the second artificial intelligence model converges.

9. A video processing apparatus, characterized in that, include: The decomposition module is used to decompose the video into multiple video segments; The first determining module is used to determine multiple action event information corresponding to the multiple video segments respectively, and to determine the target video segment from the multiple video segments based on the action event information. The action event includes an event in which an object in the video segment performs interactive movement and generates a motion trajectory. The first identification module is used to identify object description information from the target video segment; as well as The second determining module is used to determine whether a target action event has occurred in the scene described by the video, based on the object description information. The device is also used to perform filtering preprocessing on the video, filtering out still images and images with poor image quality in the video; The decomposition module is specifically used to decompose the preprocessed video to obtain multiple video segments; The number of target video segments is multiple, including: The second identification module is used to identify at least two video segments to be merged from multiple target video segments before identifying object description information from the target video segments. The multiple target video segments are identified from the start time to the end time to identify whether there is an action event in the target video segment that spans multiple video segments at the end time. If the action event spans two or more video segments, then the multiple video segments that completely contain the action event are taken as video segments to be merged. The merging module is used to merge the at least two video segments to be merged to obtain a merged video segment. The first identification module is specifically used for: The target video segment and the merged video segment are analyzed pixel by pixel, and the target object in each frame is labeled for target object tracking. The feature values ​​in each frame are extracted separately, and then multiple features are fused to obtain depth feature values. The corresponding object description information is determined based on the depth feature values ​​of the target video segment and the merged video segment.

10. The apparatus as claimed in claim 9, characterized in that, The action event information includes: action event type and start and end times of the action event; The second identification module is specifically used for: Determine the interval value of the start and end times between adjacent target video segments; If the interval value is less than or equal to the interval threshold, and the action event types of the adjacent target video segments are the same, then the adjacent target video segments are taken as the video segments to be merged.

11. The apparatus as claimed in claim 9, characterized in that, The decomposition module is specifically used for: Determine the total duration of the video; The total duration is divided into multiple duration segments; The video is segmented according to the multiple duration segments to obtain the multiple video segments.

12. The apparatus as claimed in claim 10, characterized in that, The first determining module is specifically used for: The multiple video segments are respectively input into a pre-trained temporal motion localization model to obtain multiple motion event information output by the temporal motion localization model. The motion event information also includes: recognition score value. Select the video segment corresponding to the action event information to which the identification score value is greater than or equal to the scoring threshold as the target video segment; Specifically, a first artificial intelligence model is trained in advance using sample video clips and sample action event information of the sample video clips until the first artificial intelligence model converges, and the trained first artificial intelligence model is used as the time action localization model.

13. The apparatus as claimed in claim 12, characterized in that, The time-based motion localization model includes: a video feature extraction sub-model, a time regression sub-model, and a category classification sub-model connected to the video feature extraction sub-model. The first determining module is further configured to: The multiple video segments are respectively input into a pre-trained temporal action localization model to obtain multiple action event information output by the temporal action localization model, including: The multiple video segments are respectively input into the video feature extraction sub-model to obtain multiple video features output by the video feature extraction sub-model; The time regression sub-model is used to identify the start and end times of action events for the multiple video features, thereby obtaining multiple start and end times corresponding to the multiple action events. The category classification sub-model is used to identify the types of action events in the multiple video features respectively, thereby obtaining multiple action event types corresponding to the multiple action events; Based on the action event type and the start and end times of the action event, a score is given for the identification of the action event from the corresponding video features to obtain an identification score value; The action event type, the start and end times of the action event, and the recognition score are collectively used as the action event information.

14. The apparatus as claimed in claim 9, characterized in that, The first identification module is specifically used for: Multiple target objects were identified from the target video segment; Determine the multiple object categories and multiple location information corresponding to the multiple target objects; Determine the interaction state information and relative motion information between different target objects; The multiple object categories, the multiple location information, the interaction state information, and the relative motion information are collectively used as the object description information.

15. The apparatus as claimed in claim 14, characterized in that, The second determining module is specifically used for: The multiple object categories, the multiple location information, the interaction state information, and the relative motion information are respectively input into a pre-trained action event recognition model to obtain the recognition result output by the action event recognition model, which indicates whether the target action event has occurred in the scene described by the video. Specifically, a second artificial intelligence model is initially trained using sample object description information and the sample video to which the sample object description information belongs, until the second artificial intelligence model converges, and the trained second artificial intelligence model is used as the action event recognition model.

16. The apparatus of claim 15, further comprising: The receiving module is used to receive business scenario requirement information; The generation module is used to generate a labeling and recognition result corresponding to the description information of the sample object based on the business scenario requirements information. The labeling and recognition result indicates whether the target action event has occurred in the scene described by the sample video. The annotation and recognition results are used to determine when the second artificial intelligence model converges.

17. A computer device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.