Video content information processing method and device and computer program product

By identifying and extracting the behavioral characteristics of objects in video segments within a security system, and generating statistical descriptive information, the problem of event overload in traditional security systems is solved, achieving efficient security situation awareness and decision support.

CN121888045APending Publication Date: 2026-04-17TP-LINK INT SHENZHEN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TP-LINK INT SHENZHEN CO LTD
Filing Date
2025-12-31
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Traditional security systems suffer from event overload during event detection, requiring users to review video segments one by one to obtain information, which is time-consuming and makes it difficult to quickly identify truly threatening events, thus posing security risks.

Method used

By identifying object behaviors in multiple video segments, extracting behavioral feature information, and generating statistical description information based on preset rules, meaningless events are excluded, related events are sorted and merged, and structured statistical descriptions are provided, reducing the number of video segments that users need to view.

Benefits of technology

It improves the efficiency of real-time security situation awareness and decision-making, reduces the burden of user review, and enhances information processing efficiency and the readability of critical events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121888045A_ABST
    Figure CN121888045A_ABST
Patent Text Reader

Abstract

The invention relates to an information processing method and device for video content and a computer program product. The method comprises the following steps: identifying object behaviors in a plurality of corresponding events from a plurality of video segments collected within a predetermined time period; aiming at an object behavior in each event in the plurality of corresponding events, behavior characteristic information of the object behavior is extracted, and the behavior characteristic information comprises at least one of a behavior type, behavior occurrence time, a behavior occurrence position, an action characteristic, an object identification characteristic and object sound information; and based on the behavior feature information of each object behavior in the plurality of corresponding events, according to a preset statistical rule, generating statistical description information of the plurality of corresponding events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of surveillance and security, and more specifically, to methods, apparatus and computer program products for processing video content. Background Technology

[0002] Security systems are core technologies for ensuring public safety, industrial safety, and home safety. They are widely used in various scenarios such as urban security monitoring, industrial park management, and residential community protection. Their core value lies in capturing abnormal events in real time by collecting on-site video data, providing users with decision-making support, thereby reducing security risks and improving management efficiency. With the popularization of video surveillance technology, the event detection capabilities of security systems continue to improve. Security cameras can collect large amounts of video data around the clock and trigger massive event alarms based on mechanisms such as motion detection.

[0003] Therefore, it is necessary to improve the processing of large amounts of video data in security systems to enhance the intelligence level of security systems and user experience. Summary of the Invention

[0004] In view of the above problems, this disclosure provides a method, apparatus and computer program product for processing video content information, which summarizes and describes multiple events within a time period, enabling users to grasp the overall security situation without having to view massive amounts of video, thereby improving users' efficient cognition and response efficiency regarding the security situation.

[0005] One aspect of this disclosure provides a method for processing video content information, comprising: identifying object behaviors in multiple corresponding events from multiple video segments collected within a predetermined time period; extracting behavioral feature information for each object behavior in the multiple corresponding events, wherein the behavioral feature information includes at least one of behavior type, behavior occurrence time, behavior occurrence location, action features, object identification features, and object sound information; and generating statistical description information of the multiple corresponding events based on the behavioral feature information of each object behavior in the multiple corresponding events, according to preset statistical rules.

[0006] Another aspect of this disclosure provides an information processing apparatus for video content, comprising: a processor; a memory coupled to the processor; and computer program instructions stored in the memory, the computer program instructions executing the aforementioned information processing method for video content when executed by the processor.

[0007] Another aspect of this disclosure provides a computer program product, including computer program instructions that, when executed by a processor of the video content information processing device, perform the aforementioned video content information processing method. Attached Figure Description

[0008] The aspects, features, and advantages of this disclosure will become clearer and more readily understood from the following description of embodiments in conjunction with the accompanying drawings. The drawings are provided to offer a further understanding of the embodiments of this disclosure and form part of the specification. The drawings, together with the embodiments of this disclosure, are used to explain this disclosure but do not constitute a limitation thereof. In the drawings:

[0009] Figure 1 Scene diagrams of security systems according to various embodiments of the present disclosure are shown.

[0010] Figure 2 A flowchart illustrating a method for processing video content according to various embodiments of the present disclosure is shown.

[0011] Figure 3 Example forms of statistical description information (sorted from highest to lowest importance) according to various embodiments of this disclosure are shown.

[0012] Figure 4 Example forms of merged statistical description information according to various embodiments of this disclosure are shown.

[0013] Figure 5 An example block diagram of an information processing apparatus for video content according to various embodiments of the present disclosure is shown. Detailed Implementation

[0014] The technical solutions of this disclosure will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the protection scope of this disclosure.

[0015] Furthermore, the technical features involved in the different embodiments of this disclosure described below can be combined with each other as long as they do not conflict with each other.

[0016] The terms “exemplary” and / or “example” are used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” and / or “example” is not necessarily to be construed as superior to or better than other aspects. Similarly, the term “aspects of this disclosure” does not require that all aspects of this disclosure include the features, advantages, or modes of operation discussed.

[0017] Traditional security systems commonly suffer from event overload. Typically, these systems monitor areas using methods like motion detection and area intrusion detection. Once a change matching preset rules appears in the surveillance video, the system automatically triggers an event and records a video segment. For example, a security camera might trigger over 100 events daily, many of which are caused by the behavior of various objects (such as humans, vehicles, and animals) (e.g., falls, walking). Many more events may be false alarms caused by environmental interference (such as wind blowing leaves, small animal activity, changes in lighting). Users must review the complete video segment corresponding to each event (usually 10-30 seconds long) to determine its actual security significance.

[0018] Therefore, traditional security systems suffer from high time costs. A complete video segment of a single event is approximately 10 to 30 seconds long, and users need to watch each segment to obtain complete information. For example, 100 events would take 20 to 50 minutes, far exceeding the acceptable range for actual management efficiency.

[0019] Based on this, this disclosure provides a method for generating statistical description information for multiple events. Users only need to view the statistical description information extracted and statistically analyzed from multiple video segments, without having to watch a massive number of video segments one by one, which greatly improves information processing efficiency and thus enhances the real-time perception and decision-making efficiency of the security situation.

[0020] Figure 1 A scene diagram of a security system 100 according to various embodiments of the present disclosure is shown.

[0021] The security system 100 may include a video acquisition device 101, a video content information processing device 102 (hereinafter referred to as the information processing device), and an output device 103.

[0022] The video acquisition device 101 can acquire multiple video segments, such as real-time video segments or input video segments, adapting to different security monitoring scenarios such as homes, supermarkets, and public safety. The video acquisition device 101 may include, for example, cameras. Cameras may include visible light cameras, infrared cameras, thermal imaging cameras, 3D cameras, ultraviolet cameras, etc. There are many types of cameras, and selection can be made based on security scenario requirements (e.g., lighting conditions, accuracy requirements, concealment), technical characteristics (e.g., resolution, sensor type), and system integration capabilities (e.g., artificial intelligence (AI) algorithms, network transmission).

[0023] The information processing device 102 can process multiple video segments, for example, performing a video content information processing method, including, for example, identifying object behaviors in multiple corresponding events from multiple video segments collected within a predetermined time period; extracting behavioral feature information for each object behavior in the multiple corresponding events, wherein the behavioral feature information includes at least one of behavior type, behavior occurrence time, behavior occurrence location, action characteristics, object identification characteristics, and object sound information; and generating statistical description information for the multiple corresponding events according to preset statistical rules based on the behavioral feature information of each object behavior in the multiple corresponding events, etc.

[0024] The term "event" used here refers to an event that can be detected by a security system (e.g., through motion detection, area intrusion detection, etc.) and triggers the security system to generate a video clip. Examples include detecting target movement in surveillance video via motion detection (e.g., leaves rustling in the wind, a person walking by), detecting a target entering a restricted area via area intrusion detection, detecting a target crossing a virtual warning line via boundary crossing detection (e.g., moving from area A to area B), and detecting a target lingering in an area for more than a set time via loitering detection, and so on.

[0025] The output device 103 can receive statistical description information output by the information processing device 102 and output the statistical description information to the user. The output device 103 may include a display device (e.g., a liquid crystal display (LCD) / organic light-emitting diode (OLED) display, etc.). The output device 103 may also include a speaker for broadcasting the statistical description information to the user.

[0026] The security system 100 may also include input devices (not shown), such as touch screens, keyboards / mouse, microphones, etc., for receiving user input (e.g., user-defined options).

[0027] In one embodiment, the video acquisition device 101 can function as a camera for acquiring video, while the information processing device 102 and the output device 103 are deployed on the management side to perform management functions. In another embodiment, the video acquisition device 101 and the information processing device 102 can be intelligent cameras with information processing capabilities (e.g., built-in computing power) responsible for video acquisition and analysis, while the output device 103 serves as a remote display terminal for result push and presentation. Alternatively, the video acquisition device 101, information processing device 102, and output device 103 can be integrated into a single device (e.g., a personal computer) to achieve functions such as video acquisition and analysis, result push and presentation. Those skilled in the art should understand that this disclosure does not limit the hardware form of the security system 100.

[0028] Figure 2A flowchart of a video content information processing method 200 according to various embodiments of the present disclosure is shown.

[0029] In step 210, object behaviors in multiple corresponding events can be identified from multiple video segments collected within a predetermined time period (e.g., within a day, a week, or a custom time period). For example, N video segments are collected within the predetermined time period, generated due to various reasons that trigger the security system. Corresponding events may include, for example, an elderly person falling, flames / smoke, fights, the appearance of a blacklisted person, vehicles entering or leaving a garage, swaying leaves, changes in lighting, pedestrians passing by, etc. For example, consider the following eight events: Event 1 (Video Segment 1): 08:10, an elderly person falls outside the restroom; Event 2 (Video Segment 2): 09:15, a man in a black hat paces back and forth at the office building entrance, lingering for more than 8 minutes and repeatedly looking around; Event 3 (Video Segment 3): 10:05, leaves on trees in Area A sway in the wind; Event 4 (Video Segment 4): 11:30, a black backpack is left in the activity room for more than 15 minutes; Event 5 (Video Segment 5): 14:17, a man in a black hat paces back and forth in the activity room, repeatedly looking around; Event 6 (Video Segment 6): 18:10, a cat walks past the restaurant entrance; Event 7 (Video Segment 7): 21:01, a man in red climbs over the fence into a restricted area; Event 8 (Video Segment 8): 23:26, a baby cries loudly in a room for more than 1 minute. The objects can refer to humans, vehicles, animals, plants, etc. The object behavior can refer to the actions performed by the objects in these events, for example, the elderly person falling.

[0030] In step 220, for the object behavior in each of the multiple corresponding events, behavioral feature information can be extracted. For example, behavioral feature information can be extracted using visual processing techniques (e.g., machine learning models). Behavioral feature information may include at least one of the following: behavior type, behavior occurrence time, behavior occurrence location, action characteristics, object identification characteristics, and object sound information. Behavior type can also be called event type, which may include, for example, an elderly person falling, area intrusion, abnormal wandering, normal passage, leaving items behind, or an infant crying. For example, behavioral feature information can be extracted by analyzing one or more of the following information from the video segment: visual information, audio information, camera location, clock, etc.

[0031] The behavior type can be determined based on preset behavior type rules. These preset rules can be, for example, a list of behavior types, including various preset behavior types. For instance, the behavior type can be determined by analyzing the object's behavior (e.g., behavior feature extraction and behavior classification). Preset behavior type rules can be user-defined and adjusted as needed. The behavior occurrence time refers to the specific time the object's behavior occurs. The behavior occurrence location refers to the location where the object's behavior occurs, which can be determined by locating the camera. Action characteristics refer to the action process used to describe the object's behavior; this can be dynamic (e.g., falling) or static (standing still). Object identification characteristics refer to the characteristic attributes used to identify the object, such as elderly or child, male or female, or someone wearing red. Object sound information refers to the sounds emitted by the object, such as shouting or crying, which can be obtained from collected audio information. The behavior characteristic information for the above eight events is shown in Table 1 below.

[0032] Table 1. Behavioral characteristics of the above 8 events

[0033]

[0034] Furthermore, the inventors have noted that traditional security systems lack effective event filtering and prioritization mechanisms. Truly threatening events (such as suspicious individuals loitering at night or repeatedly appearing in sensitive areas) are often mixed in with a large number of low-value events, making it difficult for users to identify and respond to core risks in a state of information overload, posing significant security risks. Based on this, this disclosure proposes excluding meaningless events (including environmental interference events and events of non-user concern, as described below) and prioritizing the statistical descriptions of events, as described below.

[0035] According to embodiments of this disclosure, the importance of each event can be determined based on behavioral characteristic information according to preset importance rules. Importance can be used to characterize the level of behavioral safety risk (which may also be called urgency) and / or the user's attention priority. The higher the level of behavioral safety risk, the higher the importance. Similarly, the higher the user's attention priority, the higher the importance. The preset importance rules can be rules for determining importance levels. For example, importance levels can include high importance, medium importance, and low importance, etc., and are not limited to this; more refined or coarser levels can also be used. The preset importance rules can include importance criteria for one or more of the following: behavior type, behavior occurrence time, behavior occurrence time, or object type. For example, events such as an elderly person falling, a baby crying, and area intrusion are of high importance; events such as abnormal wandering and leaving items behind are of medium importance; and events such as plant movement and animal movement are of low importance. For another example, if a user attaches great importance to area intrusion (its attention priority is the highest), then the area intrusion event is determined to be of high importance. Furthermore, the importance can also be determined by comprehensively considering the level of behavioral safety risk and the user's attention priority. For example, if both the behavioral security risk level and user attention priority of an area intrusion event are high, then the event is determined to be of high importance. Preset importance rules can be customized by the user and adjusted as needed.

[0036] In addition, for the same type of behavior, importance can be determined by other factors (e.g., the time when the behavior occurs). For example, the importance of abnormal loitering events between 1 p.m. and 6 p.m. is medium, while the importance of abnormal loitering events between 12 a.m. and 3 a.m. is high.

[0037] According to embodiments of this disclosure, environmental interference events (which can be considered as environmental noise) (e.g., event 3: plant movement) can be excluded from multiple corresponding events (e.g., the eight events mentioned above). This is because environmental interference events caused by environmental factors (such as light, wind, rain, etc.) are generally not very important, and users do not need to pay much attention to these events. Environmental interference events can be identified based on visual and / or audio information from video segments. For example, changes in light can be identified from the visual information of a video segment, or, for example, swaying leaves can be identified from the visual information (swaying leaves) and / or audio information (wind sound, rustling leaves) of a video segment. Environmental factors are the most significant source of false alarms in traditional security systems. By excluding environmental interference events, this disclosure can increase the proportion of valid events, improve the credibility of the security system, make users more willing to trust and continue using the security system, and significantly reduce the user's review burden.

[0038] According to embodiments of this disclosure, events of non-user concern (which can be considered as non-environmental noise) (e.g., event 3: plant movement, event 6: animal movement) can be excluded from multiple corresponding events (e.g., the eight events mentioned above). These are events that the user is not concerned with. Non-user concern events can be identified based on preset user concern behavior rules. For example, preset user concern behavior rules can indicate which behaviors the user cares about and which behaviors the user does not care about. For example, preset user concern behavior rules can include user concern level standards. Preset user concern behavior rules can be customized by the user and can be adjusted by the user as needed. Therefore, by excluding non-user concern events, this disclosure can focus on the user's real needs, improve event relevance, and significantly reduce the user's burden.

[0039] In step 230, based on the behavioral characteristic information of each object's behavior in multiple corresponding events, statistical description information for multiple corresponding events can be generated according to preset statistical rules; that is, a summary description of multiple events (e.g., an event overview) can be provided. The statistical description information can be structured information, such as text or tables. The structured information can be field-based information, such as JSON format (e.g.,...). Figure 3 As shown in the image, and other formats of information, such as simplified information like "Object XX did what at time YY location ZZ" (i.e., a summary). The structure of the statistical description information can be user-defined. Preset statistical rules can be user-defined and adjusted as needed.

[0040] For example, preset statistical rules may include preset importance ranking rules. Statistical descriptions of multiple corresponding events can be ranked based on importance (e.g., from highest to lowest importance). Figure 3 As shown, statistical descriptive information (e.g., field-based information) is sorted from highest to lowest importance, with environmental disturbance events (Event 3: Plant Movement) and events of no concern to the user (Event 6: Animal Movement) excluded. The preset importance sorting rules can be customized by the user and adjusted as needed. Therefore, sorting by importance ensures that the most important (e.g., most urgent, most dangerous, most concerning) events are presented first, allowing users to see them immediately and preventing critical information from being buried by a large number of low-value events.

[0041] For example, preset statistical rules may include preset time sorting rules. Statistical descriptions of multiple corresponding events can be sorted based on the chronological order of the events (e.g., from first to last). Preset time sorting rules can be user-defined and adjusted as needed. Therefore, sorting by time can reconstruct the temporal logic of behaviors and support causal and intent inference.

[0042] Preset statistical rules can combine importance and chronological order. For example, when importance is the same, the actions can be sorted by the order in which they occurred, or when the actions occurred at the same time, they can be sorted by importance. In addition to time and importance, other factors can be considered for sorting, such as sorting based on the location where the actions occurred, but this is not a limitation.

[0043] In addition, the inventors also noted that traditional security systems can only provide isolated video clips of events and cannot automatically identify and present the correlation between multiple events. For example, the behavioral pattern of the same person appearing in different sensitive areas multiple times in a short period of time cannot be effectively identified and alerted.

[0044] According to embodiments of this disclosure, based on preset event association rules, it can be determined whether two or more events among multiple corresponding events are related, and related events can be merged to generate corresponding merged statistical description information, that is, two or more related events are merged and described. For example, the preset event association rules can be based on at least one of the following dimensions: proximity of the time of occurrence of the behavior, proximity of the location of occurrence of the behavior, whether they belong to the same object, whether they belong to the same behavior type, and similarity of action characteristics. The preset event association rules can be customized by the user and can be adjusted by the user as needed. For example, events 2 and 5 (which in this example also belong to abnormal loitering) belonging to the same object (e.g., the man in the black hat) can be merged and described as "the man in the black hat loitered at the entrance of the office building and the activity room at 09:15 and 14:17", or field-based statistical description information such as JSON format (e.g., Figure 4 (As shown). After seeing the combined description of events 4 and 5, users can suspect that the man in the black hat is not a normal visitor or employee, and may intend to commit theft or robbery, and can then focus more attention on him. For example, combining events 4 and 5 (both in the activity room) in a similar location as "a black backpack was left in the activity room at 11:30, and a man in a black hat was loitering and looking around the black backpack in the activity room." After seeing the combined description of events 4 and 5, users can reasonably suspect that the man in the black hat intended to steal the black backpack, and is a suspected thief, and can then focus more attention on him. Therefore, this disclosure, by merging multiple related events into a single combined description, not only significantly reduces the number of events that need to be stored, effectively saving storage resources, but also transforms fragmented events into a semantically coherent and focused combined description, greatly improving the readability of event information and user cognitive efficiency.

[0045] According to embodiments of this disclosure, the statistical description information may further include total statistics on the number of corresponding events and / or classification statistics on the number of corresponding events. Classification statistics may be the result of statistics based on at least one dimension, such as behavior type, time of occurrence of the behavior, location of occurrence of the behavior, and object category. For example, for the aforementioned 8 events, the statistical description information may include "Total number of events: 8," and / or the statistical description information may include "Elderly person falls: 1; Baby cries: 1; Area intrusion: 1; Abnormal wandering: 2; Item left behind: 1," with classification statistics based on behavior type being used as an example. Therefore, by including total event counts and / or classification statistics in the statistical description information, this disclosure enables users to quickly grasp the overall event scale and type distribution within a monitoring period.

[0046] According to embodiments of this disclosure, the statistical description information may further include detailed description information of key events. A key event can be defined as an event whose importance among multiple events meets a preset threshold, such as a high-importance event, such as "area intrusion." The detailed description information can be generated by supplementing the behavioral characteristic information of the key behavior based on video segments corresponding to the key event. The detailed description information may include elements such as time, location, object characteristics, behavioral process, duration, and importance. For example, for the key event "area intrusion," its detailed description information could be: "An area intrusion event occurred at 21:01. A man in a red shirt without a work badge appeared outside the warehouse restricted area fence at 21:01, then climbed onto the top of the fence with his hands and entered the warehouse. The entire action lasted 20 seconds. He did not carry any obvious tools. High importance." As can be seen, the detailed description information includes more details than the behavioral characteristic information, helping users quickly understand the nature of the event in order to make security decisions.

[0047] For key events, the statistical description information can also include video access interfaces (e.g., hyperlinks, or any other means of jumping to video segments) for accessing the video segments corresponding to the key events from multiple video segments. For example, clicking the hyperlink for the key event "Regional Intrusion" will jump to the video segment containing "Regional Intrusion." This supports quickly retrieving the video segments corresponding to the events, eliminating the need for manual searching through massive amounts of video data and significantly improving search efficiency.

[0048] In addition to key events, the statistical description information can also include video access interfaces (e.g., hyperlinks, or any other means of jumping to video segments) for other events. This allows users to access the video segment corresponding to the specific event from among multiple video segments via these video access interfaces (e.g., hyperlinks, or any other means of jumping to video segments). For example, clicking the hyperlink for the key event "Abnormal Loitering" will jump to the video segment containing "Abnormal Loitering".

[0049] According to embodiments of this disclosure, when generating statistical description information, the information can be generated at a preset statistical granularity. The preset statistical granularity may include at least one of time granularity, behavior type granularity, location granularity, and object type granularity. Time granularity can be used to specify one or more statistical time periods, such as 12:00 AM to 6:00 AM. This statistical time period can be accurate to the hour, minute, or second. Therefore, the statistical description information may only include information about events occurring between 12:00 AM and 6:00 AM, and optionally include total quantity statistics and / or category statistics, but exclude information about other events. Behavior type granularity can be used to specify one or more types of events to be statistically analyzed, such as "area intrusion" and "abnormal loitering" events. Location granularity can be used to specify the location where one or more behaviors of the events to be statistically analyzed occur, such as "near a restricted area." Object type granularity can be used to specify one or more object types, such as "elderly" and "child." The preset statistical granularity can be user-defined; for example, a user can select the preset statistical granularity through an input device (such as a touchscreen, keyboard / mouse, microphone, etc.). This disclosure supports the generation of statistical description information with preset statistical granularity, enabling users to obtain multi-dimensional statistical data that is highly relevant to their business scenarios as needed. This not only avoids interference from irrelevant information and improves the targeting of information processing and decision-making efficiency, but also enhances the flexibility and adaptability of the system, meeting the personalized monitoring and management needs of different business scenarios.

[0050] Figure 5 An example block diagram of an information processing apparatus 500 for video content according to various embodiments of the present disclosure is shown.

[0051] like Figure 5 As shown, the information processing apparatus 500 may include a processor 510 and a memory 520. The processor 510 is communicatively coupled to the memory 520 and is configured to perform the methods described above.

[0052] A set of computer program instructions stored in memory, when executed by a processor, performs any step of the above method, including: identifying object behaviors in multiple corresponding events from multiple video segments acquired within a predetermined time period; extracting behavioral feature information for the object behavior in each of the multiple corresponding events, wherein the behavioral feature information includes at least one of behavior type, behavior occurrence time, behavior occurrence location, action characteristics, object identification characteristics, and object sound information; and generating statistical description information of the multiple corresponding events based on the behavioral feature information of each object behavior in the multiple corresponding events, according to preset statistical rules. The above relates to... Figure 2 The details described in the method shown also apply here.

[0053] Examples of processor 510 include microprocessors, microcontrollers, DSPs, FPGAs, PLDs, state machines, gated logic, discrete hardware circuits, and other suitable hardware configured to perform various functionalities throughout the present disclosure.

[0054] Processor 510 can execute software. Whether referred to as software, firmware, middleware, microcode, hardware description language, or other terms, software should be broadly interpreted to mean instructions, instruction sets, code, code segments, program code, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, etc. Software can reside on memory 520.

[0055] Memory 520 may be a non-transitory computer-readable medium. As examples, non-transitory computer-readable media include magnetic storage devices (e.g., hard disks, floppy disks, magnetic stripes), optical disks (e.g., compact discs (CDs) or digital versatile discs (DVDs)), smart cards, flash memory devices (e.g., cards, sticks, or key drives), random access memory (RAM), read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), registers, removable disks, and any other suitable medium for storing software and / or instructions that can be accessed and read by a computer. Memory 520 may reside in processor 510, be external to processor 510, or be distributed across multiple entities including processor 510. Memory 520 may be embodied in a computer program product. For example, a computer program product may include a computer-readable medium in encapsulation material. Those skilled in the art will recognize how the functionality described herein can be implemented depending on the specific application and the overall design constraints imposed on the system as a whole.

[0056] Additionally, according to another embodiment of this disclosure, a computer program product for information processing of video content is disclosed. As an example, the computer program product includes a non-transitory computer-readable storage medium having program instructions embodied therein, and the program instructions are executable by a processor. When executed, the program instructions cause the processor to perform one or more of the processes described above, and details are omitted herein for brevity.

[0057] The above description, with reference to the accompanying drawings, outlines a method and apparatus for processing video content according to embodiments of the present disclosure. This disclosure, through an automated process of event filtering (e.g., excluding environmental interference events and events of no concern to the user), behavioral feature information extraction, related event merging, and statistical description information generation, transforms the traditional inefficient security model relying on manual video playback into a highly efficient cognitive model with automated statistical description. This significantly improves information processing efficiency and enhances the readability and user cognitive efficiency of information related to key events.

[0058] Unless otherwise expressly stated, expressions such as “according to,” “based on,” “depending on,” etc., as used in this disclosure do not mean “according to only,” “based on only,” or “depending on only.” In other words, in this disclosure, such expressions generally mean “at least according to,” “at least based on,” or “at least depending on.”

[0059] Any references to elements in this disclosure, such as the names "first," "second," etc., are not intended to comprehensively limit the number or order of these elements. These expressions may be used in this disclosure as a convenient way to distinguish two or more units. Therefore, references to the first unit and the second unit do not imply that only two units may be used, or that the first unit must precede the second unit in some form.

[0060] As used in this disclosure, the term "determine" can include a variety of operations. For example, "determine," calculation, operation, processing, derivation, investigation, search (e.g., searching in a table, database, or other data structure), and ascertainment are all considered "determine." Additionally, "determine" also refers to receiving (e.g., receiving information), sending (e.g., sending information), inputting, outputting, and accessing (e.g., accessing data in memory). Furthermore, "determine" can also refer to parsing, selecting, picking, building, and comparing. In other words, several actions can be considered "determine."

[0061] As used in this disclosure, terms such as “connection,” “coupling,” or any variation thereof refer to any direct or indirect connection or combination between two or more units, which may include situations where one or more intermediate units exist between two units that are “connected” or “coupled” to each other. The coupling or connection between units may be physical or logical, or a combination of both. As used in this disclosure, two units may be considered electrically connected by means of one or more wires, cables, and / or printing, and as numerous non-limiting and non-exhaustive examples, may be “connected” or “coupled” to each other by means of electromagnetic energy in the radio frequency region, microwave region, and / or light (visible and invisible) region, etc.

[0062] When the terms “comprising,” “including,” and variations thereof are used in this disclosure or claims, these terms are open-ended, just like the term “having.” Furthermore, the term “or” as used in this disclosure or claims is not an exclusive “or.”

[0063] Those skilled in the art will understand that many changes and / or modifications can be made to the present disclosure shown in the specific embodiments without departing from the spirit or scope of the present disclosure as broadly described. Therefore, the embodiments are to be considered illustrative rather than restrictive in all respects.

Claims

1. A method for processing video content, comprising: Identify object behaviors in multiple corresponding events from multiple video segments collected within a predetermined time period; For the object behavior in each of the multiple corresponding events, its behavioral feature information is extracted, wherein the behavioral feature information includes at least one of the following: behavior type, behavior occurrence time, behavior occurrence location, action features, object identification features, and object sound information; as well as Based on the behavioral characteristic information of each object's behavior in the multiple corresponding events, statistical description information of the multiple corresponding events is generated according to preset statistical rules.

2. The method according to claim 1, further comprising: Environmental interference events are excluded from the plurality of corresponding events, which are identified based on the visual and / or audio information of the video segment.

3. The method according to claim 1, further comprising: Non-user-concerned events are excluded from the plurality of corresponding events, which are identified based on preset user-concern behavior rules.

4. The method according to claim 1, further comprising: Based on preset event association rules, determine whether two or more of the multiple corresponding events are related; as well as Related events are merged to generate corresponding merged statistical description information.

5. The method of claim 4, wherein, The event association rules are based on at least one of the following dimensions: proximity of the time of occurrence of the behavior, proximity of the location of occurrence of the behavior, whether they belong to the same object, whether they belong to the same behavior type, and similarity of the action characteristics.

6. The method according to claim 1, further comprising: According to the preset importance rules, the importance of each event is determined based on the behavioral feature information. The importance is used to characterize the level of behavioral security risk and / or the user's attention priority.

7. The method of claim 6, wherein, The preset statistical rules include preset importance ranking rules, and the statistical description information of the multiple corresponding events is ranked based on the importance.

8. The method of claim 6, wherein, The statistical description information includes: detailed description information of key events, wherein the key events are those whose importance meets a preset threshold among the multiple events, and the detailed description information is generated by supplementing the behavioral feature information of the key behavior based on the video segments corresponding to the key events.

9. The method of claim 1, wherein, The preset statistical rules include preset time sorting rules, and the statistical description information of the multiple corresponding events is sorted according to the chronological order of the occurrence of the behavior.

10. The method according to claim 1, wherein, The statistical description information includes: the total number of the multiple corresponding events and / or the classification statistics of the multiple corresponding events, wherein the classification statistics are the results of statistics performed according to at least one dimension among the behavior type, the time of occurrence of the behavior, the location of occurrence of the behavior, and the object category.

11. The method of claim 1, wherein, Generating statistical description information for the multiple corresponding events includes generating the statistical description information at a preset statistical granularity, wherein the preset statistical granularity includes at least one of time granularity, behavior type granularity, location granularity, and object type granularity.

12. The method of claim 11, wherein, The preset statistical granularity is user-defined.

13. The method of claim 8, wherein, The statistical description information also includes: a video access interface for the key event, for accessing the video segment corresponding to the key event among the plurality of video segments through the video access interface.

14. The method of claim 1, wherein, The statistical description information also includes: a video access interface for each event, for accessing the video segment corresponding to the corresponding event among the plurality of video segments through the video access interface.

15. An information processing apparatus for video content, comprising: processor; Memory coupled to the processor; as well as Computer program instructions stored in the memory, which, when executed by the processor, perform the method as described in any one of claims 1-14.

16. A computer program product comprising computer program instructions that, when executed by a processor of a network device, perform the method as described in any one of claims 1-14.