Method and apparatus for recognizing the actions of an object from a first set of video streams
By optimizing video stream assignment based on processing unit accuracy and probability, the method addresses inconsistent recognition accuracy in large-scale camera systems, enhancing action recognition efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NEC CORP
- Filing Date
- 2023-11-07
- Publication Date
- 2026-05-15
AI Technical Summary
Existing methods for recognizing object actions from multiple video streams result in inconsistent action recognition accuracy due to differing models and platforms across video processing units, leading to inefficient load distribution and suboptimal performance.
A method and system that considers the action recognition accuracy and probability of each processing unit to assign video streams, optimizing load balancing and ensuring accurate action recognition by assigning streams to higher-precision units.
Enhances action recognition accuracy and efficiency by effectively distributing video streams based on processing unit precision, ensuring optimal performance in large-scale camera systems.
Smart Images

Figure 0007859599000003 
Figure 0007859599000004 
Figure 0007859599000005
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method, apparatus, and system for recognizing object actions, and more particularly, to a method, apparatus, and system for recognizing object actions from multiple video streams via multiple processing units.
Background Art
[0002] For example, there is an increasing demand for processing multiple video streams in order to detect objects and recognize object actions. Generally, it is necessary to assign such multiple video streams to an action recognition model to perform action recognition and output a recognition result (e.g., running, riding a bicycle, two people fighting, etc.). Generally, the camera / video system is large-scale (e.g., more than 200 cameras / video streams), but when the number of action recognition models (or the number of processing units executing the models) is limited, it is necessary to ensure that the tasks of action recognition are distributed over limited resources, and to avoid unevenly overloading some action recognition models (or processing units executing the models) while leaving other action recognition (or processing units executing the models) idle, thereby making the processing more efficient.
Summary of the Invention
Problems to be Solved by the Invention
[0003] However, since the models and platforms used in different video processing units may be different, even if the same input video stream is assigned to a video processing unit, the user may obtain different action recognition accuracy results. Considering the load distribution across multiple video processing units, the action recognition accuracy results are inconsistent and strongly dependent on the assignment.
[0004] Therefore, in order to address the problems and limitations in load balancing and to achieve optimal action recognition accuracy and efficiency, it is necessary to develop methods, apparatus, and systems for recognizing object actions from multiple video streams via multiple processing units, for example by considering the action recognition accuracy of each processing unit and / or the probability of recognizing actions from the video stream.
[0005] Furthermore, other desirable features and characteristics will become apparent from the following detailed description and the attached claims, in combination with the attached drawings and this background of the disclosure. [Means for solving the problem]
[0006] In a first aspect, the Disclosure provides a method for recognizing the actions of an object from a first plurality of video streams, the method comprising detecting an object from each of the first plurality of video streams, and assigning one of the first plurality of video streams to one of a second plurality of processing units, and recognizing the actions of the object based on the accuracy of the recognition of the object's actions by one of the second plurality of processing units.
[0007] In a second aspect, the Disclosure provides a device for recognizing the actions of an object from a first plurality of video streams, the device comprising at least one processor and at least one memory containing computer program code, the at least one memory and computer program code being configured to use at least one processor to cause the device to recognize an object from at least each of the first plurality of video streams, to assign one of the first plurality of video streams to one of a second plurality of processing units, and to recognize the actions of the object based on the accuracy of the recognition of the object action by one of the second plurality of processing units.
[0008] In a third aspect, the disclosure provides a system for recognizing the actions of an object from a first plurality of video streams, comprising an apparatus according to the second aspect and a third plurality of image capture devices. [Effects of the Invention]
[0009] Further benefits and advantages of the disclosed embodiments will become apparent from the specification and drawings. Benefits and / or advantages may also be obtained individually from the various embodiments and features of the specification and drawings, and not all of them are required to be provided in order to obtain one or more of such benefits and / or advantages. [Brief explanation of the drawing]
[0010] Embodiments of this disclosure will be readily apparent to those skilled in the art, in conjunction with the drawings, from the following description, which is merely an example. [Figure 1] This is a schematic diagram illustrating the process for recognizing object actions from an input video stream. [Figure 2] This is a schematic diagram illustrating the process of recognizing object actions from multiple input video streams using a conventional load balancer. [Figure 3] This is a block diagram of a system for recognizing object actions from multiple video streams, according to various embodiments of the present disclosure. [Figure 4] This flowchart shows a method for recognizing the actions of an object from a first set of video streams, according to various embodiments of the present disclosure. [Figure 5] A system 500 for recognizing the actions of an object from a first set of video streams, according to various embodiments of the present disclosure. [Figure 6] This block diagram shows the various components of the action recognition device shown in Figure 5 according to one embodiment of the present disclosure, and the process flow between them. [Figure 7]This is a block diagram showing a lightweight estimator for a probability calculation unit and an action recognition probability estimation unit according to one embodiment of the present disclosure. [Figure 8] This is a block diagram showing the communication between an allocation unit and a probability calculation unit and a multiplex video processing unit according to one embodiment of the present disclosure. [Figure 9] This flowchart shows a process overview of a video stream input unit, a probability calculation unit, and an assignment unit of an action recognition device according to one embodiment of the present disclosure. [Figure 10] This is a flowchart of a Markov decision process used to assign each of a first plurality of video streams to one of a second plurality of VPUs, according to one embodiment of the present disclosure. [Figure 11] This is a flowchart 1100 illustrating the training process of a lightweight classifier or estimator according to one embodiment of the present disclosure. [Figure 12] This flowchart shows a process for calculating and estimating the probability of a target action from a video stream, according to one embodiment of the present disclosure. [Figure 13] This is a schematic diagram of an exemplary computing device suitable for use in implementing the method shown in Figure 5 and the apparatus shown in Figure 6. [Modes for carrying out the invention]
[0011] <Explanation of Terms> Objects - Objects can be people, pets, vehicles, things, items, devices, pillars, furniture, or anything that is stationary or moving. Objects can be living or inanimate. In the case of living or biological objects such as people and pets, objects can typically be identified based on appearance features, body parts, physical characteristics, the movement of the object, or a combination thereof. Examples of appearance features of an object (person) include the relative position, size, shape and / or contour of the eyes, nose, cheekbones, jaw and chin, as well as iris pattern, skin color, hair color, or a combination thereof. Characteristics include physical attributes such as height, build, body shape, body proportions, limb length, hair color, skin color, clothing, possessions, and other similar characteristics or combinations thereof.
[0012] Movement includes behavioral characteristics such as body movement, limb positioning, direction of movement, speed of movement, gait patterns, how an object stands, moves, or speaks, changes in physical attributes when interacting with other objects, and other similar characteristics or combinations. In the case of inanimate objects, the object can typically be based on its speed of movement, movement characteristics / patterns, and changes in physical attributes when interacting with other objects.
[0013] In various embodiments of this disclosure, an object can be detected by capturing an image of the object using an image capture device such as a camera, and identifying the object based on its appearance features, body parts, physical characteristics, and / or the movement of the object in the image. In the following various embodiments, the term “sensor” may be used to refer to an image capture device.
[0014] In various embodiments, the object at the time of detection, as well as the data acquired by a sensor that is used for and associated with the identification and detection of the object, is assigned to an object identifier for subsequent identification and tracking of the object. When the object is subsequently identified based on the same or other appearance features, body parts, body characteristics, the movement of the object, or a combination thereof, the same object ID is assigned.
[0015] The action of an action-object may refer to the type of activity performed by an object that can be recognized and classified from an image or video stream based on a series of physical features (e.g., appearance features, body parts, body characteristics) and / or motion / behavior features (e.g., movement, motion) of the object identified from the video stream. Examples of actions include sitting, talking, running, jumping, riding a bicycle, fighting, stealing.
[0016] In one example, the action of an object is recognized based on the same or similar actions of the same or similar object(s) previously recognized and stored in a database. In another example, the action of an object is recognized based on a series of physical features (e.g., appearance features, body parts, body characteristics) and / or motion / behavior features of both the object and another object in the video stream, such as the fighting and throwing of the object.
[0017] In the following various embodiments, the action of an object to be recognized from a video stream may be referred to as a "target action".
[0018] A video stream refers to the continuous transmission or input of video or image files. The video or images may be generated by a processor in connection with an image capture device, or retrieved from a database. In one example, the processor and database may be connected to a server. The transmission or input of video or image files may be via wired or wireless (e.g., via NFC, Wi-Fi, Bluetooth, etc.) or a network (e.g., the Internet).
[0019] Processing Unit – A processing unit refers to a processor configured to process a video or image file to recognize the action of an object using one or more action recognition models / algorithms stored in memory accessible by the processor. Furthermore, the processing unit may be further configured to process the video or image file to detect objects using one or more object detection models / algorithms stored in memory prior to the action recognition process.
[0020] Precision – Precision in recognizing object actions refers to the precision with which the processing unit recognizes the action of an object from a sequence of motion, using one or more action recognition models / algorithms stored in memory accessible by the processing unit. In one implementation, additional information and features that affect the performance of the processing unit, such as processing time, video resolution, and video duration, may be used to determine and adjust the precision of the processing unit or the assignment of the video stream to the processing unit in order to detect and recognize the action of objects in the video stream.
[0021] Rank - The rank of a processing unit is determined based on its accuracy in recognizing object actions compared to the accuracy of other processing units in recognizing object actions. A higher rank indicates that the processing unit is more accurate in recognizing object actions than other lower-ranked processing units. In one embodiment, such ranks are used for video stream assignment. For example, if a video stream has higher importance, it may be assigned to a processing unit with a higher rank that has a higher accuracy in recognizing object actions, so that actions can be accurately recognized from high-importance video streams. Alternatively or additionally, video streams with a lower probability of detecting and recognizing object actions may also be assigned to higher-ranked processing units that have a higher accuracy in recognizing object actions, resulting in easier identification and recognition of actions from video streams with less processing power or time. Conversely, if a video stream has lower importance, it may be assigned to a lower-ranked processing unit with a lower accuracy in recognizing object actions, so that higher-ranked processing units can be used to process more important video streams. Alternatively or additionally, video streams with a high probability of detecting and recognizing object actions can be assigned to lower-ranked processing units with higher accuracy in recognizing object actions, allowing for better identification and recognition of actions from the video stream.
[0022] Probability – Probability relates to the likelihood that an object's action will be detected and recognized from the video stream and is determined based on a set of physical and / or motion / behavioral features of the object, including but not limited to the number of objects, the location of objects, the position of parts of objects, the relative distance between objects, the timestamp when an object and / or an object's action was detected, physical attributes (or changes in physical attributes) of the object, changes in the size of objects in the video stream, the duration for which an object and / or an object's action was detected, the positional distribution of different objects in the video stream, or a combination thereof, as well as other object features identified from the video stream.
[0023] In one implementation, different sets of physical, motion / behavioral, and object features may be relied upon to detect different actions of an object and their probabilities from a video stream. For example, features such as the number of people, arm positions, object collisions, and the frequency of arm extensions and object collisions can form a set of features for detecting fighting actions, and as a result, when those features are identified from a video stream, for example, when two people are sparring and colliding with each other in a short period of time, the probability of detecting a fighting action will be higher compared to other actions (e.g., riding a bicycle, sitting, stealing).
[0024] Weight Parameters - Weight parameters are the weights of a deep learning model that are assigned to and applied to physical, motion / behavioral, or object features, or combinations of two features identified from a video stream, in order to calculate the probability of detecting one or more actions of an object from a video stream. Such weight parameters correlate with the emphasis or priority given to a feature (or combination of features) in the probability calculation.
[0025] In one implementation, weight parameters assigned to and applied to physical, motion / behavior, or object features, or combinations of two features, are updated by training a deep learning model. During training, multiple video streams containing known actions (including physical, motion / behavior, and object features, and combinations thereof, that lead to the recognition of known actions) are used to check whether such known actions can be accurately recognized from such video streams. If the action recognized from such video streams does not match a known action, an instruction may be generated. Upon receiving such an instruction, the weight parameters of the physical, motion / behavior, and object features, and combinations identified from the video streams (and / or other physical, motion / behavior, and object features and combinations that contribute to the recognition of known actions) may be updated accordingly so that the known actions can be accurately recognized from such video streams using the updated weight parameters.
[0026] Importance – Importance relates to how important it is to correctly recognize the actions of objects from the video stream. Generally, higher-importance video streams are assigned to higher-rank (higher accuracy) processing units, making it easier to accurately recognize the actions of objects.
[0027] <Exemplary Embodiment> Embodiments of this disclosure will be described by reference to the drawings, merely as examples. Similar reference numerals and letters in the drawings refer to similar elements or equivalents.
[0028] Some parts of the following description are presented, explicitly or implicitly, with respect to algorithms and functional or symbolic representations of operations on data in computer memory. These algorithmic descriptions and functional or symbolic representations are means used by those skilled in the data processing techniques to communicate the content of their work to others skilled in the techniques in the most effective way. An algorithm is considered to be a self-consistent set of steps that produce a desired result. These steps require the physical manipulation of physical quantities, such as electrical, magnetic, or optical signals, which can be stored, transferred, combined, compared, and other manipulated.
[0029] Unless otherwise specified, as will be evident from the following, any description using terms such as “receive,” “calculate,” “determine,” “update,” “generate,” “initialize,” “output,” “retrieve,” “identify,” “distribute,” and “authenticate” will be understood to refer to actions and processes of a computer system or similar electronic device that manipulate data represented as physical quantities within a computer system and convert it into other data similarly represented as physical quantities within a computer system or other information storage, transmission, or display device.
[0030] This specification also discloses apparatus for carrying out the operations of the method. Such apparatus may be specifically constructed for the required purpose, or may comprise a computer or other device that is selectively activated or reconfigured by a computer program stored within the computer. The algorithms and representations presented herein are not inherently related to any particular computer or other apparatus. Various machines may be used together with the programs taught herein. Alternatively, it may be appropriate to construct a more specialized apparatus for carrying out the required method steps. The structure of a computer will become apparent from the following description.
[0031] In addition, this specification also implicitly discloses computer programs in such a way that it will be apparent to those skilled in the art that the individual steps of the methods described herein can be carried out by computer code. The computer programs are not intended to be limited to any particular programming language and its implementation. It will be understood that various programming languages and their codings may be used to implement the teachings of this disclosure contained herein. Furthermore, the computer programs are not intended to be limited to any particular control flow. There are many other variations of the computer programs that may use different control flows without departing from the spirit or scope of this disclosure.
[0032] Furthermore, one or more steps of a computer program may be performed in parallel rather than sequentially. Such a computer program may be stored on any computer-readable medium. Computer-readable mediums may include storage devices such as magnetic disks or optical disks, memory chips, or other storage devices suitable for interfacing with a computer. Computer-readable mediums may also include hardwired media, such as those exemplified by internet systems, or wireless media, such as those exemplified by GSM mobile phone systems, Long Term Evolution (LTE) systems, and 5G mobile network systems. When a computer program is loaded and executed on such a computer, it effectively brings to life an apparatus that implements the steps of the preferred method.
[0033] Various embodiments of this disclosure relate to methods and apparatus for recognizing the actions of objects from multiple video streams. Those skilled in the art will understand that such apparatus and image capture devices may be implemented as part of a system to provide the same technical effect.
[0034] FIG. 1 shows a schematic diagram 100 illustrating an overview of a process for recognizing object actions from an input video stream 102. The input video stream is fed through a video buffer into multiple video streams (indexed from 1 to M within queue 106), and the multiple video streams 106 need to be assigned to one or more processing units to execute an action recognition model 104 (indexed from 1 to N) for recognizing object actions (e.g., running, riding a bicycle, people fighting) and output recognition results from the video streams 106. Here, in this example, each action recognition model is assumed to be executed by a processing unit or a processor. However, it is understood that the processing unit or the processor may be configured to execute two or more action recognition models.
[0035] For example, when the number of action recognition models (or particularly the number of processing units) is limited, i.e., N < M, problems occur in a large-scale camera system having more than 200 camera streams (and input video streams). Additionally, since the models and platforms used in the processing units may be different, assigning the same video stream to different video processing units may result in differences in action recognition, potentially affecting the accuracy and reliability of the load balancer.
[0036] Figure 2 shows a schematic diagram 200 illustrating the process for recognizing object actions from multiple input video streams 202 using a conventional load balancer 204. The conventional load balancer 204 is configured to allocate the multiple input video streams 202 to three video processing units (VPUs) 206a, 206b, and 206c in order to recognize target actions. Assuming that the target action appears in only one of the input video streams, the conventional load balancer 204 does not consider VPU precision when allocating video streams 202, so the video stream in which the target action appears may be allocated to a lower-precision VPU (e.g., VPU 206a). As a result, the VPU cannot detect and recognize the target action, and a recognition result indicating that the target action was not detected is output from the multiple input video streams 202. Therefore, if this target action can be detected by other VPUs, such as VPUs 206b and 206c, which have higher precision, the user could not obtain optimal action recognition accuracy.
[0037] Therefore, the objective is to consider how to allocate multiple video stream loads to a limited number of action recognition models. This disclosure provides novel action recognition methods, apparatus, and systems to address such problems and limitations in load balancing between multiple video streams from a large-scale camera system and a limited number of action recognition models (processing units) in order to achieve optimal action recognition accuracy and efficiency.
[0038] In one implementation, information such as the action recognition accuracy of the processing unit and the probability of detecting a target action within each video stream is taken into consideration in the new action recognition method, apparatus, and system. For example, as shown in Figure 1, the video stream 106 is supplied to an action recognition apparatus 108 which comprises a probability calculation unit 108a configured to calculate the probability of recognizing an object action from each video stream, and an assignment unit 108b configured to assign each video stream to an action recognition model.
[0039] Figure 3 shows a block diagram of a system 300 for recognizing object actions from multiple video streams according to various embodiments of the present disclosure.
[0040] System 300 comprises a requesting device 302, an action recognition server 308, a collaboration server 340, hosts 350A to 350N, and sensors 342A to 342N.
[0041] The requesting device 302 communicates with the action recognition server 308 and / or the cooperation server 340 via connections 316 and 321, respectively. Connections 316 and 321 may be wireless (e.g., via NFC communication, Wi-Fi communication, Bluetooth, etc.) or via a network (e.g., the Internet). Connections 316 and 321 may be network-based (e.g., the Internet).
[0042] The action recognition server 308 further communicates with the cooperation server 340 via connection 320. Connection 320 may be via a network (e.g., a local area network, a wide area network, the internet, etc.). In one configuration, the action recognition server 308 and the cooperation server 340 are combined, and connection 320 may be an interconnected bus.
[0043] Next, the coordinating server 340 communicates with hosts 350A to 350N via their respective connections 322A to 322N. Connections 322A to 322N may be a network (for example, the Internet).
[0044] Hosts 350A to 350N are servers. The term "host" is used herein to distinguish between hosts 350A to 350N and the cooperative server 340. Hosts 350A to 350N are collectively referred to as host 350 herein, and host 350 refers to one of the hosts 350. Host 350 may be combined with the cooperative server 340.
[0045] In one example, host 350 may be managed by the entity's security officer, and the coordinating server 340 is a central server that coordinates the hosts 350 and determines which of the hosts 350 will transfer data or search for data such as image input.
[0046] Sensors 342A to 342N are connected to the coordinating server 340 or the action recognition server 308 via their respective connections 344A to 344N or 346A to 346N. Sensors 342A to 342N are collectively referred to as sensor 342 in this specification. Connections 344A to 344N are collectively referred to as connection 344 in this specification, and connection 344 refers to one of the connections 344. Similarly, connections 346A to 346N are collectively referred to as connection 346 in this specification, and connection 346 refers to one of the connections 346. Connections 344 and 346 may be wireless (e.g., via NFC communication, Wi-Fi communication, Bluetooth, etc.) or network (e.g., the Internet). Sensor 342 may be one of an image capture device, an object tracking device, a video capture device, a motion sensor, and a temperature sensor, and depending on its type, it may be configured to transmit input to at least one of the action recognition servers 308.
[0047] In an exemplary embodiment, device 302 and sensorEach of 342, as well as servers 308, 340, and 350, connects to other connected devices 302 and sensor It provides an interface that enables communication with 342 and / or servers 308, 340, and 350. Such communication is facilitated by an Application Programming Interface (API). Such an API may be part of a user interface that includes a Graphical User Interface (GUI), a web-based interface, a set of program interfaces such as Application Programming Interfaces (APIs) and / or Remote Procedure Calls (RPCs) corresponding to interface elements, a messaging interface where interface elements correspond to messages of a communication protocol, and / or an appropriate combination thereof.
[0048] In this specification, the term "server" can mean a single computing device comprising a set of processors working together, or a set of interconnected computing devices working together to perform a particular function. That is, a server may be contained within a single hardware unit, or it may be distributed across multiple or many different hardware units.
[0049] <Linked Server 340> The collaboration server 340 is associated with an entity (e.g., a service company or organization or moderator). In one deployment, the collaboration server 340 is owned and operated by the entity that operates the server 308. In such a deployment, the collaboration server 340 may be implemented as part of the server 308 (e.g., a computer program module, a computing device, etc.).
[0050] The integration server 340 may also be configured to manage user registration. Registered users have an action-aware account that contains user details. The registration step is called onboarding. Users can perform onboarding to the integration server 340 using either the requesting device 302 or the host 350.
[0051] To access the functions of the integration server 340, it is not necessary to have an action recognition account on the integration server 340. However, there are functions available to registered users. For example, functions such as recognition of more complex actions or actions involving multiple objects, or an increase in the maximum number of input video streams may be exclusive to registered users.
[0052] The user onboarding process is performed by the user via one of the requesting devices 302. In one deployment, the user downloads an application (including an API for interacting with the integration server 340) to the sensor 342. In another deployment, the user accesses a website (including an API for interacting with the integration server 340) on the requesting device 302.
[0053] Registration details include, for example, the user's user identifier (ID) or appearance portrait, the user's address, contact information, or other important information, and the sensor 342 authorized to update the action recognition account.
[0054] Upon onboarding, users will have an action-aware account that remembers all their details.
[0055] <Requesting device 302> The requesting device 302 is associated with the subject (or requesting party) that is a party to an action recognition request that begins with the requesting device 302. The requesting party may be a public person or a security officer of the entity who is helping to obtain the data necessary to detect and recognize a target action (e.g., steal, fight) of a person(s) or object within the entity. The requesting device 302 may be a computing device such as a desktop computer, an interactive voice response (IVR) system, a smartphone, a laptop computer, a personal digital assistant computer (PDA), a mobile computer, or a tablet computer.
[0056] In one configuration example, the requesting device 302 is a wristwatch or similar wearable computing device equipped with a wireless communication interface.
[0057] <Action Recognition Server 308> The action recognition server 308, as described above in the glossary section, is configured to recognize the actions of objects from multiple video streams.
[0058] <Host 350> Host 350 is a server associated with an entity (e.g., a company or organization) that manages object information about objects for which actions have been recognized (e.g., establish, operate).
[0059] In one configuration, the entity is a bank. Thus, each entity operates host 350 to manage resources by that entity. In one configuration, host 350 receives an alert signal that a target action has been detected. Host 350 may then be configured to send resources to a location identified by location or camera information included in the alert signal. For example, the host may be configured to acquire relevant video or image input for processing.
[0060] In one configuration, video streams, detected objects, and recognized actions can be stored and updated in an action recognition account associated with the user. Preferably, such information is valuable to law enforcement agencies, such as security or building management staff, and to users for identifying, tracking, and monitoring objects. This reduces the time it takes to view camera footage and recognize object actions.
[0061] <Sensor 342> Sensor 342 is associated with a user associated with the requesting device 302. Sensor 342 may be one of the following: an image capture device, an object tracking device, a video capture device, a motion sensor, and a temperature sensor, and depending on its type, it may be configured to send input to at least one of the action recognition servers 308. Further details on how the sensor can be used to recognize the action of an object are provided below.
[0062] Figure 4 shows a flowchart 400 illustrating a method for recognizing object actions from a first set of video streams according to various embodiments of the present disclosure. Step 402 involves detecting objects from each of the first set of video streams. Step 404 involves assigning one of the first set of video streams to one of a second set of processing units and recognizing the object actions based on the accuracy with which the second set of processing units recognize the actions of one or more objects.
[0063] Figure 5 shows a block diagram illustrating a system 500 for recognizing object actions from a first set of video streams, according to various embodiments of the present disclosure.
[0064] In one example, the management of image and signal inputs is performed by all image capture devices 502a, 502b. System 500 includes multiplex image capture devices 502a, 502c that communicate with device 504 (for simplicity, only two image capture devices are shown). In one implementation, device 504 can generally be described as a physical device comprising at least one processor 506 and at least one memory 508 containing computer program code. The at least one memory 508 and the computer program code, together with the at least one processor 506, are configured to cause the physical device to perform the operations described in Figure 4. The processor 506 is configured to receive a first set of video streams from the image capture devices 502a, 502b or to retrieve a first set of video streams from the database 510. Alternatively or additionally, the first multiple video streams captured by the image capture devices 502a, 502b are stored in the database 510, and the processor 506 is configured to retrieve the first multiple video streams from the database 510.
[0065] The image capture devices 502a and 502b may be devices such as closed-circuit television (CCTV) that provide various data (camera data) of physical, motion / behavior, and object characteristic data that can be used by the system to detect objects under an object ID and recognize the actions of the objects. In one implementation, the data derived from the image capture devices 502a and 502b may be stored in the memory 508 of the device 504 or in a database 510 accessible by the device 504.
[0066] Furthermore, camera data such as location data regarding the position where the camera is fixed or capturing, and time data such as video / image timestamps can be received, stored, and / or retrieved to derive the location and timestamp of an action related to an object in order to recognize the action of the object.
[0067] According to this disclosure, the device 504 may be configured to communicate with image capture devices 502a, 502b, a database 510, and a multiplexing unit (not shown). In one implementation, the processing unit may be part of the device 504, and the processor 506 and They communicate. Similarly, in one implementation, the database 510 can be part of the device 504.
[0068] Device 504 may receive multiple video streams as input from the image capture device 502 or retrieve them from the database 510. 508 Call memory 508 The stored computer program code is configured to use the processor 506 to cause the device 504 or image capture devices 502a, 502b to directly detect objects from each of the multiple video streams. Detection may be based on appearance features, body parts, physical characteristics, object motion, or a combination thereof.
[0069] Memory 508 or database 510 may store the accuracy of each processing unit in recognizing the target action(s) of an object(s) (or other similar objects), and processor 506 may retrieve such accuracy from memory 508 or database 510 and assign each of the multiple video streams to a processing unit to process and detect the target action(s) of the object. Detection of the target action may be based, for example, on a set of physical features (e.g., appearance features, body parts, physical characteristics) and / or motion / behavioral features (e.g., motion, movement) of the object identified from the video stream by having the processing unit run an action recognition model on the assigned video stream.
[0070] More specifically, memory 508 and the computer program code stored in memory 508 are configured to cause the device 504 to use processor 506 to compare the accuracy of each processing unit in recognizing a specific action(s) of an object(s) (or other similar objects), and to determine which processing unit(s) have higher (or highest) accuracy in detecting the action(s) than the other(s). The determination result is then used by processor 506 to assign video stream(s) to processing units(s) to detect and recognize the action(s) of the object. In one implementation, after the accuracy comparison, a ranking table is generated showing the rank of each processing unit communicating with device 504, and the rank of the processing units is used directly to determine the assignment of video streams, for example, a set of processing units (e.g., from high-precision processing units to low-precision processing units) to be assigned to a video stream load.
[0071] In one embodiment, the memory 508 and the computer program code stored in the memory 508 are configured to use the processor 506 to cause the device 504 or image capture devices 502a, 502b to identify one or more features from each of a plurality of video streams. The memory 508 and the computer program code stored in the memory 508 are configured to use the processor 506 to cause the device 504 to calculate the probability of a target action of an object (or similar object) recognized from the plurality of video streams.
[0072] Furthermore, memory 508 or database 510 may store weight parameters for each feature identified from the video stream, and memory 508 and the computer program code stored in memory 508 are configured to use processor 506 to cause device 504 to retrieve such weight parameters from memory 508 or database 510, apply each weight parameter to one or more features identified from the video stream, and calculate the probability of an object's target action.
[0073] Memory 508 and the computer program code stored in memory 508 are configured to cause the device 504 to use the processor 506 to further compare the probabilities of target actions for objects (or similar objects) recognized from multiple video streams, and to determine which video stream has a higher (or highest) probability of having a target action recognized by the processing unit than the other video streams. The determination result is then used by the processor 506 to assign the video stream(s) to the processing unit(s). For example, the result can be used to determine a set of video streams (e.g., from high probability to low probability) to be assigned to the processing unit.
[0074] Furthermore, the memory 508 and the computer program code stored in the memory 508 are configured by the processor 506 to cause a processing unit or device communicating with the device 504 to execute a deep learning model. The device 504 may assign a video stream with known actions to a processing unit, execute an action recognition model to recognize the known actions, and receive instructions from the processing unit executing the action recognition model on whether the known actions were accurately recognized from the video stream. Alternatively, the device 504 may receive action recognition results from the processing unit executing the action recognition model, indicating the actions recognized from the video stream by the processing unit, and the memory 508 and the computer program code stored in the memory 508 are configured, using the processor 506, to cause the device 504 to determine whether the action of an object matches a known action, and whether the known action was accurately recognized by the processing unit executing the action recognition model. If it is indicated that an object's action is not recognized from the video stream, the memory 508 and the computer program code stored in the memory 508 are configured to use the processor 506 to cause the device 504 to update the weight parameters of features detected by the processing unit from the video stream and contributing to the action recognition result, so that known actions can be more accurately recognized from the video stream by the processing unit.
[0075] In an alternative embodiment, the memory 508 and the computer program code stored in the memory 508 are configured to use the processor 506 to cause the device 504 to compare the importance of each of the multiple video streams and determine which video stream has a higher (or highest) importance and from which the target action is recognized more than from the other video streams. The determination result is then used by the processor 506 to assign the video streams to the processing unit. For example, the result may determine a set of video streams (e.g., from high probability to low probability) to be assigned to the processing unit.
[0076] Figure 6 shows a block diagram 600 illustrating the various components of the action recognition device of Figure 5 according to one embodiment of the present disclosure and the process flow between them. In this embodiment, the action recognition device 600 (or the processor of the action recognition device 600) comprises a video stream input unit 602, a probability calculation unit 604, an assignment unit 606, a video processing unit 608, and an information display unit 610. As illustrated in the action recognition system of Figure 1, the video stream input unit 602, the video processing unit 608, and the information display unit 610 do not have to be part of the action recognition device 600 (or the processor of the action recognition device 600).
[0077] As shown in the exemplary method for recognizing the action of an object in Figure 4, the action recognition device 600, when in operation, Step 402, wherein the video stream input unit 602 can detect objects from each of the first multiple video streams, Step 404, in which the assignment unit assigns one of the first multiple video streams to one of the video processing units 608, and the action of the object can be recognized based on the accuracy with which the action of the object in one of the video processing units 608 is recognized. It is configured to execute.
[0078] In step 402, the video stream input unit 602 may receive or retrieve the first plurality of video streams from, for example, a multiplex image capture device or a database (not shown) before detecting objects from each of the first plurality of video streams. The video stream input unit 602 may temporarily buffer the captured video streams and use them to obtain basic camera information (e.g., resolution, frame rate, etc.) for further processing.
[0079] Furthermore, in step 402, the video stream input unit 602 may identify one or more features from each of the first plurality of video streams for further processing by the probability calculation unit 604. Alternatively, such identification step may be performed by the probability calculation unit 604 itself.
[0080] Furthermore, in step 402, the video stream input unit 602 may determine the importance of each of the first multiple video streams whose actions of the object should be recognized. Alternatively, such a determination step may be performed by the assignment unit 606 itself. Alternatively, the importance of each of the first multiple video streams may be received from the image capture device.
[0081] In step 404, before assigning one of the first multiple video streams to one of the video processing units 608, the assignment unit 606 may retrieve the average recognition accuracy of each video processing unit 608 in recognizing object actions from each of the video processing units 608 or from a database (not shown) and determine whether the average recognition accuracy of the processing unit is higher than that of each of the remaining video processing units. The assignment in step 404 will be based on such a determination. The assignment unit 606 may further determine the rank of each video processing unit relative to the remaining video processing units. The assignment in step 404 will then be based on the rank of one of the video processing units 608, for example, that the assigned video processing unit has the highest (or lowest) accuracy in recognizing object actions among all the video processing units 608.
[0082] Furthermore, in step 404, before assigning one of the first plurality of video streams to one of the video processing units 608, the probability calculation unit 604 may calculate the probability that an object action is recognized from each of the first plurality of video streams, either by itself or based on one or more features identified by the video stream input unit 602, and determine whether the probability that an object action is recognized from one of the video streams is higher than that of each of the other video streams in the first plurality of video streams. The assignment unit 606 receives the video streams from the probability calculation unit 604 with the calculated probabilities. The assignment in step 404 may then be based on such determination results, for example, that the assigned video stream has the highest (or lowest) probability of detecting and recognizing an object action.
[0083] In one implementation, the allocation unit 606 maintains a video stream allocation table based on the rank of each video processing unit, updates the video stream allocation based on the new rank of the video processing unit calculated from their updated accuracies, and allocates video streams based on the video stream allocation table, for example, assigning video streams with a higher probability of target action to video processing units with higher average action recognition accuracy.
[0084] Furthermore, in step 404, before assigning one of the first multiple video streams to one of the video processing units 608, the assignment unit 606 is It may be determined whether the importance of one of the first multiple video streams is higher than the importance of each of the other video streams in the first multiple video streams. The assignment in step 404 is then based on such determination, for example, whether the assigned video stream in which the object's action should be recognized has the highest (or lowest) importance.
[0085] Subsequently, for those video processing units 608 that receive video streams by assignment, the video processing unit 608 may then run an action recognition model to detect and recognize the target action of an object from the video stream assigned to it. If a target action is detected and recognized, the video processing unit 608 may then send an alert to the information display unit 610.
[0086] The information display unit 610 receives an alert based on the target action recognition result output from the video processing unit 608 and displays the result (e.g., probability) to the user.
[0087] Figure 7 shows a block diagram 700 illustrating a lightweight estimator 704 and an action recognition probability estimation unit 706 of a probability calculation unit according to one embodiment of the present disclosure. Each video stream 702 received from the video stream input unit is assigned to the lightweight estimator 704 of the probability calculation unit to identify one or more features. Examples of features include, but are not limited to, the number of people detected, the positions of the people, the average distance, the attributes of the people (e.g., gender, age, height), the timestamp of the video, the change in the detected bounding box size, the persistence of the detection (duration, frequency, start timestamp and end timestamp of feature detection), and the positional distribution of various objects in the video / frame. The features thus identified are then sent to the assignment unit, which assigns the video stream 702 to the video processing unit, which processes and recognizes the actions of the objects, before being used by the action recognition probability estimation unit 706 to estimate and calculate the probability that an action (e.g., fight, ride, run) is recognized from the video stream.
[0088] Figure 8 shows a block diagram 800 illustrating the communication between an allocation unit 802 and a probability calculation unit 804 and a multiplexed video processing unit 806 (VPU1, VPU2, ..., VPUN) according to one embodiment of the present disclosure. The allocation unit may receive multiplexed video streams (video stream 1, video stream 2, ..., video stream M). In this case, the number of video streams is M. The allocation unit 802 may also receive VPU information, such as action recognition accuracy, from the video processing unit 806. In this case, the number of VPUs is N, where N is less than M. Furthermore, the allocation unit 802 may receive from the probability calculation unit 804 the probability of detecting and recognizing a target action from each of the video streams. The allocation unit 802 may allocate each video stream to a VPU to process and recognize the target action based on the received probabilities and VPU information. In this embodiment, the allocation unit further comprises a segment-to-pod matching subunit 803 for dividing each video stream into video segments (segment 1, segment 2, segment M) and assigning each video segment to a VPU. In this case, the number of video segments is M, and the number of VPUs N is less than M. Based on the received probability of recognizing a target action from the video segment and VPU information, the allocation unit 802 assigns segment 2 to VPU 1, segment 1 to VPU 2, and segment M to VPUN.
[0089] Figure 9 shows a flowchart 900 illustrating the process overview of a video stream input unit, a probability calculation unit, and an assignment unit of an action recognition device according to one embodiment of the present disclosure.
[0090] The video stream input unit 902 receives a first set of video streams (stream 1, stream 2, stream 3, ..., stream M), performs a stream filtering function to filter out irrelevant video streams (at different locations), and then transmits the first set of video streams to the probability calculation unit 904. The video stream input unit 902 may also transmit basic camera information such as video segment length, video resolution, frame rate, and importance of video streams (not shown) to the assignment unit.
[0091] In the probability calculation unit 904, a lightweight estimator is assigned to each video stream to process it and calculate the probability of recognizing a target action from the video stream. In this case, as shown in Table 904a, video streams 1, 2, 3, ..., M are calculated to have probabilities of recognizing a target action of 0.82, 0.63, 0.69, ..., and 0.75, respectively.
[0092] The allocation unit 906 receives the video stream and its probability, and allocates it to the video processing units (VPU1, VPU2, VPU3, ..., VPUN). The allocation may also be based on the probability calculation unit 904, the video stream input unit 902, and the allocation unit 906, respectively, as well as the basic camera information and VPU information.
[0093] VPU information, including the average precision of the action recognition model performed by each VPU, is received from the processing unit. In this case, as shown in Table 908a, VPU1, VPU2, VPU3, ..., VPUN are calculated to have precisions of 0.73, 0.68, 0.56, ..., and 0.71, respectively. Such precisions may be ranked to form an allocation table indicating which VPUs should be given the highest / lowest priority to allocate to the video stream.
[0094] For example, the allocation unit 906 may perform stream-versus-VPU matching using a scheduling algorithm to achieve optimal or better recognition accuracy. For instance, high-probability streams are considered with higher importance and therefore allocated to VPUs with better model accuracy. Alternatively, high-probability streams may be allocated to VPUs with lower accuracy because the target action is already easily recognizable from the stream and the VPU requirements for recognizing the target action are not as stringent.
[0095] Figure 10 shows a flowchart 1000 illustrating a Markov decision process used to assign each of a first plurality of video streams to one of a second plurality of VPUs according to one embodiment of the present disclosure. The Markov decision process may be performed to assign the video streams to the VPUs based on the following equation (1):
number
number
[0096] Figure 11 shows a flowchart 1100 illustrating the training process of a lightweight classifier or estimator according to one embodiment of the present disclosure. First, a training video stream 1102 having known actions (including physical, motion / behavioral, object features, and combinations thereof that lead to the recognition of known actions) is fed to a lightweight classifier network for detecting and recognizing actions(s) from the training video stream for training a deep learning model. The lightweight classifier network may have pre-configured or existing weight parameters for each physical, motion / behavioral, or object feature, or each combination thereof. Actions(s) may be detected and recognized by a processing unit (not shown) based on several physical, motion / behavioral, and object features identified from the training video stream. The detected and recognized actions(s) may then be compared and matched against known actions(s) to verify the accuracy of the detection and recognition. If, for example, a processing unit receives a signal indicating that the detected and recognized action(s) do not match any known action(s), the existing weight parameters for each of the identified physical, motion / behavioral, object features and combinations (and / or other physical, motion / behavioral, object features and combinations that contribute to the recognition of known actions) are updated so that the known actions can be accurately detected and recognized.
[0097] Once the lightweight classifier network 1104 can identify physical, motion / behavioral, object features, and combinations thereof, and recognize known actions in all training video streams, it is deployed as a trained lightweight classifier 1106 in a probability calculation unit to estimate the probability that actual video streams 1108, 1110 will be received from the video stream input unit. In one example, the estimation may be a similarity estimation, where the similarity between the physical, motion / behavioral, and object features identified from the actual video streams 1108, 1110 and those from each training video stream is calculated and used to determine the probability 1112 of the (known) action to be recognized from the actual video stream. The estimated probabilities are transmitted to an assignment unit (not shown) to assign the actual video stream to a processing unit for action recognition.
[0098] Figure 12 shows a flowchart 1200 illustrating a process for calculating and estimating the probability of a target action from a video stream according to one embodiment of the present disclosure. In this embodiment, 11 different people are detected from the video, as shown in the rectangular boxes within video frame 1202. The distance between each person and the others (e.g., the closest person) is calculated based on the distance between each box of the person and the others. By analyzing the distance values and patterns in step 1204, the probability of a target action (e.g., wait, fight) can be estimated.
[0099] Figure 13 shows a schematic diagram of an exemplary computing device 1300, which will be referred to hereafter as computer system 1300, and one or more such computing devices 1300 may be used or suitable for use in performing the method of Figure 4 and implementing the apparatus of Figure 5. The following description of computing device 1300 is provided as an example only and is not intended to be limiting.
[0100] As shown in Figure 13, the exemplary computing device 1300 includes a processor 1304 for executing software routines. Although a single processor is shown for clarity, the computing device 1300 may also include a multiprocessor system. The processor 1304 is connected to a communication infrastructure 1306 for communicating with other components of the computing device 1300. The communication infrastructure 1306 may include, for example, a communication bus, a crossbar, or a network.
[0101] The computing device 1300 further includes main memory 1308, such as Random Access Memory (RAM), and secondary memory 1310. The secondary memory 1310 may include a storage drive 1312, which may be, for example, a hard disk drive, a solid-state drive, or a hybrid drive, and / or a removable storage drive 1314, which may include a magnetic tape drive, an optical disc drive, a solid-state storage drive (e.g., a USB flash drive, a flash memory device, a solid-state drive, or a memory card). The removable storage drive 1314 reads from and / or writes to a removable storage medium 1318 in a well-known manner. The removable storage medium 1318 may include a magnetic tape, an optical disc, a non-volatile memory storage medium, etc., which is read from and written to by the removable storage drive 1314. As will be understood by those skilled in the art, the removable storage medium 1318 includes a computer-readable storage medium storing computer executable program code instructions and / or data.
[0102] In alternative implementations, the secondary memory 1310 may additionally or alternatively include other similar means for enabling computer programs or other instructions to be loaded into the computing device 1300. Such means may include, for example, a removable storage unit 1322 and interface 1320. Examples of removable storage units 1322 and interface 1320 include program cartridges and cartridge interfaces (such as those found in video game console devices), removable memory chips (such as EPROMs or PROMs) and associated sockets, removable solid-state storage drives (such as USB flash drives, flash memory devices, solid-state drives, or memory cards), and other removable storage units 1322 and interface 1320 that enable the transfer of software and data from the removable storage unit 1322 to the computer system 1300.
[0103] The computing device 1300 also includes at least one communication interface 1324. The communication interface 1324 enables software and data to be transferred between the computing device 1300 and external devices via a communication path 1326. In various embodiments of this disclosure, the communication interface 1324 enables data to be transferred between the computing device 1300 and a data communication network, such as a public data or private data communication network. The communication interface 1324 may also be used to exchange data between different computing devices 1300, such computing devices 1300 forming part of an interconnected computer network. Examples of the communication interface 1324 include a modem, a network interface (such as an Ethernet card), a communication port (such as serial, parallel, printer, GPIB, IEEE1394, RJ45, USB), an antenna with associated circuitry, and the like. The communication interface 1324 may be wired or wireless. The software and data transferred via the communication interface 1324 are in the form of signals that can be electronic signals, electromagnetic signals, optical signals, or other signals that can be received by the communication interface 1324. These signals are provided to the communication interface via the communication path 1326.
[0104] As shown in Figure 13, the computing device 1300 further includes a display interface 1302 for performing operations to render images to an associated display 1330, and an audio interface 1332 for performing operations to play audio content via an associated speaker(s) 1334.
[0105] As used herein, the term “computer program product” may, in part, refer to a removable storage medium 1318, a removable storage unit 1322, a hard disk installed in a storage drive 1312, or a carrier wave that carries software to a communication interface 1324 via a communication path 1326 (wireless link or cable). Computer-readable storage medium refers to any non-temporary, non-volatile tangible storage medium that provides recorded instructions and / or data to the computing device 1300 for execution and / or processing. Examples of such storage mediums include magnetic tapes, CD-ROMs, DVDs, Blu-ray discs, hard disk drives, ROMs or integrated circuits, solid-state storage drives (such as USB flash drives, flash memory devices, solid-state drives or memory cards), hybrid drives, magneto-optical disks, or computer-readable cards such as PCMCIA cards, whether such devices are inside or outside the computing device 1300. Examples of temporary or intangible computer-readable transmission media that may also be involved in providing software, application programs, instructions and / or data to the computing device 1300 include wireless or infrared transmission channels, as well as network connections to other computers or networked devices, and the Internet or intranet, including information recorded in email transmissions and websites, etc.
[0106] The computer program (also referred to as computer program code) is stored in main memory 1308 and / or secondary memory 1310. The computer program may also be received via the communication interface 1324. When such a computer program is executed, it enables the computing device 1300 to implement one or more features of the embodiments discussed herein. In various embodiments, when the computer program is executed, it enables the processor 1304 to implement the features of the embodiments described above. Thus, such a computer program represents the controller of the computer system 1300.
[0107] The software is stored in a computer program product and can be loaded onto the computing device 1300 using a removable storage drive 1314, a storage drive 1312, or interface 1320. The computer program product may be a non-temporary computer-readable medium. Alternatively, the computer program product may be downloaded to the computer system 1300 via a communication path 1326. Once executed by the processor 1304, the software causes the computing device 1300 to perform the operations necessary to execute the method shown in Figure 5 and implement the apparatus shown in Figure 6.
[0108] It should be understood that the embodiment shown in Figure 13 is presented merely as an example to illustrate the operation and structure of the device. Therefore, in some embodiments, one or more features of the computing device 1300 may be omitted. Also, in some embodiments, one or more features of the computing device 1300 may be combined together. In addition, in some embodiments, one or more features of the computing device 1300 may be divided into one or more component parts.
[0109] Those skilled in the art will understand that numerous variations and / or modifications can be made to the disclosure shown in specific embodiments without departing from the spirit or scope of the broadly described disclosure. Accordingly, these embodiments should be considered illustrative and not restrictive in all respects.
[0110] This application claims priority based on Singapore Patent Application No. 10202260150W, filed on 21 November 2022, the disclosure of which is incorporated herein by reference in its entirety.
[0111] <Note> All or part of the exemplary embodiments disclosed above may be described as follows, but are not limited to:
[0112] (Note 1) A method for recognizing the actions of an object from a first set of video streams, wherein the method is The first step is to detect objects from each of multiple video streams, Assigning one of the first multiple video streams to one of the second multiple processing units, and recognizing the object's action based on the accuracy of recognizing the object's action in one of the second multiple processing units, Methods that include... (Note 2) The method involves determining whether the accuracy of recognizing the action of one of the second set of processing units is higher than the accuracy of each of the remaining processing units of the second set of processing units, and the assignment of one of the first set of video streams to one of the second set of processing units for recognizing the action of an object is based on the result of the accuracy determination. The method described in Appendix 1, further including the method described in Appendix 1. (Note 3) The rank of one of the second multiple processing units relative to the remaining processing units of the second multiple processing units is determined based on the accuracy of each in recognizing the action of an object, and the assignment of one of the first multiple video streams to one of the second multiple processing units for detecting the action of an object is based on the rank of one of the second multiple processing units. The method described in Appendix 2, further including the method described in Appendix 2. (Note 4) The process involves determining whether the probability of recognizing an object's action from one of the first multiple video streams is higher than that of each other in the first multiple video streams, and the assignment of one of the first multiple video streams to one of the second multiple processing units for recognizing an object's action is based on the result of the probability determination. The method described in any one of the appendices 1 to 3, further including the method described in any one of the appendices 1 to 3. (Note 5) Identifying one or more features from each of the first multiple video streams, The method involves calculating the probability that an action of an object is recognized from each of the first multiple video streams, based on one or more features identified from each of the first multiple video streams, wherein the determination of the probability that an action is recognized from one of the first multiple video streams is based on the calculated probabilities of one of the first multiple video streams and each of the video streams of the first multiple video streams. The method described in Appendix 4, further including the method described in Appendix 4. (Note 6) Applying weight parameters to each of one or more features in order to calculate the probability that an object's action is recognized from each of the first multiple video streams. The method described in Appendix 5, further including the method described in Appendix 5. (Note 7) Receiving instructions from each of the first multiple video streams regarding whether an object's action was recognized, Updating the weight parameters of one or more features based on instructions, The method described in Appendix 6, further including the method described in Appendix 6. (Note 8) The method according to any one of Annexes 5 to 7, wherein one or more features identified from each of the first plurality of video streams include at least one of the following: the number of objects detected from each of the first plurality of video streams, the location of an object, the position of a part of an object, the relative distance between an object and another object, physical attributes relating to an object, the movement of an object, a timestamp when an object and / or an action of an object was detected, a change in the box size of an object in each of the first plurality of video streams, the duration for which an object and / or an action of an object was detected, and the positional distribution of different objects in each of the first plurality of video streams. (Note 9) The process involves determining whether the importance of one of the first set of video streams is higher than the importance of each of the other video streams in the first set of video streams, and the assignment of one of the first set of video streams to one of the second set of video streams for recognizing the action of an object is further based on the result of the importance determination. The method described in any one of the appendices 1 to 8, further including the method described in the appendices 1 to 8. (Note 10) (i) determining the processing time of one of the second plurality of processing units, (ii) the resolution of one of the first plurality of video streams, and (iii) the duration of one of the first plurality of video streams, wherein the assignment of one of the first plurality of video streams to one of the second plurality of video streams for recognizing the action of an object is based on the result of determining at least one of the processing time, resolution, and duration of one of the first plurality of video streams. The method described in any one of the appendices 1 to 9, further including the method described in any one of the appendices 1 to 9. (Note 11) The method according to any one of the appendices 1 to 10, wherein the number of first multiple video streams is greater than the number of second multiple processing units. (Note 12) Receiving a first set of video streams from a third set of image capture devices, The method described in any one of the appendices 1 to 11, further including the method described in any one of the appendices 1 to 11. (Note 13) A device for recognizing the actions of an object from a first set of video streams, wherein the device is At least one processor, At least one memory containing computer program code, Equipped with, At least one memory and computer program code are provided to the device using at least one processor, Recognize objects from each of the first multiple video streams, Assign one of the first set of video streams to one of the second set of processing units, and have the second set of processing units recognize the action of the object based on the accuracy of recognizing the action of the object. It is configured in such a way. Device. (Note 14) At least one memory and computer program code are provided to the device using at least one processor, at least further, The system determines whether the accuracy of recognizing the action of one of the second set of processing units is higher than the accuracy of each of the remaining processing units in the second set of processing units. Assign one of the first set of video streams to one of the second set of processing units, and have the object's action recognized based on the accuracy determination result. The apparatus described in Appendix 13, configured as follows. (Note 15) At least one memory and computer program code are provided to the device using at least one processor, at least further, Based on the accuracy of each in recognizing the object's action, the rank of one of the second multiple processing units is determined relative to the remaining processing units of the second multiple processing units. Assign one of the first set of video streams to one of the second set of processing units, and have the object's action recognized based on the rank of one of the second set of processing units. The apparatus described in Appendix 14, configured as follows. (Note 16) At least one memory and computer program code are provided to the device using at least one processor, at least further, Determine whether the probability of an object's action being recognized from one of the first set of video streams is higher than the probability of each other in the first set of video streams. Assign one of the first set of video streams to one of the second set of processing units to recognize the object's action based on the results of the probability determination. The apparatus described in any one of appendices 13 to 15, configured as follows. (Note 17) At least one memory and computer program code are provided to the device using at least one processor, at least further, Identify one or more features from each of the first multiple video streams. Based on one or more features identified from each of the first multiple video streams, calculate the probability that an object action is recognized from each of the first multiple video streams. Based on the calculated probabilities of one of the first multiple video streams and each of the video streams in the first multiple video streams, the probability of an action being recognized from one of the first multiple video streams is determined. The apparatus as described in Appendix 16, configured as follows. (Note 18) At least one memory and computer program code are provided to the device using at least one processor, at least further, To calculate the probability that an object's action is recognized from each of the first multiple video streams, a weight is applied to each of one or more features. The apparatus as described in Appendix 17, configured as follows. (Note 19) At least one memory and computer program code are provided to the device using at least one processor, at least further, The system receives instructions from each of the first multiple video streams indicating whether an object action was recognized. Update the weights of one or more features based on the instructions. The apparatus as described in Appendix 18, configured as follows. (Note 20) The apparatus according to any one of Annexes 17 to 19, wherein one or more features identified from each of the first plurality of video streams include at least one of the following: the number of objects detected from each of the first plurality of video streams, the location of an object, the position of a part of an object, the relative distance between an object and another object, physical attributes relating to an object, the movement of an object, a timestamp in which an object and / or an action of an object was detected, a change in the box size of an object in each of the first plurality of video streams, the duration in which an object and / or an action of an object was detected, and the positional distribution of different objects in each of the first plurality of video streams. (Note 21) At least one memory and computer program code are provided to the device using at least one processor, at least further, Determine whether the importance of one of the first set of video streams is higher than the importance of each of the other video streams in the first set of video streams. Based on the results of the importance assessment, one of the first multiple video streams is assigned to one of the second multiple processing units in order to detect the action of the object. The apparatus described in any one of appendices 13 to 20, configured as follows. (Note 22) At least one memory and computer program code are provided to the device using at least one processor, at least further, (i) the processing time of one of the second plurality of processing units, (ii) the resolution of one of the first plurality of video streams, and (iii) the duration of one of the first plurality of video streams, Assign one of the first multiple video streams to one of the second multiple processing units, and recognize the action of an object based on the result of determining at least one of the processing time, resolution, and duration of one of the first multiple video streams. The apparatus described in any one of the appendices 13 to 21, configured as follows. (Note 23) The apparatus according to any one of the appendices 13 to 22, wherein the number of first multiple video streams is greater than the number of second multiple processing units. (Note 24) At least one memory and computer program code are provided to the device using at least one processor, at least further, The first set of video streams are received from the third set of image capture devices. The apparatus described in any one of appendices 13 to 23, configured as follows. (Note 25) A system for recognizing the actions of an object from a first set of video streams, comprising the apparatus described in any one of appendices 13 to 24, and a third set of image capture devices. [Explanation of Symbols]
[0113] 302 Requesting Device 308 Action Recognition Server 340 Linked Servers 342 sensors 350 hosts 502 Image Capture Device 504 Equipment 506 Processors 508 memory 510 Databases 602 Video Stream Input Unit 604 Probability Calculation Unit 606 allocated units 608 Video Processing Units 610 Information Display Unit 802 Assigned Units 804 Probability Calculation Unit 806 Multiplex Video Processing Unit 1302 Display Interface 1304 Processor 1306 Communication Infrastructure 1308 Main Memory 1310 Secondary Memory 1312 Storage Drives 1314 Removable Storage Drive 1318 Removable storage media 1320 Interface 1322 Removable Storage Unit 1324 Communication Interface 1330 Display 1332 Audio Interface 1334 Speakers (multiple speakers possible)
Claims
1. A method for recognizing the actions of an object from a first set of video streams, wherein the method is Detecting the object from each of the first plurality of video streams, Assigning one of the first multiple video streams to one of the second multiple processing units based on the accuracy with which one of the second multiple processing units recognizes the action of the object, The assigned second processing unit recognizes the action of the object in the video stream, Methods that include...
2. A device for recognizing the actions of an object from a first set of video streams, wherein the device is At least one processor, At least one memory containing computer program code, Equipped with, The at least one memory and the computer program code are transmitted to the device using at least one processor, The object is recognized from each of the first plurality of video streams, Based on the accuracy with which one of the second plurality of processing units recognizes the action of the object, one of the first plurality of video streams is assigned to one of the second plurality of processing units. The assigned second processing unit causes the action of the object in the video stream to be recognized. It is configured in such a way. Device.