Method and apparatus for recognizing object actions from a first plurality of video streams - Patent Application 20070122997
By optimizing video stream allocation based on processing unit accuracy and probability, the method addresses inconsistent accuracy in large-scale camera systems, enhancing action recognition efficiency and accuracy.
Patent Information
- Application Number
- JP2025528385
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-21
- Filing Date
- 2023-11-07
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2043-11-07
AI Technical Summary
In large-scale camera systems with limited action recognition models, inconsistent action recognition accuracy occurs due to differing models and platforms across video processing units, leading to uneven load distribution and inefficient resource utilization.
A method and system that considers the accuracy and probability of each processing unit to optimally allocate video streams, ensuring high-importance streams are assigned to high-accuracy units, using a probability calculation and assignment mechanism.
Enhances action recognition accuracy and efficiency by optimizing load balancing across multiple video streams, ensuring accurate and timely detection of target actions.
Smart Images

Figure 2026503829000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to object action recognition methods, devices, and systems, and more particularly to methods, devices, and systems for object action recognition from multiple video streams via multiple processing units. [Background technology]
[0002] For example, there is an increasing demand for processing multiple video streams to detect objects and recognize actions of the objects. Typically, these multiple video streams need to be assigned to action recognition models to perform action recognition and output recognition results (e.g., running, riding a bicycle, two people fighting, etc.). Typically, when a camera / video system is large (e.g., over 200 cameras / video streams) but the number of action recognition models (or the number of processing units running the models) is limited, load balancing is needed to ensure that the task of action recognition is distributed across limited resources and avoid unevenly overloading some action recognition models (or processing units running the models) while leaving other action recognition models (or processing units running the models) idle, thereby making processing more efficient. Summary of the Invention [Problem to be solved by the invention]
[0003] However, because the models and platforms used by different video processing units may differ, users may obtain different action recognition accuracy results even when the same input video stream is assigned to the video processing units. Considering the load distribution across multiple video processing units, the action recognition accuracy results are inconsistent and strongly depend on the assignment.
[0004] Therefore, to address issues and limitations in load balancing and achieve optimal action recognition accuracy and efficiency, there is a need to develop a method, apparatus, and system for recognizing actions of objects from multiple video streams via multiple processing units, for example, by considering the action recognition accuracy of each processing unit and / or the recognition probability of actions from the video streams.
[0005] Furthermore, other desirable features and characteristics will become apparent from the following detailed description and the appended claims, taken in conjunction with the accompanying drawings and this background of the present disclosure. [Means for solving the problem]
[0006] In a first aspect, the present disclosure provides a method for recognizing an action of an object from a first plurality of video streams, the method including: detecting an object from each of the first plurality of video streams; and assigning one of the first plurality of video streams to one of a second plurality of processing units, and recognizing the action of the object based on an accuracy of the one of the second plurality of processing units in recognizing the action of the object.
[0007] In a second aspect, the present disclosure provides an apparatus for recognizing an action of an object from a first plurality of video streams, the apparatus comprising at least one processor and at least one memory containing computer program code, the at least one memory and the computer program code configured to: use the at least one processor to cause the apparatus to at least recognize an object from each of the first plurality of video streams; assign one of the first plurality of video streams to one of a second plurality of processing units; and recognize the action of the object based on an accuracy of one of the second plurality of processing units in recognizing the action of the object.
[0008] In a third aspect, the present disclosure provides a system for recognizing an action of an object from a first plurality of video streams, comprising an apparatus according to the second aspect and a third plurality of image capture devices. [Effects of the Invention]
[0009] Further benefits and advantages of the disclosed embodiments will become apparent from the specification and drawings. Benefits and / or advantages may be obtained individually by various embodiments and features of the specification and drawings, and not all need be provided to obtain one or more of such benefits and / or advantages. [Brief explanation of the drawings]
[0010] Embodiments of the present disclosure will be better understood and readily apparent to those skilled in the art from the following description, given by way of example only, in conjunction with the drawings in which: [Figure 1] FIG. 1 is a schematic diagram outlining a process for recognizing object actions from an input video stream. [Figure 2] FIG. 1 is a schematic diagram illustrating a process for recognizing object actions from multiple input video streams using a conventional load balancer. [Figure 3] FIG. 1 is a block diagram of a system for recognizing object actions from multiple video streams according to various embodiments of the present disclosure. [Figure 4] 1 is a flowchart illustrating a method for recognizing an action of an object from a first plurality of video streams according to various embodiments of the present disclosure. [Figure 5] 5 is a block diagram illustrating a system 500 for recognizing an action of an object from a first plurality of video streams, according to various embodiments of the present disclosure. [Figure 6] FIG. 6 is a block diagram illustrating various components of the action recognizer of FIG. 5 and the process flow between them according to one embodiment of the present disclosure. [Figure 7]FIG. 2 is a block diagram illustrating a lightweight estimator of a probability computation unit and an action-recognition probability estimation unit according to one embodiment of the present disclosure. [Figure 8] FIG. 2 is a block diagram illustrating an allocation unit and its communication with a probability computation unit and multiple video processing units, according to one embodiment of the present disclosure. [Figure 9] 1 is a flowchart illustrating an overview of the processes of a video stream input unit, a probability calculation unit, and an allocation unit of an action recognizer according to one embodiment of the present disclosure. [Figure 10] 10 is a flowchart illustrating a Markov Decision Process used to assign each of a first plurality of video streams to one of a second plurality of VPUs, according to one embodiment of the present disclosure. [Figure 11] 11 is a flowchart 1100 illustrating a process for training a lightweight classifier or estimator, according to one embodiment of the present disclosure. [Figure 12] 1 is a flowchart illustrating a process for calculating and estimating the probability of a target action from a video stream according to one embodiment of the present disclosure. [Figure 13] 7 is a schematic diagram of an exemplary computing device suitable for use in performing the method of FIG. 5 and implementing the apparatus of FIG. 6. DETAILED DESCRIPTION OF THE INVENTION
[0011] <Terminology> Object—An object can be a person, a pet, a vehicle, an object, an item, a device, a pillar, furniture, or anything stationary or moving. An object can be animate or inanimate. For living or biological objects such as people and pets, the object can typically be detected based on appearance features, body parts, physical characteristics, object movement, or a combination thereof. Examples of appearance features of an object (person) include the relative position, size, shape, and / or contour of the eyes, nose, cheekbones, chin, and jaw, as well as iris pattern, skin color, hair color, or a combination thereof. Characteristics include physical attributes such as height, build, body type, body proportions, limb length, hair color, skin color, clothing, possessions, and other similar characteristics or combinations.
[0012] Movement includes behavioral characteristics such as body movement, limb position, direction of movement, speed of movement, walking pattern, the way the object stands, moves or speaks, changes in physical attributes when interacting with other objects, other similar characteristics or combinations. For inanimate objects, the object can typically be based on movement speed, movement characteristics / patterns, and changes in physical attributes when interacting with other objects.
[0013] In various embodiments of the present disclosure, an object may be detected by capturing an image of the object with an image capture device, such as a camera, and identifying the object based on its appearance features, body parts, body characteristics, and / or the object's movement in the image. In various embodiments below, the term "sensor" may be used to refer to the image capture device.
[0014] In various embodiments, the object at the time of detection, and the sensor-acquired data used to identify and detect the object and associated with the object's identification and detection, are assigned an object identifier for subsequent identification and tracking of the object. When the object is subsequently identified based on the same or other appearance features, body parts, body characteristics, object motion, or a combination thereof, it is assigned the same object ID.
[0015] Action—An object's action may refer to a type of activity performed by an object that can be recognized and classified from an image or video stream based on a set of physical features (e.g., appearance features, body parts, body characteristics) and / or motor / behavioral features (e.g., movement, motion) of the object identified from the video stream. Examples of actions include sitting, talking, running, jumping, biking, fighting, and stealing.
[0016] In one example, the action of the object is recognized based on the same or similar action(s) of the same or similar object(s) previously recognized and stored in a database. In another example, the action of the object is recognized based on a set of physical features (e.g., appearance features, body parts, physical characteristics) and / or movement / behavior features of both the object and another object in the video stream, such as fighting and throwing of the object.
[0017] In the various embodiments below, the action of the object to be recognized from the video stream may be referred to as a "target action."
[0018] Video Stream—Video stream refers to the continuous transmission or input of video or image files. The video or images may be generated by a processor in association with an image capture device or may be retrieved from a database. In one example, the processor and database may be connected to a server. The transmission or input of the video or image files may be wired or wireless (e.g., via NFC communication, Wi-Fi communication, Bluetooth, etc.) or over a network (e.g., the Internet).
[0019] Processing Unit—A processing unit refers to a processor configured to process a video or image file to recognize object actions using one or more action recognition models / algorithms stored in a memory accessible by the processor. Additionally, the processing unit may be further configured to process the video or image file to detect objects using one or more object detection models / algorithms stored in the memory prior to the action recognition process.
[0020] Accuracy—Accuracy in recognizing an object's action refers to the accuracy of a processing unit in recognizing an object's action from a sequence of the object's motion using one or more action recognition models / algorithms stored in a memory accessible by the processing unit. In one implementation, additional information and features that affect the performance of the processing unit, such as processing time, video resolution, and video duration, may be used to determine and adjust the accuracy of the processing unit or the allocation of video streams to the processing unit to detect and recognize the object's action in the video stream.
[0021] Rank—The rank of a processing unit is determined based on its accuracy in recognizing object actions compared to the accuracy of other processing units in recognizing object actions. A higher rank indicates that the processing unit is more accurate in recognizing object actions than other processing units with lower ranks. In one embodiment, such ranks are used for video stream allocation. For example, if a video stream has higher importance, the video stream may be assigned to a processing unit with a higher rank that has a higher accuracy in recognizing object actions so that actions can be accurately recognized from the high-importance video stream; alternatively or additionally, a video stream with a lower probability of detecting and recognizing object actions may also be assigned to a processing unit with a higher rank that has a higher accuracy in recognizing object actions, resulting in easier identification and recognition of actions from the video stream and less processing power or time. Conversely, if a video stream has lower importance, the video stream may be assigned to a processing unit with a lower rank that has a lower accuracy in recognizing object actions so that a processing unit with a higher rank can be used to process the more important video stream. Alternatively or additionally, video streams with a high probability of detecting and recognizing object actions may be assigned to lower-rank processing units with higher accuracy in recognizing object actions, which can better identify and recognize actions from the video stream.
[0022] Probability - The probability relates to the likelihood that an object's action will be detected and recognized from the video stream and is determined based on a set of physical and / or motion / behavioral characteristics of the object, such as, but not limited to, the number of objects, the location of the object, the position of parts of the object, the relative distance between the object and another object, the timestamp at which the object and / or object's action was detected, physical attributes (or changes in physical attributes) related to the object, changes in the size of the object in the video stream, the duration at which the object and / or object's action was detected, the distribution of positions of different objects in the video stream, or combinations thereof, as well as other object features identified from the video stream.
[0023] In one implementation, different sets of physical, motion / behavior, and object features may be relied upon to detect different actions of objects and their probabilities of being detected from a video stream. For example, features such as the number of people, arm positions, object collisions, and frequency of arm extensions, frequency of object collisions, etc., can form a set of features for detecting fighting actions, such that when those features are identified from a video stream, for example, when two people are sparring and bumping into each other in a short period of time, the probability of detecting the fighting action is higher compared to other actions (e.g., riding a bicycle, sitting, stealing).
[0024] Weight Parameter - A weight parameter is a weight in a deep learning model that is assigned and applied to a physical, motion / behavior, or object feature, or a combination of two features identified from a video stream, in order to calculate the probability of detecting one or more actions of an object from the video stream. Such a weight parameter correlates with the emphasis or priority given to the feature (or combination of features) in the calculation of the probability.
[0025] In one implementation, weight parameters assigned to and applied to physical, motion / behavior, or object features, or combinations of the two features, are updated by training a deep learning model. During training, multiple video streams with known actions (including physical, motion / behavior, object features, and combinations thereof that lead to the recognition of the known actions) are used to check whether such known actions can be accurately recognized from such video streams. If the actions recognized from such video streams do not match the known actions, an indication may be generated. Upon receiving such an indication, weight parameters for physical, motion / behavior, object features, and combinations identified from the video streams (and / or other physical, motion / behavior, and object features and combinations that contribute to the recognition of the known actions) may be updated accordingly so that the known actions are accurately recognized from such video streams using the updated weight parameters.
[0026] Importance - Importance relates to how important it is to correctly recognize the action of an object from a video stream. In general, video streams with higher importance are assigned to processing units with higher ranks (higher accuracy), which makes it easier to accurately recognize the action of an object.
[0027] Exemplary Embodiments Embodiments of the present disclosure will now be described, by way of example only, with reference to the following drawings in which like reference numbers and letters indicate like or equivalent elements:
[0028] Some portions of the description which follow are presented explicitly or implicitly in terms of algorithms and functional or symbolic representations of operations on data within a computer memory. These algorithmic descriptions and functional or symbolic representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is conceived to be a self-consistent sequence of steps leading to a desired result. These steps require physical manipulations of physical quantities, such as electrical, magnetic, or optical signals capable of being stored, transferred, combined, compared, and otherwise manipulated.
[0029] Unless otherwise indicated, and as will become apparent hereinafter, descriptions utilizing terms such as "receive," "calculate," "determine," "update," "generate," "initialize," "output," "retrieve," "identify," "distribute," "authenticate," and the like throughout this specification will be understood to refer to the actions and processes of a computer system or similar electronic device that manipulates and converts data represented as physical quantities within the computer system into other data similarly represented as physical quantities within the computer system or other information storage, transmission, or display device.
[0030] This specification also discloses apparatus for performing the method operations. Such apparatus may be specially constructed for the required purposes, or may comprise a computer or other device selectively activated or reconfigured by a computer program stored in the computer. The algorithms and displays presented herein are not inherently related to any particular computer or other apparatus. Various machines may be used with programs in accordance with the teachings herein. Alternatively, construction of more specialized apparatus to perform the required method steps may be appropriate. The structure of the computer will be apparent from the description below.
[0031] In addition, this specification also implicitly discloses a computer program in that it is obvious to one skilled in the art that the individual steps of the methods described herein may be implemented by computer code. The computer program is not intended to be limited to any particular programming language and its implementation form. It will be understood that various programming languages and their coding can be used to implement the teachings of the present disclosure contained herein. Furthermore, the computer program is not intended to be limited to any particular control flow. There are many other variations of the computer program that can use different control flows without departing from the spirit or scope of the present disclosure.
[0032] Furthermore, one or more of the steps of the computer program may be performed in parallel rather than sequentially. Such a computer program may be stored on any computer-readable medium. The computer-readable medium may include a storage device such as a magnetic or optical disk, a memory chip, or other storage device suitable for interfacing with a computer. The computer-readable medium may also include a hardwired medium, such as exemplified by the Internet system, or a wireless medium, such as exemplified by the GSM mobile phone system, the Long Term Evolution (LTE) system, and the 5G mobile network system. When loaded and executed on such a computer, the computer program effectively produces an apparatus that implements the steps of the preferred method.
[0033] Various embodiments of the present disclosure relate to methods and apparatus for recognizing object actions from multiple video streams. Those skilled in the art will appreciate that such apparatus and image capture devices may be implemented as part of a system to provide the same technical effect.
[0034] Figure 1 shows a schematic diagram 100 that outlines a process for recognizing object actions from an input video stream 102. The input video stream is fed through a video buffer into multiple video streams (indexed from 1 to M within queue 106), and the multiple video streams 106 need to be assigned to one or more processing units to execute an action recognition model 104 (indexed from 1 to N) for recognizing object actions (such as running, riding a bicycle, people fighting) and output recognition results from the video streams 106. Here, in this example, each action recognition model is assumed to be executed by a processing unit or a processor. However, it is understood that the processing unit or processor may be configured to execute two or more action recognition models.
[0035] For example, problems occur in a large-scale camera system having more than 200 camera streams (and input video streams) when the number of action recognition models (or particularly the number of processing units) is limited, i.e., N < M. Additionally, since the models and platforms used in the processing units may be different, assigning the same video stream to different video processing units may result in differences in action recognition, potentially affecting the accuracy and reliability of the load balancer.
[0036] FIG. 2 shows a schematic diagram 200 illustrating a process for recognizing an object action from multiple input video streams 202 using a conventional load balancer 204. The conventional load balancer 204 is configured to allocate the multiple input video streams 202 to three video processing units (VPUs) 206a, 206b, and 206c to recognize a target action. Assuming that a target action appears in only one of the input video streams, the conventional load balancer 204 does not consider VPU accuracy when allocating the video streams 202. Therefore, the video stream in which the target action appears may be allocated to a low-accuracy VPU (e.g., VPU 206a). As a result, the VPU cannot detect and recognize the target action, and a recognition result indicating that the target action was not detected from the multiple input video streams 202 is output. Therefore, if the target action can be detected by other VPUs with higher accuracy, such as VPUs 206b and 206c, the user may not achieve optimal action recognition accuracy.
[0037] Therefore, the objective is to consider how to allocate multiple video stream loads to a limited number of action recognition models. In this disclosure, a new action recognition method, apparatus, and system are provided to address such issues and limitations in load balancing between multiple video streams from a large-scale camera system and limited action recognition models (processing units) to achieve optimal action recognition accuracy and efficiency.
[0038] In one implementation, information such as the action recognition accuracy of a processing unit and the probability of detecting a target action in each video stream are taken into account in the new action recognition method, apparatus, and system. For example, as shown in Figure 1, video streams 106 are supplied to an action recognizer 108 comprising a probability calculation unit(s) 108a configured to calculate a probability of recognizing an action of an object from each video stream, and an assignment unit 108b configured to assign each video stream to an action recognition model.
[0039] FIG. 3 illustrates a block diagram of a system 300 for recognizing object actions from multiple video streams, according to various embodiments of the present disclosure.
[0040] The system 300 includes a requesting device 302, an action recognition server 308, a cooperation server 340, hosts 350A to 350N, and sensors 342A to 342N.
[0041] The requesting device 302 communicates with the action recognition server 308 and / or the coordination server 340 via connections 316 and 321, respectively. The connections 316 and 321 may be wireless (e.g., via NFC communication, Wi-Fi communication, Bluetooth, etc.) or over a network (e.g., the Internet). The connections 316 and 321 may be of a network (e.g., the Internet).
[0042] The action recognition server 308 further communicates with a coordination server 340 via a connection 320. The connection 320 may be over a network (e.g., a local area network, a wide area network, the Internet, etc.). In one arrangement, the action recognition server 308 and the coordination server 340 are combined, and the connection 320 may be an interconnected bus.
[0043] The linking server 340 then communicates with the hosts 350A-350N via the respective connections 322A-322N. The connections 322A-322N may be a network (eg, the Internet).
[0044] The hosts 350A to 350N are servers. The term host is used herein to distinguish between the hosts 350A to 350N and the linked server 340. The hosts 350A to 350N are collectively referred to as hosts 350 herein, and host 350 refers to one of the hosts 350. The host 350 may be combined with the linked server 340.
[0045] In one example, the hosts 350 may be managed by security personnel of an entity, and the coordination server 340 is a central server that coordinates the hosts 350 and determines which of the hosts 350 are to transfer data or retrieve data, such as image capture.
[0046] The sensors 342A-342N are connected to the coordination server 340 or the action recognition server 308 via respective connections 344A-344N or 346A-346N. The sensors 342A-342N are collectively referred to herein as sensors 342. The connections 344A-344N are collectively referred to herein as connections 344, where connection 344 refers to one of the connections 344. Similarly, the connections 346A-346N are collectively referred to herein as connections 346, where connection 346 refers to one of the connections 346. The connections 344 and 346 may be wireless (e.g., via NFC communication, Wi-Fi communication, Bluetooth, etc.) or over a network (e.g., the Internet). The sensor 342 may be one of an image capture device, an object tracking device, a video capture device, a motion sensor, and a temperature sensor, and may be configured to send input to at least one of the action recognition servers 308 depending on its type.
[0047] In an exemplary embodiment, each of devices 302 and 342, and servers 308, 340, and 350, provides an interface that enables communication with other connected devices 302 and 342 and / or servers 308, 340, and 350. Such communication is facilitated by an application programming interface (API). Such an API may be part of a user interface that may include a graphical user interface (GUI), a web-based interface, a programmatic interface such as a set of application programming interfaces (APIs) and / or remote procedure calls (RPCs) that correspond to interface elements, a messaging interface in which interface elements correspond to messages of a communication protocol, and / or any suitable combination thereof.
[0048] Use of the term "server" herein can refer to a single computing device with processors that work together, or to multiple interconnected computing devices that work together to perform a particular function, i.e., a server may be contained within a single hardware unit or may be distributed among multiple or many different hardware units.
[0049] <Collaboration Server 340> The coordination server 340 is associated with an entity (e.g., a company or organization or a moderator of the service). In one arrangement, the coordination server 340 is owned and operated by the entity that operates the server 308. In such an arrangement, the coordination server 340 may be implemented as part of the server 308 (e.g., a computer program module, a computing device, etc.).
[0050] The collaboration server 340 may also be configured to manage user registration. A registered user has an action-aware account that contains the user's details. The registration step is called onboarding. A user may onboard to the collaboration server 340 using either the requesting device 302 or the host 350.
[0051] It is not necessary to have an action recognition account with the coordination server 340 to access the functionality of the coordination server 340. However, some features are available to registered users. For example, features such as recognizing more complex actions, or actions involving multiple objects, or increasing the maximum number of input video streams may be exclusive to registered users.
[0052] The user onboarding process is performed by the user via one of the requesting devices 302. In one arrangement, the user downloads an application (including an API for interacting with the cooperation server 340) to the sensor 342. In another arrangement, the user accesses a website (including an API for interacting with the cooperation server 340) on the requesting device 302.
[0053] Registration details include, for example, the user's user identifier (ID) or appearance portrait, the user's address, contacts, or other important information, and the sensors 342 authorized to update the action recognition account.
[0054] Once onboarded, users have an action-aware account that remembers all their details.
[0055] <Requesting Device 302> The requesting device 302 is associated with a subject (or requestor) who is the party to an action recognition request originating at the requesting device 302. The requestor may be a public official or an entity's security officer assisting in obtaining the data necessary to detect and recognize a target action (e.g., steal, fight) of a person(s) or object(s) within the entity. The requesting device 302 may be a computing device such as a desktop computer, an Interactive Voice Response (IVR) system, a smartphone, a laptop computer, a Personal Digital Assistant Computer (PDA), a mobile computer, a tablet computer, etc.
[0056] In one example arrangement, the requesting device 302 is a wristwatch or similar wearable computing device equipped with a wireless communication interface.
[0057] <Action Recognition Server 308> The action recognition server 308 is as described above in the terminology section and is configured to recognize the actions of objects from multiple video streams.
[0058] <Host 350> Host 350 is a server associated with an entity (eg, a company or organization) that manages (eg, establishes, administers) object information regarding an object for which an action is recognized.
[0059] In one arrangement, the entities are banks. Accordingly, each entity operates a host 350 to manage resources by that entity. In one arrangement, the host 350 receives an alert signal that a target action has been detected. The host 350 may then be configured to send resources to a location identified by location or camera information included in the alert signal. For example, the host may be configured to capture relevant video or image input for processing.
[0060] In one arrangement, the video stream, detected objects, and recognized actions may be stored and updated in an action recognition account associated with the user. Advantageously, such information is valuable to law enforcement agencies and users, such as security or building management staff, who identify, track, and monitor objects. This reduces the time it takes to view camera footage and recognize object actions.
[0061] <Sensor 342> The sensor 342 is associated with a user associated with the requesting device 302. The sensor 342 may be one of an image capture device, an object tracking device, a video capture device, a motion sensor, and a temperature sensor, and may be configured to send input depending on its type to at least one of the action recognition servers 308. Further details on how sensors may be utilized to recognize actions of an object are provided below.
[0062] 4 shows a flowchart 400 illustrating a method for recognizing an action of an object from a first plurality of video streams according to various embodiments of the present disclosure. Detecting an object from each of the first plurality of video streams occurs in step 402. Assigning one of the first plurality of video streams to one of the second plurality of processing units and recognizing the action of the object based on the accuracy of the second plurality of processing units in recognizing the action of one or more objects occurs in step 404.
[0063] FIG. 5 shows a block diagram illustrating a system 500 for recognizing an action of an object from a first plurality of video streams according to various embodiments of the present disclosure.
[0064] In one example, management of image input and signal input is performed by all image capture devices 502a, 502b. System 500 includes multiple image capture devices 502a, 502c in communication with apparatus 504 (for simplicity, only two image capture devices are shown). In one implementation, apparatus 504 can be generally described as a physical device including at least one processor 506 and at least one memory 508 containing computer program code. The at least one memory 508 and the computer program code, together with the at least one processor 506, are configured to cause the physical device to perform the operations described in FIG. 4. Processor 506 is configured to receive a first plurality of video streams from image capture devices 502a, 502b or retrieve a first plurality of video streams from database 510. Alternatively or additionally, the first plurality of video streams captured by the image capture devices 502a, 502b are stored in a database 510, and the processor 506 is configured to retrieve the first plurality of video streams from the database 510.
[0065] The image capture devices 502a, 502b may be devices such as closed-circuit televisions (CCTVs) that provide various data (camera data) of physical, motion / behavior, and object feature data that can be used by the system to detect objects under object IDs and recognize object actions. In one implementation, the data derived from the image capture devices 502a, 502b may be stored in a memory 508 of the device 504 or in a database 510 accessible by the device 504.
[0066] Additionally, camera data, such as position data regarding the location where the camera is fixed or capturing, and time data, such as video / image timestamps, may be received, stored, and / or retrieved to derive the location and timestamp of an action relative to an object for recognizing the action of the object.
[0067] According to the present disclosure, the apparatus 504 may be configured to communicate with the image capture devices 502a, 502b, a database 510, and multiple processing units (not shown). In one implementation, the processing units may be part of the apparatus 504 and are in communication with the processor 508. Similarly, in one implementation, the database 510 may be part of the apparatus 504.
[0068] The apparatus 504 may receive as input multiple video streams from the image capture device 502 or retrieve them from a database 510. The memory 506 and the computer program code stored in the memory 506 are configured to cause the apparatus 504 or the image capture devices 502a, 502b, using the processor 506, to directly detect objects from each of the multiple video streams. The detection may be based on appearance features, body parts, body characteristics, object motion, or a combination thereof.
[0069] Memory 508 or database 510 may store the accuracy of each of the processing units in recognizing target action(s) of the object (or other similar objects), and processor 506 may retrieve such accuracy from memory 508 or database 510 and assign each of the multiple video streams to a processing unit to process and detect the target action of the object. The detection of the target action may be based on a set of physical features (e.g., appearance features, body parts, body characteristics) and / or motion / behavioral features (e.g., movement, motion) of the object identified from the video stream, for example, by running an action recognition model by the processing unit on the assigned video stream.
[0070] More specifically, the memory 508 and the computer program code stored therein are configured to cause the processor 506 to cause the device 504 to compare the accuracy of each of the processing units in recognizing a particular action(s) of an object (or other similar object) and determine which processing unit has a higher (or highest) accuracy in detecting the action than the other processing unit(s). The determination result is then utilized by the processor 506 to assign the video stream(s) to the processing unit(s) for detecting and recognizing the action(s) of the object. In one implementation, after the accuracy comparison, a ranking table is generated indicating the rank of each processing unit in communication with the device 504, and the processing unit ranks are directly used to assign the video streams, e.g., determine the sequence of processing units (e.g., high-precision processing units to low-precision processing units) to be assigned to the video stream load.
[0071] In one embodiment, memory 508 and the computer program code stored in memory 508 are configured to cause apparatus 504 or image capture devices 502a, 502b, using processor 506, to identify one or more features from each of the multiple video streams. Memory 508 and the computer program code stored in memory 508 are configured to cause apparatus 504, using processor 506, to calculate a probability of a target action of an object (or similar object) recognized from the multiple video streams.
[0072] Further, the memory 508 or the database 510 may store a weighting parameter for each feature identified from the video stream, and the memory 508 and the computer program code stored in the memory 508 are configured to cause the device 504, using the processor 506, to retrieve such weighting parameters from the memory 508 or the database 510 and apply the respective weighting parameters to one or more features identified from the video stream to calculate a probability of a target action of the object.
[0073] The memory 508 and the computer program code stored therein are configured to cause the device 504, using the processor 506, to further compare the respective probabilities of the target action of the object (or similar objects) recognized from the multiple video streams and determine which video stream has a higher (or highest) probability of having the target action recognized by the processing unit than the other video streams. The determination result is then utilized by the processor 506 for assigning the video stream(s) to the processing unit(s). For example, the result may determine a series of video streams (e.g., from high probability to low probability) to be assigned to the processing units.
[0074] Further, the memory 508 and the computer program code stored therein are configured, via the processor 506, to cause a processing unit or device in communication with the apparatus 504 to execute the deep learning model. The apparatus 504 may assign a video stream having known actions to the processing unit to execute an action recognition model to recognize the known actions, and receive an indication of whether the known actions were accurately recognized from the video stream by the processing unit executing the action recognition model. Alternatively, the apparatus 504 may receive action recognition results from the processing unit indicating actions recognized from the video stream by the processing unit executing the action recognition model, and the memory 508 and the computer program code stored therein are configured, via the processor 506, to cause the apparatus 504 to determine whether the action of the object matches the known action and whether the known action was accurately recognized by the processing unit executing the action recognition model. If it is indicated that the action of the object is not recognized from the video stream, the memory 508 and the computer program code stored in the memory 508 are configured to use the processor 506 to cause the device 504 to update weight parameters of features detected by the processing unit from the video stream that contributed to the result of the action recognition, so that the known action can be more accurately recognized from the video stream by the processing unit.
[0075] In an alternative embodiment, memory 508 and the computer program code stored therein are configured to cause device 504, using processor 506, to compare the importance of each of a plurality of video streams and determine which video stream has a higher (or highest) importance from which a target action can be recognized than the other video streams. The determination result is then used by processor 506 to assign video streams to processing units. For example, the result may determine a sequence of video streams (e.g., from high probability to low probability) to be assigned to the processing units.
[0076] 6 shows a block diagram 600 illustrating various components of the action recognizer of FIG. 5 and the process flow between them according to one embodiment of the present disclosure. In this embodiment, the action recognizer 600 (or the processor of the action recognizer 600) comprises a video stream input unit 602, a probability calculation unit 604, an allocation unit 606, a video processing unit 608, and an information display unit 610. As in the action recognition system illustrated in FIG. 1, the video stream input unit 602, the video processing unit 608, and the information display unit 610 may not be part of the action recognizer 600 (or the processor of the action recognizer 600).
[0077] As shown in the exemplary method for recognizing an action of an object in FIG. 4, the action recognizer 600, in operation, Step 402, in which a video stream input unit 602 may detect an object from each of a first plurality of video streams; Step 404, in which the allocation unit may allocate one of the first plurality of video streams to one of the video processing units 608 to recognize the action of the object based on the accuracy of the one of the video processing units 608 in recognizing the action of the object; is configured to execute
[0078] In step 402, the video stream input unit 602 may receive or retrieve the first plurality of video streams, for example from multiple image capture devices or a database (not shown), before detecting objects from each of the first plurality of video streams. The video stream input unit 602 may be used to temporarily buffer the captured video streams and obtain basic camera information (e.g., resolution, frame rate, etc.) for further processing.
[0079] Furthermore, in step 402, the video stream input unit 602 may identify one or more features from each of the first plurality of video streams for further processing by the probability computation unit 604. Alternatively, such an identification step may be performed by the probability computation unit 604 itself.
[0080] Furthermore, in step 402, the video stream input unit 602 may determine the importance of each of the first plurality of video streams for recognizing the action of the object. Alternatively, such a determining step may be performed by the allocation unit 606 itself. Alternatively, the importance of each of the first plurality of video streams is received from an image capture device.
[0081] In step 404, before assigning one of the first plurality of video streams to one of the video processing units 608, the assignment unit 606 may retrieve the average recognition accuracy of the video processing units 608 in recognizing the action of the object from each of the video processing units 608 or from a database (not shown), and determine whether the average recognition accuracy of the processing unit is higher than that of each of the remaining video processing units. The assignment in step 404 will be based on the result of such determination. The assignment unit 606 may further determine the rank of each video processing unit relative to the remaining video processing units. Then, the assignment in step 404 is based on the rank of one of the video processing units 608, for example, the assigned video processing unit has the highest (or lowest) accuracy in recognizing the action of the object among all the video processing units 608.
[0082] Also, in step 404, before allocating one of the first plurality of video streams to one of the video processing units 608, the probability calculation unit 604 may calculate a probability that an object action will be recognized from each of the first plurality of video streams based on one or more features identified by itself or the video stream input unit 602, and determine whether the probability that an object action will be recognized from one of the video streams is higher than that of each of the other video streams of the first plurality of video streams. The allocation unit 606 receives the video streams with the calculated probabilities from the probability calculation unit 604. The allocation in step 404 may then be based on the result of such determination, for example, the allocated video stream has the highest (or lowest) probability of detecting and recognizing an object action.
[0083] In one implementation, the allocation unit 606 maintains a video stream allocation table based on the rank of each video processing unit, updates the video stream allocation based on the new rank of the video processing units calculated from their updated accuracy, and allocates video streams based on the video stream allocation table, for example, a video stream with a higher probability of a target action is allocated to a video processing unit with a higher average action recognition accuracy.
[0084] Also, in step 404, before allocating one of the first plurality of video streams to one of the video processing units 608, the allocation unit 608 may determine whether the importance of one of the first plurality of video streams is higher than the importance of each of the other video streams of the first plurality of video streams. The allocation in step 404 is then based on the result of such determination, for example, the assigned video stream in which the action of the object should be recognized has the highest (or lowest) importance.
[0085] Thereafter, for those video processing units 608 that receive the video streams according to the assignment, the video processing units 608 may then execute the action recognition model to detect and recognize the target action of the object from the video streams assigned to it. If the target action is detected and recognized, the video processing units 608 may then send an alert to the information display unit 610.
[0086] The information display unit 610 receives the target action recognition result alert output from the video processing unit 608 and displays the result (eg, probability) to the user.
[0087] FIG. 7 illustrates a block diagram 700 illustrating a lightweight estimator 704 and an action recognition probability estimation unit 706 of the probability calculation unit according to one embodiment of the present disclosure. Each video stream 702 received from the video stream input unit is assigned to the lightweight estimator 704 of the probability calculation unit to identify one or more features. Examples of features include, but are not limited to, the number of detected people, the people's positions, the average distance, the people's attributes (e.g., gender, age, height, etc.), the video's timestamp, the change in the detected bounding box size, the persistence of detection (duration, frequency, start timestamp, and end timestamp of the feature's detection), and the location distribution of various objects within the video / frame. The features thus identified are then used by the action recognition probability estimation unit 706 to estimate and calculate the probability that an action (e.g., fighting, riding, running, etc.) will be recognized from the video stream before being sent to an assignment unit to assign the video stream 702 to a video processing unit for processing and recognizing the object's action.
[0088] FIG. 8 shows a block diagram 800 illustrating an allocation unit 802 and its communication with a probability calculation unit 804 and multiple video processing units 806 (VPU1, VPU2, . . . , VPUN) according to one embodiment of the present disclosure. The allocation unit may receive multiple video streams (video stream 1, video stream 2, . . . , video stream M), where the number of video streams is M. The allocation unit 802 may also receive VPU information, such as action recognition accuracy, from the video processing units 806. In this case, the number of VPUs is N, where N is less than M. Furthermore, the allocation unit 802 may receive a probability of detecting and recognizing a target action from each of the video streams from the probability calculation unit 804. The allocation unit 802 may assign each video stream to a VPU for processing and recognizing the target action based on the received probability and VPU information. In this embodiment, the allocation unit further includes a segment-to-pod matching subunit 803 for dividing each of the video streams into video segments (segment 1, segment 2, segment M) and assigning each video segment to a VPU. In this case, the number of video segments is M, and the number of VPUs N is less than M. Based on the reception probability of recognizing the target action from the video segment and VPU information, the allocation unit 802 assigns segment 2 to VPU1, segment 1 to VPU2, and segment M to VPUN.
[0089] FIG. 9 shows a flowchart 900 outlining the processes of the video stream input unit, the probability calculation unit, and the allocation unit of the action recognizer according to one embodiment of the present disclosure.
[0090] The video stream input unit 902 receives a first plurality of video streams (stream 1, stream 2, stream 3, ..., stream M), performs a stream filter function to filter out irrelevant video streams (different positions), and then transmits the first plurality of video streams to the probability calculation unit 904. The video stream input unit 902 may also transmit basic camera information, such as video segment length, video resolution, frame rate, and importance of the video streams (not shown), to the allocation unit.
[0091] The probability calculation unit 904 assigns a lightweight estimator to each video stream to process the video stream and calculates the probability of recognizing the target action from the video stream. In this case, as shown in Table 904a, the probabilities of recognizing the target action for video streams 1, 2, 3, ..., M are calculated to be 0.82, 0.63, 0.69, ..., 0.75, respectively.
[0092] The video streams and their probabilities are transmitted to the allocation unit 906, which allocates them to the video processing units (VPU1, VPU2, VPU3, ..., VPUN). The allocation may be based on the probabilities, basic camera information, and VPU information received from the probability calculation unit 904, the video stream input unit 902, and the allocation unit 906, respectively.
[0093] VPU information, including the average accuracy of the action recognition models executed by each VPU, is received from the processing unit. In this case, VPU1, VPU2, VPU3, . . . , VPUN are calculated to have accuracies of 0.73, 0.68, 0.56, . . . , and 0.71, respectively, as shown in Table 908a. Such accuracies may be ranked to form an allocation table indicating which VPUs should be given the highest / lowest priority for allocating video streams.
[0094] In one example, the allocation unit 906 may use a scheduling algorithm to perform stream-to-VPU matching to achieve optimal or better recognition accuracy. For example, a high-probability stream is considered with higher importance and therefore assigned to a VPU with better model accuracy. Alternatively, a high-probability stream may be assigned to a VPU with lower accuracy because the target action is already easily recognized from the stream and the VPU's requirements for recognizing the target action are less stringent.
[0095] 10 shows a flowchart 1000 illustrating a Markov Decision Process used to assign each of a first plurality of video streams to one of a second plurality of VPUs, according to one embodiment of the present disclosure. A Markov Decision Process may be performed to assign video streams to VPUs based on the following equation (1):
number
number
[0096] FIG. 11 shows a flowchart 1100 illustrating a training process for a lightweight classifier or estimator according to one embodiment of the present disclosure. First, a training video stream 1102 having known actions (including physical, motion / behavior, object features, or combinations thereof that lead to the recognition of the known actions) is fed to a lightweight classifier network for detecting and recognizing action(s) from the training video stream for deep learning model training. The lightweight classifier network may have preconfigured or pre-existing weight parameters for each physical, motion / behavior, or object feature, or each combination thereof. The action(s) may be detected and recognized by a processing unit (not shown) based on several physical, motion / behavior, and object features identified from the training video stream. The detected and recognized action(s) may then be compared against known action(s) to verify whether the detection and recognition are accurate. If an indication is received, for example from the processing unit, that the detected and recognized action(s) do not match the known action(s), then the existing weight parameters for each of the identified physical, motion / behavior, object features and combinations (and / or other physical, motion / behavior, object features and combinations that contribute to the recognition of the known action) are updated so that the known action can be accurately detected and recognized.
[0097] Once the lightweight classifier network 1104 is able to identify physical, motion / behavior, and object features, and their combinations, and recognize known actions in all training video streams, it is deployed as a trained lightweight classifier 1106 in a probability calculation unit to estimate the probability that actual video streams 1108, 1110 will be received from the video stream input unit. In one example, the estimation may be a similarity estimation, in which the similarity between the identified physical, motion / behavior, and object features from the actual video streams 1108, 1110 and those from each training video stream is calculated and used to determine the probability 1112 of the (known) action being recognized from the actual video stream. The estimated probabilities are transmitted to an allocation unit (not shown) to allocate the actual video streams to processing units for action recognition.
[0098] 12 shows a flowchart 1200 illustrating a process for calculating and estimating the probability of a target action from a video stream, according to one embodiment of the present disclosure. In this embodiment, 11 different people are detected from the video, as shown by rectangular boxes in a video frame 1202. The distance between each person and other people (e.g., the closest person) is calculated based on the distance between the person and each box of the other people. By analyzing the distance values and patterns in step 1204, the probability of the target action (e.g., wait, fight) can be estimated.
[0099] Figure 13 shows a schematic diagram of an exemplary computing device 1300, hereinafter interchangeably referred to as computer system 1300, one or more such computing devices 1300 may be used or suitable for use to perform the method of Figure 4 and implement the apparatus of Figure 5. The following description of computing device 1300 is provided by way of example only and is not intended to be limiting.
[0100] 13, the exemplary computing device 1300 includes a processor 1304 for executing software routines. While a single processor is shown for clarity, the computing device 1300 may also include a multi-processor system. The processor 1304 is connected to a communications infrastructure 1306 for communicating with other components of the computing device 1300. The communications infrastructure 1306 may include, for example, a communications bus, crossbar, or network.
[0101] The computing device 1300 further includes a main memory 1308, such as random access memory (RAM), and a secondary memory 1310. The secondary memory 1310 may include a storage drive 1312, which may be, for example, a hard disk drive, a solid state drive, or a hybrid drive, and / or a removable storage drive 1314, which may include a magnetic tape drive, an optical disk drive, a solid state storage drive (e.g., a USB flash drive, a flash memory device, a solid state drive, or a memory card), or the like. The removable storage drive 1314 reads from and / or writes to a removable storage medium 1318 in a well-known manner. The removable storage medium 1318 may include a magnetic tape, an optical disk, a non-volatile memory storage medium, or the like, which is read from and written to by the removable storage drive 1314. As will be appreciated by one or more skilled in the art, the removable storage medium 1318 includes a computer-readable storage medium having stored thereon computer-executable program code instructions and / or data.
[0102] In alternative implementations, the secondary memory 1310 may additionally or alternatively include other similar means for allowing computer programs or other instructions to be loaded into the computing device 1300. Such means may include, for example, a removable storage unit 1322 and interface 1320. Examples of removable storage units 1322 and interfaces 1320 include program cartridges and cartridge interfaces (such as those found in video game console devices), removable memory chips (such as EPROMs or PROMs) and associated sockets, removable solid-state storage drives (such as USB flash drives, flash memory devices, solid-state drives, or memory cards), and other removable storage units 1322 and interfaces 1320 that allow software and data to be transferred from the removable storage unit 1322 to the computer system 1300.
[0103] The computing device 1300 also includes at least one communication interface 1324. The communication interface 1324 allows software and data to be transferred between the computing device 1300 and external devices via a communication path 1326. In various embodiments of the present disclosure, the communication interface 1324 allows data to be transferred between the computing device 1300 and a data communication network, such as a public or private data communication network. The communication interface 1324 may be used to exchange data between different computing devices 1300, where such computing devices 1300 form part of an interconnected computer network. Examples of the communication interface 1324 may include a modem, a network interface (such as an Ethernet card), a communication port (such as serial, parallel, printer, GPIB, IEEE 1394, RJ45, USB), an antenna with associated circuitry, etc. The communication interface 1324 may be wired or wireless. Software and data transferred through communications interface 1324 are in the form of signals, which may be electronic, electromagnetic, optical, or other signals capable of being received by communications interface 1324. These signals are provided to communications interface via communications path 1326.
[0104] As shown in FIG. 13, the computing device 1300 further includes a display interface 1302 for performing operations to render images on an associated display 1330 and an audio interface 1332 for performing operations to play audio content through associated speaker(s) 1334.
[0105] As used herein, the term "computer program product" may refer, in part, to removable storage medium 1318, removable storage unit 1322, a hard disk installed in storage drive 1312, or a carrier wave that carries software over communications path 1326 (wireless link or cable) to communications interface 1324. A computer-readable storage medium refers to any non-transitory, non-volatile, tangible storage medium that provides recorded instructions and / or data to computing device 1300 for execution and / or processing. Examples of such storage media include magnetic tape, CD-ROM, DVD, Blu-ray disk, hard disk drive, ROM or integrated circuit, solid-state storage drive (such as a USB flash drive, flash memory device, solid-state drive or memory card), hybrid drive, magneto-optical disk, or computer-readable card such as a PCMCIA card, whether such device is internal or external to computing device 1300. Examples of transitory or non-tangible computer-readable transmission media that may also be involved in providing software, application programs, instructions and / or data to computing device 1300 include wireless or infrared transmission channels, as well as network connections to other computers or networked devices, and the Internet or intranets, including email transmissions and information stored on websites and the like.
[0106] Computer programs (also referred to as computer program code) are stored in main memory 1308 and / or secondary memory 1310. Computer programs may also be received via communications interface 1324. When executed, such computer programs enable computing device 1300 to implement one or more features of the embodiments discussed herein. In various embodiments, when executed, the computer programs enable processor 1304 to implement the features of the above-described embodiments. Thus, such computer programs represent controllers of computer system 1300.
[0107] The software may be stored on a computer program product and loaded into the computing device 1300 using the removable storage drive 1314, the storage drive 1312, or the interface 1320. The computer program product may be a non-transitory computer-readable medium. Alternatively, the computer program product may be downloaded to the computer system 1300 via communications path 1326. The software, when executed by the processor 1304, causes the computing device 1300 to perform the operations necessary to execute the method of FIG. 5 and implement the apparatus of FIG. 6.
[0108] It should be understood that the embodiment of Figure 13 is presented merely as an example to illustrate the operation and structure of the apparatus. Thus, in some embodiments, one or more features of computing device 1300 may be omitted. Also, in some embodiments, one or more features of computing device 1300 may be combined together. Additionally, in some embodiments, one or more features of computing device 1300 may be divided into one or more component parts.
[0109] Those skilled in the art will appreciate that numerous variations and / or modifications may be made to the present disclosure as set forth in the specific embodiments without departing from the spirit or scope of the present disclosure as broadly described. The present embodiments are, therefore, to be considered in all respects as illustrative and not restrictive.
[0110] This application claims priority to Singapore Patent Application No. 10202260150W, filed November 21, 2022, the disclosure of which is incorporated herein by reference in its entirety.
[0111] <Additional Notes> All or part of the exemplary aspects disclosed above may be described in the following appendices, but are not limited thereto.
[0112] (Appendix 1) 1. A method for recognizing an action of an object from a first plurality of video streams, the method comprising: detecting an object from each of the first plurality of video streams; assigning one of the first plurality of video streams to one of the second plurality of processing units to recognize an action of the object based on an accuracy of the one of the second plurality of processing units in recognizing the action of the object; A method comprising: (Appendix 2) determining whether an accuracy in recognizing an action of the object of one of the second plurality of processing units is higher than an accuracy of each of the remaining processing units of the second plurality of processing units, and wherein the assignment of one of the first plurality of video streams to one of the second plurality of processing units for recognizing an action of the object is based on a result of the accuracy determination; 2. The method of claim 1, further comprising: (Appendix 3) determining a rank of one of the second plurality of processing units relative to the remaining processing units of the second plurality of processing units based on their respective accuracies in recognizing the action of the object, wherein the assignment of one of the first plurality of video streams to one of the second plurality of processing units for detecting the action of the object is based on the rank of the one of the second plurality of processing units; 3. The method of claim 2, further comprising: (Appendix 4) determining whether a probability of recognizing an action of an object from one of the first plurality of video streams is higher than each other video stream of the first plurality of video streams, and assigning one of the first plurality of video streams to one of the second plurality of processing units for recognizing an action of the object is based on a result of the determination of the probability; 4. The method of any one of claims 1 to 3, further comprising: (Appendix 5) identifying one or more features from each of the first plurality of video streams; calculating a probability that an action of the object will be recognized from each of the first plurality of video streams based on the one or more features identified from each of the first plurality of video streams, wherein determining the probability that an action will be recognized from one of the first plurality of video streams is based on the calculated respective probabilities for one of the first plurality of video streams and each of the video streams of the first plurality of video streams; 5. The method of claim 4, further comprising: (Appendix 6) applying a weighting parameter to each of the one or more features to calculate a probability of recognizing an action of the object from each of the first plurality of video streams. 6. The method of claim 5, further comprising: (Appendix 7) receiving an indication of whether an action of the object was recognized from each of the first plurality of video streams; updating a weight parameter for each of the one or more features based on the instruction; 7. The method of claim 6, further comprising: (Appendix 8) 8. The method of any one of claims 5 to 7, wherein the one or more features identified from each of the first plurality of video streams include at least one of: a number of objects detected in each of the first plurality of video streams, a position of the object, a position of a part of the object, a relative distance between the object and another object, a physical attribute of the object, a movement of the object, a timestamp at which the object and / or action of the object is detected, a change in box size of the object in each of the first plurality of video streams, a duration at which the object and / or action of the object is detected, and a position distribution of different objects in each of the first plurality of video streams. (Appendix 9) determining whether an importance of one of the first plurality of video streams is higher than an importance of each of the other video streams of the first plurality of video streams, and the assignment of one of the first plurality of video streams to one of the second plurality of video streams for recognizing an action of an object is further based on a result of the importance determination; 9. The method of any one of claims 1 to 8, further comprising: (Appendix 10) determining at least one of (i) a processing time of one of the second plurality of processing units, (ii) a resolution of one of the first plurality of video streams, and (iii) a duration of one of the first plurality of video streams, wherein the allocation of one of the first plurality of video streams to one of the second plurality of video streams for recognizing an action of an object is based on a result of determining at least one of the processing time, resolution, and duration of one of the first plurality of video streams; 10. The method of any one of appendices 1 to 9, further comprising: (Appendix 11) 11. The method of any one of claims 1 to 10, wherein the number of the first plurality of video streams is greater than the number of the second plurality of processing units. (Appendix 12) receiving a first plurality of video streams from a third plurality of image capture devices; 12. The method of any one of claims 1 to 11, further comprising: (Appendix 13) 1. An apparatus for recognizing an action of an object from a first plurality of video streams, the apparatus comprising: at least one processor; at least one memory containing computer program code; Equipped with The at least one memory and the computer program code, using the at least one processor, cause the apparatus to at least: Recognizing an object from each of the first plurality of video streams; assigning one of the first plurality of video streams to one of the second plurality of processing units to recognize an action of the object based on an accuracy of the one of the second plurality of processing units in recognizing the action of the object; It is configured as follows: Device. (Appendix 14) At least one memory and computer program code are used to configure the apparatus, using at least one processor, to further include at least: determining whether an accuracy in recognizing the action of the object of one of the second plurality of processing units is higher than an accuracy of each of the remaining processing units of the second plurality of processing units; assigning one of the first plurality of video streams to one of the second plurality of processing units to recognize an action of the object based on a result of the accuracy determination; 14. The apparatus of claim 13, configured to: (Appendix 15) At least one memory and computer program code are used to configure the apparatus, using at least one processor, to further include at least: determining a rank of one of the second plurality of processing units relative to the remaining processing units of the second plurality of second processing units based on their respective accuracy in recognizing the action of the object; assigning one of the first plurality of video streams to one of the second plurality of processing units to recognize an action of the object based on a rank of the one of the second plurality of processing units; 15. The apparatus of claim 14, configured to: (Appendix 16) At least one memory and computer program code are used to configure the apparatus, using at least one processor, to further include at least: determining whether a probability that an action of the object is recognized from one of the first plurality of video streams is higher than a probability from each other video stream of the first plurality of video streams; assigning one of the first plurality of video streams to one of the second plurality of processing units to recognize an action of the object further based on the result of the probability determination; 16. The apparatus of any one of appendixes 13 to 15, configured to (Appendix 17) At least one memory and computer program code are used to configure the apparatus, using at least one processor, to further include at least: identifying one or more features from each of the first plurality of video streams; calculating a probability of recognizing an action of the object from each of the first plurality of video streams based on the one or more features identified from each of the first plurality of video streams; determining a probability that an action is recognized from one of the first plurality of video streams based on the calculated probabilities for one of the first plurality of video streams and each of the video streams of the first plurality of video streams; 17. The apparatus of claim 16, configured to: (Appendix 18) At least one memory and computer program code are used to configure the apparatus, using at least one processor, to further include at least: applying a weighting to each of the one or more features to calculate a probability of recognizing an action of the object from each of the first plurality of video streams; 18. The apparatus of claim 17, configured to: (Appendix 19) At least one memory and computer program code are used to configure the apparatus, using at least one processor, to further include at least: receiving an indication of whether an object action was recognized from each of the first plurality of video streams; updating a weighting of each of the one or more features based on the instructions; 19. The apparatus of claim 18, configured to: (Appendix 20) 20. The apparatus of any one of claims 17 to 19, wherein the one or more features identified from each of the first plurality of video streams include at least one of a number of objects detected in each of the first plurality of video streams, a position of the object, a position of a part of the object, a relative distance between the object and another object, a physical attribute of the object, a movement of the object, a timestamp at which the object and / or action of the object is detected, a change in box size of the object in each of the first plurality of video streams, a duration at which the object and / or action of the object is detected, and a position distribution of different objects in each of the first plurality of video streams. (Appendix 21) At least one memory and computer program code are used to configure the apparatus, using at least one processor, to further include at least: determining whether an importance of one of the first plurality of video streams is greater than an importance of each of the other video streams of the first plurality of video streams; assigning one of the first plurality of video streams to one of the second plurality of processing units for detecting an action of the object further based on a result of the importance determination; 21. The apparatus of any one of clauses 13 to 20, configured to (Appendix 22) At least one memory and computer program code are used to configure the apparatus, using at least one processor, to further include at least: determining at least one of (i) a processing time of one of the second plurality of processing units, (ii) a resolution of one of the first plurality of video streams, and (iii) a duration of one of the first plurality of video streams; assigning one of the first plurality of video streams to one of the second plurality of processing units to recognize an action of the object based on a result of determining at least one of a processing time, a resolution, and a duration of the one of the first plurality of video streams; 22. The apparatus of any one of clauses 13 to 21, configured to (Appendix 23) 23. The apparatus of any one of claims 13 to 22, wherein the number of the first plurality of video streams is greater than the number of the second plurality of processing units. (Appendix 24) At least one memory and computer program code are used to configure the apparatus, using at least one processor, to further include at least: receiving the first plurality of video streams from a third plurality of image capture devices; 24. The apparatus of any one of clauses 13 to 23, configured to (Appendix 25) 25. A system for recognizing object actions from a first plurality of video streams, comprising: an apparatus according to any one of clauses 13 to 24; and a third plurality of image capture devices. [Explanation of symbols]
[0113] 302 Requesting Device 308 Action Recognition Server 340 Collaboration Server 342 Sensors 350 hosts 502 Image Capture Device 504 Equipment 506 processor 508 memory 510 Database 602 Video Stream Input Unit 604 Probability Computation Unit 606 Allocation Units 608 Video Processing Unit 610 Information display unit 802 allocation units 804 Probability Computation Unit 806 Multiplexed Video Processing Unit 1302 Display Interface 1304 processor 1306 Communications Infrastructure 1308 main memory 1310 Secondary Memory 1312 storage drive 1314 Removable Storage Drive 1318 Removable storage media 1320 Interface 1322 Removable Storage Unit 1324 communication interface 1330 Display 1332 Audio Interface 1334 Speaker(s)
Claims
1. 1. A method for recognizing an action of an object from a first plurality of video streams, the method comprising: detecting the object from each of the first plurality of video streams; assigning one of the first plurality of video streams to one of a second plurality of processing units to recognize the action of the object based on an accuracy of the one of the second plurality of processing units in recognizing the action of the object; A method comprising:
2. determining whether the accuracy of the one of the second plurality of processing units in recognizing the action of the object is higher than the accuracy of each of the remaining processing units of the second plurality of processing units, and the assignment of the one of the first plurality of video streams to the one of the second plurality of processing units for recognizing the action of the object is based on a result of the determination of the accuracy; The method of claim 1 further comprising:
3. determining a rank of the one of the second plurality of processing units relative to the remaining processing units of the second plurality of processing units based on the respective accuracies in recognizing the action of the object, wherein an assignment of the one of the first plurality of video streams to the one of the second plurality of processing units for detecting the action of the object is based on the rank of the one of the second plurality of processing units; The method of claim 2 further comprising:
4. determining whether a probability of recognizing the action of the object from the one of the first plurality of video streams is higher than from each other video stream of the first plurality of video streams, and the assignment of the one of the first plurality of video streams to the one of the second plurality of processing units for recognizing the action of the object is based on a result of the determination of the probability; The method of any one of claims 1 to 3, further comprising:
5. identifying one or more features from each of the first plurality of video streams; calculating a probability that the action of the object will be recognized from each of the first plurality of video streams based on the one or more features identified from each of the first plurality of video streams, wherein the determination of the probability that the action will be recognized from the one of the first plurality of video streams is based on the calculated respective probabilities for the one of the first plurality of video streams and each of the video streams of the first plurality of video streams; The method of claim 4 further comprising:
6. applying a weighting parameter to each of the one or more features to calculate the probability that the action of the object is recognized from each of the first plurality of video streams. The method of claim 5 further comprising:
7. receiving an indication of whether the action of the object was recognized from each of the first plurality of video streams; updating the weight parameters for each of the one or more features based on the indication; and The method of claim 6 further comprising:
8. 8. The method of claim 5, wherein the one or more features identified from each of the first plurality of video streams comprises at least one of: a number of objects detected from each of the first plurality of video streams, a position of the object, a position of a part of the object, a relative distance between the object and another object, a physical attribute of the object, a movement of the object, a timestamp at which the object and / or an action of the object is detected, a change in box size of the object in each of the first plurality of video streams, a duration at which the object and / or an action of the object is detected, and a position distribution of different objects in each of the first plurality of video streams.
9. determining whether an importance of the one of the first plurality of video streams is higher than an importance of each of the other video streams of the first plurality of video streams, and the assignment of the one of the first plurality of video streams to the one of the second plurality of video streams for recognizing the action of the object is further based on a result of the determination of the importance; 9. The method of claim 1, further comprising:
10. determining at least one of (i) a processing time of the one of the second plurality of processing units, (ii) a resolution of the one of the first plurality of video streams, and (iii) a duration of the one of the first plurality of video streams, wherein the allocation of the one of the first plurality of video streams to the one of the second plurality of video streams for recognizing the action of the object is based on a result of the determination of the at least one of the processing time, the resolution, and the duration of the one of the first plurality of video streams; 10. The method of claim 1, further comprising:
11. The method of claim 1 , wherein the number of the first plurality of video streams is greater than the number of the second plurality of processing units.
12. receiving the first plurality of video streams from a third plurality of image capture devices; 12. The method of claim 1, further comprising:
13. 1. An apparatus for recognizing an action of an object from a first plurality of video streams, the apparatus comprising: at least one processor; at least one memory containing computer program code; Equipped with The at least one memory and the computer program code are used by at least one processor to cause the device to at least: recognizing the object from each of the first plurality of video streams; assigning one of the first plurality of video streams to one of a second plurality of processing units to recognize the action of the object based on an accuracy of the one of the second plurality of processing units in recognizing the action of the object; It is configured as follows: Device.
14. The at least one memory and the computer program code, using at least one processor, cause the device to further include at least: determining whether the accuracy of the one of the second plurality of processing units in recognizing the action of the object is higher than the accuracy of each of the remaining processing units of the second plurality of processing units; assigning the one of the first plurality of video streams to the one of the second plurality of processing units to recognize the action of the object based on a result of the determination of the accuracy; The device of claim 13 configured to:
15. The at least one memory and the computer program code, using at least one processor, cause the device to further include at least: determining a rank of the one of the second plurality of second processing units relative to the remaining processing units of the second plurality of second processing units based on the respective accuracy in recognizing the action of the object; assigning the one of the first plurality of video streams to the one of the second plurality of processing units to recognize the action of the object based on the rank of the one of the second plurality of processing units; The apparatus of claim 14 configured to:
16. The at least one memory and the computer program code, using at least one processor, cause the device to further include at least: determining whether a probability that the action of the object is recognized from the one of the first plurality of video streams is higher than a probability from each other video stream of the first plurality of video streams; assigning the one of the first plurality of video streams to the one of the second plurality of processing units to recognize the action of the object further based on a result of the determination of the probability; 16. The apparatus of any one of claims 13 to 15, configured to:
17. The at least one memory and the computer program code, using at least one processor, cause the device to further include at least: identifying one or more features from each of the first plurality of video streams; calculating a probability of recognizing the action of the object from each of the first plurality of video streams based on the one or more features identified from each of the first plurality of video streams; determining the probability that the action is recognized from the one of the first plurality of video streams based on the respective calculated probabilities for the one of the first plurality of video streams and each of the video streams of the first plurality of video streams; 17. The apparatus of claim 16, configured to:
18. The at least one memory and the computer program code, using at least one processor, cause the device to further include at least: applying a weighting to each of the one or more features to calculate a probability of recognizing the action of the object from each of the first plurality of video streams; 18. The apparatus of claim 17, configured to:
19. The at least one memory and the computer program code, using at least one processor, cause the device to further include at least: receiving an indication of whether the action of the object was recognized from each of the first plurality of video streams; updating the weighting of each of the one or more features based on the indication; 20. The apparatus of claim 18, configured to:
20. 20. The apparatus of claim 17, wherein the one or more features identified from each of the first plurality of video streams comprises at least one of: a number of objects detected from each of the first plurality of video streams, a position of the object, a position of a part of the object, a relative distance between the object and another object, a physical attribute of the object, a movement of the object, a timestamp at which the object and / or an action of the object is detected, a change in box size of the object in each of the first plurality of video streams, a duration at which the object and / or an action of the object is detected, and a position distribution of different objects in each of the first plurality of video streams.
21. The at least one memory and the computer program code, using at least one processor, cause the device to further include at least: determining whether an importance of the one of the first plurality of video streams is greater than an importance of each of the other video streams of the first plurality of video streams; assigning the one of the first plurality of video streams to the one of the second plurality of processing units for detecting the action of the object further based on a result of the determination of the importance; 21. The apparatus of any one of claims 13 to 20, configured to:
22. The at least one memory and the computer program code, using at least one processor, cause the device to further include at least: determining at least one of (i) a processing time of the one of the second plurality of processing units, (ii) a resolution of the one of the first plurality of video streams, and (iii) a duration of the one of the first plurality of video streams; assigning the one of the first plurality of video streams to the one of the second plurality of processing units to recognize the action of the object based on a result of the determination of at least one of the processing time, the resolution, and the duration of the one of the first plurality of video streams; 22. The apparatus of any one of claims 13 to 21, configured to:
23. 23. The apparatus of claim 13, wherein the number of the first plurality of video streams is greater than the number of the second plurality of processing units.
24. The at least one memory and the computer program code, using at least one processor, cause the device to further include at least: receiving the first plurality of video streams from a third plurality of image capture devices; 24. Apparatus according to any one of claims 13 to 23, configured to:
25. A system for recognizing an action of an object from a first plurality of video streams, the system comprising: an apparatus according to any one of claims 13 to 24; and a third plurality of image capture devices.
Citation Information
Patent Citations
Acquisition method, loading method and selection method of deep learning model
CN112114892A
Multi-functional computer-assisted genoscopy system and method using an optimized
CN115004316A
Vehicle driving assist system with enhanced data processing
US20170330067A1