Methods, apparatus, systems, and programs for allocating hardware resources to process video streams.
By configuring video capture devices to detect target attributes and allocating resources based on interaction levels and neural network layers, the method optimizes hardware resource usage and simplifies deployment for processing multiple video streams.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-13
- Publication Date
- 2026-03-25
AI Technical Summary
Existing systems face inefficiencies in allocating hardware resources for processing multiple video streams from different cameras due to varying application requirements and neural network model complexities, leading to potential resource wastage or increased deployment complexity.
A method and system for allocating hardware resources that involves configuring video capture devices to detect target attributes, selecting appropriate attribute detection modules, and distributing resources based on interaction levels and neural network layers to optimize processing.
This approach optimizes hardware resource usage by aligning resource allocation with specific video analysis tasks, reducing waste and simplifying deployment complexity while meeting diverse application requirements.
Smart Images

Figure 2026509783000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method, apparatus, and system for allocating hardware resources. In particular, the present disclosure relates to a method, apparatus, and system for allocating hardware resources for processing a plurality of video streams generated by a plurality of video capture devices.
Background Art
[0002] Visual analysis, such as person / object classification and action recognition, is an important application accompanying the widespread deployment of smart camera systems. Generally, a video stream executes a visual analysis task / function and extracts to an analysis server that executes one or more neural network models based on hardware resources (such as GPU, CPU, RAM, etc.) from a surveillance camera in order to predict the visual analysis result of an end user.
[0003] For example, when more surveillance cameras are deployed in different locations for visual analysis tasks / functions such as action recognition and object classification, different video streams generated by those cameras need to be processed by different video analysis tasks / functions. If the video streams of all cameras are processed by the same neural network model that executes the same video analysis task / function for recognizing / detecting the same set of target attributes, different application requirements may not be satisfied, or some hardware resources may be wasted. On the other hand, if the video streams of all cameras are processed by different neural network models that execute different video analysis tasks / functions for recognizing / detecting different sets of target attributes, the complexity of deployment becomes extremely high when the neural network models require different libraries / environments.
Summary of the Invention
Problems to be Solved by the Invention
[0004] Therefore, in order to address the problems of visual analysis tasks and optimize hardware resources using flexible neural network model configurations, it is necessary to develop methods, devices, and systems for allocating hardware resources to process multiple video streams generated by multiple video capture devices.
[0005] Furthermore, by combining the following detailed description and the attached claims with the accompanying drawings and the background of this disclosure, other desirable features and characteristics will become apparent. [Means for solving the problem]
[0006] In a first embodiment, the Disclosure provides a method for allocating hardware resources to process a plurality of video streams generated by a plurality of video capture devices, the method comprising: receiving instructions to configure each of the plurality of video capture devices to detect one or more target attributes of objects appearing in the video streams generated by the corresponding video capture devices; selecting an attribute detection module configured to detect the one or more target attributes of the objects in accordance with the instructions; and allocating one or more of the hardware resources to the attribute detection module for processing the video streams and detecting the one or more target attributes of the objects.
[0007] In a second embodiment, the Disclosure provides an apparatus for allocating hardware resources to process a plurality of video streams generated by a plurality of video capture devices, comprising at least one processor and at least one memory containing computer program code, wherein the at least one memory and the computer program code are configured to cause at least one processor to cause the apparatus to receive instructions to configure each of the plurality of video capture devices to detect one or more target attributes of objects appearing in video streams generated by each of the plurality of video capture devices, to cause the apparatus to select an attribute detection module configured to detect the one or more target attributes of the objects in accordance with the instructions, and to cause the apparatus to process the video streams and allocate one or more hardware resources to the attribute detection module for detecting the one or more target attributes of the objects.
[0008] In a third embodiment, the Disclosure provides a system for allocating hardware resources for processing the plurality of video streams, comprising the apparatus according to the second embodiment and a plurality of video capture devices configured to generate a plurality of video streams.
[0009] Further benefits and advantages of the disclosed exemplary embodiments will become apparent from the specification and drawings. Benefits and / or advantages may be obtained individually from the various embodiments and features of the specification and drawings, and it is not necessary for all of them to be provided in order to obtain one or more of such benefits and / or advantages. [Brief explanation of the drawing]
[0010] The exemplary embodiments described herein will be better understood and readily apparent to those skilled in the art, in conjunction with the drawings, based on the description below, which is merely an example. [Figure 1]Figure 1 shows a schematic diagram illustrating the process of allocating hardware resources to process the video stream generated by the camera. [Figure 2] Figure 2 shows a site plan where a video analysis system, including multiple cameras, is located. [Figure 3] Figure 3 shows the process of a conventional video analysis system that uses the same neural network model to detect target attributes of objects from each camera. [Figure 4] Figure 4 shows a block diagram of a system that allocates hardware resources for processing multiple video streams generated by multiple video capture devices, according to various exemplary embodiments of the present disclosure. [Figure 5] Figure 5 shows a diagram illustrating the process of detecting the target attributes of an object from multiple video streams generated by multiple cameras. [Figure 6] Figure 6 shows a flowchart illustrating the process of allocating hardware resources for processing multiple video streams generated by multiple video capture devices, according to various exemplary embodiments of the present disclosure. [Figure 7] Figure 7 shows a block diagram illustrating a system for recognizing the motion of an object from a first set of video streams, according to various exemplary embodiments of the present disclosure. [Figure 8] Figure 8 shows a diagram illustrating the process of detecting the target motion of an object from a video stream generated by a camera, according to a first exemplary embodiment of the present disclosure. [Figure 9] Figure 9 shows a diagram illustrating three different levels of interaction according to an exemplary embodiment. [Figure 10] Figure 10 shows an exemplary graphical user interface according to a first exemplary embodiment of the present disclosure. [Figure 11] Figure 11 shows a diagram illustrating a process for detecting the target motion of an object from a video stream generated by a camera, according to a second exemplary embodiment of the present disclosure. [Figure 12] Figure 12 shows a diagram illustrating a neural network layer according to an exemplary embodiment of the present disclosure. [Figure 13] Figure 13 shows an exemplary graphical user interface according to a second exemplary embodiment of the present disclosure. [Figure 14] Figure 14 shows a schematic diagram of an exemplary computing device suitable for use in carrying out the method in Figure 6 and implementing the apparatus in Figure 7. [Modes for carrying out the invention]
[0011] Glossary Objects – Objects may be people, pets, vehicles, things, items, devices, pillars, furniture, or stationery or any object in motion. Objects may be living or non-living. In the case of living or biological objects, such as people or pets, objects can usually be identified based on external features, body parts, physical characteristics, the actions of the object, or a combination thereof. Examples of external features of an object (person) include the relative position, size, shape and / or contour of the eyes, nose, cheekbones, and jaw, as well as iris pattern, skin color, hair color, or a combination thereof. Features include physical characteristics such as height, build, body type, body proportions, limb length, hair color, skin color, clothing, possessions, and other similar features or combinations thereof.
[0012] Action includes behavioral features such as body movements, limb positions, direction of movement, speed of movement, gait patterns, how an object stands, how it moves, how it speaks, changes in physical characteristics due to interaction with other objects, and other similar characteristics or combinations thereof. In the case of inanimate objects, an object can usually be described based on its speed of movement, movement characteristics / patterns, and changes in physical characteristics due to interaction with other objects.
[0013] In the following various exemplary embodiments, the object may refer to one of the objects identified based on appearance features, a part of the body, physical features, the movement of the object, or a combination thereof from a video stream, that is, a target object, and one or more attributes of such a target object are then detected and identified using an attribute detection module.
[0014] Video stream - A video stream refers to the continuous transmission or input of a video or image file. The video or image may be generated by a processor to which a video capture device is connected or obtained from a database. In one example, the processor and the database may be connected to a server. The transmission or input of the video or image file may be via wired or wireless (e.g., NFC communication, Wi-Fi communication, Bluetooth, etc.) or via a network (e.g., the Internet).
[0015] Hardware resources - Hardware resources refer to processor units or memories such as a CPU, GPU, or RAM used by an attribute detection module or model that stores one or more target attributes of an object displayed in a video stream generated by a video capture device and for detecting them.
[0016] Attribute - An attribute related to an object refers to the movement of the object or the object class.
[0017] Attribute Detection Module - The attribute detection module refers to a specific module configured to execute video analysis tasks / functions such as action recognition or object classification, and configured to execute a model (e.g., a neural network model) configured to detect object attributes from a video stream by the configured video analysis tasks / functions, and connected to the processor of a device or server (e.g., an attribute detection server). When a video stream is processed through the attribute detection module, video analysis results regarding the detection of attributes related to the video analysis tasks / functions of the pre-configured model. Further, prior to the detection of attributes, the attribute detection module or a separate object detection module can be utilized to detect objects appearing in the video stream based on such attributes or other attributes of the objects. In the present disclosure, a specific video analysis task / function may be configured to process a video stream generated by a camera, and for generating video analysis results, an attribute detection module having a model configured to process the video analysis task / function as such is selected to process the video stream generated by the camera.
[0018] In an exemplary embodiment, there are multiple attribute detection modules each requiring different hardware resources to execute video analysis tasks / functions for detecting specific or identical attributes of an object. In the present disclosure, the term "video analysis task" can be used interchangeably with the term "video analysis function".
[0019] In this disclosure, an attribute detection module assigned to run a model configured to handle an action recognition task (e.g., a neural network model) may be called an action recognition module, and the model it runs may be called an action recognition model. Similarly, an attribute detection module assigned to run a model configured to handle an object classification task (e.g., a neural network model) may be called an object classification module, and the model it runs may be called an object classification model. In an implementation, an attribute detection module may be assigned to run one or both models to perform an action recognition task and an object classification task.
[0020] Action – An object's action can refer to the type of activity performed by an object that can be recognized and classified based on a set of physical characteristics (e.g., appearance, body parts, physical features) and / or motor / behavioral characteristics (e.g., movements and actions) of the object identified from a video stream. Examples of actions include sitting, talking, running, jumping, riding a bicycle, fighting, and stealing.
[0021] Object Class – An object class associated with an object refers to the class or category to which an object belongs, based on its external appearance, body parts, physical characteristics, actions, or a combination thereof. Examples of object classes include, but are not limited to, people, adults, children, inanimate objects, devices, furniture, animals, personal belongings, and employees permitted to enter the premises.
[0022] Target Attributes - Target attributes refer to attributes of particular interest that should be detected. In some exemplary embodiments where the attribute of particular interest is behavior (referred to here as “target behavior”), the process of detecting the target attributes of objects appearing in a video stream includes the process of recognizing the target behavior of objects appearing in a video stream. In some other exemplary embodiments where the attribute of particular interest is object class (referred to here as “target object class”), the process of detecting the target attributes of objects appearing in a video stream includes the process of identifying the target object class of objects appearing in a video stream, or classifying objects appearing in a video stream into a target object class.
[0023] Interaction Level – The interaction level related to an action correlates with the number of objects required for relational analysis for the action to be detected and identified by an attribute detection module, which correlates with the amount of hardware resources required by the attribute detection module to detect and identify such action under the interaction level. In this disclosure, actions involving two or fewer objects may be pre-configured as low-interaction-level actions, actions involving three objects may be pre-configured as medium-interaction-level actions, and actions involving four or more objects may be pre-configured as high-interaction-level actions. In this disclosure, the term “interaction level” may be used interchangeably with “sequential interaction level” or “sequential level.”
[0024] For example, the action of riding a bicycle involves a person (object 1) and a bicycle (object 2), so it is pre-configured as a low-interaction level action, and therefore the amount of hardware resources required to analyze and recognize the action is small. On the other hand, gang fighting actions may involve four or more people, so they are pre-configured as a high-interaction level action, and therefore the amount of hardware resources required to analyze and recognize the action is large. In some exemplary embodiments, interaction levels can be associated with a video stream to indicate the amount of hardware resources that need to be allocated to process the video stream. Typically, one or more actions with pre-configured interaction levels are detected from the video stream, and the highest interaction level among the actions detected from the video stream is selected as the interaction level associated with the video stream. As a result, a corresponding amount of hardware resources is allocated to process the video stream to ensure that there is enough hardware resources to detect the actions of all objects from the video stream.
[0025] Processing Neural Network Layer - The processing neural network layer of an attribute detection module refers to a submodule configured to generate results or signals relating to the detection of specific attributes of an object, particularly the object class of an object, either alone or in combination with one or more other neural network processing layers. Hereinafter, the term "processing neural network layer" can be used interchangeably with "processing layer" or "network layer." Traditionally, processing layers specialized for identifying different object classes are arranged in series, and a video stream is processed through each processing layer (each combination of processing layers) to obtain detection results or signals indicating whether an object can be classified into one of the object classes.
[0026] In one implementation, for example, fewer processing layers may be required for object classes that are easy to detect and classify due to their clear characteristics. On the other hand, more processing layers may be required for object classes that are difficult to detect and classify.
[0027] In an alternative implementation, a processing layer dedicated to identifying the most important or relevant object classes in the application context is placed at the beginning of the layer sequence, ensuring that detection results or signals for objects of the important or relevant object classes are generated first. Alternatively, a processing layer dedicated to identifying the most important or relevant object classes in the application context is placed at the end of the layer sequence.
[0028] The number of processing layers correlates with the amount of hardware resources required by the attribute detection module that detects and recognizes such object classes. According to this disclosure, when the detection of a particular target object class is indicated, the video stream is processed up to a processing layer dedicated to identifying such target object class. This minimizes the hardware resources required to run the remaining processing layers.
[0029] Instructions – Instructions may refer to signals or information received from another connected device, server, or processing unit. In one exemplary embodiment, a user of a device or system for detecting target attributes of an object appearing in a video stream generated by a camera can input, select, or instruct a target camera and target attributes via a connected device or server, and such instructions relating to the task of selecting a target camera and detecting target attributes are sent to the device or system. Such communication can be facilitated by a programming interface, such as a graphical user interface (GUI), a web-based interface, an application programming interface (API), and / or a set of remote procedure calls (RPCs) corresponding to interface elements, a messaging interface in which interface elements correspond to messages of a communication protocol, and / or an appropriate combination thereof, which may be part of a user interface.
[0030] Example of an exemplary embodiment Exemplary embodiments of this disclosure are described only by reference to the drawings. Similar reference numerals and letters in the drawings refer to similar elements or equivalents.
[0031] Some of the following explanations are presented, explicitly or implicitly, in terms of algorithms and functional or symbolic representations of operations on data in computer memory. These algorithmic descriptions and functional or symbolic representations are means used by those skilled in data processing techniques to most efficiently communicate the content of their work to others skilled in the art. An algorithm can be thought of as a self-consistent set of steps to reach a desired result. These steps require the physical manipulation of physical quantities, such as electrical, magnetic, or optical signals, which can be stored, transferred, combined, compared, and otherwise manipulated.
[0032] Unless otherwise specified, and as will be evident below, throughout this specification, discussions using terms such as “receive,” “calculate,” “determine,” “update,” “generate,” “initialize,” “output,” “retrieve,” “identify,” “distribute,” and “authenticate” refer to the operations and processing of a computer system or similar electronic device for manipulating data represented as physical quantities within a computer system and converting it into other similar data represented as physical quantities within a computer system or other information storage device, transmission device, or display device.
[0033] This specification also discloses apparatus for performing the operations of these methods. Such apparatus may be configured specifically for the required purpose, or may consist of a computer or other device that is selectively activated or reconfigured by a computer program stored in the computer. The algorithms and displays shown herein are not inherently related to any particular computer or other apparatus. Various machines may be used with the program in accordance with the teachings of this specification. Alternatively, a configuration of more specialized apparatus for performing the required method steps may be appropriate. The structure of the computer will become clear from the following description.
[0034] Furthermore, this specification also implicitly discloses computer programs in such a way that it will be obvious to those skilled in the art that the individual steps in the methods described herein can be performed by computer code. The computer programs are not intended to be limited to any particular programming language and its execution. It should be understood that various programming languages and their codings may be used to implement the teachings contained in this disclosure. Also, the computer programs are not intended to be limited to any particular control flow. There are numerous variations of the computer programs that may use different control flows without departing from the spirit or scope of this disclosure.
[0035] Furthermore, one or more steps in a computer program may be executed in parallel rather than sequentially. Such a computer program can be stored in any computer-readable medium. The computer-readable medium may include storage devices such as magnetic disks or optical disks, memory chips, or other storage devices suitable for interfacing with a computer. The computer-readable medium may also include wired media, as exemplified by the Internet system, or wireless media, as exemplified by the GSM mobile telephone system, Long-Term Evolution (LTE) system, and 5G mobile network system. A computer program loaded into and executed on such a computer can effectively provide a device that performs the steps in a desired manner.
[0036] Various exemplary embodiments of this disclosure relate to methods and apparatus for recognizing the motion of an object from multiple video streams. Those skilled in the art will understand that such apparatus and image capture apparatus may be implemented as part of a system to provide similar technical effects.
[0037] Figure 1 shows a schematic diagram 100 illustrating the process of allocating hardware resources 108 for processing the video stream 104 generated by camera 102. The analysis server 106 can extract the video stream 104 from camera 102 in order to process it. Specifically, the analysis server uses hardware resources 108 such as CPU, CPU, and / or RAM to pass the video stream 104 through a neural network model 110, performs video analysis tasks on the video stream 104 such as detecting objects and attributes, recognizing the behavior of objects, and identifying object classes from the video stream 104, and generates visual analysis results 112 for the end user.
[0038] If the video streams from all cameras are processed by the same neural network model to perform the same video analysis task / function (e.g., object classification only or motion recognition only), different application requirements may not be met, or some hardware resources may be wasted. On the other hand, if the video streams from all cameras are processed by different neural network models to perform different video analysis tasks / functions that recognize / detect different sets of target attributes, the complexity of deployment increases significantly if the neural network models require different libraries / environments, and therefore hardware resources.
[0039] Figure 2 shows a site plan 200 in which a video analysis system including multiple cameras is located. The cameras are positioned in different locations within the site. Each camera can be configured for a specific video analysis task. For example, camera 1 is configured for object classification, with the target attribute being the object class "person". Camera 2 is configured for object classification, with the target attributes being the object classes "person" and "box". Camera 3 is configured for motion recognition, with the target attributes being the actions "walk", "fight", and "talk". And camera 4 is configured for motion recognition, with the target attributes being the actions "walk" and "talk". In such a case, the analysis server (not shown) can perform the configured video analysis task by extracting one or more video streams generated by each camera and passing the video streams through the corresponding attribute detection modules (e.g., cameras 3 and 4 are motion recognition modules, and cameras 1 and 2 are object classification modules).
[0040] Figure 3 shows Figure 300 illustrating the process of a conventional video analysis system that uses the same neural network model to detect target attributes of objects from each camera. In this case, the analysis server receives a video stream and runs a neural network model that can perform object classification to detect the object classes "person," "vehicle," "motorcycle," and "box" of objects appearing in the video stream. However, different video streams generated by different cameras may have different classification (application) requirements. In particular, the analysis server may request that it detect the object classes "person," "vehicle," "motorcycle," and "box" of objects appearing in video stream 1 generated by camera 1, the object classes "person" and "box" of objects appearing in video stream 2 generated by camera 2, the object classes "person" and "vehicle" of objects appearing in video stream 3 generated by camera 3, the object class "person" of objects appearing in video stream 4 generated by camera 4, and only the object class "motorcycle" of objects appearing in video stream N generated by camera N. In such a case, deploying the same neural network capable of performing object classification to detect the object classes "person," "vehicle," "motorcycle," and "box" of objects appearing in the video stream could potentially optimize GPU resource usage because it doesn't need to detect vehicles and boxes in video stream 2, motorcycles and boxes in video stream 3, vehicles, motorcycles and boxes in video stream 4, and people, vehicles, and boxes in video stream N.
[0041] Figure 4 shows a block diagram of a system 400 that allocates hardware resources for processing multiple video streams generated by multiple video capture devices, according to various exemplary embodiments of the present disclosure.
[0042] System 400 comprises a requesting device 402, an attribute detection server 408, a collaboration server 440, hosts 450A to 450N, and sensors 442A to 442N.
[0043] The requesting device 402 communicates with the attribute detection server 408 and / or the cooperation server 440 via the connection units 416 and 421, respectively. The connection units 416 and 421 may be wireless (e.g., via NFC communication, Wi-Fi communication, Bluetooth, etc.) or via a network (e.g., the Internet). The connection units 416 and 421 may also be network (e.g., Internet) connection units.
[0044] The attribute detection server 408 further communicates with the cooperating server 440 via the connection unit 420. The connection unit 420 may be accessed via a network (e.g., a local area network, a wide area network, the internet, etc.). In one configuration, the attribute detection server 408 and the cooperating server 440 are combined, and the connection unit 420 may be an interconnected bus.
[0045] The coordinating server 440 communicates with hosts 450A to 450N via their respective connection units 422A to 422N. The connection units 422A to 422N may be a network (for example, the Internet).
[0046] Hosts 450A through 450N are servers. The term "host" is used here to distinguish between hosts 450A through 450N and the collaborative server 440. Hosts 450A through 450N are collectively referred to as host 450, and host 450 refers to any of the hosts 450. Host 450 may be combined with collaborative server 440.
[0047] In one example, host 450 may be managed by the entity's security officer, and the coordinating server 440 may be a central server that coordinates with host 450 and decides whether any of the hosts 450 will transfer data or acquire data such as image input.
[0048] Sensors 442A to 442N are connected to the coordinating server 440 or attribute detection server 408 via their respective connectors 444A to 444N or 446A to 446N. Sensors 442A to 442N are collectively referred to here as sensor 442. Connectors 444A to 444N are collectively referred to here as connector 444, and connector 444 refers to any of the connectors 444. Similarly, connectors 446A to 446N are collectively referred to here as connector 446, and connector 446 refers to any of the connectors 446. Connectors 444 and 446 may be wireless (e.g., via NFC communication, Wi-Fi communication, Bluetooth, etc.) or via a network (e.g., the Internet). Sensor 442 may be an image capture device, an object tracking device, a video capture device, a motion sensor, or a temperature sensor, and depending on its type, it may be configured to transmit input to at least one of the attribute detection servers 408.
[0049] In exemplary embodiments, each of devices 402 and 442, and servers 408, 440, and 450 provides an interface that enables communication with other connected devices 402 and 442 and / or servers 408, 440, and 450. Such communication is facilitated by an application programming interface (API). Such an API may be part of a user interface that includes a program interface such as a graphical user interface (GUI), a web-based interface, an application programming interface (API) and / or a set of remote procedure calls (RPCs) corresponding to interface elements, a messaging interface in which interface elements correspond to messages of a communication protocol, and / or an appropriate combination thereof.
[0050] The term "server" as used herein may refer to a single computing device comprising a processor, or to multiple interconnected computing devices working together to perform a particular function. In other words, a server may be contained within a single hardware unit, or it may be distributed across several or many different hardware units.
[0051] Linkage server 440 The collaboration server 440 is associated with an entity (e.g., a service company, organization, or moderator). In one configuration, the collaboration server 440 is owned and operated by the entity operating server 408. In such a configuration, the collaboration server 440 may be implemented as part of server 408 (e.g., a computer program module, computing device, etc.).
[0052] The integration server 440 may be configured to manage user registration. Registered users have an action-aware account that includes user details. The registration step is called onboarding. Users may perform onboarding to the integration server 440 using either the requesting device 402 or the host 450.
[0053] To access the functions of the integration server 440, it is not necessary to have a motion recognition account on the integration server 440. However, there are functions available to registered users. For example, functions such as the recognition of more complex motions or motions involving multiple objects, or increasing the maximum number of input video streams, can be restricted to registered users only.
[0054] The user onboarding process is performed by the user via one of the requesting devices 402. In one configuration, the user downloads an application (including an API that interacts with the integration server 440) to the sensor 442. In another configuration, the user accesses a website (including an API that interacts with the integration server 440) on the requesting device 402.
[0055] Registration details include, for example, a user identifier (ID) or a portrait of the user's appearance, the user's address, contact information, or other important information, and a sensor 442 authorized to update the motion recognition account.
[0056] Upon onboarding, users will have an action-aware account that stores all their details.
[0057] Requesting device 402 The requesting device 402 is associated with the subject (or requester) that is a party to an attribute discovery request initiated by the requesting device 402. The requester may be a member of the public interest or a security officer of the entity who is helping to obtain the data necessary to detect and recognize target attributes of objects within the entity (e.g., theft, combat, animal class, person class). The requesting device 402 may be a desktop computer, an interactive voice response (IVR) system, a smartphone, a laptop computer, a personal digital assistant computer (PDA), a mobile computer, a tablet computer, etc.
[0058] In one example configuration, the requesting device 402 is a watch or a similarly wearable computing device equipped with a wireless communication interface.
[0059] Attribute detection server 408 The attribute detection server 408 is as described in the "Terminology" section above, and is configured to perform video analysis tasks / functions such as motion recognition or object classification, and to run a model (e.g., a neural network model) configured to detect object attributes from the video stream received from the sensor 442 according to the configured video analysis task / function.
[0060] Host 450 Host 450 is a server associated with an entity (e.g., a company or organization) that manages (e.g., establishes and operates) object information about objects for which attributes have been detected.
[0061] In one configuration, the entity is a bank. Therefore, each entity manages resources by operating host 450. In one configuration, host 450 receives an alarm signal indicating that target activity has been detected. Host 450 can then be configured to transmit resources to a location identified by location or camera information included in the alarm signal. For example, the host may be configured to acquire relevant video or image input for processing.
[0062] In one configuration, video streams, detected objects, and attributes may be stored and updated in an attribute detection account associated with the user. Advantageously, such information is valuable to users such as law enforcement agencies and security or building management staff who identify, track, and monitor objects. The time spent examining camera footage to detect target attributes of objects from multiple video streams generated by multiple sensors 442 can be reduced.
[0063] Sensor 442 Sensor 442 is associated with a user associated with the requesting device 402. Sensor 442 may be an image capture device, an object tracking device, a video capture device, a motion sensor, or a temperature sensor, and may be configured to generate a video stream or video file and send it to at least one of the attribute detection servers 408 for detecting target attributes of an object from the video stream. Sensor 442 may also be configured to send information about the sensor or the video stream generated by the sensor (e.g., location and resolution) to the cooperating server 440 for the allocation of hardware resources for processing the video stream generated by sensor 442.
[0064] Figure 5 shows a diagram illustrating the process of detecting target attributes of an object from multiple video streams generated by multiple cameras. The video streams may be sent to a server or module 502 for processing the corresponding video analysis results generated. Module 502 comprises a requirements analysis unit, a model configuration unit, a resource allocation unit, and a model inference unit. The requirements analysis unit is configured to collect and analyze application requirements information from different camera streams, including visual / video analysis tasks, behavior / class lists, and other possible requirements. The output of the requirements analysis unit is the type of attribute detection model (e.g., a behavior recognition model or an object classification model), the order level in the case of behavior recognition, and the number of network layers in the case of an object classification model.
[0065] The Model Configuration Unit configures the settings of the neural network model (e.g., the number of sequential interaction levels or the number of neural network layers) based on the requirements analysis results from the Requirements Analysis Unit. Specifically, the Model Configuration Unit receives instructions from the Requirements Analysis Unit regarding the model type, sequential levels, or network layers for each video stream, selects a neural network model configured for the requested task, and sets the sequential levels / network for inference and processing the video stream. Next, the Resource Allocation Unit is configured to allocate hardware resources based on the Model Configuration Results to minimize hardware resource consumption while meeting application requirements. The Model Inference Unit uses the video stream input, the configured neural network model, and the allocated hardware resources to predict the results (e.g., detection results for motion recognition or object classification).
[0066] Figure 6 shows a flowchart illustrating the process of allocating hardware resources for processing multiple video streams generated by multiple video capture devices, according to various exemplary embodiments of the present disclosure.
[0067] In step 602, instructions are received to detect one or more target attributes of objects appearing in the video streams generated by each of the multiple video capture devices. In step 604, an attribute detection module configured to detect one or more target attributes of objects according to the instructions is selected. In step 606, one or more hardware resources are allocated to the attribute detection module to process the video streams and detect one or more target attributes of objects.
[0068] Figure 7 shows a block diagram illustrating a system 700 for recognizing the motion of an object from a first set of video streams, according to various exemplary embodiments of the present disclosure.
[0069] In one example, the management of image and signal inputs is performed by each video capture device 702a, 702b. System 700 includes video capture devices 702a, 702b (for simplicity, only two video capture devices are shown) that communicate with device 704. In one implementation, device 704 may generally be described as a physical device comprising at least one processor 706 and at least one memory 708 containing computer program code. The at least one memory 708 and computer program code are configured by at least one processor 706 to cause the physical device to perform the operations described in Figure 6. Processor 706 is configured to receive a first set of video streams from video capture devices 706a, 706b or to retrieve a first set of video streams from a database 710. Alternatively or additionally, the first set of video streams acquired by video capture devices 702a, 702b are stored in a database 710, and the host processor 706 is configured to retrieve a first set of video streams from the database 710.
[0070] The video capture devices 702a and 702b are collectively referred to here as video capture device 702, and video capture device 702 refers to one of the hosts 702. Video capture device 702 may be a closed-circuit television (CCTV)-like device that provides various data (camera data), including physical, motion / behavioral, and object feature data, which can be used by a system that detects objects as well as detecting target attributes of objects. In one implementation, the data generated from video capture device 702 may be stored in the memory 708 of device 704 or in a database 710 accessible by device 704.
[0071] Furthermore, camera data such as positional data related to the fixed or acquired position of the camera, and time data such as video / image timestamps, may be received, stored, and / or acquired in order to recognize the motion of an object and to derive timestamps of the position and motion related to the object.
[0072] According to this disclosure, the device 704 may be configured to communicate with a video capture device 702, a database 710, and a plurality of processing units (not shown). In one implementation, the processing units may be part of the device 704 and communicate with a processor 708. Similarly, in one implementation, the database 710 may be part of the device 704.
[0073] The device 704 can receive multiple video streams as input from the video capture device 702 or obtain them from the database 710. The device 704 may also receive instructions to detect one or more target attributes of objects appearing in the video streams generated by the video capture device 702.
[0074] The memory 706 and the computer program code stored therein may be configured to cause the device to select an attribute detection module (not shown) configured to detect one or more target attributes of an object according to instructions, and to allocate one or more hardware resources (e.g., part of the memory 708 or a processing unit of the processor 706) to the attribute detection module in order to process the video stream and detect one or more target attributes of the object. The attribute detection module may be part of the device 704.
[0075] In one exemplary embodiment, the device 704 may receive a first input relating to the selection of each of a plurality of video capture devices, and a second input relating to the selection of a task for detecting one or more target attributes of an object. The memory 706 and the computer program code stored therein are configured to cause the processor 706 to cause the device to select an attribute detection module based on the first and second inputs.
[0076] In one exemplary embodiment, when an instruction is received to detect target behavior of an object appearing in a video stream generated by one of the video capture devices 702, where one or more target attributes are all target behaviors, the memory 706 and the computer program code stored therein are configured to cause the processor 706 to select an attribute detection module (e.g., a behavior recognition module) configured to detect such target behaviors, and to allocate the hardware resources required for such attribute detection module to perform a behavior recognition task and target behavior detection.
[0077] Furthermore, if each target action is associated with a pre-configured interaction level, and a higher pre-configured interaction level for an action indicates that more hardware resources are required by an attribute detection module to detect the action, then the memory 706 and the computer program code stored therein are configured to allow the processor 706 to cause the device to identify the interaction level associated with the video stream based on the pre-configured object interaction level corresponding to the target action. The memory 706 and the computer program code stored therein are configured to allow the processor 706 to cause the device to select an attribute detection module (e.g., an action recognition module) configured to detect such target actions, and to allocate the hardware resources required by such attribute detection module to perform an action recognition task and detect the target action based on the interaction level associated with the video stream. A list of target actions and their corresponding pre-configured interaction levels is stored in the memory 706 or database 710, and the device 704 is configured to retrieve the pre-configured interaction levels for target actions from the memory 706 or database 710.
[0078] If there are two target behaviors to be detected by two different pre-configured interaction levels, the memory 706 and the computer program code stored therein are configured to cause the processor 706 to cause the device to select the higher (or highest) pre-configured interaction level from the two pre-configured interaction levels as the interaction level associated with the video stream.
[0079] In another exemplary embodiment, when an instruction is received to detect the target object class of an object appearing in a video stream generated by one of the video capture devices 702, such that one or more target attributes are all target object classes, the memory 706 and the computer program code stored therein are configured to cause the processor 706 to cause the device to select an attribute detection module (e.g., an object classification module) and to allocate the hardware resources necessary for such an attribute detection module to perform an object classification task and the detection of target object classes. A list of target object classes and the required number of pre-configured processing layers corresponding to them is stored in the memory 706 or database 710, and the device 704 is configured to retrieve the number of pre-configured processing layers for a target object class from the memory 706 or database 710.
[0080] Furthermore, the memory 706 and the computer program code stored therein are configured to cause the device to identify the number of processing neural network layers required for the attribute detection module to process the video stream, based on the number of pre-configured processing layers required to detect each of one or more target object classes, as determined by the processor 706. The memory 706 and the computer program code stored therein are configured to cause the device to select an attribute detection module (e.g., an action recognition module) configured to allocate one or more hardware resources required for the attribute detection module to execute a second video stream through the number of processing neural network layers required to detect one or more target object classes from the video stream.
[0081] If there are two target objects to be detected by two different pre-configured numbers of processing neural network layers, the memory 706 and the computer program code stored therein are configured to cause the processor 706 to cause the device to select the higher (or highest) number of processing layers from the two different pre-configured numbers of processing neural network layers as the number of processing layers associated with the video stream.
[0082] The memory 708 and the computer program code stored therein are configured by the processor 706 to cause the device 704 to generate the results of detecting one or more target attributes of an object from multiple video streams using one or more hardware resources allocated to each of the video capture devices 702.
[0083] The following paragraphs describe a first exemplary embodiment in which the motion recognition video analysis task is configured to be performed on a video stream received from one of several cameras.
[0084] Figure 8 shows Figure 800 illustrating a process for detecting target motion of an object from a video stream generated by a camera, according to a first exemplary embodiment of the present disclosure. Instructions via APIs, which may be part of a user interface including a graphical user interface (GUI), a web-based interface, a program interface, and a messaging interface, may be received by a requirements analysis unit from another connected device (not shown) to indicate video analysis requirements for detecting target motion (i.e., performing motion recognition of such target motion) from the video stream. The video stream, consisting of multiple video frames, is passed to a scene context extraction network module in the motion recognition module, where contextual information regarding motion recognition is obtained. The model configuration unit then selects a higher-order interaction network module, a middle-order interaction network module, or a lower-order interaction network module in the motion recognition module based on the video analysis requirements and contextual information, for example, based on a pre-configured interaction level for the target motion. Each hierarchical interaction network module can be configured to detect and recognize motion at different hierarchical interaction levels and therefore require different hardware resources to operate and perform the video analysis task. The video frames are then processed by the selected hierarchical interaction network module to generate the motion recognition result.
[0085] Hierarchical interactions analyze the relationships between detected objects. Higher-order interactions or higher interaction levels indicate that more objects are involved in the relationship analysis, such as motion recognition. For example, higher-order interaction levels indicate that more objects (or more movement of objects) are required in the analysis, and therefore, generally, analyzing relationships between objects and recognizing actions performed by objects requires more powerful modules (e.g., higher-order interaction network modules) with more hardware resources than those for lower-order interaction levels.
[0086] Generally, higher-order interaction network modules require maximum hardware resources for operation and are used to perform video analysis tasks for high-level interaction behavior, while lower-order interaction network modules require minimal hardware resources for operation and are used to perform video analysis tasks for low-level interaction behavior. Deploying higher-order interaction network modules to analyze low-level interaction behavior may result in wasted hardware resources that could be used for other purposes. On the other hand, lower-order interaction network modules may not be powerful enough to detect high-level interaction target behavior and may not meet video analysis requirements.
[0087] Figure 9 shows Figure 900 illustrating three different levels of interaction according to an exemplary embodiment. Multiple objects (O1, O2, O3, O4, O5, O6, O7, ...) are detected from the video frame. In one implementation, actions involving two or fewer objects may be pre-configured as low-interaction-level actions. Actions involving three objects may be pre-configured as medium-interaction-level actions. Actions involving four or more objects may be pre-configured as high-interaction-level actions. When an instruction is received to detect a target action of objects, and the target action involves only three objects (medium-interaction level), the medium-level interaction network module is deployed to process the video stream and detect the medium-interaction-level target action. This can optimize the hardware resources allocated to processing the selected module and video stream.
[0088] Figure 10 shows Figure 1000, which illustrates an exemplary graphical user interface according to a first exemplary embodiment of the present disclosure. In this example, the video stream is shown on the display window 1002. The settings panel 1004 shows a list of cameras (cameras 1-4). In this case, input is received regarding the selection of camera 1, which is the target camera. Subsequently, the settings panel 1004 shows two video analysis tasks available for processing the video stream of the target camera: an action recognition task and an object classification task. In this case, input is received regarding the selection of the action recognition task. In response to the input regarding the portion of the action recognition task, a list of target actions is shown. In this case, a list of target actions is shown, consisting of four different target actions: "sleep," "walk," "talk," and "fight." Subsequently, input is received regarding the selection of the target action "sleep" from the list of target actions. Overall, the input from the GUI operation indicates that the analysis server (not shown) is configured to perform an action recognition task to detect the action "sleep" of an object appearing in the video stream generated by camera 1, which is the target camera. To facilitate user operation and selection, a description of each input option may be displayed at the bottom of the GUI.
[0089] Table 1 shows the estimated corresponding hardware resources (GPU memory) required by attribute discovery modules for detecting different target operations and their different pre-configured hierarchy levels, according to exemplary embodiments of the present disclosure.
[0090] [Table 1]
[0091] In this exemplary embodiment, if, in relational analysis, actions such as sleeping, walking, and suicide require the involvement of two or fewer objects, they are pre-configured as low-level actions, and a low-level network module and an estimated 2GB of GPU memory may be required to detect such actions. If, in relational analysis, a conversation action requires the involvement of three objects (e.g., two faces, hands), it is pre-configured as a medium-level action, and an medium-level network module and an estimated 4GB of GPU memory may be required to detect the conversation action. If, in relational analysis, a combat action requires the involvement of four or more objects (e.g., four people), it is pre-configured as a high-level action, and a high-level network module and an estimated 6GB of GPU memory may be required to detect the combat action.
[0092] Table 2 shows the hardware resources allocated to process video streams with different instructions to detect different target behaviors of an object, according to exemplary embodiments of the present disclosure.
[0093] [Table 2]
[0094] In this exemplary embodiment, instructions are received to detect sleep and walking behavior from video stream 1, conversation behavior from video stream 2, suicide behavior from video stream 3, and combat and suicide behavior from video stream 4. Based on Table 1, since both sleep and walking behavior are low-level behaviors, video stream 1 is associated with the low level. A low-level network module is selected to detect sleep and walking behavior of objects appearing in video stream 1, and 2GB of GPU memory is allocated to processing video stream 1. Conversation behavior is a medium-level behavior, and therefore video stream 2 is associated with the medium level. A medium-level network module is selected to detect conversation behavior of objects appearing in video stream 2, and 4GB of GPU memory is allocated to processing video stream 2. Suicide behavior is a low-level behavior, and therefore video stream 3 is also associated with the low level. A low-level network module is selected to detect suicide behavior of objects appearing in video stream 3, and 2GB of GPU memory is allocated to processing video stream 3. Suicide is a low-level behavior, but combat is a high-level behavior. In this case, since two operations with different predetermined hierarchical levels are shown, the higher of the two, i.e., the higher hierarchical level, is selected as the hierarchical level for video stream 4. A higher-order network module is selected to detect both suicide and combat of objects appearing in video stream 4, and 6GB of GPU memory is allocated to processing video stream 4.
[0095] Note that in conventional methods, the same module with 6GB of GPU memory is used to process the video stream and detect each target operation.
[0096] The following paragraphs describe a second exemplary embodiment in which the object classification video analysis task is configured to be performed on a video stream received from one of several cameras.
[0097] Figure 11 shows a diagram illustrating the process of detecting the target motion of an object from a video stream generated by a camera, according to a second exemplary embodiment of the present disclosure. The object classification module 1102 consists of multiple neural network layers arranged in series.
[0098] Each neural network layer can generate object classification results for a specific object class and is therefore used as a classifier to detect a specific object class from a video frame when the video frame is passed through the neural network layers. In some implementations, the results of two or more neural networks can be combined as a classifier to generate object classification results for a specific object class. Different combinations of neural network layers may be used as different classifiers to classify different object classes and are arranged in series so that the video frame can pass through all the different classifiers to classify different object classes from the video frame.
[0099] Figure 12 shows Figure 1200, which illustrates a neural network layer according to an exemplary embodiment of the present disclosure. The input video frame has a resolution of 768 × 768 and has seven convolutional layers, each having a Rectified Linear (ReLU) function arranged in series. The video frame passes through the seven convolutional layers to obtain different binary classification results for classifying different object classes. In particular, the video frame passes through the first two convolutional layers to obtain a first binary classification result with maximum and sigmoid activations for use as a first classifier for detecting a first class of objects appearing in the video frame (e.g., humans). Next, the video frame passes through the next two convolutional layers to obtain a second binary classification result with maximum and sigmoid activations for use as a second classifier for detecting a second class of objects appearing in the video frame (e.g., vehicles). Next, the video passes through two more convolutional layers to obtain a third binary classification result with maximum and sigmoid activations, which can be used as a third classifier to detect a third class of objects appearing in the video frames (e.g., cats).
[0100] Returning to Figure 11, instructions via an API, which may be part of a user interface including a graphical user interface (GUI), a web-based interface, a program interface, and a messaging interface, may be received by a requirements analysis unit from another connected device (not shown) to indicate video analysis requirements for detecting such target classification classes (i.e., performing object classification of such target object classes) from a video stream. Based on the video analysis requirements, the model configuration unit selects the number of neural network layers in the attribute detection module 1102 required to perform classification of the target object classes, for example, based on a pre-configured number of processing layers for the target object classes.
[0101] For example, the first and second neural network layers are pre-configured to be used as classifiers for classifying people. The third and fourth neural network layers are pre-configured to be used as another classifier for classifying vehicles. The fifth and sixth neural network layers are pre-configured to be used as yet another classifier for classifying cats. When an instruction is received to detect a vehicle object class from a video stream, the model configuration unit selects four neural network layers so that the video frame passes through the fourth neural network layer in order to generate a classification result for the vehicle object class. In a real system deployment, the selection of different layers is implemented by an inference program to determine the number of layers the image data must pass through for the final output.
[0102] Generally, processing video frames through more neural network layers requires more hardware resources. Deploying numerous neural networks and classifiers to detect many classification classes may waste hardware resources that could have been effectively used for other purposes, while deploying too few neural networks may fail to satisfy the requirements for detecting target object classes and video analysis.
[0103] Figure 13 shows Figure 1300, which illustrates an exemplary graphical user interface according to a second exemplary embodiment of the present disclosure. In this example, a video stream is shown on display window 1302. The settings panel 1304 shows a list of cameras (cameras 1-4). In this case, input is received regarding the selection of camera 1, which is the target camera. Subsequently, the settings panel 1304 shows two video analysis tasks available for processing the video stream of the target camera: an action recognition task and an object classification task. In this case, input is received regarding the selection of the object classification task. In response to the input regarding the portion of the object classification task, a list of target actions is shown. In this case, a list of target actions is shown, consisting of four different target object classes: "human," "vehicle," "cat," and "apple." Subsequently, input is received regarding the selection of the target action "human" from the list of target actions. Overall, the input from the GUI operation indicates that the analysis server (not shown) is configured to perform an object classification task to detect the class "human" of objects appearing in the video stream generated by camera 1, which is the target camera. To facilitate user operation and selection, explanations for each input option may be displayed at the bottom of the GUI.
[0104] Table 3 shows the estimated corresponding hardware resources (GPU memory) required by attribute detection modules for detecting different target object classes according to exemplary embodiments of the present disclosure.
[0105] [Table 3]
[0106] In this exemplary embodiment, each of the different object classes has a pre-configured number of network layers required to be detected. Specifically, two neural network layers with an estimated 0.5 GB of GPU memory are required to generate cat detection results, three neural network layers with an estimated 1 GB of GPU memory are required to generate box or bicycle detection results, and four neural network layers with an estimated 2 GB of GPU memory are required to generate person and bicycle detection results.
[0107] Table 4 shows the hardware resources allocated to process video streams having different instructions for detecting different target object classes of objects, according to exemplary embodiments of the present disclosure.
[0108] [Table 4]
[0109] In this exemplary embodiment, instructions are received to detect the object classes of boxes and people from video stream 1, boxes and cats from video stream 2, time-saving vehicle objects from video stream 3, and people and vehicles from video stream 4. Based on Table 3, boxes and people can be detected using three and four network layers, and therefore video stream 1 is associated with four network layers with 2GB of GPU memory, i.e., the higher of two predetermined numbers of network layers. Boxes and cats can be detected using three and two network layers, respectively, and therefore video stream 2 is associated with three network layers with 1GB of GPU memory, i.e., the higher of two predetermined numbers of network layers. Cats and bicycles can be detected using two and three network layers, respectively, and therefore video stream 3 is associated with three network layers with 1GB of GPU memory, i.e., the higher of two predetermined numbers of network layers. Both people and vehicles can be detected using four network layers, and therefore, video stream 4 is associated with four network layers having 2GB of GPU memory.
[0110] Note that conventional methods use the same module to process the video stream through all network layers (e.g., seven layers using 4GB of GPU memory) to detect each target object class.
[0111] Figure 14 shows a schematic diagram of an exemplary computing device 1400, hereafter also interchangeably referred to as the computer system 1400. Here, one or more such computing devices 1400 can be used or are suitable for use to perform the method of Figure 6 and to implement the apparatus of Figure 7. The following description of the computing device 1400 is provided for illustrative purposes only and is not intended to limit it.
[0112] As shown in Figure 14, the exemplary computing device 1400 includes a processor 1404 for executing software routines. Although a single processor is shown for clarity, the computing device 1400 can include a multiprocessor system. The processor 1404 is connected to a communication infrastructure 1406 for communicating with other elements of the computing device 1400. The communication infrastructure 1406 can include, for example, a communication bus, a crossbar, or a network.
[0113] The computing device 1400 further includes main memory 1408, such as random access memory (RAM), and secondary memory 1410. The secondary memory 1410 may include a storage drive 1412, which is, for example, a hard disk drive, a solid-state drive, or a hybrid drive, and / or a removable storage drive 1414, which includes a magnetic tape drive, an optical disc drive, a solid storage drive (such as a USB flash drive, a flash memory device, a solid-state drive, or a memory card). The removable storage drive 1414 reads from and / or writes to the removable storage media 1418 in a well-known manner. The removable storage media 1418 includes a magnetic tape, an optical disc, a non-volatile memory storage medium, etc., which is read from or written to by the removable storage drive 1414. As will be understood by those skilled in the art, the removable storage media 1418 includes a computer-readable storage medium storing program code instructions and / or data that can be executed by a computer.
[0114] In other implementations, secondary memory 1410 may additionally or alternatively include other similar means that enable computer programs or other instructions to be loaded into computing device 1400. Such means may include, for example, removable storage units 1422 and interfaces 1420. Examples of removable storage units 1422 and interfaces 1420 include program cartridges and cartridge interfaces (e.g., those found in video game console devices), removable memory chips (such as EPROM or PROM) and associated sockets, removable solid-state storage drives (such as USB flash drives, flash memory devices, solid-state drives, or memory cards), and other removable storage units 1422 and interfaces 1420 that enable software and data to be transferred from removable storage units 1422 to computer system 1400.
[0115] The computing device 1400 also includes at least one communication interface 1424. The communication interface 1424 enables software and data to be transferred between the computing device 1400 and external devices via a communication channel 1426. In various exemplary embodiments of this disclosure, the communication interface 1424 enables data to be transferred between the computing device 1400 and a data communication network, such as a public data or private data communication network. The communication interface 1424 can be used to exchange data between different computing devices 1400 that form part of an interconnected computer network. Examples of the communication interface 1424 may include a modem, a network interface (such as an Ethernet card), a communication port (such as serial, parallel, printer, GPIB, IEEE 1394, RJ45, USB, etc.), an antenna with associated circuitry, etc. The communication interface 1424 may be wired or wireless. The software and data transferred via the communication interface 1424 may be in the form of signals that are electronic, electromagnetic, optical, or other signals that can be received by the communication interface 1424. These signals are provided to the communication interface via communication channel 1426.
[0116] As shown in Figure 14, the computing device 1400 further includes a display interface 1402 that performs operations for generating images on an associated display 1430, and an audio interface 1432 that performs operations for playing audio content via one or more associated speakers 1434.
[0117] As used herein, the term “computer program product” may, in part, refer to a removable storage medium 1418, a removable storage unit 1422, a hard disk installed in a storage drive 1412, or a carrier wave that transmits software to a communication interface 1424 via a communication channel 1426 (wireless link or cable). Computer-readable storage medium refers to any non-temporary, non-volatile tangible storage medium that provides instructions and / or data recorded for execution and / or processing to the computing device 1400. Examples of such storage mediums include magnetic tape, CR-ROM, DVD, Blu-ray disc, hard disk drive, ROM or integrated circuit, solid-state storage drive (such as USB flash drive, flash memory device, solid-state drive or memory card), hybrid drive, magneto-optical disk, or computer-readable card such as a PCMCIA card, whether such devices are internal or external to the computing device 1400. Examples of temporary or intangible computer-readable transmission media that may also be involved in providing software, application programs, instructions and / or data to the computing device 1400 include wireless or infrared transmission channels and network connections to other computers or network devices, as well as the Internet or intranets, including email transmissions and information stored on websites, etc.
[0118] The computer program (also called computer program code) is stored in main memory 1408 and / or secondary memory 1410. The computer program can also be received via the communication interface 1424. When such a computer program is executed, it enables the computing device 1400 to perform one or more features of the exemplary embodiments described herein. In various exemplary embodiments, when the computer program is executed, it enables the processor 1404 to perform features of the exemplary embodiments described above. Thus, such a computer program represents the controller of the computer system 1400.
[0119] The software may be stored in a computer program product and loaded into the computing device 1400 using a removable storage drive 1414, a storage drive 1412, or interface 1420. The computer program product may be a non-temporary computer-readable medium. Alternatively, the computer program product may be downloaded to the computer system 1400 via a communication channel 1426. Once executed by the processor 1404, the software causes the computing device 1400 to perform the operations necessary to carry out the method of Figure 6 and implement the apparatus of Figure 7.
[0120] It should be understood that the exemplary embodiment in Figure 14 is presented as an example to illustrate the operation and structure of a single device. Therefore, in some exemplary embodiments, one or more features of the computing device 1400 may be omitted. Also, in some exemplary embodiments, one or more features of the computing device 1400 may be combined together. Furthermore, in some exemplary embodiments, one or more features of the computing device 1400 may be divided into one or more components.
[0121] Those skilled in the art will understand that many variations and / or modifications may be made to the present disclosure shown in certain exemplary embodiments without departing from the spirit or scope of the broadly described disclosure. Therefore, these exemplary embodiments are considered illustrative and not restrictive in all respects.
[0122] This application is based on Singapore Patent Application No. 10202300627Q, filed on 8 March 2023, and claims priority from said application, the disclosure of which is incorporated herein by reference in its entirety.
[0123] For example, some, but not limited to, examples of exemplary embodiments disclosed above may be described in whole or in part as follows: (Note 1) A method for allocating hardware resources to process multiple video streams generated by multiple video capture devices, Receiving instructions to configure each of the plurality of video capture devices to detect one or more target attributes of objects appearing in a video stream generated by the corresponding video capture device, Select an attribute detection module configured to detect one or more target attributes of the object in accordance with the instructions above, A method comprising allocating one or more hardware resources to an attribute detection module that processes the video stream and detects one or more target attributes of the object. (Note 2) Receiving the instruction for detecting one or more target attributes of the object appearing in the video stream generated by the corresponding video capture device is: Receiving a first input relating to the selection of the corresponding video capture device, The process includes receiving a second input relating to the selection of a task for detecting one or more target attributes of the object, The selection of the attribute detection module is based on the first input and the second input, The method described in Appendix 1. (Note 3) Each of the one or more target attributes is a target action, and a first instruction is received to detect one or more target actions of the object appearing in a first video stream generated by a first video capture device among the plurality of video capture devices, and the selection of the attribute detection module is based on the first instruction, and the method is, Identifying interaction levels associated with the first video stream based on predetermined object interaction levels corresponding to the one or more target actions, wherein a higher pre-configured interaction level for an action indicates that more hardware resources are required by the attribute detection module to detect the action, and the allocation of the one or more hardware resources for processing the video stream allocates the one or more hardware resources to the attribute detection module to process the video stream and detect the one or more target actions from the first video stream, based on the interaction levels associated with the first video stream. The method described in Appendix 1 or 2, further including the method described in Appendix 1 or 2. (Note 4) Selecting the highest pre-configured interaction level from among two or more pre-configured interaction levels for the one or more target actions as the interaction level. The method described in Appendix 3, which further includes the method described in Appendix 3. (Note 5) Each of the one or more target attributes is a target object class, and a second instruction is received to detect one or more target object classes of the object appearing in a second video stream generated by a second video capture device among the plurality of video capture devices, and the selection of the attribute detection module is based on the second instruction, and the method is, Identifying the number of processing neural network layers of the attribute detection module required to process the second video stream based on the number of pre-configured processing layers required to detect each of the one or more target object classes, wherein the allocation of the one or more hardware resources for processing the video stream is such that the allocation of the one or more hardware resources processes the second video stream through the number of processing neural network layers and detects the one or more target object classes from the second video stream. The method described in any one of the appendices 1 to 4, further including the method described in any one of appendices 1 to 4. (Note 6) Selecting a higher number of processing neural network layers from among two or more pre-configured processing neural network layers to be used as the number of processing neural network layers, corresponding to detecting one or more attributes from the second video stream. The method described in Appendix 5, which further includes the method described in Appendix 5. (Note 7) To generate the result of detecting the one or more target attributes of the object from the multiple video streams using the one or more hardware resources allocated to each of the multiple video capture devices, The method described in any one of the appendices 1 to 6, further including the method described in any one of the appendices 1 to 6. (Note 8) A device that allocates hardware resources to process multiple video streams generated by multiple video capture devices, At least one processor, A device comprising at least one memory containing computer program code, The at least one memory and the computer program code are used by at least one processor to power the device, Instructions are received to configure each of the plurality of video capture devices to detect one or more target attributes of objects appearing in the video stream generated by the corresponding video capture device. In accordance with the above instructions, select an attribute detection module configured to detect one or more target attributes of the object, A device configured to allocate one or more hardware resources to an attribute detection module that processes the video stream and detects one or more target attributes of the object. (Note 9) The at least one memory and the computer program code are used by at least one processor to power the device, The system receives a first input for configuring the corresponding video capture device. A second input is received indicating a task to detect one or more target attributes of the object. Based on the first and second inputs, the system is configured to select the attribute detection module configured to detect the one or more target attributes of the object in accordance with the instructions, The apparatus described in Appendix 8. (Note 10) Each of the one or more target attributes is a target action, and a first instruction is received to detect one or more target actions of the object appearing in a first video stream generated by a first video capture device among the plurality of video capture devices, and the selection of the attribute detection module is based on the first instruction. The at least one memory and the computer program code are further controlled by at least one processor to enable the device to function at least further. Based on a predetermined interaction level corresponding to each of the one or more target actions, the interaction level associated with the first video stream is identified, and a higher pre-configured interaction level for an action indicates that the attribute detection module requires more hardware resources to detect the action. The system is configured to process the video stream based on the interaction level associated with the first video stream and to cause the attribute detection module to allocate the one or more hardware resources in order to detect the one or more target actions from the first video stream. The apparatus described in Appendix 8 or 9. (Note 11) The at least one memory and the computer program code are further controlled by at least one processor to enable the device to function at least further. The system is configured to select the highest pre-configured interaction level from among two or more pre-configured interaction levels for the one or more target actions described above. The apparatus described in Appendix 10. (Note 12) Each of the one or more target attributes is a target object class, and a second instruction is received to detect one or more target object classes of the object appearing in a second video stream generated by a second video capture device among the plurality of video capture devices, and the selection of the attribute detection module is based on the second instruction. The at least one memory and the computer program code are further controlled by at least one processor to enable the device to function at least further. Based on the number of pre-configured processing neural network layers required to detect each of the one or more target object classes, the number of processing neural network layers required for processing the second video stream of the attribute detection module is identified. The system is configured to process the second video stream through the aforementioned number of processing neural network layers and to allocate the one or more hardware resources to detect the one or more target object classes from the second video stream. The apparatus described in any one of the appendices 8 to 11. (Note 13) The at least one memory and the computer program code are further controlled by at least one processor to enable the device to function at least further. The system is configured to select the number of processing neural network layers that is higher than the number of pre-configured processing neural network layers corresponding to detecting the one or more attributes from the second video stream. The apparatus described in Appendix 12. (Note 14) The at least one memory and the computer program code are further controlled by at least one processor to enable the device to function at least further. The system is configured to generate results for detecting one or more target attributes of an object from the multiple video streams using the one or more hardware resources allocated to each of the multiple video capture devices. The apparatus described in any one of the appendices 8 to 13. (Note 15) A system comprising the apparatus described in any one of the appendices 8 to 14 and a plurality of video capture devices configured to generate a plurality of video streams, and for allocating hardware resources for processing the plurality of video streams. [Explanation of Symbols]
[0124] 102 Cameras 104 Video Streams 106 Analysis Server 108 Hardware Resources 110 Neural Network Models 112 Visual Analysis Results 400 System 402 Requesting Device 408 Attribute Detection Server 440 Coordinated Serve 450A-450N Host 442A-442N Sensor 502 Server or Module 700 System 702a, 702b video capture devices 704 Equipment 706 Processor 708 memory 710 Databases 1002 Display Window 1004 Settings Panel 1102 Object Classification Module 1302 Display Window 1304 Settings Panel 1400 computing devices 1402 Display Interface 1404 Processor 1406 Communication infrastructure 1408 Main Memory 1410 Secondary Memory 1412 Storage Drives 1414 Removable Storage Drive 1418 Removable Storage Media 1420 Interface 1422 Removable Storage Unit 1424 Communication Interface 1426 Communication Channel 1430 Display 1432 Audio Interface 1434 Speaker
Claims
1. A method for allocating hardware resources to process multiple video streams generated by multiple video capture devices, Receiving instructions to configure each of the plurality of video capture devices to detect one or more target attributes of an object appearing in a video stream generated by the corresponding video capture device, Select an attribute detection module configured to detect one or more target attributes of the object in accordance with the above instructions, A method comprising allocating one or more hardware resources to an attribute detection module that processes the video stream and detects one or more target attributes of the object.
2. Receiving the instruction for detecting one or more target attributes of the object appearing in the video stream generated by the corresponding video capture device means Receiving a first input relating to the selection of the corresponding video capture device, This includes receiving a second input relating to the selection of a task for detecting one or more target attributes of the object, The selection of the attribute detection module is based on the first input and the second input. The method according to claim 1.
3. Each of the one or more target attributes is a target action, and a first instruction is received to detect one or more target actions of the object appearing in a first video stream generated by a first video capture device among the plurality of video capture devices, and the selection of the attribute detection module is based on the first instruction, and the method is, Identifying interaction levels associated with the first video stream based on predetermined object interaction levels corresponding to the one or more target actions, wherein a higher pre-configured interaction level for an action indicates that more hardware resources are required by the attribute detection module to detect the action, and the allocation of the one or more hardware resources for processing the video stream allocates the one or more hardware resources to the attribute detection module to process the video stream and detect the one or more target actions from the first video stream, based on the interaction levels associated with the first video stream. The method according to claim 1 or 2, further comprising:
4. Selecting the highest pre-configured interaction level from among two or more pre-configured interaction levels for the one or more target actions as the interaction level. The method according to claim 3, further comprising:
5. Each of the one or more target attributes is a target object class, and a second instruction is received to detect one or more target object classes of the object appearing in a second video stream generated by a second video capture device among the plurality of video capture devices, and the selection of the attribute detection module is based on the second instruction, and the method is, Identifying the number of processing neural network layers of the attribute detection module required to process the second video stream based on the number of pre-configured processing layers required to detect each of the one or more target object classes, wherein the allocation of the one or more hardware resources for processing the video stream is such that the allocation of the one or more hardware resources processes the second video stream through the number of processing neural network layers and detects the one or more target object classes from the second video stream. The method according to any one of claims 1 to 4, further comprising:
6. Selecting a higher number of processing neural network layers from among two or more pre-configured processing neural network layers to be used as the number of processing neural network layers, corresponding to detecting one or more attributes from the second video stream. The method according to claim 5, further comprising:
7. To generate the result of detecting the one or more target attributes of the object from the multiple video streams using the one or more hardware resources allocated to each of the multiple video capture devices, The method according to any one of claims 1 to 6, further comprising:
8. A device that allocates hardware resources to process multiple video streams generated by multiple video capture devices, At least one processor, A device comprising at least one memory containing computer program code, The at least one memory and the computer program code are used by at least one processor to power the device, Instructions are received to configure each of the plurality of video capture devices to detect one or more target attributes of an object appearing in a video stream generated by the corresponding video capture device. The system will select an attribute detection module configured to detect one or more target attributes of the object in accordance with the above instructions. A device configured to allocate one or more hardware resources to an attribute detection module that processes the video stream and detects one or more target attributes of the object.
9. The at least one memory and the computer program code are used by at least one processor to power the device, The system receives a first input for configuring the corresponding video capture device. A second input is received indicating a task to detect one or more target attributes of the object. Based on the first and second inputs, the system is configured to select the attribute detection module configured to detect the one or more target attributes of the object in accordance with the instructions. The apparatus according to claim 8.
10. Each of the one or more target attributes is a target action, and a first instruction is received to detect one or more target actions of the object appearing in a first video stream generated by a first video capture device among the plurality of video capture devices, and the selection of the attribute detection module is based on the first instruction. The at least one memory and the computer program code are further controlled by at least one processor to enable the device to function at least further. Based on a predetermined interaction level corresponding to each of the one or more target actions, the interaction level associated with the first video stream is identified, and a higher pre-configured interaction level for an action indicates that the attribute detection module requires more hardware resources to detect the action. The system is configured to process the video stream based on the interaction level associated with the first video stream and to cause the attribute detection module to allocate the one or more hardware resources in order to detect the one or more target actions from the first video stream. The apparatus according to claim 8 or 9.
11. The at least one memory and the computer program code are further controlled by at least one processor to enable the device to function at least further. The system is configured to select the highest pre-configured interaction level from among two or more pre-configured interaction levels for the one or more target actions described above as the interaction level. The apparatus according to claim 10.
12. Each of the one or more target attributes is a target object class, and a second instruction is received to detect one or more target object classes of the object appearing in a second video stream generated by a second video capture device among the plurality of video capture devices, and the selection of the attribute detection module is based on the second instruction. The at least one memory and the computer program code are further controlled by at least one processor to enable the device to function at least further. Based on the number of pre-configured processing neural network layers required to detect each of the one or more target object classes, the number of processing neural network layers required for processing the second video stream of the attribute detection module is identified. The system is configured to process the second video stream through the aforementioned number of processing neural network layers and to allocate the one or more hardware resources to detect the one or more target object classes from the second video stream. The apparatus according to any one of claims 8 to 11.
13. The at least one memory and the computer program code are further controlled by at least one processor to enable the device to function at least further. The system is configured to select a higher number of processing neural network layers from among two or more pre-configured processing neural network layers, corresponding to detecting one or more attributes from the second video stream, as the number of processing neural network layers. The apparatus according to claim 12.
14. The at least one memory and the computer program code are further controlled by at least one processor to enable the device to function at least further. The system is configured to generate results for detecting one or more target attributes of an object from the multiple video streams using the one or more hardware resources allocated to each of the multiple video capture devices. The apparatus according to any one of claims 8 to 13.
15. A system comprising the apparatus according to any one of claims 8 to 14 and a plurality of video capture devices configured to generate a plurality of video streams, for allocating hardware resources for processing the plurality of video streams.
Citation Information
Patent Citations
Image recognition device
JP2012068965A
System and method for subject re-identification
JP2016072964A
Method and device for image recognition and program
JP2016162232A
Content selection device, content selection metho, content selection system and program
JP2020115159A
Estimating device, learning device, estimating method, and program
JP2021165984A