Apparatus, method, and program
The imaging device addresses the challenge of unknown resource allocation by embedding metadata in video data to signal processing constraints, ensuring efficient object recognition and preventing failures.
Patent Information
- Application Number
- JP2025196722
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-02-10
AI Technical Summary
Existing imaging devices lack the ability for clients to accurately determine the available processing resources for object recognition, leading to potential recognition failures and inefficiencies due to unknown resource constraints.
An imaging device that generates video data with embedded metadata indicating the availability of processing resources for object recognition, using color-coded bounding boxes to signal resource constraints, allowing clients to adjust settings accordingly.
Enables clients to manage resource allocation effectively, preventing recognition failures and delays by providing real-time feedback on processing capacity, thus optimizing network camera performance.
Smart Images

Figure 2026021615000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an imaging device, a control method for an imaging device, and a program. [Background technology]
[0002] The Annotated Regions SEI (ARSEI) standard in H.265 makes it possible to add information such as the type and position of objects within the field of view to the stream as metadata. Generally, a large amount of processing resources is required to recognize objects within video and obtain tracking result information. The processing resources possessed by network cameras vary from camera to camera, but the ARSEI standard specifies a maximum number of recognizable objects of 255. Many network cameras do not have sufficient processing resources to update the 255 objects required by the ARSEI standard, and the number of objects a network camera can update also varies depending on the number of streams being distributed and the network camera settings.
[0003] On the other hand, the frequency of requests for object recognition varies depending on the use case. Patent Document 1 discloses a technology in which a client notifies a network camera of a reservation for processing resources to be used for object recognition by the network camera. Patent Document 2 also discloses a method in which, when the number of objects recognized by a network camera reaches the maximum allowable number, the minimum detection area for object recognition is increased to a larger size so that the number remains within the maximum allowable number. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2007-221510 [Patent Document 2] Japanese Patent Application Laid-Open No. 2012-242970 Summary of the Invention [Problem to be solved by the invention]
[0005] However, the conventional techniques disclosed in the above-mentioned Patent Documents 1 and 2 have a problem in that the client side cannot know the surplus processing resources that can be allocated to object recognition.
[0006] In view of the above-mentioned problems, an object of the present invention is to enable a client to recognize processing resources on the imaging device side that can be allocated to object recognition. [Means for solving the problem]
[0007] The imaging device of the present invention is an imaging device that generates video data and has: a recognition means for recognizing objects from video corresponding to the video data; a generation means for generating object information related to the objects recognized by the recognition means and storing it in the video data; and an output means for outputting the video data in which the object information generated by the generation means is stored to an external device; wherein the generation means generates predetermined information and stores it in the video data when the number of objects recognized by the recognition means is equal to or greater than a predetermined ratio of the maximum number of objects that the recognition means can recognize, and the generation means generates metadata as the predetermined information for displaying the frames of the recognized objects in a predetermined color. [Effects of the Invention]
[0008] According to the present invention, it is possible for the client side to recognize the processing resources of the image capture device that can be allocated to object recognition. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 2 is a block diagram showing an example of the internal configuration of a network camera. [Figure 2]10 is a flowchart illustrating an example of a processing procedure when the number of recognizable objects reaches an upper limit. [Figure 3] FIG. 10 is a diagram illustrating the maximum number of recognizable objects. [Figure 4] FIG. 10 is a diagram for explaining a method of notifying the degree of pressure on processing resources. DETAILED DESCRIPTION OF THE INVENTION
[0010] A network camera 100 according to an embodiment of the present invention will be described below with reference to FIG. 1 is a block diagram showing an example of the internal configuration of a network camera 100 according to this embodiment. The network camera 100 according to this embodiment includes an imaging unit 110 and a controller unit 120, and transfers video data from the imaging unit 110 to the controller unit 120. Communication also takes place between the network camera 100 and a network device 130.
[0011] The imaging unit 110 includes an optical lens 111 , an imaging element 112 , a signal processing circuit 113 , an imaging control circuit 114 , and a memory transfer circuit 115 . The image sensor 112 converts light focused through the optical lens 111 into an electric charge to generate an image signal. The signal processing circuit 113 receives the image signal generated by the image sensor 112 and digitizes it to generate a captured image.
[0012] The imaging control circuit 114 controls the imaging element 112 at the same cycle as the image output cycle. Furthermore, if the accumulation time is longer than the image output cycle, the imaging control circuit 114 controls the signal processing circuit 113 to hold the captured image in the frame memory of the signal processing circuit 113 during the period when the imaging element 112 cannot output an imaging signal. The memory transfer circuit 115 transfers the captured image digitized by the signal processing circuit 113 to the memory 122 in the controller unit 120.
[0013] The controller unit 120 includes a CPU 121 , a memory 122 , a network I / F 123 , a non-volatile memory 124 , and an encoding circuit 125 . The captured image transferred to the memory 122 is subjected to object recognition processing and encoding processing by the encoding circuit 125. The object recognition processing and encoding processing are performed by the encoding circuit 125, but may be performed by a module separate from the encoding circuit 125. The encoding circuit 125 performs object recognition processing on the captured image to generate information about the object. The information about the object includes object information such as the object type and position. The encoding circuit 125 then performs encoding processing on the captured image in accordance with the ARSEI standard, which is a standard for H.265, and stores the information about the object in the metadata section of the video data. Under the control of the CPU 121, the video data with the information about the object stored in the metadata section is transmitted to the outside of the network camera 100 via the network I / F 123.
[0014] The network device 130 communicates with the network camera 100 via the network I / F 123. An example of the network device 130 is a network hub. The information processing device 140 can perform various settings on the network camera 100 via the network device 130 and the network I / F 123, and can acquire data stored in the non-volatile memory 124. An example of the information processing device 140 is a personal computer (PC) or a smartphone.
[0015] Fig. 2 is a flowchart showing an example of a processing procedure when the number of recognizable objects reaches an upper limit. The processing procedure shown in Fig. 2 is a processing procedure when the number of objects recognized in the captured image obtained by the imaging unit 110 approaches the number of objects recognizable by the network camera. Note that each process shown in Fig. 2 is performed under the control of the CPU 121, and is performed by the CPU 121 reading out a program stored in the non-volatile memory 124. Processing begins when the network camera 100 is powered on and begins capturing and distributing images with a predetermined setting. Then, in S201, the CPU 121 uses the encoding circuit 125 to calculate the maximum number of objects that can be recognized while images are being captured and distributed.
[0016] An example of a method for calculating the maximum number of recognizable objects in S201 will be described using Fig. 3(a). Here, in order to calculate the maximum number of recognizable objects, it is assumed that a test video showing a large number of objects (15 in the case of Fig. 3(a)) is stored in the network camera 100. For example, when the network camera 100 operates at a frame rate of 30 fps, the time for one frame is 1 / 30 seconds. If ten objects 301 to 310 can be recognized in 1 / 30 seconds, and recognition of the eleventh object is not completed within 1 / 30 seconds, the maximum number of recognizable objects in the network camera is calculated to be 10.
[0017] 3(b), another example of the method for calculating the maximum number of recognizable objects shown in S201 will be described. Here, in order to calculate the maximum number of recognizable objects, it is assumed that a test video showing five objects 311 to 315 is stored in the network camera 100. If the time for one frame is 1 / 30 second and recognition of the five objects 311 to 315 is completed in 1 / 60 second, it is assumed that 10 objects can be recognized in 1 / 30 second. From this result, the maximum number of recognizable objects in the network camera is calculated to be 10.
[0018] Next, in S202, the CPU 121 acquires the number of objects recognized by the encoding circuit 125 in the captured image. Subsequently, in S203, the CPU 121 compares the maximum number of recognizable objects calculated in S201 with the number of objects acquired in S202. Then, it determines whether there is a surplus in the number of recognizable objects. In this embodiment, for example, if the number of recognized objects is less than 80% of the maximum number of recognizable objects, it is determined that there is a surplus in the number of recognizable objects. If this determination shows that there is not a surplus in the number of recognizable objects, the process proceeds to S204. Then, in S204, under the control of the CPU 121, the encoding circuit 125 inserts information indicating that processing resources for recognizing objects are tight into the metadata of the video data. Then, in S205, the encoding circuit 125 transmits the video data to the network device 130 via the network I / F 123.
[0019] On the other hand, if the result of the determination in S203 is that there is a surplus in the number of recognizable objects, there is no need to insert information indicating that the processing resources for recognizing objects are tight into the metadata. Therefore, the process proceeds to S205, and the video data is transmitted to the network device 130 via the network I / F 123.
[0020] Next, in S206, the CPU 121 determines whether or not an event has occurred that requires the recalculation of the maximum number of recognizable objects. If the result of this determination is that an event has occurred that requires the recalculation of the maximum number of recognizable objects, the process returns to S201; if not, the process returns to S202. Here, an event that requires the recalculation of the maximum number of recognizable objects is an event that changes the maximum number of recognizable objects of the network camera 100. Examples include when the settings (frame rate, etc.) of the network camera 100 are changed, when an add-on application is added, or when settings are changed to draw an additional stream.
[0021] 4(a) and 4(b), an example of the process of inserting information indicating that processing resources are constrained into metadata in S204 will be described. For example, assume that the maximum number of recognizable objects calculated in S201 is 10, and the number of objects acquired in S202 is 8. In this case, since the ratio to the maximum number of recognizable objects is approaching 80%, it is determined in S203 that there is no room for the number of recognizable objects. In this case, by setting 'bb color yellow', for example, in ar_label[ar_label_idx[i]], which is the syntax of ARSEI, information indicating that processing resources are constrained is inserted into the metadata of the video data. As a result, on the information processing device 140 side, the bounding boxes surrounding objects 401 to 408 are displayed in yellow instead of the usual white, as shown in FIG. 4(a).
[0022] Also, if the number of objects acquired in S202 is 10, it is possible that not all objects have been recognized, and it is determined in S203 that there is not enough recognizable objects. In this case, ar_label[ar_label_idx[i]] is set to, for example, 'bb color red'. As a result, on the information processing device 140 side, the bounding boxes surrounding the objects 411 to 420 are displayed in red, as shown in FIG. 4(b).
[0023] As described above, in this embodiment, information indicating that processing resources are constrained is inserted into the metadata of the video data. Then, by checking the ARSEI syntax ar_label[ar_label_idx[i]], the client side can determine the degree of constrained processing resources required for object recognition. In other words, the client side can recognize that a lack of processing resources in the network camera may cause various problems, such as object recognition not being completed within one frame, transmission delays, and frame drops. As shown in Figure 4, by coloring when the number of recognized objects exceeds a certain percentage (maximum number of recognizable objects), the client side can recognize the degree of constrained processing resources in the network camera.
[0024] The degree of resource congestion of the network camera may be indicated not only by color information but also by other visual information. Specifically, the encoding circuit 125 of the network camera 100 can provide other visual information, such as blinking the bounding box, depending on the content set in ar_label[ar_label_idx[i]]. Furthermore, when the number of recognized objects reaches the maximum number of recognizable objects, the blinking speed of the bounding box may be changed, for example, by blinking at a high speed.
[0025] As described above, according to this embodiment, it is possible to notify the client side of the degree of pressure on the processing resources required for object recognition by the network camera, which allows the client side to make appropriate setting changes to the camera side, thereby avoiding various inconveniences caused by a lack of processing resources in the network camera.
[0026] (Other embodiments) Although preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments and various modifications and variations are possible within the scope of the gist of the present invention. The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or storage medium, and having one or more processors in the computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., an ASIC) that realizes one or more functions. [Explanation of symbols]
[0027] 121 CPU, 125 encoding circuit
Claims
1. An imaging device that generates video data, recognition means for recognizing an object from an image corresponding to the image data; a generating means for generating object information relating to the object recognized by the recognizing means and storing the object information in the video data; an output means for outputting the video data in which the object information generated by the generation means is stored to an external device; and the generating means generates predetermined information and stores it in the video data when the number of objects recognized by the recognizing means is equal to or greater than a predetermined ratio of the maximum number of objects that can be recognized by the recognizing means; The imaging device, wherein the generating means generates, as the predetermined information, metadata for displaying a frame of the recognized object in a predetermined color.
2. 2. The imaging device according to claim 1, wherein the generation means generates the metadata for displaying the frames of the recognized objects in different colors when the number of objects recognized by the recognition means is equal to or greater than a predetermined ratio of the maximum number of objects recognizable by the recognition means and has not reached the maximum number of objects recognizable by the recognition means, and when the number of objects recognized by the recognition means has reached the maximum number of objects recognizable by the recognition means.
3. An imaging device that generates video data, recognition means for recognizing an object from an image corresponding to the image data; a generating means for generating object information relating to the object recognized by the recognizing means and storing the object information in the video data; an output means for outputting the video data in which the object information generated by the generation means is stored to an external device; and the generating means generates predetermined information and stores it in the video data when the number of objects recognized by the recognizing means is equal to or greater than a predetermined ratio of the maximum number of objects that can be recognized by the recognizing means; The imaging device, wherein the generating means generates, as the predetermined information, metadata for causing a frame of the recognized object to blink.
4. 4. The imaging device according to claim 3, wherein the generation means generates the metadata for displaying the frames of the recognized objects at different blinking speeds when the number of objects recognized by the recognition means is equal to or greater than a predetermined ratio of the maximum number of objects recognizable by the recognition means and has not yet reached the maximum number of objects recognizable by the recognition means, and when the number of objects recognized by the recognition means has reached the maximum number of objects recognizable by the recognition means.
5. 5. The imaging device according to claim 1, wherein the generating means stores the predetermined information in metadata included in the video data.
6. A method for controlling an imaging device that generates video data, comprising: a recognition step of recognizing an object from an image corresponding to the image data; a generating step of generating object information relating to the object recognized in the recognizing step and storing the object information in the video data; an output step of outputting the video data in which the object information generated in the generation step is stored to an external device; and In the generating step, when the number of objects recognized in the recognizing step is equal to or greater than a predetermined ratio with respect to the maximum number of objects recognizable in the recognizing step, predetermined information is generated and stored in the video data; A method for controlling an imaging device, wherein in the generating step, metadata for displaying a frame of the recognized object in a predetermined color is generated as the predetermined information.
7. A method for controlling an imaging device that generates video data, comprising: a recognition step of recognizing an object from an image corresponding to the image data; a generating step of generating object information relating to the object recognized in the recognizing step and storing the object information in the video data; an output step of outputting the video data in which the object information generated in the generation step is stored to an external device; and In the generating step, when the number of objects recognized in the recognizing step is equal to or greater than a predetermined ratio with respect to the maximum number of objects recognizable in the recognizing step, predetermined information is generated and stored in the video data; A method for controlling an imaging device, wherein in the generating step, metadata for causing a frame of the recognized object to blink is generated as the predetermined information.
8. A program for causing a computer to function as each of the means of the imaging device according to any one of claims 1 to 5.
Citation Information
Patent Citations
Network camera and monitoring system using the same
JP2007221510A
Image processing device and control method therefor
JP2012242970A