System and method for visual search

JP7901564B2Active Publication Date: 2026-08-06MOTOROLA SOLUTIONS INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
MOTOROLA SOLUTIONS INC
Filing Date
2023-07-19
Publication Date
2026-08-06

Smart Images

  • Figure 0007901564000001
    Figure 0007901564000001
  • Figure 0007901564000002
    Figure 0007901564000002
  • Figure 0007901564000003
    Figure 0007901564000003
Patent Text Reader

Abstract

To provide an appearance search system, a method and a computer readable medium including one or more cameras, where a video having an image of an object.SOLUTION: A video capture / playback system includes one or more processors and a memory containing a computer program code, and implements a method having the one or more processors. A method 400 comprises the steps of: identifying, by a camera, one or more objects in images of objects; generating a signature for an object identified by a server and mounting a learning machine to generate the signature for an object of interest; transmitting the images of the objects from the camera to the one or more processors via a network; comparing the signature of the identified object with the signature of the object of interest to generate a similarity score for the identified object; and transmitting, on the basis of the similarity score, an instruction for presenting one or more images of the objects on a display.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Related Applications This specification claims the benefit of U.S. Provisional Patent Application No. 62 / 430,292, filed Dec. 5, 2016, and U.S. Provisional Patent Application No. 62 / 527,894, filed Jun. 30, 2017, and incorporates by reference the entirety of both documents herein.

[0002] Technical Field The present subject matter relates to video surveillance, and more particularly, to identifying objects of interest in video of a video surveillance system.

Background Art

[0003] Computer-executed visual object classification, also known as object recognition, involves classifying visual representations of real-life objects found in still images or videos captured by a camera. By performing visual object classification, each visual object found in a still image or video is classified according to its type (e.g., human, vehicle, or animal, etc.).

[0004] In automated security and surveillance systems, image data such as video or video footage is typically collected using a video camera or other imaging device or sensor. In the simplest systems, the images represented by the image data are displayed by a security officer for simultaneous viewing and / or recorded for later review after a security breach. In such systems, the task of detecting and classifying visually interesting objects is performed by a human observer. Significant progress would occur if the system itself had the ability to perform object detection and classification, either in part or in whole.

[0005] Conventional surveillance systems often focus on detecting objects such as people, vehicles, and animals moving within their environment. However, if, for example, a child gets lost in a large shopping mall, it can be extremely time-consuming for security personnel to manually review video footage to locate the lost child. Computer-driven detection of objects within images captured by cameras significantly simplifies the process of security personnel reviewing relevant video segments, allowing for the timely discovery of the lost child.

[0006] That being said, computer-based video analysis for detecting and recognizing objects, and identifying similar objects, requires considerable computing resources, especially as the desired accuracy increases. Distributing the processing to optimize resource utilization makes it easier for computers to execute. [Overview of the Initiative]

[0007] A first aspect of the present disclosure provides an appearance search system comprising one or more cameras configured to capture video of a scene, wherein the video has images of objects. The system comprises one or more processors and memory, which include computer program code stored in the memory, and which, when executed by the one or more processors, are configured to perform a certain method. The method includes identifying one or more objects within an image of objects. The method further includes generating signatures for the identified objects and implementing a learning machine configured to generate signatures for objects of interest. The system further includes a network configured to transmit images of objects from the cameras to one or more processors. The method further includes comparing the signatures for the identified objects with signatures for objects of interest to generate similarity scores for the identified objects and transmitting instructions to present one or more images of objects on a display based on the similarity scores.

[0008] This system may also include a storage system for storing signatures generated from identified objects and images. The implemented learning machine may be a second learning machine, and the discrimination may be performed by a first learning machine implemented by one or more processors.

[0009] The first and second learning machines may include neural networks. The neural networks may include convolutional neural networks. The neural networks or convolutional neural networks include training models.

[0010] The system may further include one or more graphics processing units that power the first and second learning machines. One or more cameras may be further configured to capture images of objects using video analysis.

[0011] One or more cameras may be configured to sort images of objects by object classification. One or more cameras may be configured to identify one or more images containing human objects, and the network may be configured to send only the identified images to one or more processors.

[0012] The image of an object may include a portion of the image frame of the video. The portion of the image frame may include a first image portion of the image frame, and the first image portion includes at least an object. The portion of the image frame may include a second image portion of the image frame, and the second image portion is larger than the first image portion. The first learning machine may be configured to draw the outlines of one or more objects, or all of them, within the second image portion for the second learning machine.

[0013] One or more cameras may be configured to generate reference coordinates so that images of objects can be extracted from the video. A memory system may be configured to store the reference coordinates.

[0014] One or more cameras may be further configured to select one or more images from video captured over a certain period of time to obtain one or more images of an object. Object identification may include drawing the outlines of one or more objects in the image.

[0015] Identification may include identifying multiple objects within at least one image, and dividing at least one image into multiple segmented images, with each segmented image containing at least a portion of one of the identified objects. The method may further include determining the confidence level for each identified object, and, if the confidence level does not meet the confidence requirements, having the first learning machine perform identification and segmentation, or if the confidence level does meet the confidence requirements, having the second learning machine perform identification and segmentation.

[0016] One or more cameras may further include one or more video analysis modules for determining reliability. In yet another aspect of the present disclosure, a method is provided which includes capturing video of a scene and having images of objects. The method further includes identifying one or more objects within the images of objects. The method further includes generating signatures of the identified objects and signatures of objects of interest using a learning machine. The method further includes generating similarity scores for the identified objects by comparing the signatures of the identified objects with first signatures of the objects of interest. The method further includes presenting one or more images of the objects on a display based on the similarity scores.

[0017] The method may further include performing any of the above steps or operations in combination with the first aspect of the present disclosure. In yet another aspect of the present disclosure, a computer-readable medium is provided which stores computer program code executable by one or more processors and is configured such that when executed by one or more processors, one or more processors perform a certain method. The method includes capturing a video of a scene, wherein the video has images of objects. The method further includes identifying one or more objects within the images of objects. The method further includes using a learning machine to generate signatures of the identified objects and signatures of objects of interest. The method further includes generating similarity scores for the identified objects by comparing the signatures of the identified objects with first signatures of objects of interest. The method further includes presenting one or more images of objects on a display based on the similarity scores.

[0018] A method performed by one or more processors may further include performing any of the above steps or operations in combination with the first aspect of this disclosure. In yet another aspect of the present disclosure, a system is provided comprising one or more cameras configured to capture video of a scene. The system further comprises one or more processors and memory containing computer program code stored in memory, which is configured to perform a certain method when executed by one or more processors. The method comprises extracting a chip from the video, the chip containing an image of an object. The method further comprises identifying a plurality of objects within at least one chip. The method further comprises dividing at least one chip into a plurality of segmented chips, each segmented chip containing at least a portion of one of the identified objects.

[0019] The method may further include implementing a learning machine configured to generate signatures of objects of interest by generating signatures of identified objects. The learning machine may be a second learning machine, and identification and segmentation may be performed by a first learning machine implemented by one or more processors. The method may further include determining the confidence level for each identified object, and if the confidence level does not meet the confidence requirements, having the first learning machine perform identification and segmentation, or if the confidence level does meet the confidence requirements, having the second learning machine perform identification and segmentation. One or more cameras may be equipped with one or more video analysis modules for determining confidence levels.

[0020] At least one chip may contain at least one padded chip. Each padded chip may contain a first image portion of the image frame of the video. At least one chip may further contain at least one unpadded chip. Each unpadded chip may contain a second image portion of the image frame of the video, the second image portion being smaller than the first image portion.

[0021] In yet another aspect of the present disclosure, a computer-readable medium is provided, which stores computer program code executable by one or more processors and is configured such that when executed by one or more processors, one or more processors perform a certain method. The method includes obtaining an image of a scene. The method further includes extracting a chip from the image, the chip containing an image of an object. The method further includes identifying a plurality of objects within at least one chip. The method further includes dividing at least one chip into a plurality of segmented chips, each segmented chip containing at least a portion of one of the identified objects.

[0022] A method performed by one or more processors may further include performing any of the above steps or operations directly in conjunction with the above system. In yet another aspect of the present disclosure, there is provided an appearance search system including a camera that captures an image of a scene, the image having an image of an object, a processor including a learning machine that generates a signature from the image of the object associated with the image and generates a first signature from a first image of an object of interest, a network for transmitting the image of the object from the camera to the processor, and a storage system that stores the generated signature of the object and the associated image. The processor further generates a similarity score by comparing the signature from the image with the first signature of the object of interest, and further prepares an image of an object having a higher similarity score and presents it to the user on a display.

[0023] According to some exemplary embodiments, the learning machine is a neural network. According to some exemplary embodiments, the neural network is a convolutional neural network.

[0024] According to some exemplary embodiments, the neural network is a trained model. According to some exemplary embodiments, a graphics processing unit is used to operate the learning machine.

[0025] According to some exemplary embodiments, the image of the object is captured by a camera and processed by the camera using video analysis. According to some exemplary embodiments, the image of the object is screened by classifying the type of the object with the camera before being transmitted to the processor.

[0026] According to some exemplary embodiments, the type of the object transmitted to the processor is a human. According to some exemplary embodiments, a camera that captures an image of an object from a video further includes capturing reference coordinates of an image in the video and extracting the image of the object from the video based on the reference coordinates.

[0027] According to some exemplary embodiments, the image extracted from the video is deleted, and the storage system stores the signature, the reference coordinates, and the video. According to some exemplary embodiments, video analysis selects one or more images of an object over a certain period of time and represents the images of the object captured during that period.

[0028] In yet another aspect of the present disclosure, there is provided a computer-executable method for performing an appearance search on an object of interest in a video captured by a camera, the method comprising extracting an image of the object from the video captured by the camera, transmitting the image of the object and the video to a processor over a network, generating, by the processor, a signature from the image of the object using a learning machine, storing, in a storage system, the signature of the object and the video associated with the object, generating, by the processor, a signature from an image of any object of interest using a learning machine, comparing, by the processor, the signature from the image in the storage system with the signature of the object of interest to generate a similarity score for each comparison, and presenting, on a display, to a user, an image of an object having a higher similarity score.

[0029] In yet another aspect of the present disclosure, a computer-executable method for visually searching for an object of interest in a video captured by a camera is provided, the method comprising: extracting an image of an object from a video captured by a camera; transmitting the image and video of the object to a processor over a network; using a learning machine, the processor generating a signature from the image of the object, the image of the object including an image of the object of interest; storing the object's signature and associated video in a memory system; searching for an instance of the image of the object of interest via the memory system; retrieving a signature of the object of interest from the memory for the instance of the image of the object of interest; and having the processor compare the signature from the image in the memory system with the signature of the object of interest to generate a similarity score for each comparison; and preparing and presenting images of objects with higher similarity scores to the user on a display.

[0030] In yet another aspect of the present disclosure, a non-transient, computer-readable storage medium is provided, which stores instructions causing the processor to perform a method for visually retrieving an object of interest in a video captured by a camera when executed by the processor, the method comprising: extracting an image of an object from a video captured by a camera; transmitting the image and video of the object to the processor over a network; using a learning machine, having the processor generate a signature from the image of the object; storing in a storage system that the image of the object includes an image of the object of interest, the signature of the object, and the video associated with the object; retrieving an instance of the image of the object of interest from the storage unit for the instance of the image of the object of interest; and having the processor compare the signature from the image in the storage system with the signature of the object of interest to generate a similarity score for each comparison; and preparing and presenting to the user on a display an image of the object with a higher similarity score.

[0031] For a detailed explanation, please refer to the following diagram. [Brief explanation of the drawing]

[0032] [Figure 1] This is a block diagram of connected devices in a video capture and playback system according to an exemplary embodiment. [Figure 2A] This is a block diagram of a series of operating modules for a video capture and playback system according to one exemplary embodiment. [Figure 2B] This is a block diagram of a set of operating modules in one particular exemplary embodiment, in which a video analysis module 224, a video management module 232, and a storage device 240 are fully implemented on one or more image acquisition devices 108. [Figure 3] This is a flowchart illustrating an exemplary embodiment of a method for performing video analysis on one or more image frames of video captured by a video acquisition device. [Figure 4] This is a flowchart illustrating an exemplary embodiment of a method for performing appearance matching to identify the location of an object of interest on one or more image frames of video captured by a video acquisition device (camera). [Figure 5] Figure 4 is a flowchart illustrating an exemplary embodiment of a visual search that performs visual matching on the client to locate the location of an object of interest in recorded video. [Figure 6] Figure 4 is a flowchart illustrating an exemplary embodiment of a time-specified appearance search that performs appearance matching on client 420 to identify the location of video footage in which an object of interest was recorded either before or after a selected time. [Figure 7] This is an illustrative block diagram of metadata for an object profile before it is stored and for an object profile that has been reduced in size for storage. [Figure 8] Figure 4 shows a scene and trimming boundary box of an exemplary embodiment. [Figure 9] This is a block diagram of a series of operational submodules of a video analysis module according to one exemplary embodiment. [Figure 10A] This is a block diagram of a process for generating feature vectors according to one exemplary embodiment. [Figure 10B] This is a block diagram of an alternative process for generating feature vectors according to another exemplary embodiment. [Figure 11] This is a flowchart of an exemplary embodiment for generating a trimming boundary box. [Figure 12] This figure shows an example of the image seen by the camera, a padded cropped bounding box, and a cropped bounding box generated by the analysis module. [Modes for carrying out the invention]

[0033] For the sake of simplification and clarity, it should be understood that the elements shown in the drawings are not necessarily depicted to actual size. For example, the dimensions of some elements may be exaggerated compared to others for clarity. Furthermore, where necessary, reference numerals may be repeated throughout the drawings to indicate corresponding or identical elements.

[0034] Numerous specific details are provided to ensure a thorough understanding of the exemplary embodiments described herein. However, those skilled in the art will understand that the embodiments described herein can be carried out without these specific details. In other instances, known methods, procedures, and components are omitted to avoid obscuring the embodiments described herein. Furthermore, this description should not be considered as limiting the scope of the embodiments described herein, but rather as simply describing the various embodiments described herein.

[0035] When the words "a" or "an" are used in a claim and / or specification with the terms "comprising" or "including," they may mean "one," but unless otherwise specified, they can also mean "one or more," "at least one," and "more than one." Similarly, the word "another" may mean at least the second and subsequent, unless otherwise specified.

[0036] The terms “coupled,” “coupling,” and “connected” as used herein may have several different meanings depending on the context in which they are used. For example, the terms coupled, coupling, and connected may have mechanical or electrical connotations. For instance, as used herein, the terms coupled, coupling, and connected may, depending on the specific context, refer to two elements or devices being directly connected to each other or connected to each other by one or more intermediary elements or devices via electrical elements, electrical signals, or mechanical elements.

[0037] In this specification, an image may encompass multiple consecutive image frames, where the image frames together form images captured by a video camera. Each image frame may be represented by a matrix of pixels, where each pixel has a pixel image value. For example, the pixel image value may be a grayscale number (e.g., 0 to 255) or, in the case of a color image, a number of values. Examples of color spaces used to represent the pixel image values ​​of image data include RGB, YUV, CYKM, YCBCR4:2:2, and YCBCR4:2:0.

[0038] As used herein, “metadata” or its derivatives refer to information obtained through computer-executed analysis of images, such as images within video. For example, video processing may include, but is not limited to, image processing, analysis, management, compression, encoding, storage, transmission, and / or playback of video data. Video analysis may include segmenting image frame regions and detecting visual objects, tracking, and / or classifying visual objects located within the captured scene represented by the image data. Processing of image data may also produce additional information about the image data or visual objects captured within the image. For example, such additional information is generally understood to be metadata. Metadata may also be used for other processing of image data, such as drawing bounding boxes around objects detected within the image frame.

[0039] As those skilled in the art will understand, the various exemplary embodiments described herein can be embodied as methods, systems, or computer program products. Thus, the various exemplary embodiments can take the form of entirely hardware embodiments, entirely software embodiments (including firmware, resident software, microcode, etc.), or embodiments combining software and hardware aspects, all of which may collectively be referred to herein as “circuits,” “modules,” or “systems.” Furthermore, the various exemplary embodiments can take the form of computer program products on a computer-usable storage medium having computer-usable program code embedded in the medium.

[0040] Any suitable computer-usable or computer-readable medium may be used. This computer-usable or computer-readable medium may, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, devices, or propagation media. In the context of this specification, a computer-usable or computer-readable medium may be any medium that can contain, store, communicate, propagate, or carry programs for use by, or in connection with, an instruction execution system, apparatus, or device.

[0041] Computer program code for performing the operations of various exemplary embodiments may be written in an object-oriented programming language such as Java, Smalltalk, C++, or Python. However, computer program code for performing the operations of various exemplary embodiments may also be written in a conventional procedural programming language, such as the C programming language or a similar programming language. The program code may run entirely on a computer, partially on a computer, as a standalone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server. In the last scenario, the remote computer may be connected to the computer via a local area network (LAN) or wide area network (WAN), or this connection may be to an external computer (for example, via the Internet using an Internet service provider).

[0042] Various exemplary embodiments are described below with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It will be understood that each block in the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be executed by computer program instructions. These computer program instructions may be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device for the purpose of manufacturing a machine, thereby creating means for executing the instructions via the computer's processor or other programmable data processing device to perform the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams.

[0043] These computer program instructions can also be stored in computer-readable memory, which can instruct a computer or other programmable data processing device to function in a specific way, thereby generating manufacturing items that include instructions to perform functions / operations specified in one or more blocks of a flowchart and / or block diagram.

[0044] Computer program instructions can also be loaded into a computer or other programmable data processing device to perform a series of operations on the computer or other programmable device, thereby creating a computer-executable process, in which the instructions executed by the computer or other programmable device provide a process that performs a function / operation specified in one or more blocks of a flowchart and / or block diagram.

[0045] Referring to Figure 1, the diagram shows a block diagram of the connection devices of a video capture and playback system 100 according to an exemplary embodiment. For example, the video capture and playback system 100 may be used as a video surveillance system. The video capture and playback system 100 comprises hardware and software that perform the processes and functions described herein.

[0046] The video capture and playback system 100 includes at least one video capture device 108 that operates to capture multiple images and generate image data representing the multiple captured images. The video capture device 108 or camera 108 is an image capture device and includes a security video camera.

[0047] Each video acquisition device 108 includes at least one image sensor 116 for acquiring multiple images. The video acquisition device 108 may be a digital video camera, and the image sensor 116 may output the acquired light as digital data. For example, the image sensor 116 may be a CMOS, NMOS, or CCD. In some embodiments, the video acquisition device 108 may be an analog camera connected to an encoder.

[0048] At least one image sensor 116 may operate to capture light in one or more frequency ranges. For example, at least one image sensor 116 may operate to capture light in a range substantially corresponding to the visible light frequency range. In other examples, at least one image sensor 116 may operate to capture light outside the visible light range, for example, light in the infrared and / or ultraviolet range. In other examples, the image acquisition device 108 may be a multi-sensor camera having two or more sensors that operate to capture light in separate frequency ranges.

[0049] At least one video acquisition device 108 may be equipped with a dedicated camera. In this specification, a dedicated camera refers to a camera whose primary feature is the acquisition of images or video. In some exemplary embodiments, the dedicated camera may perform functions related to the acquired images or video, such as processing image data generated by the camera or another video acquisition device 108, but is not limited to these. For example, the dedicated camera may be a surveillance camera, such as a pan-tilt-zoom camera, dome camera, ceiling camera, box camera, or bullet camera.

[0050] In addition to, or instead of, the video capture device 108 may include at least one embedded camera. It will be understood herein that an embedded camera refers to a camera embedded within a device that operates to perform a function unrelated to the captured image or video. For example, an embedded camera may be a camera found in any one of the following: a laptop, tablet, drone device, smartphone, or video game console or controller.

[0051] Each video capture device 108 comprises one or more processors 124, and one or more memory devices 132 connected to these processors and one or more network interfaces. The memory devices may include local memory (e.g., random access memory and cache memory) used in the execution of program instructions. The processors execute computer program instructions (e.g., operating system and / or application programs), and these instructions may be stored in the memory devices.

[0052] In various embodiments, the processor 124 may be implemented by any suitable processing circuit having one or more circuit units, including a digital signal processor (DSP), a processor with an embedded graphics processing unit (GPU), and any suitable combination thereof, and these processors may operate separately, in parallel, or redundantly. Such processing circuit may be implemented by one or more integrated circuits (ICs), including monolithic integrated circuits (MICs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or any suitable combination thereof. In addition to or instead, such processing circuit may be implemented as, for example, a programmable logic controller (PLC). The processor may have circuits for storing memory, such as digital data, and may include memory circuits or communicate with memory circuits via wires, for example.

[0053] In various exemplary embodiments, the memory device 132 connected to the processor circuit operates to store data and computer program instructions. Typically, the memory device is all or part of a digital electronic integrated circuit, or is formed from multiple digital electronic integrated circuits. The memory device may be implemented as, for example, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, one or more flash drives, a universal serial bus (USB) connected to the memory unit, a magnetic storage device, an optical storage device, a magneto-optical storage device, or any combination thereof. The memory device may operate to store memory such as volatile memory, non-volatile memory, dynamic memory, or any combination thereof.

[0054] In various exemplary embodiments, multiple components of the image acquisition device 108 may be implemented together within a system-on-a-chip (SOC). For example, the processor 124, memory device 116, and network interface may be implemented within the SOC. Furthermore, in such an implementation, the general-purpose processor and one or more GPUs and DSPs may be implemented together within the SOC.

[0055] Continuing in Figure 1, each of the at least one video acquisition device 108 is connected to the network 140. Each video acquisition device 108 outputs the image data represented by the image it acquires and operates to transmit the image data over the network.

[0056] It will be understood that network 140 may be any suitable communication network for sending and receiving data. For example, network 140 may be a local area network, an external network (e.g., a WAN, or the internet), or a combination thereof. In other examples, network 140 may include a cloud network.

[0057] In some examples, the video capture and playback system 100 includes a processing unit 148. The processing unit 148 operates to process image data output by the video capture device 108. The processing unit 148 also includes one or more processors and one or more memory devices connected to the processors (CPUs). The processing unit 148 may also include one or more network interfaces. For the sake of explanation, only one processing unit 148 is shown, but it will be understood that the video capture and playback system 100 may include any appropriate number of processing units 148.

[0058] For example, as shown in the figure, the processing device 148 is connected to a video acquisition device 108 which may not have memory 132 or CPU 124 for processing image data. The processing device 148 may also be connected to a network 140.

[0059] According to one exemplary embodiment, and as shown in Figure 1, the video capture and playback system 100 comprises at least one workstation 156 (e.g., a server), each having one or more processors, including a graphics processing unit (GPU). At least one workstation 156 may also have storage memory. The workstation 156 receives image data from at least one video capture device 108 and performs processing on the image data. The workstation 156 may further send commands to manage and / or control one or more image capture devices 108. The workstation 156 may receive unprocessed image data from the video capture device 108. Alternatively, or in addition to this, the workstation 156 may receive image data that has already undergone some intermediate processing, such as processing at the video capture device 108 and / or processing equipment 148. The workstation 156 may receive metadata from the image data and perform further processing on the image data.

[0060] Although Figure 1 shows a single workstation 156, it should be understood that workstations can be implemented as a collection of multiple workstations. The video capture and playback system 100 further comprises at least one client device 164 connected to the network 140. The client device 164 is used by one or more users to interact with the video capture and playback system 100. Thus, the client device 164 comprises at least one display device and at least one user input device (e.g., a mouse, keyboard, or touch screen). The client device 164 operates to display a user interface on its display device that displays information, receives user input, and plays back video. For example, the client device may be any one of the following: a personal computer, laptop, tablet, personal digital assistant (PDA), mobile phone, smartphone, game console, or other mobile device.

[0061] The client device 164 operates to receive image data on the network 140 and to play back the received image data. The client device 164 may also have the ability to process the image data. For example, the processing capabilities of the client device 164 can be limited to processing related to the ability to play back the received image data. In another example, the image processing capabilities may be shared between the workstation and one or more client devices 164.

[0062] In some examples, the image acquisition and playback system 100 may be implemented without a workstation 156. Therefore, the image processing function may be performed entirely by one or more video acquisition devices 108. Alternatively, the image processing function may be shared by two or more of the video acquisition device 108, processing equipment 148, and client device 164.

[0063] Next, referring to Figure 2A, the diagram shows a block diagram of a set of 200 operating modules for a video capture and playback system 100 according to one exemplary embodiment. The operating modules may be implemented in hardware, software, or both on one or more devices of the video capture and playback system 100 as shown in Figure 1.

[0064] The set of operating modules 200 includes at least one video acquisition module 208. For example, each video acquisition device 108 may implement a video acquisition module 208. The video acquisition module 208 operates to control one or more components of the video acquisition device 108 (e.g., a sensor 116) to acquire an image.

[0065] The set of operational modules 200 includes a subset 216 of image data processing modules. For example, as shown in the figure, the subset 216 of image data processing modules includes a video analysis module 224 and a video management module 232.

[0066] The video analysis module 224 receives image data, analyzes the image data, and determines the characteristics or features of the captured image or video, and / or the characteristics or features of objects seen in the scene represented in the image or video. Based on the determination, the video analysis module 224 may further output metadata that provides information about the determination. Examples of decisions made by the video analysis module 224 include one or more of the following: foreground / background segmentation, object detection, object tracking, object classification, virtual tripwire, anomaly detection, face detection, face recognition, license plate recognition, identification of "behind" or "deleted" objects, and business intelligence. However, it will be understood that other video analysis functions known in the prior art may also be implemented by the video analysis module 224.

[0067] The video management module 232 receives image data and performs processing functions on the image data related to the transmission, playback, and / or storage of the video. For example, the video management module 232 can process the image data so that it can be transmitted according to bandwidth requirements and / or capacity. The video management module 232 may also process the image data according to the playback capabilities of the client device 164 that plays the video, such as the processing power and / or resolution of the display of the client device 164. The video management module 232 may also process the image data according to the storage capacity in the video acquisition and playback system 100 in order to store the image data.

[0068] According to some exemplary embodiments, it will be understood that a subset of the video processing module 216 may include only one of the video analysis module 224 and the video management module 232.

[0069] The set of operational modules 200 further includes a subset of storage modules 240. For example, as shown in the figure, the subset of storage modules 240 includes a video storage module 248 and a metadata storage module 256. The video storage module 248 stores image data, which may be image data processed by the video management module. The metadata storage module 256 stores information data output from the video analysis module 224.

[0070] Although the video storage module 248 and the metadata storage module 256 are shown as separate modules, it will be understood that they may be implemented within the same hardware storage device, thereby implementing logical rules for separating the stored video from the stored metadata. In other exemplary embodiments, the video storage module 248 and / or the metadata storage module 256 may be implemented within a plurality of hardware storage devices that may implement a distributed storage scheme.

[0071] The set of operating modules further includes at least one video playback module 264, which operates to receive image data and play the image data as video. For example, the video playback module 264 may be implemented in a client device 164.

[0072] The operating modules of set 200 may be implemented in one or more of the image acquisition device 108, processing equipment 148, workstation 156, and client device 164. In some exemplary embodiments, the operating modules may be fully implemented in a single device. For example, the video analysis module 224 may be fully implemented in the workstation 156. Similarly, the video management module 232 may be fully implemented in the workstation 156.

[0073] In other exemplary embodiments, some functions of the operation module of set 200 may be partially implemented on a first device, and the remaining functions of the operation module may be implemented on a second device. For example, the video analysis function may be shared among one or more of the image acquisition device 108, processing equipment 148, and workstation 156. Similarly, the video management function may be shared among one or more of the image acquisition device 108, processing equipment 148, and workstation 156.

[0074] Referring now to Figure 2B, the diagram shows a block diagram of a set of operating modules 200 for a video acquisition and playback system 100 according to one particular exemplary embodiment, in which the video analysis module 224, the video management module 232, and the storage device 240 are fully implemented in one or more image acquisition devices 108. Alternatively, the video analysis module 224, the video management module 232, and the storage device 240 are fully implemented in a processing device 148.

[0075] It will be understood that by enabling the implementation of a subset 216 of the image data (video) processing module in a single device or various devices of the video acquisition and playback system 100, the system 100 can be configured flexibly.

[0076] For example, one may choose to use a specific device that has a certain function together with another device that does not have the same function. This can be useful when integrating devices from different parties (e.g., manufacturers) or when installing an existing video capture and playback system.

[0077] Referring now to Figure 3, the diagram shows a flowchart of an exemplary embodiment of a method 350 for performing video analysis on one or more image frames of video captured by the video capture device 108. The video analysis is performed by the video analysis module 224 to determine the characteristics or features of the captured image or video, and / or the characteristics or features of visual objects seen in the scene captured within the video.

[0078] In version 300, at least one image frame of the video is segmented into a foreground region and a background region. This segmentation separates the regions of the image frame corresponding to moving objects (or objects that have moved beforehand) in the captured scene from the stationary regions of that scene.

[0079] In 302, one or more foreground visual objects within a scene represented by an image frame are detected based on the segmentation of 300. For example, some randomly adjacent foreground regions or "blobs" may be identified as foreground visual objects in the scene. For example, only adjacent foreground regions larger than a certain size (e.g., number of pixels) may be identified as foreground visual objects in the scene.

[0080] Further metadata may be generated for one or more detected foreground regions. The metadata may define the foreground visual object, or the object's position and reference coordinates within the image frame. For example, position metadata may be used to generate a bounding box (e.g., when encoding or playing back video) that shows the contour of the detected foreground visual object. The image within the bounding box is extracted and included in the metadata, called a trimmed bounding box (also called a "tip"), which may be further processed on other devices, such as workstation 156 on network 140, along with the associated video. In short, a trimmed bounding box, or tip, is a portion of the image frame of the video containing the detected foreground visual object. The extracted image is the trimmed bounding box and may be either smaller or larger than the content within the bounding box. The size of the extracted image should be, for example, close to the actual boundary of the detected object, but not exceed it. The bounding box is usually rectangular, but may be an irregular shape that is approximately the same as the object's contour. The bounding box may, for example, roughly follow the boundary (contour) of a human object. In yet another embodiment, the size of the extracted image is larger than the actual boundary of the detected object, and is referred to herein as a padded cropped bounding box (also called a “padding chip”). The padded cropped bounding box may be twice the size of the bounding box, for example, to include all or part of an object that is close to or overlaps with the detected foreground visual object. To clarify further, the padded cropped bounding box has an image larger than the cropped bounding box of the image of the object within the bounding box (referred to herein as an unpadding cropped bounding box). To clarify further, the cropped bounding boxes used herein include padded cropped bounding boxes and unpadding cropped bounding boxes. It will be understood that the image size of the padded cropped bounding box may vary from slightly larger (e.g., 10% larger) to considerably larger (e.g., 1000% larger).

[0081] In embodiments of this specification, a padded cropping bounding box is described as an enlarged unpadded cropping bounding box containing extra pixels but still maintaining the reference coordinates of the original unpadded cropping bounding box, although this enlargement or extra pixels may be further added along the horizontal axis instead of the vertical axis. Furthermore, the enlargement of the extra pixels may be symmetrical or asymmetrical with respect to the axis with respect to the object. The object in the unpadded cropping bounding box may be at the center of both the padded cropping bounding box and the unpadded cropping bounding box, although in some embodiments such an object may be off-center.

[0082] In some embodiments, the trimming bounding box, which includes padded and unpadded trimming bounding boxes, may be the reference coordinates of the image frame of the video, rather than an image actually extracted from the image frame of the video. In this case, the image of the trimming bounding box may be extracted from the image frame when necessary. Examples of the image seen by camera 108, the padded trimming bounding box, and the trimming bounding box resulting from the padded trimming bounding box are sent to the video analysis module 224, which may process the trimming bounding box on a server, for example.

[0083] To visually identify each of the one or more foreground visual objects detected, visual indicators may be added to the image frame. The visual indicators may be bounding boxes surrounding each of the one or more foreground visual objects within the image frame.

[0084] In some exemplary embodiments, the video analysis may further include in 304 classifying the foreground visual objects (or objects) detected in 302. For example, pattern recognition may be performed to classify the foreground visual objects. Foreground visual objects may be classified into classes such as people, cars, or animals. In addition to or instead of this, visual objects may be classified by actions such as the movement and direction of movement of the visual objects. Other classification elements such as color, size, and orientation may be determined. In more specific examples, the classification of visual objects may include person identification based on face detection and character recognition such as license plates. Visual classification may be performed in accordance with the systems and methods described in U.S. Patent No. 8,934,709, jointly owned, which is incorporated herein by reference in its entirety.

[0085] Video analysis in 306 may further include detecting whether an event has occurred and what type of event it is. Event detection may be based on comparing the classification of one or more foreground visual objects with one or more predetermined rules. Events may be anomaly detection or business intelligence, and may include, for example, whether a tripwire in the video has been activated, the number of people in one area, whether an object in the scene is "behind," or whether an object in the scene has been removed.

[0086] One example of video analysis in 306 may be a setting to detect only humans, and when a human is detected, the trimming bounding box of the human object is extracted and included in the metadata along with the reference coordinates of each trimming bounding box, and this metadata may be further processed on other devices such as a workstation 156 on network 140 along with the associated video 310.

[0087] Referring now to Figure 4, the diagram illustrates a flowchart of an exemplary embodiment of a method 400 for performing appearance matching to locate an object of interest in one or more image frames of video captured by a video acquisition device 108 (camera 108). The video is captured by camera 108 over a period of time. This period may span several hours, several days, or several months, and may span multiple video files or segments. As used herein, “video” includes information suggesting time and encompasses video files and video segments containing relevant metadata that identifies which camera 108 it is, if there are two or more cameras. The processing of the video is divided into several stages and distributed to optimize resource utilization and indexing for later retrieval of objects (or people) of interest. Video in which such a person of interest is found during the search may then be reviewed by the user.

[0088] The video of scene 402 is captured by camera 108. Scene 402 is within the field of view of camera 108. The video is processed by the video analysis module 224 in camera 108 to generate metadata including a cropping bounding box 404. The video analysis module 224 performs object detection and classification and also generates an image (cropping bounding box) from the video that best represents the objects in scene 402. In this example, images of objects classified as people or human are extracted from the video and included in the metadata as a cropping bounding box 404 for further identification processing. The metadata including the cropping bounding box 404 and the video are sent to server 406 on network 140. Server 406 may be a workstation 156 or a client device 164.

[0089] Server 406 has far more resources to further process the trimmed bounding box 108 and the generated feature vector (or "signature" or "binary representation") 410 to represent the object in scene 408 402. Processing 408 is known, for example, in the prior art as a feature descriptor.

[0090] In computer vision, feature descriptors are commonly known as algorithms that capture an image and output a feature description or feature vector through image transformation. A feature descriptor acts as a numerical "fingerprint" that encodes information, i.e., an image, into a series of numbers that can be used to distinguish features from one another. Ideally, this information should be immutable under image transformation so that features can be found again in other images of the same object. Examples of feature descriptor algorithms include SIFT (Scale-invariant feature transform), HOG (histogram of oriented gradients), and SURF (Speeded Up Robust Features).

[0091] A feature vector is an n-dimensional vector of numerical features (numbers) that represent an image of an object that can be processed by a computer. By comparing the feature vector of one image of an object with the feature vector of another image, a computer can determine whether two images are of the same object. An image signature (or feature vector, or embedding, or representation, etc.) is a multi-dimensional vector computed by a neural network (e.g., a convolutional one).

[0092] By calculating the Euclidean distance between two feature vectors of two images captured by camera 108, a computer-operable process can determine a similarity score indicating how similar the two images may be. The neural network is trained so that the feature vectors it computes for images are close (short Euclidean distance) for similar images and farther (longer Euclidean distance) for dissimilar images. To retrieve relevant images, the feature vector of the query image is compared to the feature vector of an image in database 414. The search results may be presented in ascending order of their distance (a value between 0 and 1) to the query image. The similarity score may be a percentage converted from a value between 0 and 1, for example.

[0093] In this exemplary embodiment, process 408 processes the trimmed bounding box 404 using a learning machine to generate image feature vectors or signatures of objects captured in the video. The learning machine is a neural network, such as a convolutional neural network (CNN) running on a graphics processing unit (GPU). The CNN may be trained using a training dataset containing an infinite number of pairs of similar and dissimilar images. The CNN is, for example, a Siamese network architecture trained with a symmetric loss function to train the neural network. An example of a Siamese network is described in Bromley, Jane, et al., "Signature verification using a "Siamese" time delay neural network," International Journal of Pattern Recognition and Artificial Intelligence 7.04(1993):669-688, the entire content of which is incorporated herein by reference.

[0094] Process 408 leverages a training model known as batch learning, where all training is performed before the appearance search system is used. In this embodiment, the training model is a convolutional neural network learning model with one possible set of parameters. For a given learning model, there are infinitely many possible sets of parameters. Optimization methods (such as stochastic gradient descent) and numerical gradient methods (such as backpropagation) may be used to find the set of parameters that minimizes the objective function (AKA loss function). The contrasting loss function is used as the objective function. This function is defined to take a high value when the current training model has low accuracy (assigning long distances to similar pairs or short distances to dissimilar pairs) and a low value when the current trained model has high accuracy (assigning short distances to similar pairs and long distances to dissimilar pairs). The training process is thus reduced to a minimum problem. The process of finding the model with the highest accuracy is the training process, the finished model with the set of parameters is the trained model, and the set of parameters remains unchanged after being deployed to the appearance search system.

[0095] An alternative embodiment of the processing unit 408 is to utilize a learning machine using what is known as an online machine learning algorithm. The learning machine is utilized in the processing unit 408 with the initial set of parameters, but the appearance search system continuously updates the model's parameters based on some source of truth (e.g., user feedback when selecting an image of an object of interest). Such a learning machine may include other types of neural networks as well as convolutional neural networks.

[0096] The trimming bounding box 404 of a human object is processed by the processing unit 408 to generate a feature vector 410. The feature vector 410 is indexed 412 and stored in the database 414 along with the video. The feature vector 410 is also associated with reference coordinates indicating where the trimming bounding box 404 of the human object may be located in the video. Storing in the database 414 includes storing the video along with the feature vector 410 of the trimming bounding box 404 and the reference coordinates indicating where the trimming bounding box 404 may be located in the video, as well as timestamps and camera identification information.

[0097] To identify the location of a specific person in the video, a feature vector of the person of interest is generated. Similar feature vectors 416 to the feature vector of the person of interest are extracted from the database 414. The extracted feature vectors 416 are compared to a threshold similarity score 418, and those exceeding the threshold are provided to the client 420 for presentation to the user. The client 420 also has a video playback module 264 so that the user can view the video associated with the extracted feature vectors 416.

[0098] More specifically, the trained model is trained using a predetermined distance function that is used to compare it with computed feature vectors. The same distance function is used when the trained model is utilized in the appearance search system. The distance function is the Euclidean distance between feature vectors that have been normalized so that each feature vector has a unit norm, so all feature vectors lie on a hypersphere of unit norm. After computed and stored the feature vectors of detected objects in the database, a search for similar objects is performed using exact nearest neighbor search, thoroughly evaluating the distance from the queried feature vector (the feature vector of the object of interest) to all other vectors in the target time frame. The search results are returned ranked in descending order of distance to the queried feature vector.

[0099] In another embodiment, approximate nearest neighbor search may be used. Approximate nearest neighbor search is similar to nearest neighbor search "itself," but it retrieves the most similar result without looking at all the results. This is faster, but it may lead to false positives. One example of approximate nearest neighbor search is the use of hashing indexing of feature vectors. Approximate nearest neighbor search can be faster when there are many feature vectors, such as when the search time frame is long.

[0100] More precisely, it will be understood that "objects of interest" encompass "people of interest," and "people of interest" encompass "objects of interest." Next, referring to Figure 5, which illustrates a flowchart of an exemplary embodiment of Figure 4, detailing an appearance search 500 that performs appearance matching on client 420 to locate the location of the recorded video of the object of interest. To initiate an appearance search for the object of interest, the feature vector of the object of interest is required to search database 414 for similar feature vectors. Two exemplary methods for initiating an appearance search 500 are shown.

[0101] In the first method of initiating an appearance search 500, the client 420 receives an image of the object of interest 502, which then sends it to the processing unit 408 to generate a feature vector of the object of interest 504. In the second method, the user searches the database 414 for an image of the object of interest 514 and retrieves a feature vector of the object of interest that was previously generated when the image was being processed for storage in the database 414 516.

[0102] Next, using either the first or second method, a search 506 is performed on database 414 for candidate feature vectors that have a similarity score exceeding a threshold, such as 70%, compared to the feature vector of the object of interest. Images of the candidate feature vectors are received 508 and then presented to the user by client 420, and the user selects images of candidate feature vectors that belong to the object of interest or that may belong to the object of interest 510. Client 420 tracks the selected images in a list. The list containing the images selected by the user belongs to the object of interest. Optionally, at selection 510, the user may remove images from the list that they later deemed inappropriate.

[0103] Each time a new image (or more images) of the object of interest is selected in selection 510, the feature vector of the new image is retrieved from the database 414 506, and new candidate images of the object of interest are presented to the user in client 420, allowing the user to select a new image that is either of the object of interest or potentially of the object of interest again in selection 510. This visual search loop 500 may continue until the user determines that they have identified enough images of the object of interest and terminates the search 512. The user may then, for example, view or download a video associated with the images in the list.

[0104] Referring now to Figure 6, which illustrates a flowchart of an exemplary embodiment of Figure 4, detailing a time-specified appearance search 600 that performs appearance matching on client 420 to identify the location of recorded video of an object of interest either before or after a selected time. This type of search is useful, for example, to locate a lost bag by identifying an image close to the present time and then tracing it back in time to determine who might have left the bag behind.

[0105] To initiate an appearance search of an object of interest, the feature vector of the object of interest is required to search the database 414 for similar feature vectors. Two exemplary methods for initiating a time-specified appearance search 600 are shown, similar to the appearance search 500. In the first method for initiating the appearance search 600, the client 420 receives an image of the object of interest 602, which then sends it to the processing unit 408 to generate a feature vector of the object of interest 604. In the second method, the user searches the database 414 for an image of the object of interest 614 and retrieves a feature vector of the object of interest that was previously generated when the image was processed before being stored in the database 414 616.

[0106] By either the first or second method, the time-specified appearance search 600 is configured to search either forward or backward in time 618. In the first method, the user may manually set the search time. In the second method, the search start time is set to the time the image was captured by the camera 108. In this example, the time-specified appearance search 600 is configured to search forward in time, for example, to locate a lost person closer to the current time. In another example, the time-specified appearance search 600 may be configured to search backward in time, for example, if the user wants to find out who left their bag (object of interest) behind.

[0107] Next, a database search 606 is performed in forward time from the search time for candidate feature vectors that have a similarity score exceeding a threshold, such as 80%, compared to the feature vector of the object of interest. The images of the candidate feature vectors are received 608 and then presented to the user by the client 420, and the user selects one image from the candidate feature vector images that are of the object of interest or that may be of the object of interest 610. The client 420 tracks the selected image in a list. The list includes the image selected by the user as the object of interest. Optionally, at selection 610, the user may remove any images from the list that they later deemed inappropriate.

[0108] Each time a new image of the object of interest is selected in selection 610, the feature vectors of the new image are searched in the database 414 in a forward time order from the search time 606. The search time is the time when the new image was captured by the camera 108. The new candidate images of the object of interest are presented to the user in client 420, and the user selects another new image that is either of the object of interest or potentially of the object of interest 610. This time-specified appearance search loop 600 may continue until the user determines that they have identified enough images of the object of interest and terminates the search 612. The user may then, for example, view or download videos related to the images in the list. This example searches forward time, but searching backward time would result in almost the same outcome, except that the search of the database 414 is filtered to include matches that occurred before the search time, or matches that occurred before the search time.

[0109] Next, referring to Figure 7, the diagram shows an example of metadata for an object profile 702 containing the trimming bounding box 404 as transmitted by camera 108 to server 406, and an example of an object profile 704 containing the feature vector 708 of the trimming bounding box 404 for storage in database 414, instead of the image 706 (trimming bounding box 404). Since the file size of image 706 is larger than the file size of feature vector 708, some storage space can be saved by storing object profile 704 containing feature vector 708 instead of image 706. As a result, a significant amount of data storage space can be saved because trimming bounding boxes are often quite large and numerous.

[0110] The data 710 for object profiles 702 and 704 includes, for example, a timestamp, frame number, pixel-level resolution based on the width and height of the scene, a segmentation mask for this frame based on the width and height of the pixels, as well as a stride based on the row width in bytes, classification (person, vehicle, etc.), confidence level of the classification in percent, a box based on the width and height in normalized sensor coordinates (a bounding box surrounding the outlined object), the image width and height in pixels and the image stride (row width in bytes), the image segmentation mask, orientation, and the xy coordinates of the image box. The feature vector 708 is, for example, a 48-dimensional, i.e., 48 floating-point binary representations of the image 706 (binary in the scene, consisting of 0s and 1s). The number of dimensions may be larger or smaller depending on the learning machine used to generate the feature vector. Higher dimensions generally result in higher accuracy, but can also require extremely high computational resources.

[0111] Since the trimming bounding box 404 or image 706 can be extracted again from the recorded video using reference coordinates, there is no need to add and save the trimming bounding box 404 to the video. The reference coordinates may include, for example, a timestamp, frame number, and box. For example, the reference coordinates may simply be a timestamp containing the relevant video file, and an image frame close to the original image frame may suffice, depending on whether the timestamp is accurate enough to trace back to the original image frame or not, because temporally close image frames in a video generally look alike.

[0112] In this exemplary embodiment, the object profile 704 has feature vectors replaced by an image, but in other embodiments, it may have an image compressed using a conventional method. Next, referring to Figure 8, the illustration shows scene 402 and the cropping bounding box 404 of the exemplary embodiment of Figure 4. Scene 402 shows three detected people. Their images 802, 806, and 808 are extracted by camera 108 and sent to server 406 as the cropping bounding box 404. Images 802, 806, and 808 are representative images of the three people in the video over a period of time. The three people in the video are moving, and consequently their images captured will be different over a period of time. To sort the images into a manageable number, one (or more) representative images are selected as the cropping bounding box 404 for further processing.

[0113] Referring now to Figure 9, the diagram shows a block diagram of a set of operational submodules for a video analysis module 224 according to one exemplary embodiment. The video analysis module 224 includes several modules that perform various tasks. For example, the video analysis module 224 includes an object detection module 904 that detects objects appearing in the field of view of the video acquisition device 108. The object detection module 904 may use any known object detection method, such as motion detection and blob detection. The object detection module 904 may use the detection methods described in U.S. Patent No. 7,627,171, entitled "Methods and Systems for Detecting Objects of Interest in Spatio-Temporal Signals," which are incorporated herein by reference in their entirety.

[0114] The video analysis module 224 also includes an object tracking module 908 connected to or linked to the object detection module 904. The object tracking module 908 operates to associate instances of objects detected by the object detection module 908 with time. The object tracking module 908 is called "Object The detection method described in U.S. Patent No. 8,224,029, entitled “Matching for Tracking, Indexing, and Search,” may be used, and is incorporated herein by reference to the entirety of that patent. The object tracking module 908 generates metadata corresponding to the visual objects it tracks. The metadata may correspond to the visual object signatures that represent the appearance or other characteristics of the objects. The metadata is transmitted to the server 406 for processing.

[0115] The video analysis module 224 also includes an object classification module 916 that classifies objects detected by the object detection module 904 and connects them to the object tracking module 908. The object classification module 916 may internally include an instantaneous object classification module 918 and a temporary object classification module 912. The instantaneous object classification module 918 determines the type of visual object (e.g., human, vehicle, or animal) based on a single instance of the object. Preferably, the input to the instantaneous object classification module 916 is a sub-region of the image (e.g., within the bounding box) where the object of visual interest is located, rather than the entire image frame. The advantage of inputting a sub-region of the image frame to the classification module 916 is that it requires less processing power because the entire scene does not need to be analyzed for classification. The video analysis module 224 may further process any object type other than, for example, a human.

[0116] The temporary object classification module 912 may maintain object class information (e.g., human, vehicle, or animal) over a certain period of time. The temporary object classification module 912 averages the instantaneous class information of the object provided by the instantaneous object classification module 918 over a certain period of time while the object exists. In other words, the temporary object classification module 912 determines the type of object based on the appearance of the object in multiple frames. For example, analyzing a person's gait may be useful for classifying people, or analyzing a person's feet may be useful for classifying cyclists. The temporary object classification module 912 may combine information about the object's trajectory (e.g., whether the trajectory is smooth or disordered, or whether the object is moving or stationary) with confidence information of the classification performed by the instantaneous object classification module 918, which is averaged over multiple frames. For example, the classification confidence value determined by the object classification module 916 may be adjusted based on the smoothness of the object's trajectory. The temporary object classification module 912 may assign objects to an unknown class until the visual objects have been classified a sufficient number of times by the instantaneous object classification module 918 and a predetermined number of statistics have been collected. When classifying objects, the temporary object classification module 912 may also take into account how long the object has been in the field of view. The temporary object classification module 912 may make a final decision regarding the object's class based on the information described above. The temporary object classification module 912 may use hysteresis techniques to change the object's class. More specifically, a threshold may be set to move an object's classification from an unknown class to a confirmed class, and this threshold may be greater than the threshold for moving it in the reverse direction (e.g., from human to unknown). The object classification module 916 may generate metadata about the object's class, which may be stored in the database 414.The temporary object classification module 912 may aggregate the classifications performed by the instantaneous object classification module 918.

[0117] In an alternative configuration, the object classification module 916 is positioned after the object detection module 904 and before the object tracking module 908, such that object classification occurs before object tracking. In yet another alternative configuration, the object detection module, tracking module, temporary classification module, and classification modules 904, 908, 912, and 916 are correlated as described above. In yet another alternative embodiment, the video analysis module 224 may detect faces in a human image using (prior art known) face recognition and provide a corresponding confidence level. The appearance search system in such an embodiment may include using a feature vector or cropped bounding box of a face image instead of the whole person, as shown in Figure 8. Such a face feature vector may be used alone or in combination with a feature vector of the whole object. Similarly, a feature vector of a part of an object may be used alone or in combination with a feature vector of the whole object. For example, a part of an object may be an image of a human ear. Ear recognition for individual identification is known in the prior art.

[0118] In each image frame of the video, the video analysis module 224 detects objects and extracts images of each object. The image selected from these images is referred to as the final object. The final object is intended to select the best representation of the visual appearance of each object while it is present in the scene. Signature / feature vectors can be extracted using the final objects, and these signature / feature vectors can then be used to query other final objects to find the closest match when setting up an appearance search.

[0119] Ideally, the final form of an object should be generated for each frame of its existence. However, if this were done, even one second of video would contain many image frames, making the computational requirements for practical use of appearance search too high. The following is an example of selecting possible final forms of an object that represent the object over a certain period of time, or selecting one image from a set of possible images of the object, in order to reduce the computational requirements.

[0120] When an object (a person) enters scene 402, it is detected as an object by the object detection module 904. Next, the object classification module 916 classifies the object as a person or human if it has a confidence level of certainty that the object is a person. The object is tracked within scene 402 by the object tracking module 908 through each image frame of the video captured by camera 108. The object may be identified by a tracking number when it is being tracked.

[0121] In each image frame, the image of the object within the bounding box surrounding the object is extracted from the image frame, and the image is a cropped bounding box. The object classification module 916 provides, for example, a confidence level for each image frame that the object is a human. In yet another exemplary embodiment, if the object classification module 916 provides a relatively low confidence level for classifying the object as (for example) a human, the padded cropped bounding box is extracted, thereby allowing a more computationally powerful object detection and classification module (e.g., process 408) on the server to resolve the padded cropped bounding box of the object before a feature vector is generated. The more computationally powerful object detection and classification module may be another neural network that resolves or extracts an object from another object that is overlapping or closely adjacent. A relatively low confidence level (e.g., 50%) may be used to indicate whether the cropped bounding box or the padded cropped bounding box should be further processed to resolve issues such as other objects within the bounding box before a feature vector is generated. The video analysis module 224 maintains a list of a specific number of cropping bounding boxes and tracks, for example, the top 10 cropping bounding boxes with the highest confidence level as objects within scene 402. When the object tracking module 908 fails to track an object, or when an object moves out of the scene, the cropping bounding box 404 is selected from the list of 10 cropping bounding boxes that represent the object with the most foreground pixels (or object pixels). The cropping bounding box 404, along with its metadata, is sent to the server 406 for further processing. The cropping bounding box 404 represents the image of the object over this tracking period. Confidence level is used to reject cropping bounding boxes where the object may not have a good image, such as when the object straddles a shadow. Alternatively, two or more cropping bounding boxes may be selected from the list of the top 10 cropping bounding boxes and sent to the server 406.For example, you may also send another trimming bounding box selected based on the highest confidence level.

[0122] The list of the top 10 trimming bounding boxes is one embodiment. Alternatively, this list could be as few as 5 trimming bounding boxes or as many as 20 trimming bounding boxes, as yet another example. Furthermore, selecting a trimming bounding box as trimming bounding box 404 from the list of trimming bounding boxes may be done periodically, not just for traces that have been missed. Alternatively, the selection of a trimming bounding box from the list may be based on the highest confidence level instead of the maximum number of object pixels. Alternatively, the video analysis module 224 may be located on a server 406 (workstation 156), a processing unit 148, a client device 164, or other device outside the camera.

[0123] The above selection criteria for trimming bounding boxes represent a possible solution to the problem of representing the existence of an object with a single trimming bounding box. The following are alternative selection criteria.

[0124] Alternatively, the top 10 out of n trimming bounding boxes can be selected using the information provided by the height estimation algorithm of the object classification module 916. The height estimation module creates a homology matrix based on the head (upper part) and feet (lower part) observed over a certain period of time. The period during which homology is learned is referred to herein as the learning phase. The obtained homology is further used to estimate the height of the actual object appearing at a particular location and is compared with the height of the object observed at that location. Once learning is complete, the top n trimming bounding boxes can be selected using the information provided by the height estimation module by comparing the height of the trimming bounding box with the expected height of the object at the location where the trimming bounding box was captured. This selection method is intended to be a criterion for rejecting trimming bounding boxes that may be false positives with high confidence as reported by the object classification module 916. The resulting selected trimming bounding boxes can then be further ranked by the number of foreground pixels captured by the object. This multi-stage selection criterion ensures that the final object not only has high classification reliability, but also conforms to the dimensions of the object expected at that location, and furthermore, has a good number of foreground pixels as reported by the object detection module 904. The trimmed bounding box obtained from the multi-stage selection criterion may give the object a better appearance over its duration in the frame compared to the trimmed bounding box obtained from any of the aforementioned criteria applied individually. In this specification, the machine learning module includes machine learning algorithms known in the prior art.

[0125] Referring now to Figure 10A, the diagram shows a block diagram of the process 408 in Figure 4 according to another exemplary embodiment. The image of the object (a trimmed bounding box including a padded trimmed bounding box) 404 is received by the processing unit 408, where it is processed by the first neural network 1010 to detect, classify, and outline objects within the trimmed bounding box 404. The first neural network 1010 and the second neural network 1030 are, for example, convolutional neural networks. The first neural network 1010 detects, for example, 0, 1, 2, or more people (as classified) for a given trimmed bounding box of the clip 404. If it is 0, no human objects were detected, the initial classification (at camera 108) was incorrect, and a feature vector 410 should not be generated for that given trimmed bounding box (end 1020). If one human object is detected, the given trimmed bounding box needs to be processed further. If a given trimming bounding box is a padded trimming bounding box, the image of the object in that given trimming bounding box is optionally resized to fit within the object's bounding box, just like other unpaded trimming bounding boxes. If two or more (2+) human objects are detected in a given trimming bounding box, in this embodiment, the image of the object closest to (or closest to) the center coordinates of the "object" in the image frame is extracted from the image frame of a new trimming bounding box that replaces the given trimming bounding box in trimming bounding box 404 and processed further.

[0126] The first neural network 1010 outputs an image (trimmed bounding box) 1040 outlining the object's contours, which is then processed by the second neural network 1030 to generate a feature vector 410, which is associated with the trimmed bounding box 404. An example of the first neural network 1010 is the single-shot multi-box detector (SSD) known in the prior art.

[0127] Next, referring to Figure 10B, the diagram shows a block diagram of the processing unit 408 of Figure 4 according to yet another exemplary embodiment. The image of the object (a trimmed bounding box including a padded trimmed bounding box) 404 is received by the processing unit 408, and the comparator 1050 determines the confidence level associated with the trimmed bounding box 404. The trimmed bounding box 404 from the camera 108 has associated metadata (such as confidence level) as determined by the video analysis module of the camera 108.

[0128] If the confidence level of a given trimming bounding box is relatively low (e.g., less than 50%), the given trimming bounding box is processed according to the embodiment of Figure 10A, which begins with the first neural network 1010 and ends with the feature vector 410. If the confidence level of a given trimming bounding box is relatively high (e.g., 50% or more), the given trimming bounding box is processed directly by the second neural network 1030, which generates the feature vector 410 without passing through the first neural network 1010.

[0129] Embodiments describing the extraction of padded trimmed bounding boxes from camera 108 include extracting the entire image of the object as a padded trimmed bounding box, while other embodiments extract only the padded trimmed bounding box when the confidence level for the classified related object is relatively low. Note that the first neural network 1010 may process both padded and unpadded trimmed bounding boxes to improve accuracy, and in some embodiments, the first neural network may process all trimmed bounding boxes if computational resources are available. The first neural network 1010 may process all padded trimmed bounding boxes, or it may process only a portion of the unpadded trimmed bounding boxes with low confidence levels. The threshold confidence level set by comparator 1050 may be lower than the threshold confidence level set to extract padded trimmed bounding boxes from camera 108. In some embodiments, some of the padded trimmed bounding boxes may be processed directly by the second neural network 1030, skipping processing by the first neural network 1010, especially when computing resources are tied to other functions of the server 406. Therefore, the number of trimmed bounding box processes handled by the first neural network may be set according to the amount of computing resources available on the server 406.

[0130] Next, referring to Figure 11, the diagram shows a flowchart of the processing unit 408 in Figures 11A and 11B according to another exemplary embodiment. Given a trimming bounding box 1110 (which may be unpaded or padded), if there are three human objects, the first neural network 1010 detects each of the three human objects and draws the outlines of each image of the three human objects to form trimming bounding boxes 1120, 1130, and 1140. The second neural network 1030 then generates feature vectors for the trimming bounding boxes 1120, 1130, and 1140. The trimming bounding boxes 1120, 1130, and 1140, along with their associated feature vectors, replace the given trimming bounding box 1110 in the trimming bounding box 404 in index 412 and database 414. In an alternative embodiment where the image contains multiple objects, only the object with the greatest overlap is retained (cropping bounding box 1130), and the other cropping bounding boxes are discarded.

[0131] Therefore, in one embodiment, object detection is performed in the following two stages: (1) Camera 108 performs low-accuracy but power-efficient object detection and sends the padded trimmed bounding box of the object to Server 406. Padding the trimmed bounding box provides the server-side algorithm with more pixel background for performing object detection and allows the server-side algorithm to reconstruct parts of the object that were cut off by the camera-side algorithm. Next, (2) Server 406 performs object detection on the padded trimmed bounding box using a high-accuracy but power-intensive algorithm.

[0132] This provides a compromise during network bandwidth usage, as network streams with trimmed bounding boxes for objects can have very low bandwidth. Sending all frames at a high frame rate is impractical in such environments unless a video codec is used (which requires decoding the video on server 406).

[0133] If server-side object detection is performed on an encoded video stream (such as the one used to record the video), the video must be decoded before the object detection algorithm can be executed. However, the computational requirements for decode multiple video streams may be too high to be practical.

[0134] Therefore, in this embodiment, camera 108 performs "approximate" object detection and sends the relevant padded trimmed bounding box to the server using a relatively low-bandwidth communication channel. Thus, camera 108 creates a padded trimmed bounding box that is likely to contain the object of interest using a less computationally intensive algorithm.

[0135] While the above description provides examples of embodiments where human objects are the primary object of interest, it will be understood that the basic method of extracting a trimming bounding box from an object, calculating a representation of feature vectors from it, and then using these feature vectors as a basis to compare opposing feature vectors with other objects, does not definitively determine the class of the object under consideration. Sample objects could be, for example, bags, backpacks, or suitcases. Therefore, visual retrieval systems for locating vehicles, animals, and inanimate objects can be implemented using the features and / or functions described herein, without departing from the spirit and principles of the operation of the embodiments described.

[0136] While the above description provides examples of embodiments, it will be understood that some features and / or functions of the described embodiments can be modified without departing from the spirit and principle of operation of the described embodiments. Therefore, it will be understood that the above is intended to be non-limiting, and other modifications and variations can be made without departing from the scope of the invention as described in the claims appended to this specification. Furthermore, any feature of any embodiment described herein may be appropriately combined with any other feature of any other embodiment described herein.

Claims

1. An appearance search system, At least one camera configured to capture video of a scene, wherein the video has images of objects, and the at least one camera is further configured to identify one or more objects in the images of the objects using a first learning machine of the at least one camera, One or more processors and memory, the one or more processors and memory including computer program code stored in the memory, A network configured to transmit images from at least one camera, including at least a portion of one or more identified objects, to one or more processors, When the computer program code is executed by one or more processors, the one or more processors The output from the second learning machine is to generate one or more signatures related to at least a portion of one or more identified objects and signatures of at least a portion of the objects of interest, To generate one or more similarity scores for at least one individual part of the one or more identified objects by comparing one or more signatures associated with at least one individual part of the one or more identified objects with signatures of at least one individual part of the object of interest, A command is transmitted to a client device having a display to present on the display one or more images related to at least a portion of the one or more identified objects, based on the one or more similarity scores. An appearance search system configured to implement a method including the following.

2. The appearance search system according to claim 1, wherein at least one of the first and second learning machines includes a neural network.

3. The appearance search system according to claim 2, wherein at least one of the first and second learning machines includes a convolutional neural network.

4. The appearance search system according to claim 2, wherein the first and second learning machines include neural networks.

5. The appearance search system according to claim 4, wherein the first and second learning machines include convolutional neural networks.

6. An appearance search system, At least one camera configured to capture video of a scene, wherein the video has images of objects, and the at least one camera is further configured to identify one or more objects in the images of the objects using a first learning machine of the at least one camera, One or more processors and memory, the one or more processors and memory including computer program code stored in the memory, A network configured to transmit images from at least one camera, including at least a portion of one or more identified objects, to one or more processors, When the computer program code is executed by one or more processors, the one or more processors Further identification of at least a portion of the one or more objects in the image using a second learning machine, The output from the third learning machine is to generate one or more signatures related to at least a portion of one or more further identified objects and signatures of at least a portion of the objects of interest, To generate one or more similarity scores for at least one of the individual parts of the one or more identified objects by comparing one or more signatures associated with at least one individual part of the one or more further identified objects with signatures of at least one part of the object of interest, A command is transmitted to a client device having a display to present on the display one or more images related to one or more identified objects or at least a portion of the further identified objects, based on the one or more similarity scores. An appearance search system configured to implement a method including the following.

7. The appearance search system according to claim 6, wherein at least one of the first, second, and third learning machines includes a neural network.

8. The appearance search system according to claim 7, wherein at least one of the first, second, and third learning machines includes a convolutional neural network.

9. The appearance search system according to claim 7, wherein the first, second, and third learning machines include a neural network.

10. The appearance search system according to claim 9, wherein the first, second, and third learning machines include a convolutional neural network.

Citation Information

Patent Citations

  • System And Method For Object And Event Identification Using Multiple Cameras

    US20140333775A1