A method and system for video analysis
The method and system efficiently filter video segments using metadata, motion detection, and semantic search to reduce computational demands, enabling effective identification of events of interest in large video datasets.
Patent Information
- Application Number
- GB2024007741
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-31
- Publication Date
- 2026-01-28
AI Technical Summary
Manual review of large amounts of video footage is time-consuming and computationally demanding, while automated methods are impractical due to high resource requirements, leading to inefficiencies in identifying events of interest.
A method and system that utilizes a multi-stage sequential analysis involving metadata filtering, motion detection, object classification, and semantic video search to identify relevant video segments, reducing computational resources and time by applying these techniques in a sequential pipeline.
Minimizes computational resources and time required for video analysis, providing a smaller amount of footage for manual review by efficiently filtering out irrelevant segments using metadata, motion, and semantic search techniques.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Field of the Invention
[0001] The present application relates to a method and system for video analysis. Background to the Invention
[0002] It is common to use video footage such as stored or live video from CCTV cameras, surveillance cameras, and the like, to identity and investigate events of interest. For example, such video footage is commonly reviewed and analysed in order to investigate or prevent accidents and / or crime. The gathering and storing of video footage is generally becoming cheaper overtime, while the quality of the video footage is generally improving, so that the availability, value and potential usefulness of video footage is increasing.
[0003] It is well known to manually review video footage for events of interest. However, this approach is time consuming and cognitively demanding, making it inconvenient, expensive and time consuming to review large amounts of video footage because of the number of human reviewers and length of time required. Although human reviewers can generally review video footage at higher speeds than the "real time" represented in the video footage, there is generally a trade-off between the speed of the review and its quality. Further, manual review of large amounts of video footage often suffers from quality or reliability issues due to the human reviewers becoming bored and inattentive.
[0004] A number of methods of automatically searching video footage for events of interest using trained Al and / or ML algorithms have been proposed and used, such as the CLIP neural network. However, such automated methods are highly computationally demanding, so that the required large amount of computing resources and cost make it impractical or undesirable to use these automated methods to review large amounts of video footage.
[0005] The inventors have devised the claimed invention in light of the above considerations. The embodiments described below are not limited to implementations which solve any or all of the disadvantages of the known approaches described above. Summary of Invention
[0006] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter; variants and alternative features which facilitate the working of the invention and / or serve to achieve a substantially similar technical effect should be considered as falling into the scope of the invention.
[0007] The invention is defined as set out in the appended set of claims.
[0008] In a first aspect of the present invention, there is provided a computer implemented method for analysing video, comprising: obtaining a search query comprising a semantic description of an event of interest and an object definition identifying one or more objects associated with the event of interest: obtaining first video segments; applying one or more image classification algorithms to the first video segments to identify parts of the first video segments containing at least one object of the one or more objects, and based on this identification, defining object searched video segments each containing at least one of the identified parts; and applying one or more semantic video search models to the object searched video segments to identify parts of the object searched video segments matching the semantic description, and based on this identification, defining semantic searched video segments each containing at least one of the identified parts; and outputting the identities of the semantic searched video segments. This may provide the advantages of minimizing, or reducing, the computational resources required to analyse video to provide a smaller amount of video footage which may contain one or more events of interest, for example for subsequent manual review.
[0009] In some embodiments, the search query further comprises a range of metadata types and values to be searched; and wherein the method further comprises; obtaining second video segments and associated metadata; carrying out metadata searching of the second video segments and associated metadata using the range of metadata types and values to identify parts of the second video segments associated with metadata corresponding to the range of metadata types and values, and based on this identification, defining metadata searched video segments each containing at least one of the identified parts; and using the metadata searched video segments as the first video segments. This may provide the advantages of further minimizing, or reducing, the computational resources required to analyse video.
[0010] In some embodiments, the method further comprising; obtaining third video segments; applying one or more movement detection algorithms to the third video segments to identify parts of the third video segments containing movement, and based on this identification, defining movement searched video segments each containing at least one of the identified parts; and using the movement searched video segments as the first video segments. This may provide the advantages of further minimizing, or reducing, the computational resources required to analyse video.
[0011] In some embodiments, the search query further comprises obtaining a range of metadata types and values to be searched; and wherein the method further comprises: obtaining fourth video segments and associated metadata; carrying out metadata searching of the fourth video segments and associated metadata using the range of metadata types and values to identify parts of the fourth video segments associated with metadata corresponding to the range of metadata types and values, and based on this identification, defining metadata searched video segments each containing at least one of the identified parts; applying one or more movement detection algorithms to 3 the metadata searched video segments to identity parts of the metadata searched video segments containing movement, and based on this identification, defining movement searched video segments each containing at least one of the identified parts; and using the movement searched video segments as the first video segments. This may provide the advantages of further minimizing, or reducing, the computational resources required to analyse video.
[0012] In some embodiments, the method further comprising displaying the semantic searched video segments to a user for manual review. This may provide the advantages of minimizing, or reducing, the manual input required to analyse video.
[0013] In some embodiments, the one or more semantic video search models include one or more multi-modal text-vision models, and optionally include CLIP. This may provide the advantages of improving the quality of the video analysis.
[0014] In some embodiments, obtaining an object definition comprises applying one or more natural language processing algorithms to the semantic description of an event of interest to expand the semantic description into a list of one or more names of objects. This may provide the advantages of improving the quality of the video analysis.
[0015] In some embodiments, the range of metadata types and values define a range of locations. This may provide the advantages of improving the quality of the video analysis.
[0016] In some embodiments, the range of metadata types and values define a range of times. This may provide the advantages of improving the quality of the video analysis.
[0017] In some embodiments, the obtained video segments are stored video segments. This may provide the advantage of improved flexibility by enabling non-real time video analysis.
[0018] In some embodiments, the obtained video segments are real time video segments.
[0019] In a second aspect of the present invention, there is provided a system for analysing video, the system comprising: a data store arranged to store a search query comprising a semantic description of an event of interest and an object definition identifying one or more objects associated with the event of interest; an object module arranged to apply one or more image classification algorithms to first video segments to identity parts of the first video segments containing at least one object of the one or more objects, and based on this identification, defining object searched video segments each containing at least one of the identified parts; a semantic video module arranged to apply one or more semantic video search models to the object searched video segments to identify parts of the object searched video segments matching the semantic description, and based on this identification, defining semantic searched video segments each containing at least one of the identified parts; and an output module arranged to output the identities of the semantic searched video segments. . This may provide the advantages of minimizing, or reducing, the computational 4 resources required to analyse video to provide a smaller amount of video footage which may contain one or more events of interest, for example for subsequent manual review.
[0020] In some embodiments, the search query further comprises a range of metadata types and values to be searched; and the system further comprising a metadata module arranged to carry out metadata searching of second video segments and associated metadata using the range of metadata types and values to identify parts of the second video segments associated with metadata corresponding to the range of metadata types and values, and based on this identification, defining metadata searched video segments each containing at least one of the identified parts; wherein the object module is arranged to use the metadata searched video segments as the first video segments. This may provide the advantages of further minimizing, or reducing, the computational resources required to analyse video.
[0021] In some embodiments, the system further comprises: a movement module arranged to apply one or more movement detection algorithms to third video segments to identify parts of the third video segments containing movement, and based on this identification, defining movement searched video segments each containing at least one of the identified parts; wherein the object module is arranged to use the movement searched video segments as the first video segments. This may provide the advantages of further minimizing, or reducing, the computational resources required to analyse video.
[0022] In some embodiments, the search query further comprises a range of metadata types and values to be searched; and the system further comprising: a metadata module arranged to carry out metadata searching of fourth video segments and associated metadata using the range of metadata types and values to identify parts of the second video segments associated with metadata corresponding to the range of metadata types and values, and based on this identification, defining metadata searched video segments each containing at least one of the identified parts; and a movement module arranged to apply one or more movement detection algorithms to the metadata searched video segments to identify parts of the metadata searched video segments containing movement, and based on this identification, defining movement searched video segments each containing at least one of the identified parts; wherein the object module is arranged to use the movement searched video segments as the first video segments. This may provide the advantages of further minimizing, or reducing, the computational resources required to analyse video.
[0023] In some embodiments, the system further comprising an output device arranged to display the semantic searched video segments to a user for manual review. This may provide the advantages of minimizing, or reducing, the manual input required to analyse video.
[0024] In some embodiments, the one or more semantic video search models include one or more multi-modal text-vision models, and optionally include CLIP. This may provide the advantages of improving the quality of the video analysis.
[0025] In some embodiments, the range of metadata types and values define a range of locations and / or a range of times. This may provide the advantages of improving the quality of the video analysis.
[0026] In some embodiments, the video segments are stored video segments or real time video segments. This may provide the advantage of improved flexibility.
[0027] In a third aspect of the present invention, there is provided a computer-readable medium comprising instructions which, when executed by one or more processors, cause the one or more processor to carry out the method of the first aspect.
[0028] The features and embodiments discussed above may be combined as appropriate, as would be apparent to a person skilled in the art, and may be combined with any of the aspects of the invention except where it is expressly provided that such a combination is not possible or the person skilled in the art would understand that such a combination is self-evidently not possible. Brief Description of the Drawings
[0029] Embodiments of the present invention are described below, by way of example, with reference to the following drawings.
[0030] Figure 1 shows a schematic diagram of an example of a video surveillance system and a video analysis system according to a first embodiment;
[0031] Figure 2 shows a flow chart of a method useable by the video surveillance system of figure 1;
[0032] Figure 3 shows a flow chart of a method useable by the video analysis system of figure 1;
[0033] Figure 4 shows a schematic diagram of a search query useable by the video analysis system of figure 1;
[0034] Figure 5 shows a schematic diagram of an example of a combined video surveillance and analysis system according to a second embodiment;
[0035] Figure 6 shows a flow chart of a method useable by the video surveillance and analysis system of figure 5; and
[0036] Figure 7 shows a schematic diagram of a search query useable by the video analysis system of figure 5.
[0037] Common reference numerals are used throughout the figures to indicate the same or similar features. Detailed Description
[0038] Embodiments of the present invention are described below by way of example only. These examples represent the best mode of putting the invention into practice that are currently known to the Applicant although they are not the only ways in which this could be achieved, the description sets forth the functions of the example and the sequence of steps for constructing and operating the example. However, the same or equivalent functions and sequences may be accomplished by different examples.
[0039] Figure 1 shows a schematic diagram of a video surveillance system 100 arranged to capture and store surveillance video, and a video analysis system 200 arranged to analyse the stored surveillance video to identify events of interest according to a first embodiment. The events of interest may be of any type, as appropriate to the purpose of the video surveillance and video analysis. In some examples the video surveillance and video analysis may be intended for security reasons, or to investigate or prevent accidents and / or crime, or for any other reason.
[0040] As shown in figure 1, the video surveillance system 100 comprises a plurality of closed circuit television (CCTV) video cameras 102a to 102n which each capture a respective video stream 104a to 104n from a respective field of view together with associated metadata. Accordingly, each video stream 104a to 104n comprises a series of images or frames together with the associated metadata. Optionally, one, some, or all of the video cameras 104a to 104n may also capture sound so that their respective video streams 104a to 104n also comprise sound data. The video surveillance system 100 further comprises a central store 106 communicatively connected to the video cameras 102a to 102n and comprising a video data store 108. The use of CCTV video cameras is not essential, and other examples may use alternative types of video cameras.
[0041] The metadata associated with each video stream 104a to 104n may typically comprise at least the time at which the different images of the series of images of the video stream 104a to 104n were captured, and the location of the capturing video camera 102a to 102n. The metadata may further comprise the facing and / or the field of view of the capturing video camera 102a to 102n. The metadata may further comprise the identity and / or type of the capturing video camera 102a to 102n. The metadata may comprise additional information as required in any specific implementation.
[0042] Figure 2 shows a flow chart of a method 300 carried out by the video surveillance system 100 of the first embodiment in operation. In the method 300, in a first video capture block 302, the plurality of video cameras 102a to 102n capture respective video streams 104a to 104n and send these video streams 104a to 104n to the central store 106. Then, in a storage block 304, the central 7 store 106 stores the received video streams 104a to 104n in the video data store 108 in association with the associated metadata as a plurality of stored video segments 110. The metadata is stored as metadata tags associated with the images or frames of the stored video segments 110.
[0043] The video surveillance system 100 may be arranged to monitor an area, such as an urban area, with the video cameras 102a to 102n being distributed across and around the monitored area. In the illustrated embodiment of figure 1 the video cameras 102a to 102n are CCTV video cameras having fixed locations. However, some, or all, of the video cameras 102a to 102n may be able to change orientation to scan their respective fields of view across respective areas.
[0044] The plurality of video cameras 102a to 102n may be the same as one another, or may comprise a plurality of different types of video camera having different characteristics. Without wishing to be bound by theory it is expected that, in practice, many video surveillance systems 100 will comprise a plurality of different types of video camera. The plurality of video cameras 102a to 102n may be a fixed group of video cameras. Alternatively, the number and identities of the video cameras 102a to 102n comprised in the video surveillance system 100 may change overtime as video cameras are switched on and off, or connected to and disconnected from, the central store 106.
[0045] The video analysis system 200 is communicatively connected to the central store 106 through a communications network 112, so that the video analysis system 200 can access the video segments 110 stored in the video data store 108 and analyse the stored video segments 110 to identify event(s) of interest.
[0046] The video analysis system 200 comprises at least one processor 202, and a memory 204. The at least one processor 202 comprises a metadata module 206, a motion module 208, an object module 210, a semantic video module 212, and an output module 214 of the video analysis system 200. The different modules 206 to 214 are provided by different software executed by the at least one processor 202. In some examples each module, or some modules, may be provided by software executed by specific dedicated processor(s) of the at least one processor 202. The memory 204 may be any form of data store.
[0047] The video analysis system 200 further comprises a user input device 218 and an output device 216. The user input device 218 may be any suitable user input device. For example, the user input device 218 may be a keyboard, a mouse, a touchscreen, a graphical user interface (GUI), or a microphone and speech to text system, or any other input device, or combinations thereof. The output device 216 may comprise a monitor or display screen, or any other output device, or combinations thereof. In some examples the user input device 218 and the output device 216 may be combined into a single input / output device. As will be explained in more detail below, the video analysis system 200 may only be able to identify a specific plurality of objects. Accordingly, the video analysis system 200 may invite the user to select objects from the plurality of objects which the video 8 analysis system is able to identify, for example by presenting the names of this plurality of objects as a drop down menu.
[0048] In the illustrated first embodiment of figures 1 and 2 it is assumed that when the video analysis system 200 carries out video analysis the video surveillance system 100 has already been in operation previously so that a plurality of video segments 110 are stored in the video data store 108 corresponding to a number of time segments of video streams 104a to 104n stored in association with their associated metadata. These stored video segments 110 may be referred to as stored video footage. The video surveillance system 100 may continue operating while the video analysis system 200 carries out video analysis. However, this is not essential. It will be understood that the video analysis by the video analysis system 200 is asynchronous with the capture of the video streams 104a to 104n by the video surveillance system 100.
[0049] Figure 3 shows a flow chart of a video analysis method 400 carried out by the video analysis system 200 of the first embodiment in operation.
[0050] In the method 400, in an obtain search query block 402, the video analysis system 200 obtains a search query which defines the objective of the video analysis to be carried out by the video analysis system 200 on the stored video streams 104a to 104n. In the illustrated embodiment of figures 1 to 3, the search query is input by a user into the video analysis system 200 using the user input device 218, and then stored in the memory 204.
[0051] Figure 4 shows a schematic diagram of a search query 500 according to the first embodiment. As shown in figure 4, the search query 500 comprises three elements, a semantic definition 502 of an event of interest, a range definition 504 of metadata types and values to be searched, and an object definition 506 of one or more objects expected to be found in association with the event of interest. The semantic definition is a definition in words. In the illustrated example the metadata types and values to be searched comprises a range of locations and / or times to be searched. The range of locations to be searched may, for example, be a specific location or an area.
[0052] The search query 500 may be entered by the user via the user input device 218 in any convenient manner, such an via a keyboard or a microphone and speech to text system. In some examples the semantic definition 502 may be entered as a natural language statement describing the event of interest, such as "people fighting" or "a man in a red shirt getting into a vehicle". The range definition 504 of metadata types and values to be searched, which comprises locations and / or times to be searched in the illustrated example, may be entered as text, for example by naming a location, defining an area, and / or defining a time range or limits. In some examples the locations to be searched may be entered by selecting a location or drawing an area on a map, for example using a mouse or touchscreen. The object definition 506 of one or more objects expected to be found in association with the event of interest may be entered by the user selecting objects from a list of names of objects identifiable by the video analysis system 200, for example by selection from a drop 9 down menu. In some examples the video analysis system 200 may alternatively, or additionally, comprise one or more natural language processing algorithms, and use these to expand the semantic definition 502 of the event of interest into a list of one or more names of objects identifiable by the video analysis system 200 which are likely to be found in association with the event of interest. In some examples the video analysis system 200 may alternatively, or additionally, comprise one or more object association algorithms, and use these to expand the list of one or more names of objects entered by a user to include one or more additional names of objects identifiable by the video analysis system 200 which are similar to objects in the list of one or more names of objects entered by a user. There are a number of natural language processing algorithms and object association algorithms known to the skilled person, and new algorithms are regularly produced. Any suitable algorithms may be used.
[0053] In the illustrated example of the first embodiment of figures 1 to 4, the search query 500 is input by a law enforcement user wishing to find video surveillance imagery relating to a number of incidents in an area of a town the previous evening in which it is believed that a group of men drinking in public were involved in a number of street fights. In this example, the semantic definition 502 of the event of interest input by the user may be "men fighting in the street". Further, in this example, the range definition 504 of locations and / or times to be searched input by the user may be a time range and a geographical area expected to include all of the reported or suspected incidents. Further, in this example the object definition 506 of one or more objects may comprise a bottle and a glass.
[0054] In the method 400, the video analysis system 200 then obtains, in an obtain video data block 404, the stored video segments 110 from the central store 106. To do this, the video analysis system 200 sends a request for stored video segments 110 to the central store 106, and the central store 106 responds to the request by sending the requested plurality of stored video segments 110 to the video analysis system 200. The video analysis system 200 stores the received video image segments 110 in the memory 204.
[0055] Then, the video analysis system 200 carries out a multi-stage sequential analysis 406 of the plurality of received video segments 110 based on the input search query 500.
[0056] In a first stage of the multi-stage sequential analysis 406, in a metadata filtering block 408, the metadata module 206 of the video analysis system 200 uses the metadata associated with the stored video segments 110, that is, the metadata in the metadata tags associated with the images of the stored video segments 110, to carry out metadata searching to identity parts of the stored video segments 110 which were captured by cameras at locations within a predetermined distance of the range of locations to be searched as defined in the range definition 504 of the search query 500 at times within the range of times to be searched which were captured by cameras at locations within a predetermined distance of the range of locations to be searched as defined in the range definition 504 of the search query 500 at times within the range of times to be searched defined in the range definition 504. This may be done by defining metadata value ranges corresponding to locations within a predetermined distance of the range of locations to be searched and the range of times to be searched and searching for, or filtering out, parts of the video segments with associated metadata values falling within these ranges. The metadata filtering block 408 then defines metadata searched video segments each comprising at least one of the identified parts of the stored video segments 110, and produces data identifying the metadata searched video segments. Accordingly, the data produced by the metadata filtering block 408 identifies a plurality of metadata searched video segments which were captured by cameras at locations within a predetermined distance of the range of locations to be searched at times within the range of times to be searched, defining a smaller pool of metadata searched stored video segments from the original stored video segments 110. The metadata filtering block 408 may be regarded as filtering out irrelevant video segments which do not show images from the locations and times of interest from the original pool of stored video segments 110.
[0057] Accordingly, in the illustrated example of the first embodiment, the metadata searched stored video segments will comprise a plurality of video segments of the stored video segments 110 which were captured in the area of interest on the previous evening.
[0058] The predetermined distance may be a fixed distance based on the maximum distance at which a video camera is expected to be able to capture useable images. In some examples the predetermined distance may be based on the type of the capturing video camera as identified in the associated metadata, so that video cameras of different types may have different predetermined distances.
[0059] In some examples, when the necessary metadata regarding the facing and / or the field of view of the capturing video cameras 102a to 102n is available, the metadata filtering block 408 may conduct a more detailed search to identity video segments which were captured by cameras at locations within a predetermined distance of the range of locations to be searched and also have fields of view showing all or part of the range of locations to be searched defined in the range definition 504 of the search query 500. In order to do this, the metadata module 206 will usually need to analyse the metadata and convert the location and facing of the video camera indicated by the metadata into a geographical field of view so that it can be determined whether or not the field of view includes all or part of the defined range of locations to be searched.
[0060] Then, in a motion filtering block 410, the motion module 208 of the video analysis system 200 conducts a motion search of the metadata searched stored video segments identified in the data produced by the metadata filtering block 408. It will be understood that these metadata searched stored video segments are smaller (that is, comprise less total video data or footage), than the stored video segments 110. In the motion filtering block 410, the motion module 208 uses one or 11 more motion detection algorithms to identify parts of the metadata searched stored video segments containing movement. The motion filtering block 410 then defines motion searched video segments each corresponding at least one ofthe identified parts of the metadata searched video segments, and produces data identifying the motion searched video segments. Accordingly, the data produced by the motion filtering block 410 identifies a plurality of motion searched video segments which were captured by cameras at locations within a predetermined distance ofthe range of locations to be searched at times within the range of times to be searched, and also contain movement, defining a smaller pool of motion searched stored video segments from the metadata searched stored video segments. The motion filtering block 410 may be regarded as filtering out irrelevant video segments which do not contain any movement from the metadata searched stored video segments. It will be understood that video segments containing no movement, or in other words showing no changes between images, are of no interest because nothing is happening in these video segments.
[0061] In general, motion detection algorithms operate by comparing successive images of a video segment and identifying changes in the image which may correspond to movement in the imaged scene shown in the image. There are a large number of computationally efficient motion detection algorithms known to the skilled person in the technical field of image processing, and new motion detection algorithms are regularly produced. Any suitable motion detection algorithm may be used. In some examples, a motion detection algorithm may be selected from a group of multiple different motion detection algorithms based on the type ofthe capturing video camera as identified in the associated metadata, so that video cameras of different types may have their captured video segments filtered using different motion detection algorithms. In some examples, a motion detection algorithm may be selected from a group of multiple different motion detection algorithms based on characteristics ofthe image, such as light level, image quality, or the like. In examples where some or all ofthe video cameras 102a to 102n are able to change orientation a suitable motion detection algorithm able to distinguish movement in the imaged scene and changes in the image caused by the changing orientation.
[0062] Accordingly, in the illustrated example ofthe first embodiment, the motion searched stored video segments will comprise video segments of the stored video segments 110 which were captured in the area of interest on the previous evening and include movement.
[0063] Then, in an object detection block 412, the object detection module 210 ofthe video analysis system 200 conducts a search ofthe motion searched stored video segments identified in the data produced by the motion filtering block 410 for the object or objects identified in the object definition 506 ofthe search query 500. It will be understood that these motion searched stored video segments are smaller (that is, comprise less total video data or footage), than the metadata searched stored video segments. In the object detection block 412, the object detection module 210 uses one or more image classification algorithms to identify parts ofthe motion searched video segments containing images of one or more ofthe object(s) identified in the object definition 506. The object 12 detection block 412 then defines object searched video segments each comprising at least one of the identified parts of the motion searched video segments, and produces data identifying the object searched video segments. Accordingly, the data output by the object detection block 412 identifies a plurality of object searched video segments which were captured by cameras at locations within a predetermined distance of the range of locations to be searched at times within the range of times to be searched, contain movement, and also contain at least one of the object(s) identified in the object definition 506, so defining a smaller pool of object searched stored video segments from the motion searched stored video segments. The object detection block 412 may be regarded as filtering out irrelevant video segments which do not contain any of the objects listed in the object definition 506 from the movement searched stored video segments.
[0064] A number of image classification algorithms able to detect specific objects shown in images are known to the skilled person, and new image classification algorithms are regularly produced. Any suitable image classification algorithm may be used. One known suitable image classification algorithm is the YOLO system, which is a neural networked based system. Another known suitable image classification algorithm is EfficientDet. Other suitable image classification algorithms are also known. YOLO, and similar neural network or machine learning based systems must be trained to identify specific objects. The image classification algorithm used by the object detection module 210, in the illustrated example, YOLO, must have been previously trained to enable the neural network(s) to identify a number of different objects, which objects correspond to the list of one or more names of objects identifiable by the video analysis system 200. Accordingly, the video analysis system 200 can only identify a plurality of objects which the one or more image classification algorithms have been trained to identify, and it is for this reason that the object definition 506 is limited to the names of objects identifiable by the video analysis system 200.
[0065] In some examples, an image classification algorithm may be selected from a group of multiple different image classification algorithms based on the type of the capturing video camera as identified in the associated metadata, so that video cameras of different types may have their captured video segments filtered using different image classification algorithms. In some examples, an image classification algorithm may be selected from a group of multiple different image classification algorithms based on characteristics of the image, such as light level, image quality, or the like.
[0066] Accordingly, in the illustrated example of the first embodiment, the object searched stored video segments will comprise video segments of the stored video segments 110 which were captured in the area of interest on the previous evening, include movement, and include a bottle and / or a glass.
[0067] Then, in a semantic video search block 414, the semantic video module 212 of the video analysis system 200 conducts a search of the object searched stored video segments identified in the data produced by the object detection block 412 for the events defined by the semantic definition 502 of the search query 500. It will be understood that these object searched stored video segments are smaller (that is, comprise less total video data or footage), than the movement searched stored video segments. In the semantic video search block 414, the semantic video module 212 applies one or more multi-modal text-vision models to each of the video segments of the object searched stored video segments to identify parts of the object searched video segments containing images matching the semantic definition 502 of the event of interest in the search query 500. The semantic video search block 414 then defines semantic searched video segments each comprising at least one of the identified parts of the object searched video segments, and produces data identifying the semantic searched video segments. Accordingly, the data output by the semantic video search block 414 identifies a plurality of semantic searched video segments which were captured by cameras at locations within a predetermined distance of the range of locations to be searched at times within the range of times to be searched, contain movement, contain at least one of the object(s) identified in the object definition 506, and contain the event of interest defined by the semantic definition 502, so defining a smaller pool of semantic searched stored video segments from the object searched stored video segments. The semantic video search block 414 may be regarded as filtering out irrelevant video segments which do not contain events corresponding to the semantic definition 502 from the object searched stored video segments. It will be understood that the semantic searched stored video segments are smaller (that is, comprise less total video data or footage), than the object searched stored video segments.
[0068] In general, multi-modal text-vision models search the object searched stored video segments by using a multi-modal text and vision neural network to link text based queries to video segments matching the meaning of the text query. This is referred to as a semantic search. This may be understood as the semantic search comprising embedding each of the video segments into the same latent space as the semantic definition and returning the closest matching segments to the semantic definition from within that latent space. A number of multi-modal text-vision models able to search images and / or video images and to return images matching a semantic definition are known to the skilled person, and models are regularly produced. Any suitable multi-modal text-vision model may be used. One known suitable multi-modal text-vision model is CLIP, which is a neural networked based system. Other known suitable multi-modal text-vision models include ALIGN and DALL-E. Other suitable multi-modal text-vision models are also known. CLIP, and similar neural network or machine learning based systems, must be trained to associate image content with semantic definitions. The multi-modal text-vision models may be used "off the shelf in a standard format, or may be trained further from their standard format to be better suited to the specific conditions of the video surveillance system, such as the types of video cameras used, and / or the types of video images produced in practice, which may include noise, blurring, and the like.
[0069] In some examples, the one or more multi-modal text-vision models may search each video segment as a number of separate images, separately searching each image for image content 14 corresponding to the events defined by the semantic definition. In such examples, a search for the event "people fighting" may return images of a two or more people in close proximity, or a search for the event "a man in a red shirt getting into a vehicle" may return images of a man in a red shirt near, or inside, a vehicle. The illustrated example using CLIP is an example of this type. In other examples, the one or more multi-modal text-vision models may be more sophisticated, and may search each video segment as a series of images, searching each image for image content corresponding to the events defined by the semantic definition, and also taking into account movement of and / or interactions between the image content over sequences of images for the events defined by the semantic definition. In such examples, a search for the event "people fighting" may return images of a two or more people in close proximity and moving in particular ways, ora search for the event "a man in a red shirt getting into a vehicle" may return images of a man in a red shirt entering a vehicle.
[0070] In some examples, a multi-modal text-vision model may be selected from a group of multiple different multi-modal text-vision models based on the type of the capturing video camera as identified in the associated metadata, so that video cameras of different types may have their captured video segments filtered using different multi-modal text-vision model. In some examples, an multi-modal text-vision model may be selected from a group of multiple different multi-modal textvision models based on characteristics of the image, such as light level, image quality, or the like. In some examples, an multi-modal text-vision model may be selected from a group of multiple different multi-modal text-vision models based on the semantic definition, Some models may be better at searching for particular semantic definitions than others.
[0071] Accordingly, in the illustrated example of the first embodiment, in examples where a multi-modal text-vision model searching the video segments as separate images is used, the semantic searched stored video segments will comprise video segments of the stored video segments 110 which were captured in the area of interest on the previous evening, include movement, include a bottle and / or a glass, and include a group of men in close proximity to one another. In examples where a multi-modal text-vision model searching the video segments as a series of images is used, the semantic searched stored video segments will comprise video segments of the stored video segments 110 which were captured in the area of interest on the previous evening, include movement, include a bottle and / or a glass, and include a group of men in close proximity to one another and moving in a particular manner associated with fighting.
[0072] The data identifying semantic searched video segments produced by the semantic video search block 414 is output by the output module 214 in a final output stage 416 as the output of the multi-stage sequential analysis 406. In the first embodiment this output data identifying semantic searched video segments is then stored by the video analysis system 200 in the memory 204 for future manual review. The video analysis system 200 may optionally display the semantic searched video segments identified by the output using the output device 216 for manual review by a user.
[0073] In the first embodiment the output data is stored in the memory 204 by the video analysis system 200. Alternatively and / or additionally, the video analysis system 200 may be arranged to send the output data elsewhere for storage and / or manual review. Further, in some examples the video analysis system 200 may be arranged to store the semantic searched video segments identified by the output data for future manual review and / or to send the semantic searched video segments identified by the output data elsewhere for manual review.
[0074] In each of the blocks 408 to 414 of the multi-stage sequential analysis 406, the identified video segments may have a predetermined minimum length. This predetermined minimum length may, for example, be defined in terms of time or number of images. Setting such a predetermined minimum length may ensure that the identified semantic searched video segments are sufficiently long for their content to be properly understood by a human reviewer carrying out a manual review.
[0075] The disclosed systems and methods may provide the advantage of minimizing, or reducing, the computational resources and time required to automatically review video footage, such as surveillance video footage, in order to provide a smaller amount of video footage which may contain one or more events of interest, typically for subsequent manual review. The disclosed systems and methods reduce the computational resources and time required (compared to a "brute force" approach of reviewing all of the video footage with a multi-modal text-vision model) by applying a sequenced series or pipeline of different search or filtering techniques to the video footage with the final stage using a multi-modal text-vision model to reduce the amount of video footage reviewed using the a multi-modal text-vision model. Metadata based filtering of video segments is straightforward, can be carried out very quickly, and requires very little computational resources. The use of motion detection algorithms to identify video segments containing movement, or motion based filtering, is also straightforward, can be carried out quickly, and requires little computational resources, although motion detection generally requires more time and computational resources than metadata based filtering. Object detection in video segments using image classification algorithms requires more time and computational resources than motion based filtering, in part because neural networks or other machine learning approaches are required. Semantic video searching of video segments using multi-modal text-vision models requires even more time and computational resources than object detection. Accordingly, by conducting the different filtering operations of the different blocks or stages of the multi-stage sequential analysis in the specified order the starting pool of surveillance video segments can be sequentially reduced in stages to smaller pools of video segments in an efficient manner whereby the more time and computational resource consuming filtering operations are applied to the smaller pools of video segments, or, to put it another way, larger pools of video segments (or larger amounts of video data) are subject to filtering operations which require less time and computational resources. Accordingly, the system and method set out above may provide the advantage of automatically filtering video footage, such as surveillance video footage, in an efficient 16 manner minimizing, or reducing, the computational resources and time required, to provide a smaller amount of video footage containing one or more events of interest for manual review.
[0076] In some examples of the first embodiment, the motion filtering block 410 and the motion module 208 may be omitted. In some examples, the metadata filtering block 408 and the metadata module 206 may be omitted. In some examples, all of these may be omitted.
[0077] In the first embodiment the system is described as comprising only a single video surveillance system 100 and a single video analysis system 200, for clarity. In other examples, the system may comprise one or more, such as a plurality, of video surveillance systems 100 connected through a communications network to one or more, such as a plurality, of video analysis systems 200. In such examples the one or more video analysis systems 200 may be arranged to obtain video segments stored in any of the one or more video surveillance systems 100, as necessary.
[0078] In the first embodiment the video data store 108 is comprised in the video surveillance system 100 and the memory 204 is comprised in the video analysis system 200. In other examples the video surveillance system 100 and video analysis system 200 may also use external data stores, such as remote servers.
[0079] In the first embodiment the stored video segments are sent through a communications network 112 from the central store 106 of the video surveillance system 100 to the video analysis system 200. In other examples the stored video segments may be obtained in other ways. For example, the stored video segments may be stored in a portable data storage device by the video surveillance system 100 and the portable data storage device transported to the video analysis system 200 so that the stored video segments can be accessed. This may provide improved security.
[0080] The first embodiment describes how the video analysis system 200 analyses stored video segments to identify video segments showing an event of interest. It will be understood that the video analysis system 200 can analyse the stored video segments to identity video segments showing multiple different events of interest by carrying out the method 400 multiple times for different search queries. Such multiple carrying out of the method 400 may be done consecutively or simultaneously.
[0081] The order of the steps of the methods described herein is exemplary, but the steps may be carried out in any suitable order, or simultaneously where appropriate. In particular, in some examples, motion filtering may be carried out at by the video surveillance system 100, either by the central store 106, or by the video cameras 104a to 104n themselves, so that the stored video segments 110 stored in the video data store 108 are motion filtered to only include video segments including movement. In such examples the video analysis system 200 will not need to carry out the motion filtering block 410, so that the motion module 208 will not be required. In such examples the motion filtering block 410 will be carried out first (by the video surveillance system 100) before the other stages of the multi-stage sequential analysis 406.
[0082] In the illustrated example of the first embodiment the multi-stage sequential analysis 406 is carried out by the video analysis system 200. In alternative examples, the metadata based filtering of the metadata filtering block 408 may be carried out by the video surveillance system 100. In some examples, the video analysis system 200 may send a request for stored video segments 110 to the central store 106 together with data identifying the metadata values of interest. The central store 106 can then search the stored video segments 110 for video segments with associated metadata values corresponding to the identified values of interest, and send only the identified video segments to the video analysis system 200.
[0083] In the illustrated example of the first embodiment the multi-stage sequential analysis 406 is applied to stored video segments. In other examples the multi-stage sequential analysis may be applied to "live" or real time video segments.
[0084] Figure 5 shows a schematic diagram of a combined video surveillance and analysis system 600 according to a second embodiment. In the illustrated second embodiment the combined video surveillance and video analysis system 600 is intended to carry out real time monitoring of surveillance video captured by a plurality of closed circuit television (CCTV) video cameras. The events of interest may be of any type, as appropriate to the purpose of the video surveillance and video analysis. In some examples the video surveillance and video analysis may be intended for security reasons, or to investigate or prevent accidents and / or crime, or for any other reason. The use of CCTV video cameras is not essential, and other examples may use alternative types of video cameras.
[0085] As shown in figure 5, the video surveillance and analysis system 600 comprises a plurality of closed circuit television (CCTV) video cameras 602a to 602n which each capture a respective video stream 604a to 604n from a respective field of view together with associated metadata. The video cameras 602a to 602n and the video streams 604a to 604n correspond to the video cameras 102a to 102n and video streams 104a to 104n of the first embodiment, and accordingly will not be described in detail to avoid unnecessary repetition.
[0086] In the illustrated second embodiment of figure 5, the video cameras 602a to 602n are communicatively connected to a video analysis system 606 so that the video analysis system 606 can analyse the "live" video streams 604a to 604n in real time to identify event(s) of interest.
[0087] The video analysis system 606 comprises at least one processor 608, and a memory 610. The at least one processor 606 comprises a motion module 612, an object module 614, a semantic video module 616, and an output module 618 of the video analysis system 602. The different modules 612 to 618 are provided by different software executed by the at least one processor 608. In some examples each module, or some modules, may be provided by software executed by specific dedicated processor(s) ofthe at least one processor 608. The memory 610 may be any form of data store. The video analysis system 606 further comprises a user input device 622 18 and an output device 620. The user input device 622 and an output device 620 correspond to the user input device 218 and the output device 216 of the first embodiment, and accordingly will not be described in detail.
[0088] Figure 6 shows a flow chart of a video analysis method 700 carried out by the video analysis system 606 of the video surveillance and analysis system 600 of the second embodiment in operation.
[0089] In the method 700, in an obtain search query block 702, the video analysis system 606 obtains a search query which defines the objective of the video analysis to be carried out by the video analysis system 606 on the received video streams 604a to 604n. In the illustrated second embodiment, the search query is input by a user into the video analysis system 606 using the user input device 622, and then stored in the memory 610.
[0090] Figure 7 shows a schematic diagram of a search query 800. As shown in figure 7, the search query 800 comprises two elements, a semantic definition 502 of an event of interest, and an object definition 506 of one or more objects expected to be found in association with the event of interest. The search query 800 of the second embodiment is similar to the search query 500 of the first embodiment, but without the range definition 504 of locations and / or times to be searched. It will be understood that the range definition is not required in the second embodiment because the received video streams 604a to 604n are analysed as they are received in real time, and the range of locations to be searched is determined by the locations of video cameras 602a to 602n.
[0091] The search query 800 of the second invention may be entered by the user via the user input device 622 in any convenient manner, such an via a keyboard ora microphone and speech to text system, in a similar manner to the search query 500 of the second invention. The semantic definition 802 and the object definition 804 of the second embodiment correspond to the semantic definition 502 and the object definition 504 of the first embodiment, and accordingly will not be described in detail.
[0092] In the method 700, in a receive video streams block 704, the video analysis system 606 receives the video streams 604a to 604n. The video analysis system 602 then carries out a multistage sequential analysis 706 of the received video streams 604a to 604n based on the input search query 800. In practice, the video surveillance and analysis system 600 may have been capturing video streams 604a to 604n before the search query 800 was obtained, but it will be understood that the video analysis system 602 cannot analyse the received video streams 604a to 604n based on the search query 800 without the search query 800.
[0093] In a first stage of the multi-stage sequential analysis 706, in a motion filtering block 708, the motion module 612 of the video analysis system 606 conducts motion searches of the received video streams 604a to 604n.
[0094] In the motion filtering block 708, the motion module 612 uses one or more motion detection algorithms to identify parts of the received video streams 604a to 604n containing movement. The motion filtering block 708 then defines motion searched video segments each corresponding at least one of the identified parts of the received video streams 604a to 604n. The motion filtering block 708 outputs the motion searched video segments and discards other parts of the received video streams 604a to 604n. Accordingly, the motion filtering block 708 outputs motion searched video segments which contain movement. Similarly to the first embodiment, the motion filtering block 708 may be regarded as filtering out irrelevant video segments which do not contain any movement from the received video streams 604a to 604n.
[0095] Similarly to the first embodiment, any suitable motion detection algorithm may be used. In some examples, a motion detection algorithm may be selected from a group of multiple different motion detection algorithms based on the type of the capturing video camera as identified in the associated metadata, so that video cameras of different types may have their captured video segments filtered using different motion detection algorithms. In some examples, a motion detection algorithm may be selected from a group of multiple different motion detection algorithms based on characteristics of the image, such as light level, image quality, or the like. In examples where some or all of the video cameras 602a to 602n are able to change orientation a suitable motion detection algorithm able to distinguish movement in the imaged scene and changes in the image caused by the changing orientation.
[0096] Then, in an object detection block 710, the object detection module 614 of the video analysis system 606 conducts a search of the motion searched video segments produced by the motion filtering block 708 for the object or objects identified in the object definition 804 of the search query 800. It will be understood that these motion searched stored video segments are smaller (that is, comprise less total video data or footage), than the received video streams 604a to 604n. In the object detection block 710, the object detection module 614 uses one or more image classification algorithms to identify parts of the movement searched video segments containing images of one or more of the object(s) identified in the object definition 804. The object detection block 710 then defines object searched video segments each comprising at least one of the identified parts of the motion searched video segments. The object detection block 710 outputs the output searched video segments, and discards other parts of the motion searched video segments. The object detection block 710 may be regarded as filtering out irrelevant video segments which do not contain any of the objects listed in the object definition 804 from the movement searched video segments.
[0097] Similarly to the first embodiment, any suitable image classification algorithm may be used. One known image classification algorithm which may be used is the YOLO system, which is a neural networked based system. In some examples, an image classification algorithm may be selected from a group of multiple different image classification algorithms based on the type of the capturing video camera as identified in the associated metadata, so that video cameras of different 20 types may have their captured video segments filtered using different image classification algorithms. In some examples, an image classification algorithm may be selected from a group of multiple different image classification algorithms based on characteristics of the image, such as light level, image quality, or the like.
[0098] Then, in a semantic video search block 712, the semantic video module 616 of the video analysis system 606 conducts a search of the object searched video segments produced by the object detection block 710 for the events defined by the semantic definition 802 of the search query 800. It will be understood that these object searched video segments are smaller (that is, comprise less total video data or footage), than the movement searched video segments. In the semantic video search block 712, the semantic video module 616 applies one or more multi-modal text-vision models to each of the video segments of the object searched video segments to identify any video segment sections matching the semantic definition 802 of the event of interest in the search query 800. The semantic video search block 710 then outputs the identified video segment sections and discards other parts of the object searched video segments, It will be understood that some identified video segment sections may comprise only parts of respective video segments of the object searched video segments. The semantic video search block 712 may be regarded as filtering out irrelevant video segments which do not contain events corresponding to the semantic definition 802 from the object searched video segments. It will be understood that these semantic searched video segments are smaller (that is, comprise less total video data or footage), than the object searched video segments.
[0099] Similarly to the first embodiment, any suitable multi-modal text-vision model may be used. One known multi-modal text-vision model is CLIP, which is a neural networked based system.
[0100] Similarly to the first embodiment, in some examples, the one or more multi-modal text-vision models may search each video segment as a number of separate images, separately searching each image for image content corresponding to the events defined by the semantic definition. The illustrated example using CLIP is an example of this type In other examples, the one or more multi-modal text-vision models may be more sophisticated, and may search each video segment as a series of images, searching each image for image content corresponding to the events defined by the semantic definition, and also taking into account movement of and / or interactions between the image content over sequences of images for the events defined by the semantic definition.
[0101] Similarly to the first embodiment, in some examples, a multi-modal text-vision model may be selected from a group of multiple different multi-modal text-vision models based on the type of the capturing video camera as identified in the associated metadata, so that video cameras of different types may have their captured video segments filtered using different multi-modal text-vision model. In some examples, an multi-modal text-vision model may be selected from a group of multiple different multi-modal text-vision models based on characteristics of the image, such as light level, 21 image quality, or the like. In some examples, an multi-modal text-vision model may be selected from a group of multiple different multi-modal text-vision models based on the semantic definition, Some models may be better at searching for particular semantic definitions than others.
[0102] The semantic searched video segments produced by the semantic video search block 712 are output by the output module 618 in a final output block 714 as the output of the multistage sequential analysis 706. In the second embodiment these output semantic searched video segments are then stored by the video analysis system 606 in the memory 610. Optionally, the semantic searched video segments produced by the semantic video search block 712 are displayed to the user for manual review via the output device 620. The display to the user may be carried out in real time using semantic searched video segments output by the output block 714. Alternatively, the semantic searched video segments output by the output block 714 may be stored in the memory 610 and the display to the user may be carried out using the semantic searched video segments stored in the memory 610.
[0103] In the second embodiment the output semantic searched video segments are stored in the memory 610. Alternatively and / or additionally, the video analysis system 606 may be arranged to send the semantic searched video segments elsewhere for storage and / or manual review.
[0104] In each of the blocks 708 to 714 of the multi-stage sequential analysis 706, the identified video segments may have a predetermined minimum length. This predetermined minimum length may, for example, be defined in terms of time or number of images.
[0105] The second embodiment may provide similar advantages to the first embodiment.
[0106] In alternative examples of the second embodiment the motion filtering may be carried out at by the video cameras 602a to 602n themselves, so that the video stream 604a to 604n are motion filtered to only include video segments including movement. In such examples the video analysis system 606 will not carry out the motion filtering block 708, so that the motion module 608 will not be required. In such examples the motion filtering block 708 will be carried out by the respective video cameras 602a to 602n.
[0107] In some examples of the second embodiment, the motion filtering block 708 and the motion module 608 may be omitted.
[0108] In the second embodiment, video data is discarded in each of the motion filtering block 708, the object detection block 710, and the semantic video search block 712. In other examples, some or all of this discarded video data may be stored for subsequent processing, for example being stored in the memory 610.
[0109] In the second embodiment the memory 610 is comprised in the video analysis system 606. In other examples the video analysis system 606 may also use external data stores, such as remote servers.
[0110] In the illustrated example of the second embodiment the multi-stage sequential analysis 706 is applied to stored real time video segments. In other examples the multi-stage sequential analysis may be applied to stored video segments
[0111] In the embodiments above, the video analysis system has a single user input device and output device. In other examples the system may have multiple user input devices and / or output devices to enable use of the system by multiple users. In particular, the video analysis system may have multiple output devices to enable manual review of identified video segments by multiple users.
[0112] The embodiments include one or more processors comprising modules. These modules may comprise software and / or dedicated processing hardware, as appropriate in any particular implementation.
[0113] In the embodiments above, the metadata searched comprises locations and / or times. In other examples, different metadata types may be searched in addition, or alternatively, to locations and times, as appropriate to the event to be identified and the metadata associated with the video streams from the video cameras.
[0114] In the embodiments above, the video cameras are at fixed locations. In other examples some or all of the video cameras may be mobile, for example being carried by ground, marine, or aerial vehicles such as autonomous or remote operated vehicles.
[0115] Features of the first and second embodiments set out above may be exchanged between the embodiments.
[0116] The embodiments described above are fully automatic. In some alternative examples a user or operator of the system may manually instruct some steps of the method to be carried out.
[0117] The acts described herein may comprise computer-executable instructions that can be implemented by one or more processors and / or stored on a computer-readable medium or media. The computer-executable instructions can include routines, sub-routines, programs, threads of execution, and / or the like. Still further, results of acts ofthe methods can be stored in a computer-readable medium, displayed on a display device, and / or the like.
[0118] The methods described herein may be performed by software in machine readable form on a tangible storage medium e.g. in the form of a computer program comprising computer program code means adapted to perform all the steps of any ofthe methods described herein when the program is run on a computer and where the computer program may be embodied on a computer 23 readable medium. Examples of tangible (or non-transitory) storage media include disks, thumb drives, memory cards etc. and do not include propagated signals. The software can be suitable for execution on a parallel processor or a serial processor such that the method steps may be carried out in any suitable order, or simultaneously. This application acknowledges that firmware and software can be valuable, separately tradable commodities. It is intended to encompass software, which runs on or controls "dumb" or standard hardware, to carry out the desired functions. It is also intended to encompass software which "describes" or defines the configuration of hardware, such as HDL (hardware description language) software, as is issued for designing silicon chips, or for configuring universal programmable chips, to carry out desired functions.
[0119] Various functions described herein can be implemented in hardware, software, or any combination thereof. If implemented in software, the functions can be stored on or transmitted over as one or more instructions or code on a computer-readable medium. Computer-readable media may include, for example, computer-readable storage media. Computer-readable storage media may include volatile or non-volatile, removable or non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. A computer-readable storage media can be any available storage media that may be accessed by a computer. By way of example, and not limitation, such computer-readable storage media may comprise RAM, ROM, EEPROM, flash memory or other memory devices, CD-ROM or other optical disc storage, magnetic disc storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Disc and disk, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray (RTM) disc (BD). Further, a propagated signal is not included within the scope of computer-readable storage media. Computer-readable media also includes communication media including any medium that facilitates transfer of a computer program from one place to another. A connection, for instance, can be a communication medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fibre optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of communication medium. Combinations of the above should also be included within the scope of computer-readable media.
[0120] Alternatively, or in addition, the functionality described herein can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, hardware logic components that can be used may include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs). Complex Programmable Logic Devices (CPLDs), etc.
[0121] It will be understood that the benefits and advantages described above may relate to one embodiment or may relate to several embodiments. The embodiments are not limited to those 24 that solve any or all of the stated problems orthose that have any or all of the stated benefits and advantages. Variants should be considered to be included into the scope of the invention.
[0122] Any reference to 'an' item refers to one or more of those items. The term 'comprising' is used herein to mean including the method steps or elements identified, but that such steps or elements do not comprise an exclusive list and a method or apparatus may contain additional steps or elements.
[0123] Further, to the extent that the term "includes" is used in either the detailed description or the claims, such term is intended to be inclusive in a manner similar to the term "comprising" as "comprising" is interpreted when employed as a transitional word in a claim.
[0124] The order of the steps of the methods described herein is exemplary, but the steps may be carried out in any suitable order, or simultaneously where appropriate. Additionally, steps may be added or substituted in, or individual steps may be deleted from any of the methods without departing from the scope of the subject matter described herein. Aspects of any of the examples described above may be combined with aspects of any of the other examples described to form further examples without losing the effect sought.
[0125] It will be understood that the above description of a preferred embodiment is given by way of example only and that various modifications may be made by those skilled in the art. What has been described above includes examples of one or more embodiments. It is, of course, not possible to describe every conceivable modification and alteration of the above devices or methods for purposes of describing the aforementioned aspects, but one of ordinary skill in the art can recognize that many further modifications and permutations of various aspects are possible. Accordingly, the described aspects are intended to embrace all such alterations, modifications, and variations that fall within the scope of the appended claims.
Claims
1. A computer implemented method for analysing video, comprising:obtaining a search query comprising a semantic description of an event of interest and an object definition identifying one or more objects associated with the event of interest:obtaining first video segments;applying one or more image classification algorithms to the first video segments to identify parts of the first video segments containing at least one object of the one or more objects, and based on this identification, defining object searched video segments each containing at least one of the identified parts; andapplying one or more semantic video search models to the object searched video segments to identify parts of the object searched video segments matching the semantic description, and based on this identification, defining semantic searched video segments each containing at least one of the identified parts; andoutputting the identities of the semantic searched video segments.
2. The method of claim 1, wherein the search query further comprises a range of metadata types and values to be searched; andwherein the method further comprises;obtaining second video segments and associated metadata;carrying out metadata searching of the second video segments and associated metadata using the range of metadata types and values to identify parts of the second video segments associated with metadata corresponding to the range of metadata types and values, and based on this identification, defining metadata searched video segments each containing at least one of the identified parts; andusing the metadata searched video segments as the first video segments.
3. The method of claim 1, and further comprising;obtaining third video segments;applying one or more movement detection algorithms to the third video segments to identify parts of the third video segments containing movement, and based on this identification, defining movement searched video segments each containing at least one of the identified parts; andusing the movement searched video segments as the first video segments.
4. The method of claim 1, wherein the search query further comprises obtaining a range of metadata types and values to be searched; andwherein the method further comprises:obtaining fourth video segments and associated metadata;carrying out metadata searching of the fourth video segments and associated metadata using the range of metadata types and values to identify parts of the fourth video segments associated with metadata corresponding to the range of metadata types and values, and based on this identification, defining metadata searched video segments each containing at least one of the identified parts;applying one or more movement detection algorithms to the metadata searched video segments to identify parts of the metadata searched video segments containing movement, and based on this identification, defining movement searched video segments each containing at least one of the identified parts; andusing the movement searched video segments as the first video segments.
5. The method of any preceding claim, and further comprising displaying the semantic searched video segments to a user for manual review.
6. The method of any preceding claim, wherein the one or more semantic video search models include one or more multi-modal text-vision models, and optionally include CLIP.
7. The method of any preceding claim, wherein the obtaining an object definition comprisesapplying one or more natural language processing algorithms to the semantic description of an event of interest to expand the semantic description into a list of one or more names of objects.
8. The method of any preceding claim, wherein the range of metadata types and values define a range of locations.
9. The method of any preceding claim, wherein the range of metadata types and values define a range of times.
10. The method of any preceding claim, wherein the obtained video segments are stored video segments.
11. The method of any one of claims 1 to 9, wherein the obtained video segments are real time video segments.
12. A system for analysing video, the system comprising:a data store arranged to store a search query comprising a semantic description of an event of interest and an object definition identifying one or more objects associated with the event of interest;an object module arranged to apply one or more image classification algorithms to first video segments to identify parts of the first video segments27 containing at least one object of the one or more objects, and based on this identification, defining object searched video segments each containing at least one of the identified parts;a semantic video module arranged to apply one or more semantic video search models to the object searched video segments to identify parts of the object searched video segments matching the semantic description, and based on this identification, defining semantic searched video segments each containing at least one of the identified parts; andan output module arranged to output the identities of the semantic searched video segments.
13. The system of claim 12, wherein the search query further comprises a range of metadata types and values to be searched; andthe system further comprising a metadata module arranged to carry out metadata searching of second video segments and associated metadata using the range of metadata types and values to identify parts of the second video segments associated with metadata corresponding to the range of metadata types and values, and based on this identification, defining metadata searched video segments each containing at least one of the identified parts;wherein the object module is arranged to use the metadata searched video segments as the first video segments.
14. The system of claim 12, wherein the system further comprises:a movement module arranged to apply one or more movement detection algorithms to third video segments to identify parts of the third video segments containing movement, and based on this identification, defining movement searched video segments each containing at least one of the identified parts;wherein the object module is arranged to use the movement searched video segments as the first video segments.
15. The system of claim 12, wherein the search query further comprises a range of metadata types and values to be searched; andthe system further comprising:a metadata module arranged to carry out metadata searching of fourth video segments and associated metadata using the range of metadata types and values to identify parts of the second video segments associated with metadata corresponding to the range of metadata types and values, and based on this identification, defining metadata searched video segments each containing at least one of the identified parts; anda movement module arranged to apply one or more movement detection algorithms to the metadata searched video segments to identify parts of the metadata searched video segments containing movement, and based on this identification, defining movement searched video segments each containing at least one of the identified parts;wherein the object module is arranged to use the movement searched video segments as the first video segments.
16. The system of any one of claims 12 to 15, the system further comprising an output device arranged to display the semantic searched video segments to a user for manual review.
17. The system of any one of claims 12 to 16, wherein the one or more semantic video search models include one or more multi-modal text-vision models, and optionally include CLIP.
18. The system of any one of claims 12 to 17, wherein the range of metadata types and values define a range of locations and / or a range of times.
19. The system of any one of claims 12 to 18, wherein the video segments are stored video segments or real time video segments.
20. A computer-readable medium comprising instructions which, when executed by one or more processors, cause the one or more processor to carry out the method of any of claims 1 to 11.
Citation Information
Patent Citations
Searching recorded video
US20120173577A1
Video search apparatus and method
US20140355823A1
Generating a summary video sequence from a source video sequence
US20170337429A1
Systems and methods for identifying events within video content using intelligent search query
US20210256061A1
Large scale video search using queries that define relationships between objects
WO2016081880A1