Method and system for composing video materials

By receiving the position indications associated with the sequence and timestamp input by the user, video records are collected from multiple video cameras and video data is automatically formed, which solves the problem that operators in the prior art is difficult for tracking and combining video data, and achieves more efficient video data generation.

CN112732978BActive Publication Date: 2025-08-01AXIS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011140123.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-28
Filing Date
2020-10-22
Publication Date
2025-08-01
Estimated Expiration
2040-10-22

AI Technical Summary

Technical Problem

In areas monitored by multiple video cameras, it is difficult for an operator to track specific action processes and combine related video materials, and the prior art requires a large amount of manual input.

Method used

By receiving the sequence of user input, using the position indication associated with timestamps, video records are collected from multiple video cameras, and video data is automatically composed based on user input, and the next user input is guided by map navigation and historical data.

Benefits of technology

The combinatorial process of video data is simplified, manual input is reduced, and the accuracy and efficiency of video data is improved, especially when tracking specific trajectories in the monitoring area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112732978B_ABST
    Figure CN112732978B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for composing video material, and more particularly to a method and system for composing video material of an action process along a trajectory in an area monitored by a plurality of video cameras. A first sequence of user inputs defining a trajectory in an area monitored by a plurality of video cameras is received. Each user input in the first sequence is associated with a timestamp and receives an indication of a position in a map of the area being monitored by the plurality of video cameras. For each user input in the first sequence, video recordings are collected from those cameras among the plurality of video cameras having a field of view covering the position indicated by the user input. The collected video recordings are recorded for a period starting at the timestamp associated with the user input and ending at the timestamp associated with the next user input in the first sequence or when an indication to stop collection is received. Then, video material is composed based on the video recordings collected for the user inputs in the first sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video surveillance of an area by a plurality of cameras. In particular, the present invention relates to a method and system for composing video material of an action process along a trajectory in an area monitored by a plurality of video cameras. Background Art

[0002] Video cameras are commonly used for surveillance purposes. A video surveillance system typically includes a plurality of video cameras installed in the area to be monitored and a video management system. The video recorded by the cameras is sent to the video management system for storage and display to an operator. For example, an operator can display the video recordings from one or more selected video cameras via the video management system in order to track events and incidents occurring in the monitored area. Further, if an incident occurs, video material that can be used as forensic evidence can be composed based on the video recordings stored in the video management system.

[0003] However, as the number of cameras in a video surveillance system increases, it becomes challenging to obtain an overview of the recorded video from all the cameras. Video surveillance devices with hundreds of cameras are not uncommon. For example, it is difficult for an operator to track a specific action process in a scene such as when a person or a moving object moves in the monitored area. Further, a large amount of manual input is required to combine the video material of a specific incident that has occurred in the monitored area, which becomes a cumbersome task.

[0004] Therefore, there is a need for methods and systems that make it easier to combine the video material of an action process in a monitored area. Summary of the Invention

[0005] In view of this, the object of the present invention is to alleviate the above problems and simplify the process of composing video material of an action process in an area monitored by a plurality of video cameras.

[0006] According to a first aspect, there is provided a method for composing video material of an action process along a trajectory in an area monitored by a plurality of video cameras, comprising:

[0007] receiving a first sequence of user inputs, the first sequence of user inputs defining a trajectory in an area monitored by a plurality of video cameras,

[0008] wherein each user input in the first sequence is associated with a timestamp and is received as an indication of a position in a map of the area being monitored by the plurality of video cameras;

[0009] For each user input in the first sequence, video recordings are collected from those of the plurality of video cameras having a field of view that covers the location indicated by the user input, and the video recordings collected are recorded during a period that starts at the timestamp associated with the user input and ends at the timestamp associated with the next user input in the first sequence or when an indication to stop collection is received; and

[0010] A video material is composed based on the video recordings collected for each user input in the first sequence.

[0011] The first sequence of user inputs is generally received and processed sequentially. Thus, when a user input is received, the video recordings for that user input can be collected. Then the next user input is received, and then the video recordings for that next user input are collected. Then the process can be repeated until all user inputs in the first sequence have been received and processed.

[0012] Using this method, the video material is automatically composed according to the trajectory defined via user inputs. Accordingly, the user does not have to browse through all the video materials recorded by multiple cameras to identify the video recordings related to the event of interest. Instead, the user only needs to define a trajectory in the monitored area, and the relevant video recordings depicting the course of action along the trajectory are collected and included in the video material.

[0013] This method further allows the user to freely select a desired trajectory in the monitored area. This is superior to methods in which the relevant video recordings are simply identified by analyzing the recorded video content.

[0014] Video material refers to a collection of video files. The video material can be in the form of an output file in which multiple video files are included.

[0015] At least one of the user inputs can further indicate a portion around the location in the map of the area being monitored, and the size of the portion reflects the degree of uncertainty of the indicated location. When collecting video recordings for the at least one user input, then video recordings are collected from those of the plurality of video cameras having a field of view that overlaps with the portion around the location indicated by the user input. The degree of uncertainty can also be considered the precision of the user input. In this case, a smaller size of the portion reflects a higher precision, and vice versa.

[0016] In this way, the user can indicate one or more surrounding regions in the user input, and videos from those cameras having a field of view overlapping with the region are collected. Since potentially more cameras will have a field of view overlapping with a larger region compared to a smaller region, a larger region generally results in more video recordings being collected. This can be advantageously used when the user is unsure of the location of the next user input. For example, the user may attempt to track an object in a monitored area and be unsure whether the object will turn to a position on the right or a position on the left. The user can then indicate a position between the position on the left and the position on the right, and further indicate a region large enough to cover both the position on the left and the position on the right. As another example, the user may attempt to track a group of objects through a monitored area. The user can then indicate a region around the indicated position such that all objects in the group fall within the region.

[0017] One or more of the plurality of cameras can have a variable field of view. For example, there can be one or more pan cameras, tilt cameras, zoom cameras. In response to receiving a user input in a first sequence of user inputs, the method can further point one or more of the plurality of video cameras to the indicated position in a map of the region being monitored. In this way, cameras pointing in another direction can be redirected to the indicated position so that they capture videos of events at the indicated position.

[0018] In response to receiving a user input in a first sequence of user inputs, the method can further display video recordings from those video cameras among the plurality of video cameras having a field of view covering the position indicated by the user input, starting from the timestamp associated with the user input. This allows the user to view the video recordings collected for the current user input. The user can use the displayed video recordings as a guide for the next user input. For example, the user can see in the video recording that an object is turning in a certain direction in the monitored area. In response, the user can locate the next user input in that direction on a map of the monitored area.

[0019] The method can further give guidance regarding the next user input via a map of the monitored area. Specifically, in response to receiving a user input in a first sequence of user inputs, the method can display one or more suggestions for the location of the next user input in the map of the monitored area. This guidance saves time and simplifies the user's decision. This is also advantageous in cases where there are blind spots in the monitored area that are not covered by any of the cameras. If an object enters a blind spot, it cannot be inferred from the video data where the object will appear after passing through the blind spot. In such cases, the suggested location can indicate to the user where the object will typically reappear after passing through the blind spot. For example, if the currently indicated location on the map is at the start of an unmonitored corridor leading in several directions, the suggested location can indicate to the user at which monitored location the object will typically appear after passing through the corridor.

[0020] One or more suggestions for the location of the next user input can be determined based on the location indicated by the most recently received user input in the first sequence of user inputs and the locations of multiple video cameras in the monitored area. Alternatively or additionally, the suggestion can be based on statistical data regarding common trajectories in the monitored area. Such statistical data can be collected from historical data. The statistical data can be used to calculate one or more most likely next locations given the current location along the trajectory. Then, one or more most likely next locations can be presented to the user in the map of the monitored area as suggestions. In this way, prior knowledge of the locations of multiple video cameras and / or prior knowledge of typical trajectories can be used to guide the user in making a decision regarding the next user input.

[0021] The trajectory that has been input by the user can be stored for later use. Specifically, the method can further include: storing the first sequence of user inputs; and accessing the stored first sequence of user inputs at a later time point to perform the steps of collecting video records and composing video materials. In this way, for example, when forensic video materials need to be generated and output, the user can return to the stored trajectory and use it later to compose the video materials. It can also happen that, at a later time point, additional video records that were recorded but not available when the trajectory was input by the user become available. For example, video records from portable cameras carried by objects in the monitored area are not available until the cameras have uploaded their videos. In such cases, the stored trajectory can be accessed when the additional video records are available for composing video materials that also include some of the additional video records.

[0022] Another advantage of using the stored trajectory is that the trajectory can be modified before composing the video material. More specifically, the method may include modifying the user input in the first sequence of user inputs stored, before performing the steps of collecting video recordings and composing the video material. In this way, the user can adjust the stored trajectory so that the resulting composed video material better reflects the course of action in the monitored area.

[0023] For example, the user input in the first sequence of user inputs stored can be modified by adjusting the position indicated by the user input in a map of the area being monitored. The modification may also include adding or removing one or more user inputs to / from the first sequence, and / or modifying the timestamps associated with one or more user inputs. All timestamps of the user inputs in the first sequence can also be offset by a certain value. The latter can advantageously be used to compose a video material that reflects the course of action along the trajectory at a time point before or after the time point indicated by the timestamp, such as the course of action along the stored trajectory 24 hours before or 24 hours after the timestamp associated with the trajectory.

[0024] The trajectory can be defined in real time via user input, that is, while the video recording is being recorded. In this case, the timestamp associated with the user input corresponds to the time point at which the user input is made.

[0025] Alternatively, the trajectory can be defined via user input after the video has been recorded. More specifically, the method may include receiving and storing video recordings recorded by a plurality of video cameras during a first time period, wherein the step of receiving the first sequence of user inputs is performed after the first time period, and wherein each user input is associated with a timestamp corresponding to a time within the first time period. Accordingly, in this case, the timestamp of the user input does not correspond to the time point at which the user input is made. Instead, the timestamp associated with the user input can be generated by offsetting the time point at which the user input is made by a certain user-specified value. For example, the user can specify an appropriate timestamp for the first user input in the trajectory, and the timestamps of the additional user inputs in the trajectory can be set relative to this first timestamp.

[0026] In addition to multiple video cameras, data sources of other data types can be arranged in the monitoring area. This can include sensors and / or detectors such as microphones, radar sensors, door sensors, temperature sensors, thermal cameras, face detectors, license plate detectors, etc. The method can further include: for each user input in the first sequence, collecting data from other data sources arranged within a predetermined distance from the location indicated by the user input, the collected data from the other data sources being generated during a period starting from the timestamp associated with the user input and ending at the timestamp associated with the next user input in the first sequence or when an indication to stop collection is received; and adding the data from the other data sources to the video material. Thus, the composed video material includes not only video recordings but also data from other types of sensors and detectors that can provide forensic evidence about the course of actions along a trajectory in the monitoring area.

[0027] Sometimes, two trajectories in the monitoring area can overlap. For example, two objects of interest can first follow a common trajectory and then they separate, forming two branches of the trajectory. Conversely, two objects can first follow two separate trajectories but then join each other along a common trajectory. In such cases, it may be of interest to compose a single video material that includes video recordings of both trajectories. To this end, the method can further include:

[0028] receiving a second sequence of user inputs that define a second trajectory in the area monitored by the multiple video cameras,

[0029] wherein the first sequence of user inputs and the second sequence of user inputs overlap because they share at least one user input;

[0030] for each user input in the second sequence that is not shared with the first sequence of user inputs, collecting video recordings from those of the multiple video cameras that have a field of view covering the location indicated by the user input, the collected video recordings being recorded during a period starting from the timestamp associated with the user input and ending at the timestamp associated with the next user input in the second sequence or when an indication to stop collection is received; and

[0031] including in the video material the video recordings collected for each user input in the second sequence that is not shared with the first sequence of user inputs.

[0032] As an alternative, if there is user input in the first sequence and there is user input in the second sequence, and the positions of these user inputs are covered by the field of view of the same camera during an overlapping period, the first sequence of user inputs and the second sequence of user inputs can be considered to overlap. In this case, when there is no overlap with the first sequence, it may be sufficient to collect the video recordings of the second sequence for the user input and the time period.

[0033] According to a second aspect, there is provided a system for composing video material of an action process along a trajectory in an area monitored by a plurality of video cameras, comprising:

[0034] A user interface, arranged to receive a first sequence of user inputs, the first sequence of user inputs defining a trajectory in an area monitored by a plurality of video cameras, wherein the user interface is arranged to receive each user input in the first sequence as an indication of a position in a map of the area being monitored by the plurality of video cameras, and is arranged to associate each user input with a timestamp;

[0035] A data storage, arranged to store video recordings from the plurality of video cameras; and

[0036] A processor, arranged to:

[0037] Receive a first sequence of user inputs from the user interface;

[0038] For each user input in the first sequence, collect from the data storage video recordings from those video cameras among the plurality of video cameras that have a field of view covering the position indicated by the user input, the collected video recordings being recorded in a time period starting from the timestamp associated with the user input and ending at the timestamp associated with the next user input in the first sequence or when an indication to stop collection is received; and

[0039] Compose video material based on the video recordings collected for each user input in the first sequence.

[0040] According to a third aspect, there is provided a computer program product comprising a non-transitory computer-readable medium having computer code instructions stored thereon, the computer code instructions, when executed by a processor, causing the processor to perform the method according to the first aspect.

[0041] The second aspect and the third aspect generally may have the same features and advantages as the first aspect. It should also be noted that unless otherwise explicitly stated, the present invention relates to all possible combinations of features. Description of the Drawings

[0042] The above and additional objects, features, and advantages of the present invention will be better understood from the following illustrative and non - limiting detailed description of embodiments of the present invention with reference to the accompanying drawings, in which like reference numerals will be used for like elements, wherein:

[0043] Figure 1 Schematically illustrates a video surveillance system according to an embodiment.

[0044] Figure 2 Is a flowchart of a method for composing video material according to a first set of embodiments.

[0045] Figures 3a to 3d Schematically illustrates a sequence of user inputs received as an indication of a position in a map of a monitored area.

[0046] Figure 4 Illustrates for Figures 3a to 3d The video records collected for the user inputs in the sequence illustrated in

[0047] Figure 5 Is a flowchart of a method for composing video material according to a second set of embodiments.

[0048] Figure 6 Is a flowchart of a method for composing video material according to a third set of embodiments.

[0049] Figure 7 Schematically illustrates two overlapping sequences of user inputs. Detailed Description of the Invention

[0050] The present invention will now be described more fully hereinafter with reference to the accompanying drawings, in which embodiments of the invention are shown.

[0051] Figure 1 Illustrates a video surveillance system 1 and a video management system 100. The video surveillance system 1 includes a plurality of video cameras 10 installed to monitor an area 12. The video management system 100 will also be referred to herein as a system for composing video material of an action process along a trajectory in the area 12 monitored by the plurality of cameras 10.

[0052] The monitored area 12 is illustrated in the form of a planned map of the area 12. In this example, it is assumed that the area 12 is an indoor area, where walls separate different parts of the area 12 to form rooms, corridors, and other spaces. However, it should be understood that the concepts described herein are equally applicable to other types of areas, including outdoor areas. A plurality of video cameras 10 (illustrated herein by the twelve cameras 10-1 to 10-12 listed) are arranged in the monitored area 12. Each of the cameras 10 has a field-of-view that covers a portion of the monitored area. Some of the cameras 10 may be fixed cameras, meaning they have a fixed field-of-view. Other cameras may have a variable field-of-view, meaning the cameras can be zoomed and / or controlled to move in a pan direction or a tilt direction such that their field-of-view covers different parts of the area 12 at different points in time. As a special case of a camera with a variable field-of-view, there may be cameras carried by objects moving around in the area 12, such as mobile phone cameras, wearable cameras, or airborne cameras. Preferably, the plurality of cameras 10 are arranged in the monitored area 12 such that each point in the monitored area 12 lies within or can lie within the field-of-view of at least one of the plurality of cameras 10. However, this is not necessary to implement the concepts described herein.

[0053] In addition to the video cameras 10, a plurality of data sources 14 may be arranged in the area 12. The data sources 14 can generally generate any type of data that provides evidence of actions or events that have occurred in the monitored area 12. This includes sensors and / or detectors such as microphones, radar sensors, door sensors, temperature sensors, thermal cameras, face detectors, license plate detectors, etc. The data sources 14 may also include point-of-sale systems that register sales and returns of purchases made in the area 12. By configuration, the data sources 14 can be associated with the cameras 10 or another data source. For example, the data source 14-1 may be associated with the camera 10-1, or the data source 14-2 may be associated with the data source 14-1. In addition, chains of such associations can be formed. For example, the data source 14-2 may be associated with the data source 14-1, which in turn is associated with the camera 10-1. Such associations and chains of associations can be used when collecting data from the cameras or the data sources 14. For example, if video is to be collected from a camera over a period of time, data can also be automatically collected from the associated data source during that period of time.

[0054] Multiple cameras 10 and additional data sources 14 (if available) communicate with a video management system 100 via a communication link 16. The communication link 16 can be provided by any type of network, such as any known wired or wireless network. For example, the multiple cameras 10 can send recorded video to the video management system 100 via the communication link 16 for display or storage. Further, the video management system 100 can send control instructions to the multiple cameras 10 to start and stop recording or to redirect or change the zoom level of one or more of the cameras 10.

[0055] The video management system 100 includes a user interface 102, a data storage 104, and a processor 106. The video management system 100 may also include a non-transitory type of computer-readable memory 108, such as non-volatile memory. The computer-readable memory 108 can store computer code instructions that, when executed by the processor 106, cause the processor 106 to perform any of the methods described herein.

[0056] The user interface 102 can include a graphical user interface through which an operator can view video recorded by one or more of the multiple cameras 10. The user interface 102 can also display a map of the monitored area, similar to the map of the area shown at the top of Figure 1 . As will be explained in more detail later, the operator can interact with the map, for example, by clicking on a location in the map with a mouse cursor to indicate a location in the map. If the operator sequentially indicates several locations on the map, the indicated locations will define a trajectory in area 12.

[0057] The data storage 104, which can be a database, stores video recordings received from the multiple cameras via the communication link 16. Through interaction with the user interface 102, the data storage 104 can further store one or more trajectories that have been defined by the operator.

[0058] The processor 106 interacts with the user interface 102 and the database 104 to compose video material of an action process along such a trajectory. This will now be explained in more detail with reference to the Figure 2 flowchart of Figure 2 which shows a first set of embodiments of a method for composing video material. Figure 2 in (and Figure 5 and Figure 6 in) the dashed lines illustrate optional steps.

[0059] In Figure 2 the first set of embodiments shown, it is assumed that the operator provides input regarding the trajectory in area 12 in real time, that is, the video is recorded simultaneously.

[0060] The method starts at step S102 by receiving a first user input via the user interface 102. The user input is received in the form of an indication of a position in the map of the monitoring area 12. This is illustrated in more detail in Figure 3a . On the user interface 102, a map 32 of the monitoring area 12 can be displayed. Via the user interface 102, the user can input an indication 34-1 of a position in the map 32, for example, by clicking on a desired position in the map 32 with a mouse cursor. Here, the indication 34-1 is graphically represented by a star icon, where the center of the icon represents the indicated position. However, it should be understood that this is just one of many possibilities.

[0061] Optionally, the user input can also indicate a sector around the position. The purpose of the sector is to associate the position indication with a degree of uncertainty. The degree of uncertainty reflects the user's certainty about the exact position of the input. In other words, the size of the sector indicates the precision of the input. For example, a larger sector can indicate a more uncertain or less precise position indication compared to a smaller sector. To specify the sector, the user can input a graphical icon such as a circle or a rectangle or Figure 3a a star as shown in

[0062] , where the center of the icon indicates the desired position and the size of the icon reflects the degree of uncertainty. Figure 3a . In the first set of embodiments where the user input is made while the video is being captured, the timestamp corresponds to the time when the user input is made. In the example of Figure 3a , the first user input identifying the position 34-1 is associated with the timestamp T1.

[0063] In some cases, especially when there is a video camera 10 with a variable field of view, the processor 106 can control one or more of the video cameras 10 to point to the indicated position 34-1 in step S103. For example, assuming that the video camera 10-2 is a camera with a pan-tilt-zoom function, the processor 106 can control the video camera 10-2 to point to the indicated position 34-1. It should be understood that in step S103, the processor 106 does not have to redirect all the cameras with a variable field of view, but only those cameras whose fields of view can cover the position 34-1 when redirected or zoomed. If an uncertain area around the position 34-1 has been provided by the user input, it may be sufficient if the field of view of the camera overlaps with the identified area when redirected or zoomed. For example, the processor 106 will not need to redirect the camera 10-3 to the position 34-1 because there is a wall between the camera 10-3 and the position 34-1. The processor 106 can identify candidate cameras to be redirected based on the position of the camera 10 relative to the indicated position 34-1 and using the knowledge of the layout of the area (such as the location of walls or other obstacles).

[0064] In step S104, the processor 106 then collects video recordings associated with the user input received in step S102. The video recordings are collected during a period starting from the timestamp associated with the user input. To do this, the processor 106 first identifies those cameras among the plurality of cameras 10 that have a field of view covering the position 34-1 indicated by the user input. The processor 106 can identify those cameras by using the information about the positions where the cameras 10 are installed in the area 12 and the information about the layout of the area (such as the location of walls and obstacles). Such information is typically provided when the cameras are installed in the area and can be stored in the data storage 104 of the video management system 100. In Figure 3a the example illustrated, the processor 106 identifies the camera 10-1 and the camera 10-2 (after redirection as described above). These cameras are represented by having a black fill in the figure.

[0065] In the case where the user input further defines the area around the position 34-1, the processor 106 can more generally identify video cameras having a field of view overlapping with the area. Thus, more cameras may be identified by the processor 106 when a larger area is indicated by the user input. In the case where a wall in the area 12 divides the area associated with the indicated position into two parts, video cameras 10 located on the other side of the wall compared to the indicated position can be excluded from being identified.

[0066] Furthermore, in the event that there are additional data sources 14 in the area 12, the processor 106 may also identify data sources 14 that are within a predetermined distance from the indicated location 34-1. The predetermined distance may be different for different types of data sources 14 and may also vary depending on the location of the data source in the area 12. Figure 3a In the example of FIG. 3 , the processor 106 recognizes that the sensor 14 - 1 is within a predetermined distance from the location 34 - 1 .

[0067] After the cameras 10 and possibly additional data sources 14 have been identified as described above, the processor 106 collects video and data from these cameras 10 and additional data sources 14. Figure 4 Further diagram in Figure 4 A timeline and video records 43 and data 44 generated by cameras 10-1 to 10-12 and additional data sources 14-1 to 14-4 are shown. Timestamps associated with user inputs are identified along the timeline, such as timestamp T1 associated with a first user input indicating a location 34-1 on a map. Processor 106 collects video records from the identified cameras and from the data sources (if available) starting at timestamp T1 associated with the first user input. Video and data are collected until another user input is received or an indication that no more user inputs are to be received. Thus, in the illustrated example, video records are collected for cameras 10-1 and 10-2 and data source 14-1 until the next user input associated with timestamp T2 is received. Figure 4 , the collected records are indicated by the shaded area.

[0068] Optionally, in step S105, the processor 106 may display the collected video recording on the user interface 102. This allows the user to track the current action at the indicated location 34-1. This also facilitates the user's decision regarding the next user input. For example, the user may see from the video that an object is moving in a certain direction and may then decide to indicate a location in that direction on a map in order to track the object.

[0069] As an option, the processor 106 may also provide one or more suggestions regarding the location of the next user input to the user via the user interface 102. The one or more suggestions may be provided in the map 32 using predefined graphical symbols. Figure 3aIn the example, the processor 106 suggests location 35-1 as a possible location for the next user input. This suggestion guides the user to select the next location. The processor 106 can base its suggestion on multiple factors. For example, it can be based on the current location of the user input 34-1 and the locations of the video cameras 10 and / or additional data sources 14 in area 12. In this way, the processor 106 can suggest the next location covered by one or more cameras. The suggestion can further be based on the layout of area 12. The layout of area 12 provides useful input on the possible trajectories that an object can take given the locations of the given walls and other obstacles in area 12. Additionally or alternatively, the processor 106 can also utilize historical data obtained by tracking objects in area 12. Based on the statistical data of the historical object trajectories in area 12, the processor 106 can infer along which trajectory an object typically moves through area 12. Assuming that the current user input 34-1 is along such a trajectory, the processor 106 can suggest the next location along that trajectory. The location suggested along the trajectory can be selected such that at least one of the cameras 10 has a field of view covering the suggested location.

[0070] Then, the processor 106 waits for further user input via the map 32 shown on the user interface 102. [[ID=,5]]

[0071] If further user input is received, the processor 106 repeats steps S102, S104, and optionally also steps S103, S105, S106 for the new user input.

[0072] Returning to this example, Figure 3b A second user input indicating location 34-2 is illustrated. The processor 106 associates the second user input with a timestamp T2 corresponding to the time point at which the second user input is received. The second user input can be provided by accepting the suggested location 35-1 (e.g., by clicking on the suggested location 35-1 with a mouse cursor). Alternatively, the user input can be provided by simply indicating the desired location in the map 32. In this case, the user input defines a region around location 34-2 that is larger than the corresponding region of location 34-1. This is illustrated by the star icon for location 34-2 being larger than the star icon for location 34-1. Thus, the user input reflects that the uncertainty of the indicated location 34-2 is greater than the uncertainty of the indicated location 34-1.

[0073] In response to the second user input, the processor 306 may optionally continue to direct one or more of the cameras 10 to location 34-2, as described above in connection with step S103. Further, the processor 106 may identify which of the cameras 10 has a field of view that covers the indicated location 34-2 or at least overlaps with the area surrounding the indicated location 34-2 as defined by the second user input. In this case, cameras 10-2, 10-4, 10-5, 10-8, 10-9, 10-12 are identified. Further, the processor 106 may identify whether any of the data sources 14 are within a predetermined distance of the indicated location 34-2. In this case, data source 14-1 is identified. Then, in step S104, the processor 106 collects video recordings from the identified video cameras and from the identified data source (if any). As Figure 4 shown, this collection begins at timestamp T2 and continues until another user input with timestamp T3 is received. Optionally, the collected video recordings may be displayed on the user interface 102 to allow the user to track the action at location 34-2 in real time.

[0074] Further, as Figure 3b shown, the processor 106 suggests multiple locations 35-2 as candidates for the next user input.

[0075] As Figure 3c and Figure 3d shown, the processor 106 repeats the above process for a third user input indicating location 34-3 and a fourth user input indicating location 34-4, respectively. The third user input is associated with timestamp T3, and the fourth user input is associated with timestamp T4. After the fourth user input, the processor 106 receives an indication via the user interface 102 that this is the last user input. For the third user input, and as Figure 4 shown, video recordings are collected from cameras 10-4, 10-11, 10-12 between timestamps T3 and T4. Further, data from data source 14-4 is collected. Additionally, candidate location 35-3 for the next user input is suggested in the map 32. For the fourth user input, between timestamp T4 and the time when the indication that the fourth user input is the last user input is received (the time which is Figure 4 represented by "stop" in

[0076] As in Figure 3dBest seen in, the received sequence of user inputs defines a trajectory 36 in the area 12 monitored by multiple cameras 10. Specifically, such a trajectory 36 is defined by the positions 34-1, 34-2, 34-3, 34-4 indicated by these user inputs. Further, the video recordings collected by the processor 106 as described above show the course of action along the trajectory 36 in the area 12.

[0077] In step S107, the processor 106 then composes video material based on the collected video recordings. Further, data collected from the data source 14 can be added to the video material. The video material can be in the form of an output file to which the collected recordings are added. The video material can be output, for example, to constitute forensic evidence. The video material can also be stored in the data storage 104 for future use.

[0078] The video material can also include a first sequence of user inputs that define the trajectory 36 in the monitored area. Specifically, the positions 34-1, 34-2, 34-3, 34-4 and the associated timestamps T1, T2, T3, T4 can be included in the video material. The video material can further include a representation of the map 32 of the area 12. This allows the recipient of the video material to not only playback the video included in the video material, but also display the map in which the trajectory is indicated simultaneously.

[0079] The video material can also include metadata associated with the cameras 10. The metadata can include an indication of the field of view of the cameras 10 and possibly also how the field of view changes over time. Specifically, the field of view of the camera from which the video recording is collected can be included as metadata. Having such metadata in the video material allows the field of view of the cameras 10 and how it changes over time to be displayed in the map 10. In other words, the metadata can be used to animate the map 10. For portable cameras such as mobile phone cameras or wearable cameras, the metadata included in the video material can relate to the position of the camera and how the position changes over time.

[0080] In a similar manner, the video material can also include metadata associated with the additional data source 14. In this case, the metadata can relate to how the values of the additional data source 14 change over time. The metadata of the data source 14 can be used to animate the map 10, for example, by animating the opening and closing of a door in the map 10 according to the value of a door sensor.

[0081] A signature can be provided for the video material, which prevents the video material from being edited and enables detection of whether the data in the video material has been tampered with. This is advantageous in cases where the video material will be used as forensic evidence.

[0082] In other cases, the video material is editable. In such a case, an edit history can be provided for the video material so that changes made to the video material after it is created can be easily tracked.

[0083] Optionally, in step S108, the memory 108 may also store the user input sequence. For example, the indicated positions 34-1, 34-2, 34-3, 34-4 may be stored in the data storage 104 together with their associated timestamps.

[0084] In a second set of embodiments, the input regarding the trajectory in region 12 is not performed in real time, that is, the video is not recorded simultaneously. More specifically, it is assumed that the video camera 10 records video during a first time period, and the input regarding the trajectory in region 12 is received after the first time period. In other words, the operator wishes to generate video material of an action process that occurred during the first time period along a trajectory in the region. However, the trajectory is not specified until after the first time period. Thus, the second set of embodiments allows the user to generate video material of an action process along a specific trajectory from pre-recorded video.

[0085] Now reference will be made to Figure 5 the flowchart of

[0086] In step S201, the processor 106 receives and stores the video recording captured by the plurality of cameras 10 during the first time period. Such a video recording may be stored by the processor 106 in the data storage 104.

[0087] Then, the processor 106 continues to receive user input in step S202 and collects the video recording for the user input in step S204. Optionally, the processor 106 may also display the video recording collected for the user input in step S205 and display the suggested positions for the next user input in step S206. These steps correspond to steps S102, S104, S105, S106 of the first set of embodiments. However, it is worth noting that since the method operates on previously recorded video data, it is not possible to perform Figure 2 the step S103 of pointing the camera. Further, contrary to the first set of embodiments, steps S202, S204, S205, S206 are performed after the first time period during which the video is recorded.

[0088] To track the course of actions that occur during a first time period, the trajectory defined by the user input sequence needs to be associated with time points during the first time period. Thus, the timestamp associated with the user input should not correspond to the time point at which the user input is received. Instead, the processor 106 associates the user input with a timestamp corresponding to a time point within the first time period. For example, the processor 106 may receive a user input that specifies a time point during the first time period, which should be the timestamp of the start of the defined trajectory of the first user input. Then the timestamps of subsequent user inputs can be set relative to the timestamp of the first user input. In effect, this can correspond to the user observing the recorded material, finding an event of interest, and then starting to track the object of interest as if it were live video. Figure 1 Of course, the difference is that when locating the next appropriate user input, the user can quickly browse through a large amount of material. Another difference is that the user can also track the object backward in time, but the relevant timestamps will be coupled to the time of recording rather than the time of the user input.

[0089] When the processor 106 receives an indication that no further user input will be received, it proceeds to step S207 to compose video material based on the collected video recordings. Optionally, it can also store the received user input sequence in the data storage 104. These steps are performed in the same manner as steps S107 and S108 described in conjunction with Figure 2 the description.

[0090] In Figure 6 the third set of embodiments illustrated in the flowchart, the method operates on the stored user input sequence. More specifically, in step S302, the processor 106 receives a stored user input sequence that defines a trajectory in the area monitored by the plurality of cameras 10. For example, the processor 106 can access the stored user input sequence from the data storage 104 (such as stored previously in step S108 of the first set of embodiments or in step S208 of the second set of embodiments).

[0091] Optionally, the processor 106 can modify one or more of the user inputs in the received user input sequence in step S303. The modification can be in response to a user input. For example, the processor 106 can display the received user input sequence and a map 32 of the monitored area on the display 102. Then the user can adjust one of the indicated positions, for example, by using a graphical representation of the position of the mouse cursor to move it.

[0092] Then, the processor 106 can continue to collect video records for each user input in the sequence at step S304, and compose a video profile based on the collected video records at step S307. Optionally, the processor 106 can also display the video records collected for the user input at step S305, and store the possibly modified sequence of user inputs at step S308. Steps S304, S305, S307, and S308 are performed in the same manner as the corresponding steps of the first and second sets of embodiments, and thus will not be described in detail.

[0093] The embodiments described herein can be advantageously used to compose a video profile of an object moving through a monitored area, such as a person walking through the area. In some cases, the trajectories of two objects may overlap, which means they share at least one location indicated by the user input. For example, two objects can first move together along a common trajectory and then separate, such that the trajectory splits into two sub-trajectories. This is illustrated in Figure 7 where the first trajectory 36-1 is defined by positions 34-1, 34-2, 34-3, 34-4, and the second overlapping trajectory 36-2 is defined by positions 34-1, 34-2, 34-5, 34-6. Alternatively, two objects can first move along two separate trajectories and then combine with each other to move along a common trajectory. In this case, it may be of interest to compose a common video profile for the overlapping trajectories. This can be achieved by receiving a second sequence of user inputs that define a second trajectory 36-2 in addition to the first sequence of user inputs that define the first trajectory 36-1. Then, the processor 106 can identify the user inputs of the second trajectory 36-2 that do not overlap with the user inputs of the first trajectory 36-1. In Figure 7 the example of, the processor 106 will then identify the user inputs corresponding to positions 34-5 and 34-6. Then, the processor 106 can proceed in the same manner as explained in connection with steps S104, S204, and S304 to collect video records from those cameras among the plurality of cameras 10 that have a field of view of the identified user input positions of the second trajectory 36-2. In Figure 7 the example of, the processor 106 can collect video records from the video camera 10-9 for the user input indicating position 3-5. The video can be collected between the timestamp associated with the user input 34-5 and the timestamp associated with the next user input 34-6. Additionally, data can be collected from the data source 14-2 within a predetermined distance of the indicated position 34-5. For the user input indicating position 34-6, the processor 106 can start collecting video records from the video camera 10-10 at the timestamp associated with the user input 34-6 and end when an indication to stop collecting is received. Further, data can be collected from the data source 14-3 within a predetermined distance of the indicated position 34-6.

[0094] It should be understood that those skilled in the art can modify the embodiments described above in various ways and still use the advantages of the present invention as shown in the above embodiments. Therefore, the present invention should not be limited to the embodiments shown, but is only defined by the appended claims. Additionally, as understood by those skilled in the art, the embodiments shown can be combined.

Claims

1. A method for video material of an action process along a trajectory in an area monitored by a plurality of video cameras, comprising: Receiving a first sequence of user inputs, the first sequence of user inputs defining a trajectory in the area monitored by the plurality of video cameras, wherein each user input in the first sequence is associated with a timestamp and is received as an indication of a position in a map of the area being monitored by the plurality of video cameras; In response to receiving a user input in the first sequence of user inputs, displaying video records from those of the plurality of video cameras having a field of view covering the position indicated by the user input to provide guidance to the user for positioning the next user input in the first sequence, the video records starting from the timestamp associated with the user input; For each user input in the first sequence, collecting video records from those of the plurality of video cameras having a field of view covering the position indicated by the user input, the collected video records being recorded in a period starting from the timestamp associated with the user input and ending at the timestamp associated with the next user input in the first sequence or when an indication to stop the collection is received; and Composing video material based on the video records collected for each user input in the first sequence.

2. The method according to claim 1, wherein, At least one of the user inputs further indicates a portion around the position in the map of the area being monitored, the size of the portion reflecting the degree of uncertainty of the indicated position, and wherein, when collecting video records for the at least one user input, collecting video records from those of the plurality of video cameras having a field of view overlapping the portion around the position indicated by the user input.

3. The method according to claim 1, further comprising: In response to receiving a user input in the first sequence of user inputs, pointing one or more of the plurality of video cameras to the indicated position in the map of the area being monitored.

4. The method according to claim 1, further comprising: In response to receiving a user input in the first sequence of user inputs, displaying one or more suggestions for the position of the next user input in the map of the area being monitored.

5. The method according to claim 4, wherein, The one or more suggestions for the position of the next user input are determined based on the position indicated by the most recently received user input in the first sequence of user inputs and the positions of the plurality of video cameras in the area being monitored.

6. The method according to claim 1, further comprising: Storing the first sequence of user inputs, and Accessing the stored first sequence of user inputs at a later time point to perform the steps of collecting video records and composing video material.

7. The method according to claim 6, further comprising: Before performing the steps of collecting video recordings and composing video material, modify the user inputs in the first sequence of stored user inputs.

8. The method according to claim 7, wherein, The user inputs in the first sequence of stored user inputs are modified by adjusting the positions indicated by the user inputs in the map of the area being monitored.

9. The method according to claim 1, wherein The timestamps associated with the user inputs correspond to the time points at which the user inputs are made.

10. The method according to claim 1, further comprising: Receiving and storing video recordings recorded by the plurality of video cameras during a first time period, wherein the step of receiving the first sequence of user inputs is performed after the first time period, and wherein each user input is associated with a timestamp corresponding to a time within the first time period.

11. The method according to claim 1, further comprising: For each user input in the first sequence, collecting data from other data sources arranged within a predetermined distance of the position indicated by the user input, the collected data from the other data sources being generated during a time period starting at the timestamp associated with the user input and ending at the timestamp associated with the next user input in the first sequence or when an indication to stop the collection is received; and Adding the data from the other data sources to the video material.

12. The method according to claim 1, further comprising: Receiving a second sequence of user inputs, the second sequence of user inputs defining a second trajectory in the area monitored by the plurality of video cameras, wherein the first sequence of user inputs and the second sequence of user inputs overlap because they share at least one user input; For each user input in the second sequence that does not share with the first sequence of user inputs, collecting video recordings from those of the plurality of video cameras having a field of view covering the position indicated by the user input, the collected video recordings being recorded during a time period starting at the timestamp associated with the user input and ending at the timestamp associated with the next user input in the second sequence or when an indication to stop the collection is received; and Including in the video material the video recordings collected for each user input in the second sequence of user inputs that does not share with the first sequence of user inputs.

13. A system for composing video material of an action process along a trajectory in an area monitored by a plurality of video cameras, comprising: A user interface arranged to receive a first sequence of user inputs, the first sequence of user inputs defining a trajectory in the area monitored by the plurality of video cameras, wherein the user interface is arranged to receive each user input in the first sequence as an indication of a position in a map of the area being monitored by the plurality of video cameras and to associate each user input with a timestamp; A data storage arranged to store video recordings from the plurality of video cameras; and A processor, arranged to: Receive a first sequence of the user input from the user interface; In response to receiving a user input in the first sequence of the user input, display video recordings from those of the plurality of video cameras having a field of view covering the position indicated by the user input on the user interface, to provide guidance to the user for positioning the next user input in the first sequence, the video recordings starting from the timestamp associated with the user input; For each user input in the first sequence, collect from the data storage video recordings from those of the plurality of video cameras having a field of view covering the position indicated by the user input, the collected video recordings being recorded in a period starting from the timestamp associated with the user input and ending at the timestamp associated with the next user input in the first sequence or when an indication to stop the collection is received; and Compose video material according to the video recordings collected for each user input in the first sequence.

14. A non-transitory computer-readable medium having computer code instructions stored thereon, the computer code instructions, when executed by a processor, cause the processor to perform the method according to claim 1.

Citation Information

Patent Citations

  • Supervision video extraction method and device

    CN104717462A