Video analytics system

By defining a sampling model in 3D vector space and applying machine learning algorithms, the reliability and efficiency issues of existing image and video analysis systems in detecting fraudulent entry are solved, achieving efficient and accurate detection of scenes in real 3D space.

CN115836321BActive Publication Date: 2026-05-08AWAAIT ARTIFICIAL INTELLIGENCE SL (AWAAIT)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
AWAAIT ARTIFICIAL INTELLIGENCE SL (AWAAIT)
Filing Date
2021-06-01
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing image and video analytics systems struggle to achieve reliable, accurate, fast, and real-time automated detection and alerts when processing image and video data from different scenes, viewpoints, and perspectives, especially in specific situations such as detecting fraudulent entry.

Method used

By defining a sampling model in 3D vector space, image frames are sampled and analyzed to extract data of regions of interest, and machine learning algorithms are applied for detection.

Benefits of technology

It improves the accuracy and efficiency of scene detection in real 3D space, simplifies processing complexity, and can reliably detect fraudulent entry from multiple angles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115836321B_ABST
    Figure CN115836321B_ABST
Patent Text Reader

Abstract

The invention relates to a computer-implemented method for sampling and analyzing data of at least one image frame (531) from at least one series of image frames captured by at least one sensor, comprising: defining at least one sampling model (501), wherein the sampling model (501) is defined in a virtual 3D vector space (521) and is based on one or more predetermined shapes (505) in the virtual 3D vector space (521); applying the at least one sampling model (501) to at least one portion of at least one image frame (531) of the at least one series of image frames, wherein the application of the at least one sampling model defines at least one region (529) of the at least one image frame (531) from which data is extracted; extracting (423) data from the at least one region (529) of the at least one image frame (531) defined by the sampling model (501); and analyzing the extracted data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a computer-implemented method for sampling and analyzing data from at least one image frame, a computer-readable storage medium, and a video analysis system. Background Technology

[0002] Image and / or video analytics systems are frequently used to investigate or monitor scenes, locations, objects and / or subjects, people, or crowds of interest to detect and alert on the occurrence of specific situations, patterns, movements, or behaviors. In the following text, the detection of such specific or categorizable situations, patterns, movements, or behaviors can be simply referred to as the detection problem or the problem to be detected by the image and / or video analytics system.

[0003] For example, image and / or video analytics systems can use captured image and / or video data to detect fraudulent or unauthorized access events, such as tailgating or riding at a control gate or doorway. A typical example of such fraudulent access is one or more individuals attempting to exploit the security delay caused by the closing of a gate after the previous individual has successfully passed through, to pass through a control gate (such as a subway station ticket gate). Another example is an individual jumping over a ticket gate. In the case of other access control gates (such as three-legged turnouts), such unauthorized access events can be more diverse, where fraudulent patterns may include fare evaders jumping over the turnout, passing underneath, passing through the same turnout with another passenger (in what is often called a 2x1 pattern), or swinging the upper arm of a three-legged turnout back and forth to use the gaps in movement within certain turnouts (typically present in those designed for entry and exit passages) to enter the payment area without verifying the fare.

[0004] However, providing reliable, accurate, fast, real-time, and automated detection and alerts remains a challenge for current systems and technologies, especially when dealing with large amounts of image and video data, such as those captured by multiple different cameras from different scenes, viewpoints, and perspectives, where the cameras may be stationary or they may be moving on their own. Summary of the Invention

[0005] question

[0006] Therefore, one object of the present invention is to improve computer-implemented video analysis methods and systems for analyzing image and / or video data to detect specific situations, patterns, movements, and behaviors of interest. For example, this may include improvements to computer-implemented video analysis methods and systems, particularly in terms of automation, speed, efficiency, reliability, and simplicity.

[0007] Solution

[0008] According to the present invention, this objective is achieved by a computer-implemented method, a computer-readable storage medium, and a video analysis system for sampling and analyzing data from at least one image frame.

[0009] An exemplary computer-implemented video analysis method according to the present invention for detecting desired specific (e.g., categorizable) patterns or problems in image and / or video data may include a computer-implemented method for sampling and analyzing data from at least one image frame of at least one series of image frames captured by at least one sensor, and may include one or more or all of the following steps:

[0010] Define at least one sampling model, wherein the sampling model is defined in a 3D vector space or a virtual 3D vector space and is based on one or more predetermined shapes in the 3D vector space or virtual 3D vector space.

[0011] • Applying at least one sampling model to at least a portion of at least one image frame in at least one series of image frames, wherein the application of the at least one sampling model defines at least one region of the at least one image frame from which data is extracted.

[0012] • Extract data from at least one region of at least one image frame defined by the sampling model.

[0013] And analyze the extracted data.

[0014] In other words, the exemplary steps described above for sampling and analyzing data from at least one image frame of at least one series of image frames captured by at least one sensor can be used in computer-implemented video analysis methods or video analysis systems, such as video analysis systems configured to detect problems.

[0015] A series or sequence of image frames can be understood in particular as a video stream or a series or sequence of image frames extracted from a video stream, wherein the image frames are captured by a sensor. A sensor should be understood in this document in particular as a device capable of generating data suitable for representation as an image. The term camera should be understood in particular as referring to all such devices or sensors.

[0016] Furthermore, it should be understood that an image or image frame is / can have a digital format, such as an array of digital pixels, or has been / can be converted from an analog format to a digital format.

[0017] The description of the extracted data analyzed herein can be understood as including analyzing extracted data from at least one region of at least one image frame defined by the sampling model for detecting desired specific (e.g., predefined or categorizable) issues or situations of interest, such as the occurrence of certain situations, patterns, movements, actions, etc., in captured image and / or video data of scenes, locations, objects and / or subjects, people or crowds of interest, i.e., in data extracted from at least one image frame in at least one series of image frames captured by at least one sensor (e.g., a camera).

[0018] In other words, extracted data from at least one region of at least one image frame defined by the sampling model can be used as input data for analyzing and detecting the exemplary situation or problem of interest.

[0019] At least one image frame from the at least one series of image frames captured by the at least one sensor can be understood as representing or including image data from a scene observed in real physical three-dimensional space or 3D space.

[0020] The application of the term in the exemplary step of applying at least one sampling model to at least a portion of at least one image frame can be understood in particular to include the step of mapping or projecting at least one sampling model onto at least a portion of at least one image frame.

[0021] The term virtual 3D vector space or 3D vector space as used herein can be understood in particular as an abstract or digital three-dimensional vector space that can be mapped to or extracted from data in at least one image frame captured by the at least one sensor.

[0022] In other words, virtual 3D vector space or 3D vector space can refer to an abstract or numerical representation or approximation of the physical real three-dimensional space or 3D space observed and captured in an image frame by at least one sensor.

[0023] As described above, the sampling model can be defined in a 3D vector space or a virtual 3D vector space, and can be based on one or more (e.g., a set) predetermined shapes in the 3D vector space or virtual 3D vector space.

[0024] These exemplary one or more predetermined shapes in the exemplary virtual 3D vector space may be selected from at least one of the following shapes: three-dimensional shape (i.e., 3D shape), two-dimensional shape (i.e., 2D shape), one-dimensional shape (i.e., 1D shape), or zero-dimensional shape (i.e., 0D shape).

[0025] Specifically, a 3D shape can have a volume and surface oriented and positioned in a 3D vector space; a 2D shape can have a surface oriented and positioned in a 3D vector space; a 1D shape can have spatial extension or length and can be oriented and positioned in a 3D vector space; and a 0D shape can be a point or point-like object positioned in a 3D vector space. In other words, the shape can have a defined orientation and / or position in a (virtual) 3D vector space.

[0026] Specifically, for example, 3D shapes can be parallelepipeds, such as cuboids and / or polyhedra and / or spheres and / or partial spheres and / or cylinders and / or partial cylinders; and / or 2D shapes can be planes or curved surfaces, and / or parallelograms; and 1D shapes can be straight or curved lines or straight or curved line segments. However, it should be emphasized that other 3D or 2D shapes different from the types described above can also be used in the sampling model.

[0027] Using a parallelepiped as a 3D shape as part of one or at least one sampling model (which will become apparent from the further examples provided below) can lead in particular to enhanced and faster processing of the image frames to be analyzed. Last but not least, this is due to the computational ease of extracting data from the image frames using the 3D shape as part of the sampling model, and the computational ease of storing and processing the data extracted from the image frames using a sampling model containing a parallelepiped in the form of a multidimensional array (e.g., a 3D tensor).

[0028] Furthermore, the parallelepiped can be considered a suitable and versatile shape for approximating a large number of real-life objects with sufficient accuracy for most applications / problems detected through video analysis, such as doors, boxes, houses, and cars.

[0029] Furthermore, parallelepipeds are easily defined using a set of three spatial vectors, which simplifies parameterization (i.e., modeling such shapes) and improves computational efficiency when parallelepipeds are used as 3D shapes in sampling models.

[0030] A convenient 2D shape for orientation and positioning in 3D vector space is, for example, a parallelogram. Similarly, the computational ease of defining such a 2D shape using two spatial vectors can benefit the parameterization and computational efficiency of sampling models that include such 2D shapes. As for parallelepipeds, another advantage of using parallelograms is that many real-world objects (e.g., streets, squares, signs) are rectangular objects or reasonably described by rectangles, and therefore can be represented and approximated with high accuracy by such 2D shapes in sampling models.

[0031] At least one sampling model defined in a 3D vector space or virtual 3D vector space and used and applied according to the steps described above and herein can be based on any combination and any number of any of the 3D and / or 2D and / or 1D and / or 0D shapes identified above as defined in the (virtual) 3D vector space. In other words, the exemplary sampling model defined in the 3D vector space or virtual 3D vector space can consist of multiple combinations and any number of predetermined shapes that can be understood to form a sampling model space that can be transformed or mapped to or applied to or projected onto at least a portion of at least one image of at least one series of image frames, which may be referred to as the image frame space or projected 3D space.

[0032] When at least one sampling model is applied, mapped, or projected onto at least a portion of at least one image in at least one series of image frames, the at least one sampling model may be associated with one or more reference points in at least one image frame of at least one series of image frames.

[0033] Specifically, associating a sampling model with one or more reference points in at least one image frame may include performing a mapping transformation or projection, such as parallel projection, between one or more points of at least one sampling model and one or more reference points in at least one image of at least one series of image frames.

[0034] The exemplary reference point may be an easily identifiable feature, such as a vertex or marker or other architectural feature in furniture or equipment or an object, whose coordinates may be correlated between 3D real space and 2D projection, especially when their true 3D coordinates (absolute or relative) are known or can be inferred from the physical true size or other architectural features of such furniture or equipment.

[0035] Alternatively, a dedicated object or marker can be placed temporarily or permanently in a real 3D scene to serve as a reference point.

[0036] For example, since reference points for vertices or markers or other architectural features in furniture or equipment already exist and can be well identified in real 3D space and its 2D projection space (video stream), the coordinates of vertices or markers or other architectural features in furniture or equipment can be correlated between 3D real space and 2D projection, especially when their real 3D coordinates (absolute or relative) are known or can be inferred from the physical real size or other architectural features of such furniture or equipment.

[0037] However, it should be noted that the computational projection or mapping of a sampling model defined in 3D space (e.g., including 3D shapes) onto an image frame does not need to be based on the following highly complex mathematical projection procedure for perfect 3D to 2D projection: approximate projection, such as simple parallel projection (e.g., parallel lines in 3D space maintain parallelism in 2D projection) may be sufficient to obtain acceptable results.

[0038] In exemplary cases where the image frame to be analyzed has an unfavorable viewpoint (such as the vanishing point being not far from the center of the image) or where the image has some aberrations (such as spherical aberration associated with wide-angle lenses), multiple local and simple transformations can be used instead of developing a fitting transformation for the entire scene of the image frame.

[0039] Therefore, a mapping transformation or conversion or projection between a 3D vector space or virtual 3D vector space and a two-dimensional projection of a real physical 3D space captured in at least one image frame and representing the real physical 3D space can be established.

[0040] At least one sampling model may be based on a predetermined shape described above in a (virtual) 3D vector space, which may be divided into one or more elements or blocks constituting the shape. Specifically, the shape in the virtual 3D vector space on which the sampling model is based may be uniformly or non-uniformly divided into one or more elements or blocks or sub-shapes constituting the predetermined shape in any or all of its geometric dimensions.

[0041] The elements or blocks that can constitute shapes in (virtual) 3D vector space can also be defined as the smallest indivisible unit of a shape and can also be called shape atoms, primitives, or voxels.

[0042] Extracting data from at least a portion of at least one image of at least one series of image frames to which a sampling model is applied, mapped, or projected may include extracting data from image frame pixels located in an image frame region contained or covered by the shape of the sampling model applied to at least a portion of the at least one image.

[0043] In other words, applying, mapping, or projecting at least one sampling model can define at least one region of interest or block of interest from at least one image frame from which data is to be extracted, such as the periphery of the region of interest or block of interest.

[0044] Specifically, extracting data from at least a portion of at least one image frame of at least one series of image frames to which the sampling model is applied, mapped, or projected may include extracting data from image frame pixels located in an image frame region contained or covered by elements or blocks of the shape of the sampling model applied to at least a portion of the at least one image.

[0045] In other words, each element or block of the shape of the sampling model applied to at least a portion of at least one image can define a projection region on at least a portion of at least one image frame.

[0046] More specifically, for example, each indivisible smallest unit of a shape, i.e. each shape atom or primitive or voxel, can define a corresponding projection region, i.e., a projection shape atom region, a projection primitive region, or a projection voxel region, on at least a portion of at least one image frame on which the sampling model is applied, thereby defining one or more regions, such as pixel regions, on an image frame on which one or more series of image frames captured from data by at least one sensor can be extracted or read out.

[0047] Data extracted or read from the projection region (i.e., the image frame pixels within the projection region) on an image frame defined by at least one sampling model, as exemplarily described above, can then be saved or stored in one or more arrays, such as in one or more multidimensional arrays, or in one or more tensors.

[0048] It should be noted that projected regions from different shapes or from different indivisible minimum units (i.e. from atoms of different shapes or different primitives or different voxels) can overlap.

[0049] As described above, once the sampling model has been projected onto an image frame, each of its projected element shapes, shape blocks, shape atoms, shape primitives, or shape voxels can define a region and / or perimeter on the image frame. For example, in the case of projecting a (virtual) one-dimensional shape of the sampling model, the projected region and perimeter may become the same, resulting in the projected voxel being just a line.

[0050] In the case of a (virtual) 0D shape in a projection sampling model, the projection area and perimeter may be folded into a single point or a single pixel or a portion of a pixel in the image frame to be analyzed.

[0051] Extracting information from the image frame to be analyzed, i.e., extracting data contained in the image frame defined by the projected shape region, such as data contained in the corresponding projected voxel region or contained in its periphery or contained in a combination of both, may in particular include extracting each projected shape, i.e. each projected element or block of the shape, especially one or more image pixel data values ​​of each shape atom or shape primitive or shape voxel.

[0052] In particular, extracting data from at least one region of at least one image frame defined by the sampling model may include extracting pixel values, such as brightness and / or color (e.g., color in a color space model such as the RGB color model), from pixels identified by the sampling model, such as image frame pixels located in the region of the projection shape and / or along or located on the projection perimeter of the projection shape (i.e., the shape projected onto the image frame).

[0053] It is important to emphasize here that the term can refer to the region or block on the image frame covered by the shape or predetermined shape of the sampling model applied to / mapped onto / projected onto the image frame, or it can refer to a pixel line or simply a single pixel or a portion of pixels in the image frame. In other words, the projection of the 1D or 0D shape of the sampling model can define one or more regions on the image frame to define a set of pixels or one or more pixels from which data is to be extracted.

[0054] Furthermore, extracting data from at least a portion of at least one image frame to which the sampling model has been applied may include transformed data.

[0055] For example, extracted data (e.g., pixel data) in image pixels within a projected shape region (e.g., within a projected shape atomic region or a projected shape primitive region), or, i.e., extracted data in image pixels covered or defined by at least one sampling model, can be transformed by calculating the maximum or minimum value or average value or pattern of pixel values ​​within the projected shape region and / or along the periphery of the projection.

[0056] Another exemplary transformation of the extracted data (i.e., data extracted from image pixels within a projected shape region, such as pixel data) may include applying a function to some or all of the extracted data. For example, applying a weighted average of some or all data values ​​extracted from pixels of an image frame within a projected shape region (e.g., within and / or around the projected voxel region or along the periphery of the projected shape region or a portion thereof).

[0057] This optional data transformation can facilitate further processing and analysis of the extracted data, and can, for example, reduce the computational burden by compressing or densifying the data when it is used to detect the aforementioned situations and problems.

[0058] Alternatively or additionally, it may be conceivable that the entire image frame to which the sampling model is applied / mapped / projected, or at least a portion thereof, is preprocessed or pre-processed before data is extracted from the image frame.

[0059] In this context, an image frame that has not undergone any preprocessing or pre-processing can also be referred to as a raw image frame or a raw image.

[0060] For example, digital processing can be applied to at least a portion of an image frame to determine edges and / or contours and / or motion flow and / or blur at least a portion of the image frame and / or modify the contrast or lighting or other properties of the original image through digital processing, or computationally segment the original image or a portion thereof, or a combination of these and / or other digital image processing procedures.

[0061] This optional and exemplary preprocessing or pre-processing can particularly facilitate the extraction and analysis of image data for a given situation or problem to be detected, because such preprocessing or pre-processing can increase the signal-to-noise ratio of the data signal to be detected for the situation or problem to be detected.

[0062] In the video analysis methods and systems described herein by way of example, it is particularly likely that the same sampling model can be applied to different portions of at least one image in at least one series of image frames.

[0063] Alternatively or additionally, it is possible that the same sampling model can be applied to multiple images in at least one series of image frames, or the same sampling model can be applied to all images in at least one series of image frames.

[0064] Furthermore, the same sampling model can be applied to multiple images from multiple different series of image frames (such as multiple different series of image frames captured by one or more sensors from different perspectives of the same scene or situation to be analyzed).

[0065] It should also be noted that the expression "the same sampling model" can refer to the same sampling model and / or a sampling model with the same topology (e.g., a predetermined shape with the same topology).

[0066] When the position of a sensor (such as a camera) used to capture images from a scene or situation in real 3D space changes, the sensor’s viewpoint and / or field of view also changes, which in turn causes the appearance size, shape and even color of different elements, objects or subjects within its view to change.

[0067] Therefore, different sensors (such as cameras) in the same or different locations may have different perspectives on the problem or situation to be detected.

[0068] Even though a single sensor (such as a single camera) may capture various instances of the same problem, these instances are located in different areas of the captured scene, and therefore at different distances and directions from the sensor's or camera's perspective.

[0069] As mentioned earlier, processing and analyzing image data from different perspectives of the same scene or situation is a challenge for current state-of-the-art video analytics systems and technologies.

[0070] For example, the parameters of a solution model for detecting and analyzing a problem or situation from image data with multiple different viewpoints need to be independently tuned and adjusted for each individual instance of the problem, scene, or situation to be analyzed (e.g., for each individual different viewpoint of the same problem, scene, or situation in the real 3D space to be analyzed) in order to, for example, try to take into account the variations in the appearance size, shape, and even color of different elements, objects, or subjects in the same scene or situation captured from different viewpoints by one or more sensors.

[0071] In particular, for example, when using machine learning algorithms and techniques to analyze data extracted from images to detect specific problems or situations, specific training requirements are needed for each instance of the problem to be detected (e.g., for each single different viewpoint of the scene or situation in the real 3D space where the specific problem or situation is to be detected and analyzed), requiring specific training or specific training datasets and / or specific different sampling models.

[0072] Alternatively or additionally, for example, it may be necessary to solve for more parameters of the model to correctly analyze the different perspectives between problem instances, which would require increasing or expanding the labeled sample set of training data (e.g., training image frames) from different perspectives so that the machine learning model can learn appropriately.

[0073] Surprisingly and surprisingly, the video analysis system method described above and herein for sampling and analyzing data from at least one image frame to detect a specific problem or situation in a scene or situation in real 3D space captured by at least one sensor in the at least one image frame, with the same training or the same sampling model, can be used to effectively and efficiently train machine learning algorithms to reliably detect desired problems or situations in scenes or situations in real 3D space for all or most possible different instances of the problem, i.e. for all or most possible different viewing angles of images taken by the at least one sensor from different perspectives.

[0074] In other words, the same sampling model (wherein, as described above and herein, the sampling model is defined in a virtual 3D vector space and is based on one or more predetermined shapes in the virtual 3D vector space) can be applied to one or more instances of the same problem or condition to be detected in a video stream (e.g., a series of image frames obtained from a sensor) and / or can be applied to other different instances of the same problem or condition to be detected in multiple different video streams (e.g., a series of image frames obtained from multiple different sensors) from the same or different scenes in real 3D space, and the same detection and / or analysis method or algorithm (e.g., the same machine learning algorithm) can be used to reliably detect the desired problem in all or most instances of said problem or in all or most instances of similar problems.

[0075] In other words, the sampling techniques exemplarily described in this paper greatly simplify and reduce the complexity and computational burden of analyzing video streams to detect desired problems or situations in a scene in real 3D space captured by one or more sensors.

[0076] The sampling technique exemplarily described herein for extracting and analyzing data from a series of image frames captured by at least one sensor from at least one sensor observing or monitoring a scene in real 3D space provides a more accurate, efficient, and effective representation in a (virtual) 3D vector space of a real physical object or subject for analysis in video analytics, particularly real three-dimensional objects or subjects that exist in the image frames captured by the sensor. This contrasts with conventional sampling techniques for extracting and analyzing data from image frames of a video stream, which do not take into account the three-dimensional spatial information present in the image frames of a scene in real 3D space, but instead use a planar two-dimensional method only when sampling, extracting, and analyzing data from image frames of a video stream.

[0077] As mentioned above, data extracted from at least one region of at least one image frame defined by the sampling model can be used to detect specific problems, patterns, or situations.

[0078] For example, the problem, pattern, or situation to be detected may include a predetermined situation and / or movement and / or behavior and / or action of an object and / or subject in a real 3D scene, represented / present in at least a portion of at least one image of at least one series of image frames captured by at least one sensor, such as fraudulent entry at a control gate or ticket gate.

[0079] For example, if the task of a video analytics system is to detect problems or situations at control gates or ticket gates, the goal of video analytics detection might be to calculate the number of passengers passing through and / or to detect and count fare evaders.

[0080] Another example of the type of problem, pattern, or situation to be detected might be a door, where the goal might be to count people crossing from one direction or another and trigger an alarm if they cross the door in the wrong way or in groups instead of individually.

[0081] Another example could be monitoring or supervising a room, or even a small space like the inside of an elevator, or an open space, or an area within an open space that is somehow designated, where the purpose is to detect leftover objects, or panic situations (such as people moving at an unusual speed), or fights, or loitering, or oversized objects, or speed monitoring, or detecting people or vehicles or other objects intruding into or moving within a space, or estimating or determining occupancy rates, or other situations.

[0082] When a video analytics system detects the specific problem, pattern, or situation using the sampling techniques described herein for extracting and analyzing data from image frames, it may, for example, provide a notification or alert to the user of the video analytics system or another software component.

[0083] To analyze the extracted data in order to detect specific problems, patterns, or situations using the sampling techniques described herein through a video analysis system, a machine learning system may be used, wherein, for example, the extracted data may be used as input data for the machine learning system to train the machine learning system to detect one or more desired patterns, problems, or situations represented or present in at least a portion of at least one image from at least one series of image frames captured by at least one sensor. For example, any of the aforementioned patterns, problems, or situations may include, for example, predetermined situations and / or movements and / or actions of objects and / or subjects in a real 3D scene, such as fraudulent entry at a control door or ticket gate, or any other type of desired problem, pattern, or situation to be detected.

[0084] Here, detecting one or more desired patterns, problems, or situations represented / existing in at least a portion of at least one image from at least one series of image frames captured by at least one sensor can particularly include: detecting multiple patterns, problems, or situations present / represented in said at least one image, or multiple patterns, problems, or situations of different types, categories, or classifications. For example, in the case of fraudulent entry at a control door or ticket gate (e.g., a three-legged revolving door), analysis of the extracted data can not only simultaneously detect whether fraudulent entry, such as a fare evasion, has occurred, but also the analysis of the same extracted data can detect / determine / classify what kind / what type of fraudulent entry occurred, such as a subject jumping over a revolving door, a subject passing under a revolving door, a subject swinging a revolving door, or two subjects passing through together.

[0085] Once a possible exemplary machine learning system is trained to detect one or more desired patterns, problems, or conditions represented / present in at least a portion of at least one image from at least one series of image frames captured by at least one sensor, such as one or more predetermined situations and / or movements and / or actions of objects and / or subjects in a real 3D scene as exemplified above, the extracted data can be used as input data for the trained machine learning system to analyze the data and detect the presence or absence of one or more desired patterns, problems, or conditions with reliable accuracy.

[0086] Further, applying, mapping, or projecting at least one sampling model onto at least a portion of at least one series of image frames may include taking into account the movement of at least one sensor during the capture of image frames from at least one series of image frames.

[0087] For example, such rotational motion of the sensor can be considered if the sensor is rotating, such as when a camera is rotating to measure or monitor a wider area of ​​space. For instance, the exemplary rotational motion of the sensor can be synchronized with / applied to the rotation of the sampling model, such that the application, mapping, or projection of the sampling model includes numerical computation steps corresponding to different rotational positions of the sensor when defining the image frame region from which data is to be extracted; that is, each different rotational position of the sensor corresponds to a different mapping or projection of the sampling model onto the image frame.

[0088] Alternatively or additionally, linear motion of the sensor can be considered by applying, mapping, or projecting a sampling model to each linear spatial location before extracting data from one or more image frames. Other, more complex motions of the sensor can also be considered.

[0089] Additionally or alternatively, as described above, image frames from multiple different series of image frames taken from different perspectives by multiple sensors used to capture image frames can also be sampled and analyzed, wherein applying or mapping or projecting at least one sampling model onto the image frames taken by the multiple sensors can take into account the different viewpoints of the multiple sensors, for example, using the same or simulated reference points in the image frames taken from different viewpoints.

[0090] For example, in a scenario where two different sensors (e.g., two cameras) are observing the same scene at a series of ticket gates (e.g., a ticket gate with multiple revolving doors), one sensor can observe from the left side of the ticket gate array while the other observes from the right side. This allows the behavior of each door in the ticket gate array, i.e., the behavior of each revolving door, to be monitored and sampled simultaneously from both sensors. Data extracted from the images of both sensors can be combined; for example, before analyzing the extracted data, the data from the different extractions can be concatenated to form a single multidimensional data array or tensor. The extracted data can then be analyzed as an analysis of a single problem or situation to be detected, observed simultaneously from two different perspectives, thereby improving the accuracy of the analysis.

[0091] The steps described above and herein for sampling and analyzing data from at least one image frame of at least one series of image frames captured by at least one sensor can be stored as instructions on one or more computer-readable storage media, wherein, when executed by one or more processors, the instructions can instruct the one or more processors to perform any of the steps described herein for sampling and analyzing data from at least one image frame of at least one series of image frames.

[0092] An exemplary video analysis system according to the present invention may include:

[0093] At least one sensor, such as a camera, is configured to capture image frames.

[0094] • And at least one computing system, including one or more processors, the processors being configured to implement a method for sampling and analyzing data from the at least one image frame, according to any of the steps described above and herein for sampling and analyzing data from the at least one image frame.

[0095] Possible exemplary camera types may include, in particular, analog and digital surveillance cameras, Internet Protocol (IP) cameras, 3D cameras, such as time-of-flight cameras or thermal imagers.

[0096] An exemplary processor for an exemplary video analytics system may include one or more central processing units (CPUs) and / or one or more graphics processing units (GPUs). It should be noted that the sampling model described herein is computationally efficient and can meet the computational resource requirements of a typical personal computer (PC). Attached Figure Description

[0097] The following figure illustrates this example:

[0098] Figure 1 Example image frame

[0099] Figure 2 Exemplary 3D shape

[0100] Figure 3 Exemplary sampling model

[0101] Figure 4 An exemplary diagram illustrating the relationship between spaces.

[0102] Figure 5a Example (First) Ticket Gate Problem

[0103] Figure 5b Example (Second) Ticket Gate Issue Detailed Implementation

[0104] Figure 1 An example of an image frame 100, which is captured from at least one exemplary series of image frames from at least one exemplary sensor (not shown), is illustrated schematically, wherein the image frame 100 has an exemplary digital format and includes a plurality of exemplary pixels 103.

[0105] Figure 1 An exemplary projection 104 onto an image frame 100 is further illustrated by an exemplary predetermined shape, wherein the exemplary predetermined shape is, for example, a 2D shape in the form of a parallelogram defined in a (virtual) 3D vector space that has been projected onto the image frame 100.

[0106] In other words, the 2D shape can define an exemplary sampling model for sampling and analyzing data from image frame 100, wherein the application or projection of the sampling model, i.e., the projection of the 2D shape, i.e., the parallelogram defined in (virtual) 3D vector space, defines an exemplary region 107, i.e., an exemplary region of interest, of at least one image frame from which data is extracted.

[0107] In other words, in order to sample and analyze data from image frame 100, data is extracted only from image frame pixels 105 located within region 107 and / or on or within the periphery 106 of the projection 104 of the sampling model (i.e., the periphery 106 of the projection 104 of the exemplary parallelogram 2D ​​shape).

[0108] For completeness, note that reference numerals 101 and 102 in the appendix exemplarily represent possible coordinate axes of an image frame, such as the X-axis 101 and the Y-axis 102.

[0109] Figure 2 An example of a possible exemplary sampling model 210 is schematically shown, which includes a possible 3D shape 200 in the form of an exemplary cuboid 211 defined in an exemplary (virtual) 3D vector space 212 spanned by exemplary orthogonal coordinate axes X, 205, Y, 206, Z, 207.

[0110] An exemplary sampling model 210 or an exemplary 3D shape 200, namely an exemplary cuboid, is defined by, for example, by an exemplary set 209 of four points, such as four reference points P1, 201, P2, 202, P3, 203, P4, 204, with coordinates P1(x1...). m ,y1 m ,z1 m P2(x2) m ,y2 m z2 m P3(x3) m y3 m z3m ) and P4(x4 m y4 m z4 m ), where x, y, z are the coordinates of the orthogonal coordinate axes X, 205, Y, 206, Z, 207, m in the superscript index represents the exemplary sampling model 210, and the subscript indices 1, 2, 3 and 4 represent the reference point numbers.

[0111] In other words, the coordinates 209 of the exemplary reference points P1, 201, P2, 202, P3, 203, P4, 204 are provided exemplary in the coordinates of an exemplary (virtual) 3D vector space spanned by exemplary orthogonal coordinate axes X, 205, Y, 206, Z, 207.

[0112] In the exemplary case shown, P1(x1) m ,y1 m ,z1 m ) is located at the origin of the exemplary (virtual) 3D vector space, i.e., P1(x1) m ,y1 m ,z1 m =P1(0,0,0), and other points lie on the example coordinate axes, i.e., P2(x2) = P1(0,0,0), and P2(x2) = P1(0,0,0). m ,y2 m z2 m )=P2(x2 m ,0,0), where x2 m ≠0, P3(x3) m y3 m z3 m )=(0,y3 m ,0), where x3 m ≠0, and P4(x4) m y4 m z4 m )=(0,0,z4 m ), where x4 m ≠0.

[0113] The exemplary sampling model 210 or the exemplary 3D shape 200 may, for example, represent a model of an object in real 3D space, such as, for example, a control door or ticket gate or corridor or volume in real 3D space.

[0114] The dimensions of the exemplary sampling model 210 or the exemplary 3D shape 200 can be adjusted in particular to better match or approximate the specific size and scale of different instances or implementations of objects in the real 3D space that the sampling model 210 or the exemplary 3D shape 200 should represent.

[0115] In this way, the same sampling model can be used to sample similar objects in real 3D space, and / or the same model can be applied to different viewing perspectives of the same object in real 3D space, such as when captured in image frames from different sensors with different viewpoints of objects or scenes in real 3D space. In this context, the expression "same sampling model" can be understood in particular as a sampling model with the same topology (i.e., including a predetermined shape with the same topology).

[0116] Furthermore, the same sampling model and / or the same model can be applied to different objects in real 3D space that have the same or similar shapes or topologies, such as different implementations or instances of ticket gate control doors in different physical locations captured by different sensors in different series of image frames, such as different subway stations.

[0117] As generally stated above, at least one sampling model may be based on a predetermined shape in a (virtual) 3D vector space, which itself may be divided into one or more elements or blocks or sub-shapes constituting said shape.

[0118] For example, the sampling model may be based on a shape in a 3D vector space that can be uniformly or non-uniformly divided into one or more elements or blocks constituting the shape in any or all of its geometric dimensions.

[0119] The elements or blocks that can constitute shapes in (virtual) 3D vector space can also be defined as the smallest indivisible unit of a shape and can also be called shape atoms, primitives, or voxels.

[0120] exist Figure 2 In the exemplary case shown, the exemplary 3D shape 200, namely the exemplary cuboid 211, is exemplary divided into voxels by using cuts of four equal slices along an exemplary longitudinal direction (i.e., along the X-axis 205), for example, the X-axis 205 may represent the direction of movement of the body or passenger; using cuts of three equal slices along an exemplary transverse direction (i.e., perpendicular to the longitudinal direction, but in the same horizontal plane), i.e., along the Y-axis 206; and using cuts of two equal slices along an exemplary vertical direction, i.e., along the Z-axis 207.

[0121] An exemplary division or exemplary slice of the exemplary 3D shape 200 then generates 4*3*2=24 smaller 3D shapes, namely smaller 3D cuboids, namely exemplary voxels 208.

[0122] Each of the voxels 208 can then be associated, for example, with an element and / or value in a data structure such as a multidimensional array (e.g., a tensor of dimension (4,3,2)).

[0123] It should be emphasized that the partitioning or slicing described here and above for the exemplary 3D shape 200 is just an example, and other partitioning or slicing schemes can also be applied to partitioning or slicing the exemplary shape of the exemplary sampling model. For example, the exemplary 3D shape can be partitioned into voxels, which can be associated with elements and / or values ​​in a data structure such as a multidimensional array (e.g., a tensor of dimension (i,j,k), where i,j,k are integers greater than 0).

[0124] This also applies to other shapes, such as two-dimensional and / or one-dimensional shapes of the sampled model.

[0125] In order to apply, map, or project the exemplary sampling model 210, namely the exemplary 3D shape 210, namely the exemplary 3D cube 211 and its voxels 208, onto a two-dimensional image frame, the exemplary process may include, for example, identifying four image reference points I1, I2, I3, I4 in the image frame, which are easily identifiable and can be replicated in different scenes in real 3D space, and then performing a mathematical transformation from the (virtual) 3D vector space to the 2D image frame space.

[0126] The exemplary image frame space can, for example, be defined by an exemplary orthogonal coordinate axis X. f Y f "Spanning" refers to the span of a frame.

[0127] For example, a parallel projection can be easily defined using four image reference points located at the vertices of the ticket gate frame (see also...). Figure 5a and Figure 5b One of them is located outside the planes defined by the other three.

[0128] Once these four exemplary image reference points are identified and located on the two-dimensional image frame to be sampled and analyzed, their pixel coordinates can be obtained or identified.

[0129] For example, let us represent the coordinates of the exemplary image reference points I1, I2, I3, I4 as follows: Here, "f" again refers to the frame, and the coordinates are the exemplary orthogonal coordinate axis X relative to the image frame. f Y f Provided.

[0130] If we associate the coordinates of these exemplary image reference points in the image frame with the corresponding reference point coordinates of the exemplary sampling model 210, then the exemplary sampling model 210 is the exemplary 3D shape 200, the exemplary 3D cuboid 211, or P1(x1) m ,y1 m ,z1 m P2(x2) m,y2 m z2 m P3(x3) m y3 m z3 m ) and P4(x4 m y4 m z4 m The parallel projection from the (virtual) 3D vector space of sampling model 210 to a two-dimensional image frame can be defined, for example, by solving the following linear equations using the four pairs of reference points mentioned above. This equations define a system of eight equations that allow the determination of the mapping or projection transformation coefficients a. j and b j The value of .

[0131]

[0132] Here, "j" is an integer from 0 to 3, "i" is an integer from 1 to 4, "f" refers to the image frame, and "m" refers to the sampling model or predetermined shape.

[0133] Therefore, based on the eight equations generated from the four pairs of corresponding points in the 3D vector space and 2D image frame space of the sampling model, eight variables or unknowns can be determined.

[0134] In this example, once the value a is determined... i and b i The sampling model 210, i.e., the 3D shape 200, i.e., the 3D cuboid 211 and its voxels 208, can be drawn / projected onto a given image frame to define at least one region of the image frame from which data / data values ​​are to be extracted, and the sampling model 210 can be assigned, for example, to the values ​​of corresponding elements of a multidimensional array (e.g., a tensor).

[0135] Specifically, each voxel of a predetermined shape, such as each voxel 208 of cuboid 211, can be associated with a projected voxel on the image frame, and each data / data value extracted from the image frame from the image frame pixel covered by the projected voxel can be assigned to the value of the corresponding multidimensional array element, i.e., the value of the corresponding tensor element.

[0136] In other words, voxels can represent elements or be associated with elements of multidimensional arrays (such as tensors), where data extracted from image frames is stored and further processed for further data analysis.

[0137] As previously mentioned, the same or similar sampling model 210, i.e. the same or similar 3D shape 200, can be used to sample and analyze other gates of the same gate type or similar geometry within the same real scene (e.g., video streams from the same sensor / camera) or from other real scenes (video streams from other sensors / cameras).

[0138] Data extracted from image frames with the different ticket gates can then be considered to face the same detection problem and can be solved with a single model / single modeling method (e.g., using the same neural network in the case of machine learning), thus bringing a general solution for many ticket gates of the same type and with the same functionality, without the need to generate (and train) additional specific solution models for additional ticket gates.

[0139] This can significantly accelerate and facilitate the resolution of detection problems in video analytics, particularly those mentioned above.

[0140] Figure 3 Further examples of possible exemplary sampling models 300 are schematically shown, which include a set 305 of various 2D shapes oriented and positioned in an exemplary (virtual) 3D vector space, such as a first two-dimensional shape 301, a second two-dimensional shape 302, a third two-dimensional shape 303, and a fourth two-dimensional shape 304.

[0141] All the exemplary two-dimensional shapes are exemplary parallelograms, wherein two-dimensional shapes 301 and 303 are exemplary rectangles.

[0142] However, the number and form of the two-dimensional shapes 301, 302, 303, and 304 are merely exemplary. Any other number and form of 2D shapes that are orientable and locatable in the exemplary (virtual) 3D vector space may also be used to define / build the exemplary sampling model 300.

[0143] The exemplary 3D vector space in which these predetermined shapes 301, 302, 303 and 304 reside is exemplarily represented by reference numeral 310.

[0144] Similar to Figure 3 In the previous example, 2D shapes 301, 302, 303 and 304 can be split or divided into one or more elements or blocks or subshapes that constitute the shape.

[0145] The element or block can also be defined as the smallest indivisible unit of shape and can also be referred to as a shape atom, primitive, or voxel.

[0146] In the example shown here, each shape 301, 302, 303, and 304 is exemplaryly divided into voxels of the same size as the given shape. For example, 2D shape 301 is divided into 2*2 voxels, i.e., 4 voxels of the same size 306; 2D shape 302 is divided into 4*4 voxels, i.e., 16 voxels of the same size 307; 2D shape 303 is divided into 2*4 voxels, i.e., 8 voxels of the same size 308; and 2D shape 304 is divided into 2*4 voxels, i.e., 8 voxels of the same size 308.

[0147] The exemplary sampling model 300 based on 2D shapes can in particular be used as a model based on one or more 3D shapes. Figure 2 An alternative to the exemplary sampling model 210.

[0148] For example, the exemplary 2D shape of the exemplary sampling model 300 can be projected onto a surface that is considered to be related to the detection of a specific problem, situation, or behavior, such as the surface of a control door or ticket gate, in order to detect tailgating or other fare evasion.

[0149] It is important to note here and generally that the surface or object in the real physical scene observed by the sensors of the video analytics system onto which the sampling model is applied / mapped / projected does not necessarily have to be limited to the actual physical surface.

[0150] For example, one can imagine that artificial surfaces, objects, or lines in a scene can be defined based on their relationship to physical surfaces, objects, or lines in the scene. For example, a plane that includes or is parallel to the plane in which an exemplary sliding door of a ticket gate moves, or a line that is aligned with or parallel to the axis of an exemplary three-legged revolving door of a ticket gate.

[0151] It is also conceivable that a sampling model or a predetermined shape of a sampling model can be projected onto the scene captured by the sensor without having a direct relationship with physical objects, objects or lines in the observed scene.

[0152] Depending on the complexity of the problem or situation to be detected or the specific geometry, the exemplary 2D shape-based sampling model 300 may be superior to the 3D shape-based sampling model 210 because, compared to 3D shapes, 2D shape processing typically generates smaller multidimensional arrays, such as tensors, for a given region of interest (region of interest) covering the data.

[0153] Therefore, in the context of the video analysis method steps and systems described in this paper, the processing of 2D shapes can be implemented faster and with fewer computational resources than the processing of 3D shapes.

[0154] Figure 4Examples of some aspects of the relationship 400 between the different spaces described above and here are illustrated.

[0155] For example, an exemplary real scene 410 in a real 3D space 401 may include real objects such as a street 405, a house 408 with a garage 407 and a tree 409 and a driveway 406 leading to the garage 407.

[0156] A sensor (e.g., a camera) can capture the exemplary real scene 410 from at least one image frame 412 in at least one series of image frames. The exemplary image frame 412 is a two-dimensional image frame representing a two-dimensional projection space (also referred to as a projected 3D space) of projected reality, such as a projection of the real scene 410 or a projection of at least a portion of the real scene 410. The image frame 412 is an exemplary digital format comprising a plurality of image pixels (not shown).

[0157] Reference numeral 413 indicates an exemplary sampling model in the form of an exemplary 3D shape 414 (e.g., a 3D cube 415) defined in an exemplary (virtual) 3D vector space 403.

[0158] The exemplary 3D shape 414, namely the exemplary 3D cuboid 415, is exemplary subdivided, sliced, or divided into multiple voxels 416 along the axis of the 3D vector space 403, for example, into 4*2*4=32 voxels 416.

[0159] Sampling model 413 is exemplarily applied to / mapped to / projected onto image frame 412. This projection 417 can be implemented, for example, according to any of the steps described above, specifically, for example, by identifying reference points in the image frame to be matched with reference points of sampling model 413 and solving a system of linear equations to determine the corresponding transformation coefficients for establishing the transformation between sampling model 413 and image frame 412.

[0160] In the exemplary case shown, the same sampling model 413 is applied twice, i.e., to two different parts of image frame 413. In this case, two independent systems of linear equations are computed to determine the corresponding transformation coefficients for establishing a transformation between sampling model 413 and the two different parts of image frame 412.

[0161] Therefore, two instances or implementations 418, 419 of the sampling model 413 applied to image frame 412 are created and two exemplary regions 421, 422 or regions of interest from which data are extracted are defined for analysis to detect specific problems or situations.

[0162] In the current situation shown, sampling model 413 or two instances or implementations of sampling model 413, 418, 419, can be used to sample and extract data from two exemplary corridors or spatial volumes or road segments 426, 427 along street 405.

[0163] The data to be extracted can be extracted, for example, by projected voxels or projected voxel regions 420, and can be extracted 425 into one or more multidimensional arrays 424, 425, such as tensors, in an abstract or digital data space 404 and having a format suitable for digital or computational processing by a processor (e.g., a graphics processing unit (GPU) or a central processing unit (CPU)).

[0164] In other words, the data of the image frame pixels covered by the projected voxels 420 of the sampled model 403 can be extracted 424 and stored in multidimensional arrays 424, 425, for example, tensors.

[0165] The extracted data can then be analyzed, for example, by applying the aforementioned machine learning techniques to detect specific problems, situations, or behaviors.

[0166] For example, for each instance 418, 419 of the sampling model 403, the analysis of the extracted data can be performed separately for different regions or blocks 421, 422 of the image frame 412 covered by the sampling model 403, or the analysis of the extracted data can be performed jointly for all regions or blocks 421, 422 of the image frame 412 covered by the sampling model 403.

[0167] The extracted data can then be analyzed, for example, to detect the presence or passage of people and / or vehicles in the analysis areas 421, 422 of image frame 412, the direction and / or speed of such passage, whether they are alone or in groups (e.g., to determine traffic flow, deceleration, congestion and blockage), to detect objects left behind, to detect panic situations, chaos or disturbances (people or vehicles moving at abnormal speeds, in abnormal directions or in abnormal numbers) or fighting, to detect loitering or oversized objects, or to monitor speed, or to estimate and / or determine occupancy levels, or other situations.

[0168] As noted in the general section above, the problems, situations, or behaviors that can be detected based on the extracted data can be diverse and are not limited to the examples given here.

[0169] It should also be noted that, taking into account the data extracted according to the steps and techniques described herein, those skilled in the art are fully capable of defining appropriate testing criteria for a selected specific problem, situation, or behavior.

[0170] Examples of problems, situations, or behaviors that can be detected based on the extracted data are only used to illustrate how the steps and video analysis techniques described herein for sampling and analyzing data from at least one image frame can be used to construct a more accurate representation of a physically three-dimensional object compared to state-of-the-art techniques that do not consider the three-dimensionality of physical reality and are limited to representations of reality based on the smooth extraction of data from the image.

[0171] For completeness, it should also be noted that, for example, a single sampling model containing a set of two 3D shapes (e.g., a set of two 3D cuboids) can be defined, and then the sampling model can be applied to the image frame only once, i.e., only a single system of linear equations can be established to determine the corresponding transformation coefficients for establishing the transformation between sampling model 413 and image frame 412. It is also conceivable to store the extracted data in a single multidimensional array or tensor.

[0172] Figure 5a An exemplary and schematic depiction of a real-world scene 508 of a ticket gate system 532 (e.g., a subway station ticket gate system) captured in exemplary image frame 529 from a series of image frames by a sensor such as a camera, wherein the ticket gate 524 is monitored, for example, by a video analytics system to detect fare evaders or fraudulent entry at the ticket gate.

[0173] What is superimposed on image frame 531 is an exemplary sampling model 510 or an exemplary instance, implementation, application, or projection of a sampling model applied to image frame 531.

[0174] In this example, the sampling model 510 is based on / defined by an exemplary 3D shape, which takes the form of an exemplary 3D cube 505 defined in an exemplary (virtual) 3D vector space 521.

[0175] As previously described in general and / or specific terms, exemplary reference points P4, 509, P1, 510, P2, 511, P3, and 512 of the sampling model 510 have been exemplary associated with or matched to the geometry of the exemplary ticket gate 524 and are projected onto the image frame 531, for example, by establishing and solving a system of linear equations between the reference points of the sampling model and the reference points in the image frame to determine the corresponding mapping or projection transformation coefficients. For better readability, the exemplary reference points in the image frame 531 are not explicitly shown, but can be assumed, for example, to be located at the positions marked by the exemplary reference points P4, 509, P1, 510, P2, 511, P3, and 512 of the sampling model 510.

[0176] The reference numeral 504 in the attached diagram represents an exemplary possible voxel or projected voxel 504 of the sampling model 510. In order to... Figure 3For easier reading, other voxels or other projected voxels of the sampled model 510 are not explicitly drawn, but can be assumed to exist in the sampled model 510, i.e., in the 3D cuboid 505.

[0177] Similarly, the image pixels of image frame 531 are not explicitly drawn or labeled, but it can be assumed that image frame 531 has a digital format that includes multiple pixels.

[0178] The sampling model 510, applied to / mapped to / projected onto image frame 531, then exemplarily defines region 529 of image frame 531 from which data is extracted.

[0179] As also described in general and / or by example above, data can then be extracted from image pixels covered by the sampled model 510 (e.g., covered by the projected voxel 504).

[0180] The extracted data can be saved into a multidimensional array (such as a tensor).

[0181] The extracted data can then be used to perform analysis to detect specific problems, situations, or behaviors.

[0182] For example, analytics can be implemented to detect fraudulent entry at ticket gate 524, such as fare evaders who tailgate.

[0183] Figure 5b An alternative real-world scenario 527 is illustrated, exemplarily and schematically, of a ticket gate system 533 with an exemplary number of ticket gates (e.g., a ticket gate system at a subway station).

[0184] The exemplary ticket gate system 533 can be used with Figure 5a The exemplary ticket gate system 532 is the same as or similar to that of the other system, although it is controlled by sensors (e.g., cameras) to capture images. Figure 5a The exemplary ticket gate system 533 is captured in image frame 530 from a different viewpoint than the ticket gate system itself.

[0185] like Figure 5a As shown, exemplary problems, situations, or behaviors that can be detected by a video analytics system that analyzes image frames 530 captured by sensors (e.g., cameras) as part of a series of image frames such as a video stream can be used to detect fare evaders or fraudulent entry at ticket gate system 533, particularly at ticket gates 525 and 526.

[0186] The reference numerals 502 and 503 in the accompanying drawings represent exemplary sampling models, instances, implementations, applications, or projections of the sampling model applied to image frame 533.

[0187] Exemplary sampling models 502 and 503 exemplarily include exemplary 3D cuboids 506 and 507 as exemplary 3D shapes, which are defined in exemplary (virtual) 3D vector spaces 522 and 523 and are exemplarily displayed as superimposed on image frame 530.

[0188] As previously described in general and / or specific terms, the exemplary reference point P1 of sampling models 502 and 503 1 513, P2 1 514, P3 1 515, P4 1 516, P1 2 517, P2 2 518, P3 2 519, P4 2 520 has been exemplary correlated or matched with the geometry of exemplary ticket gates 525 and 526, and projected onto image frame 530, for example, by establishing and solving one or more corresponding linear equations between the reference points of the sampling models and reference points in the image frame to determine the corresponding mapping or projection transformation coefficients. For better readability, the exemplary reference points in image frame 530 are not explicitly shown, but can be assumed, for example, to be located at the exemplary reference point P1 of sampling models 502 and 503. 1 513, P2 1 514, P3 1 515, P4 1 516, P1 2 517, P2 2 518, P3 2 519, P4 2 At the location marked 520.

[0189] for Figure 5b Better readability Figure 5b Elements, blocks, or voxels that may constitute the exemplary predetermined 3D shapes 506, 507 are not shown. However, as described above, it can be assumed that the 3D shapes or 3D cubes 506, 507 can be divided, subdivided, or sliced ​​into multiple elements, blocks, or voxels.

[0190] The image pixels of image frame 530 are not explicitly drawn or labeled, but it can be assumed again that image frame 530 has a digital format that includes multiple pixels.

[0191] Sampling models 502, 503 are applied to / mapped to / projected onto image frame 530, and one or more regions 528 of image frame 530 are defined exemplarily to extract data from image frame 530.

[0192] As described above in general and / or by example, data can then be extracted from image pixels covered by the sampled models 502, 503 (e.g., covered by projection elements, blocks, or voxels of the sampled models).

[0193] The extracted data can be saved into a multidimensional array (such as a tensor).

[0194] The extracted data can then be used to perform analysis to detect specific problems, situations, or behaviors.

[0195] For example, analysis can be implemented to detect fraudulent entry at the two ticket gates 525 and 526, such as fare evaders who tailgate.

[0196] As mentioned above, it is also conceivable to define a single sampling model based on a set of predetermined shapes (e.g., based on two predetermined 3D shapes 502, 503), rather than treating the 3D shapes 502, 503 as separate sampling models.

[0197] Note again that the same sampling model can be used for the exemplary real-world scenes 508 and 527 depicted in image frames 531 and 530, where "same" can mean the same and / or have the same topology.

[0198] Next are five sheets of paper, including Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5a and Figure 5b The reference numerals in the figures represent the following exemplary components.

[0199] 100 Exemplary Image Frames

[0200] 101 An exemplary first coordinate axis, such as the X-axis of an exemplary image frame coordinate system.

[0201] 102 An exemplary second coordinate axis, such as the Y-axis of an exemplary image frame coordinate system.

[0202] Example pixels of 103 image frames

[0203] 104. An exemplary projection of a predetermined shape onto an image frame, an exemplary projection of an exemplary sampled model onto the image frame, and an exemplary region of interest.

[0204] Exemplary pixels of the image frame covered by the projection of the predetermined shape of 105

[0205] 106 Exemplary periphery of the projection of the predetermined shape 104 onto the image frame

[0206] 107. An exemplary region of an image frame from which data is to be extracted, an exemplary block of interest.

[0207] Exemplary 3D shapes, exemplary parallelepipeds, and exemplary cuboids in a 200 (virtual) 3D vector space

[0208] 201 An exemplary (first) reference point with exemplary reference coordinates

[0209] 202 An exemplary (second) reference point with exemplary reference coordinates

[0210] 203 Exemplary (third) reference point with exemplary reference coordinates

[0211] 204 Exemplary (fourth) reference point with exemplary reference coordinates

[0212] 205 Exemplary (first) coordinate axis, exemplary first (virtual) 3D vector space axis 206 Exemplary (second) coordinate axis, exemplary second (virtual) 3D vector space axis

[0213] 207 Exemplary (Third) Coordinate Axes, Exemplary Third (Virtual) Three-Dimensional Vector Space Axes

[0214] Exemplary elements, blocks, or voxels of exemplary 3D shapes in a 208 (virtual) 3D vector space

[0215] 209 Exemplary reference point sets, exemplary (vertices) of 3D shapes in (virtual) 3D vector space

[0216] 210 Exemplary Sampling Model

[0217] 211 Exemplary cuboid, Exemplary 3D cuboid

[0218] 212 Exemplary (Virtual) 3D Vector Space

[0219] 300 Exemplary Alternative Sampling Models

[0220] 301 Exemplary (first) two-dimensional shape, exemplary parallelogram, exemplary rectangle 302 Exemplary (second) two-dimensional shape, exemplary parallelogram

[0221] 303 Exemplary (Third) Two-dimensional shape, exemplary parallelogram, exemplary rectangle 304 Exemplary (Fourth) Two-dimensional shape, example parallelogram

[0222] 305 A set of exemplary predetermined shapes, a set of exemplary two-dimensional shapes

[0223] 306 (First) Exemplary voxel 301 of 2D shape

[0224] 307 (Second) Exemplary voxel 302 of 2D shape

[0225] 308 (Third) Exemplary voxel 303 of 2D shape

[0226] 309 (Fourth) Exemplary voxel 304 of 2D shape

[0227] 310 Exemplary (Virtual) 3D Vector Space

[0228] Exemplary schematic relationship between 400 spaces

[0229] 401 Exemplary Real 3D Space, Exemplary Real Physical 3D Space, Exemplary Real Space

[0230] 402 Exemplary 2D Projection Space / Exemplary Projection 3D Space

[0231] 403 Exemplary Virtual 3D Vector Space, Exemplary 3D Vector Space

[0232] 404 Exemplary Abstract or Numerical Multidimensional Data Space

[0233] 405 Exemplary Street

[0234] 406 Exemplary Lane

[0235] 407 Exemplary Garage

[0236] 408 Exemplary House

[0237] 409 Exemplary Tree

[0238] 410 Exemplary Real-World Scenarios

[0239] 411 Exemplary actions / steps of an exemplary real-world scene captured or recorded by, for example, a camera's sensor. 412 Exemplary image frames, such as image frames in a digital format containing multiple pixels.

[0240] 413 Exemplary Sampling Model

[0241] 414 Exemplary Predefined Shape, Exemplary 3D Shape

[0242] 415 Exemplary 3D cuboid, Exemplary cuboid

[0243] 416 Exemplary element, exemplary block, or exemplary voxel of shape 414

[0244] 417 Exemplary action / step of applying / mapping / projecting the sampling model onto / on an image frame 418 Exemplary (first) instance or implementation of the sampling model 403 applied to / mapped onto / projected onto an image frame

[0245] 419 An exemplary (second) instance or implementation of sampling model 403 applied to / mapped onto / projected onto image frames

[0246] 420 Exemplary projection element, projection block, or projection voxel; Exemplary projection area of ​​projection element, projection block, or projection voxel.

[0247] 421 Exemplary sampling model covers an exemplary (first) region of the image frame.

[0248] 422 Exemplary sampling model covers an exemplary (second) region of the image frame.

[0249] 423 Exemplary actions / steps for extracting data from image frames 424 Exemplary (first) multidimensional data structures, such as multidimensional arrays, such as tensors

[0250] 425 Exemplary (Second) Multidimensional data structures, such as multidimensional arrays, such as tensors

[0251] 426 Exemplary (first) corridor or space volume or section along street 405

[0252] 427 Exemplary (second) corridor or space volume or section along street 405

[0253] 501 Exemplary Sampling Model, Example Instance of Sampling Model

[0254] 502 Exemplary Sampling Model, Exemplary First Instance of Sampling Model

[0255] 503 Exemplary Sampling Model, Exemplary Second Instance of Sampling Model

[0256] 504 Exemplary voxel, Exemplary projection voxel

[0257] 505 Exemplary 3D Shape, Exemplary 3D Cuboid

[0258] 506 Exemplary (First) 3D Shape, Exemplary 3D Cuboid

[0259] 507 Exemplary (Second) 3D Shape, Exemplary 3D Cube

[0260] Exemplary real-world scene at Gate 508, captured exemplary image frames.

[0261] 509 Exemplary Reference Point P4

[0262] 510 Exemplary reference point P1

[0263] 511 Exemplary reference point P2

[0264] 512 Exemplary reference point P3

[0265] 513 Exemplary reference point P1 1

[0266] 514 Exemplary reference point P2 1

[0267] 515 Exemplary Reference Point P3 1

[0268] 516 Exemplary reference point P4 1

[0269] 517 Exemplary reference point P1 2

[0270] 518 Exemplary reference point P2 2

[0271] 519 Exemplary Reference Point P3 2

[0272] 520 Exemplary reference point P4 2

[0273] 521 Exemplary (Virtual) 3D Vector Space

[0274] 522 Exemplary (Virtual) 3D Vector Space

[0275] 523 Exemplary (Virtual) 3D Vector Space

[0276] 524 Exemplary ticket gates to be monitored

[0277] 525 Exemplary ticket gates to be monitored

[0278] 526 Exemplary ticket gates to be monitored

[0279] 527 Exemplary real-world scene at the ticket gate, captured exemplary image frames.

[0280] 528 is defined as an example region or block in the image frame of the sampling model.

[0281] 529 is defined as an example region or block in an image frame of a sampling model.

[0282] 530 Exemplary Image Frames

[0283] 531 Exemplary Image Frame

[0284] 532 Exemplary ticket gate system, having an exemplary plurality of ticket gates

[0285] 533 Exemplary ticket gate system, having an exemplary plurality of ticket gates

Claims

1. A computer-implemented method for sampling and analyzing data from at least one image frame (531), said at least one image frame (531) being derived from at least one series of image frames captured by at least one sensor, said method comprising: Define at least one sampling model (501), wherein the sampling model (501) is defined in a virtual 3D vector space (521) and based on one or more predetermined shapes (505) in the virtual 3D vector space (521), wherein the shape (200) in the virtual 3D vector space (212) on which the sampling model (210) is based is divided into elements or blocks (208) constituting the shape (200). The at least one sampling model (501) is applied to at least a portion of the at least one image frame (531) in the at least one series of image frames, wherein the application of the at least one sampling model defines at least one region (529) from which data is extracted from the at least one image frame (531). Data (423) is extracted from at least one region (529) of the at least one image frame (531) defined by the sampling model (501). And the extracted data is analyzed.

2. The method according to claim 1, wherein, The one or more predetermined shapes in the virtual 3D vector space (521, 212, 310) are selected from at least one of the following shapes: 3D shape (200, 505, 506, 507), 2D shape (301, 302, 303, 304), 1D shape or 0D shape.

3. The method according to claim 2, wherein, The 3D shapes (200, 505, 506, 507) are parallelepipeds and / or polyhedra and / or spheres and / or cylinders, and / or wherein, The 2D shape is a plane or curved surface and / or a parallelogram (301, 302, 303, 304), and / or the 1D shape is a line segment, and / or the 0D shape is a point.

4. The method according to any one of claims 1 to 3, wherein, Applying the at least one sampling model (501) to at least one image frame (531) in at least one series of image frames comprises associating the at least one sampling model (501) with one or more reference points in the at least one image frame (531) in at least one series of image frames.

5. The method of claim 4, wherein the association further comprises performing a mapping transformation (417) between one or more points of the at least one sampling model (413) and one or more reference points in the at least one image frame of the at least one series of image frames.

6. The method according to claim 5, wherein, The mapping transformation is a parallel projection.

7. The method according to claim 1, wherein, The shape (200) in the virtual 3D vector space (212) on which the sampling model (210) is based is divided uniformly or non-uniformly in any or all of its geometric dimensions into one or more elements or blocks (208) that constitute the shape (200).

8. The method according to claim 1, wherein, Extracting (423) data from at least a portion of at least one image frame in at least one series of image frames to which the sampling model (413) is applied includes: extracting data from image frame pixels in image frame regions (421, 422), said image frame regions (421, 422) being contained in or covered by the shape of the sampling model applied to at least a portion of said at least one image; and storing the extracted data in an array.

9. The method according to claim 8, wherein, The array is a multidimensional array (424, 425).

10. The method of claim 8, comprising extracting data from image frame pixels in image frame regions (421, 422), said image frame regions (421, 422) being contained in or covered by elements or blocks (420) of the shape of the sampling model applied to said at least a portion of said at least one image; and storing the extracted data in at least one array.

11. The method according to claim 10, wherein, The at least one array is a multidimensional array (424, 425).

12. The method according to claim 1, wherein, The same sampling model (413) is applied to different portions of the at least one image frame in the at least one series of image frames, and / or wherein, The same sampling model is applied to multiple images in the at least one series of image frames, or the same sampling model is applied to all images in the at least one series of image frames.

13. The method according to claim 1, wherein, The extraction (423) of the data from at least a portion of the at least one image frame (412) to which the sampling model is applied includes transforming the data.

14. The method according to claim 1, wherein, The at least one image frame (412) to which the sampling model is applied, or wherein at least a portion of the at least one image frame to which the sampling model is applied, is preprocessed prior to the extraction of data.

15. The method according to claim 1, wherein, The data extracted for analysis includes: The extracted data is analyzed to detect desired patterns, wherein the patterns can include predetermined states and / or movements and / or behaviors and / or actions of objects and / or subjects within a real 3D scene (410, 508, 527), represented in at least a portion of at least one image frame in at least one series of image frames captured by the at least one sensor; and a notification or alarm is provided when the pattern is detected, and / or The extracted data is used as input to train the machine learning system for detecting desired patterns, wherein the patterns are capable of including predetermined states and / or movements and / or behaviors and / or actions of objects and / or subjects within a real 3D scene (410, 508, 527), represented in at least a portion of at least one image frame in at least one series of image frames captured by the at least one sensor. and / or The extracted data is used as input to a trained machine learning system to detect desired patterns, wherein the patterns are capable of including predetermined states and / or movements and / or behaviors and / or actions of objects and / or subjects within a real 3D scene (410, 508, 527), which are represented in at least a portion of at least one image frame in at least one series of image frames captured by at least one sensor; and a notification or alert is provided when the pattern is detected.

16. The method according to claim 15, wherein, The predetermined situation and / or movement and / or behavior and / or action of the object and / or subject in the real 3D scene constitutes fraudulent entry at the control gate.

17. The method according to claim 1, wherein, Applying the at least one sampling model to at least a portion of the at least one image frame in the at least one series of image frames takes into account the movement of the at least one sensor during the capture of image frames from the at least one series of image frames, and / or the method includes: Data from multiple series of image frames acquired by multiple sensors with different viewpoints for capturing image frames are sampled and analyzed, and the different viewpoints of the multiple sensors are taken into account when the at least one sampling model is applied to the image frames acquired by the multiple sensors.

18. One or more computer-readable storage media storing instructions therein that, when executed by one or more processors, instruct the one or more processors to perform the method according to any one of claims 1 to 17.

19. A video analysis system, comprising: At least one sensor, the at least one sensor being configured to capture image frames, and At least one computing system, the at least one computing system comprising one or more processors, the one or more processors being configured to implement the method according to any one of claims 1 to 17.