A method of visual recognition and tracking of objects using virtual sensor
Patent Information
- Application Number
- PCT/CZ2025/000009
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-24
- Filing Date
- 2025-04-24
- Publication Date
- 2025-12-18
AI Technical Summary
Existing object detection and tracking technologies require high computational power, extensive expert intervention, and are prone to errors due to the need for complex neural network training, making them inefficient and costly for real-time applications.
A method utilizing virtual sensors that process video streams from cameras, employing pre-trained convolutional neural networks (CNNs) and human annotators to rapidly deploy and adjust object detection systems, reducing computational workload and enhancing accuracy through a human-in-the-loop scheme.
Enables rapid deployment and high accuracy object detection and tracking with minimal computational resources, eliminating the need for data scientists and reducing setup time to seconds or minutes, while achieving 98-99.9% detection accuracy.
Smart Images

Figure CZ2025000009_18122025_PF_FP_ABST
Abstract
Description
Title: A method of visual recognition and tracking of objects using virtual sensorTECHNICAL FIELD
[0001] The disclosed invention relates to a method of visual recognition and tracking of objects and events in 3D space with high accuracy, reduced computational workload and fast setup, wherein the method employs use of virtual sensors which are adapted to ensure high reliability and reduced costs by maximizing their potential of accurate recognition of the object for the given visual input data.
[0002] The main role of the virtual sensor is to provide cost effective and reliable data and substitute usage of human and / or hardware sensors to detect and recpgnize objects or events of interest which are relevant for the subject, such as for example an industrial company and various businesses.BACKGROUND ART
[0003] Collection of relevant data has become one of the most important aspects for running a successful business. The data are usually used for optimization of products, processes, surveillance, security, etc. While the HW sensors might be cheap to purchase, their planning, installation, integration, and validation, processing, and interpretation of the data obtained thereof is quite burdensome and costly and requires plenty of resources and time to set up and run. HW sensors also cannot read the data unless they are installed. Furthermore, its customization is even more problematic. In case of use of conventional computer vision sensors, the integration is even more challenging and usually takes days and weeks to implement in order to receive first acceptable results.
[0004] To virtually represent an object reliably, one of the technically relevant and conceptually simplest fully computer vision-based approaches would be 3D modeling the space using photogrammetry. However, that would require camera images from a number of views, filling the observed area with a high number of cameras and costing a high amount of computation.[0Q05] Existing technical solutions designed to detect and track objects are too complex to obtain the desired result efficiently and at low computational costs. Typically, a data scientist is required to train the mathematical (deep CNN) model of the neural network for recognition of the object in the image, for example through YOLO, which is a real-time object detection algorithm that divides an image into a grid system, and each grid detects objects within itself.[0006} However, these models alone have substantially lower accuracy rate due to their general usage and are intended to recognize most common objects which are mostly pre-defined. These limitations make those models prone to error whilst, due to their robustness and volume of data they rely on, such models are substantially more difficult to implement and use and they must be trained additionally by a data scientist for intended purpose. This requires a considerable amount of time and resources to train the model appropriately according to the needs.
[0007] It is one object of this invention to largely mitigate the need for high computing power and expert data scientist work while achieving best possible results and accuracy of the recognition and digital time trace capturing the movement of tracked objects in a specified space.SUMMARY OF THE INVENTION
[0008] The above disadvantages are removed by the implementation of sequence of method steps described further and by use of technical means adapted to carry out this method in a required manner. In other words, the present method is working only with the amount of information that is not complete to reproduce a full 3D model photogrammetrically, but is sufficient for a person to identify individual objects, thereby reducing the abundant data to be processed. Human intervention is needed to some extent due to the variability in the production that cannot be fully planned ab initio, which requires additional corrections when accuracy level is unsatisfying for a given scenario. The method comprises use of computer vision to turn one or more video streams obtained from a camera into a verifiable signal with an interpretation similar to an industrial loT sensor.
[0009] The method according to this invention allows for very rapid deployment due to the presence of physical elements adapted for this purpose and the involvement of a human annotator in error removal and fast system training increasing its accuracy on tne fly.
[0010] It is required that the physical setup of the elements used to carry out this method meets certain criteria and prerequisites, whilst the present invention is not limited to physical elements available at present, but may suitably take advantage of the availability of technical alternatives thereof in the future.
[0011] This method is particularly suitable for detection and tracking of objects, which have predefined shape or appearance which does not change substantially over time and which is typical for controlled environments, and especially when there are multiple objects of the same shape or appearance. In this framework, the shape and form can be anything, but ideally in a specific sensor, observed images can be clustered into well-defined shapes. However, the virtual sensor can be trained to recognize any object of interest. For the purpose of this invention, the terms detect, recognize and identify with regards to the action towards the object or event are considered equivalent.
[0012] This method can be used to detect physically interpretable events through a discretization of the time series signal from a virtual sensor, i.e. defining times of for example opening of a door, loading or unloading of an object, ending of a machine cycle, etc.
[0013] The method brings large benefits to the user due to the fact that it can be rapidly implemented on site and does not require high computational power for detection and tracking. The training of the virtual sensor itself can be done in a short period of time within the range of seconds to minutes. Most importantly, no data scientist is required to train the model of the neural network, which is used in the state-of-the-art solutions, such as the already mentioned YOLO algorithm.
[0014] On the other hand, this method relies on the use of several models, results of which are compared to each other, and in case of uncertainty are escalated to a more complex model or even to a human annotator. The training of this system of models is automated, using a more complex model to score the data and in case of uncertainty escalates a few images to a human annotator. As a result, the model requires for training or retraining max seconds of annotator’s time instead of hours of a data scientist, and the user enjoys higher overall accuracy. The system is designed for stationary cameras, but in case of any unwanted movement, the camera can be recalibrated and the sensor automatically adjusted, for non-stationary cameras, their positions can be updated using a SLAM or visual odometry method.
[0015] The method employs the use of optical devices such as cameras, which provide an image data source for a virtual sensor which is assigned to the specific section of the image obtained by the camera. The term camera refers to any device with an optical unit that pan produce a video image. It may be a more complex device, including a mobile phone which comprises a camera, but in most cases the camera is a utility camera per definition.
[0016] In a minimalistic scenario, where the object does not have to be tracked but only recognized from the image, at least one camera would be required for obtaining the image. However, the present invention is more powerful due to the fact that it employs a tracking function in 3D space which strengthens the ability to recognize the object or event with the utmost accuracy. Due to such feature, use of multiple cameras is advantageous. The cameras shall be preferably positioned as stationary in order to collect data from the same place and angle in order to receive correct signal from the area.
[0017] The image from each camera is divided into sections of interest, wherein each section corresponds to the working area of the virtual sensor. Thus, each image can be processed with the plurality of virtual sensors, whose total number is driven by the quality of the image and complexity of the object to be recognized and possibly also tracked. For instance, the large objects with simple shape can be detected with lower resolution and quality image, whilst smaller and detailed objects require much better resolution and image quality to recognize the object with sufficient accuracy. The accuracy itself is one of the important aspects and key goals of this method. Theaccuracy is the result of correctly defining the probability of occurrence of the object in the section of image controlled by the corresponding virtual sensor. The schematic demonstration of dividing the entire image from camera to virtual sensors is shown on Fig. 1. The single object is shown on the right in dashed line, whereas a stack of objects on the left is a pre-trained virtual sensor capable of recognizing multiple units thereof.
[0018] More importantly, in addition to video consisting of sequence of images can be processed by the virtual sensors in real time, it is also possible to verify whether the objects were detected and tracked accurately ex post after re-training the virtual sensor. Thus, even the old video recordings can be repeatedly processed to obtain the most accurate results, which is impossible to achieve with HW sensors, which can collect data only after they are installed. This advantage results in no important data being lost, and the tracking can be improved by simple amendments to the virtual sensor(s). This leaves room for efficient correction of the tracking based on the original video recordings used to capture the image of the area representing the 3D space.
[0019] Further advantage is linked with the human-in-the-loop scheme, whereby a human annotator is involved in the process of training and re-training of the virtual sensor. The virtual sensor can be advantageously pre-trained and / or its accuracy is strengthened by the human annotator who decides whether the object has been rightly captured by the virtual sensor or not. Therefore, the virtual sensor does not need large learning models, but rather relies on the specific data from the given area where the method is used. This saves substantially both the computational workload and data flow between the relevant physical devices, such as servers and network components, etc.
[0020] The virtual sensor scheme looks at actors in terms of their reliability to identify various complex phenomena displayed in the video, the cost of processing a single image, and their availability.
[0021] In the basic scenario, the human intervention is based on following actors, their roles and level of skills needed to complete tasks involved in implementation of this method. As a prerequisite, it is assumed that each video signal from the camera is processed by a computer. This is usually the device on site, it has the lowest computational capacity, the lowest ability to identify problems. The edge computingmay be processed by a single board computer, which runs in real time a less complex tasks employing CNN models and / or conventional computer vision algorithms. Further processing can be made on various powerful computers or human annotators, hereinafter both defined also as “executors”, that can be used hierarchically: o The human annotator has various level of experience, he shall be at least a person who can apply "common sense" to define the similarity of a picture with the object. Annotators can value their experience by learning from experts. Annotators are available 24 / 7. o The expert may not have experience with annotations, but for a given situation he is a "user" of the signal. He does not need to be skilled in programming, he only understands what a given signal value means. His capacity is very limited and he is only available at certain times of the day / week. The expert can be a line operator, production management, consultant, etc. o The developer is the technical supervisor of the whole process who identifies toe type of problem and solves exceptions.
[0022] The minimum human-in-the-loop implementation involves one computer, one person selected from an expert or annotator. Optimally, all roles described above are used.
[0023] The system according to this method requires communication between the actors based on their reie and a method to estimate the reliability of their work. The general principle is that it is advantageous to minimize the need for less available actors, such as developers and experts, minimize the time burden on more costly actors, minimize the latency of the system, while maximizing the availability and reliability of the system.
[0024] In the initial stage, the method comprises first installation of the virtual sensor according to the following sequence:« The first image obtained from a video footage made by the camera is analyzed. The person installing the virtual sensor (usually the expert) determines the positions in the image to be monitored;« The images of the monitored area (representing the 3D space in which objects and / or events shah be tracked) are extracted from the video recording obtained from the camera. These images can be automatically sorted on the basis of their mutual similarity, for example by using clustering methods. The state images, represented by a set of annotated images used for training of the CNN model, are transferred to a person (can be an annotator) who efficiently assigns them to each sensor or just confirms or modifies the automatic sorting. Each virtual sensor is defined by a camera, region of interest (Rol) and the CNN used. This person can be the original person installing the virtual sensor or any other person, e.g. a so-called professional annotator; In other words, the person (such as annotator) assigns the Rol cut for a sample of past images from the camera to separate states. For example, given the object pertains to an industrial complex, the states can be e.g. open press, closed press, night (darkness), changeover (no mould), or e.g. open press with die 1, closed press with die 1 , open press with die 2, closed press with die 2, or| e.g. empty truck slot, truck closed, truck open, loading, crane in view, etc; o The sorted images are then used to retrain the pre-trained CNN model. The training can be fully automatic; The CNN is pre-trained by setting a region of interest (ROI), which is a part of image, usually a rectangle as shown on Fig. 1 , which contains the best possible interpretation of the object and / or its various states. Adding additional states (such as different positions, exceptional or common) for expressing the presence of the object in the image improve accuracy of the CNN model for a given object. o The obtained retrained CNN model is then passed to an executor, usually a computer, which processes each video image and transforms the image into an array of probabilities of individual states, the so-called virtual sensor value;« The values of the virtual sensors, or their changes, can be transformed into determination of a specific phenomenon, e.g. opening a door, ending a machine cycle, triggering a light signal, etc.
[0025] The implementation of the method can be also simplified by introduction of following steps;o The executor (computer) applies the pretrained CNN modei and generates a signal o The reliability of the signal is evaluated and in case of a dispute it is moved to a more accurate executor, such as human annotator, whereby using more robust CNN or other automatic models running on cloud are skipped, thus saving substantial computational workforce. o A higher level executor (a deep net model or an annotator) checks some of the selected situations and thus corrects the data. In addition, it can also perform a so- called patrol control / monitoring of some data that have not been evaluated as questionable; o In case the annotator does not know how to proceed, he will pass the problem on to an expert or the developer.
[0025] The continuous improvement of accuracy of the virtual sensor and the method as such is made as follows: o Each escalation (a dispute) from a lower level reliability executor to a higher level reliability executor is collected into an "annotations database”. o Escalations are once again checked by higher level executors with higher confidence thresholds to avoid potential errors and classify the disputed image as either o one of the states o a new state that needs to be defined o an escalation - a significant change in the image (e.g. the observed machine is no longer in the image, the camera was moved etc.) o an escalation to a developer or an expert o In case new states are defined or images are clearly assigned to one of existing states, the disputed image is added to the training set and the model is retrained.
[0027] The method advantageously comprises use of a so-called edge-cloud hybrid, wherein large datasets are processed locally on the edge device comprising computational unit such as processor(s), which is in proximity of the optical device, typically not exceeding tens of meters, and only relevant and already processed data output data are sent on cloud servers. Running the entire data processing on cloudcomputers would be technically possible, but requires transmission of substantial amounts of data. Thus, processing the computation using pre-trained models for recognition of the object on local edge devices eliminates the need for centralized processing and avoids the need for implementing robust infrastructure.
[0028] The output from the virtual sensor is the probability of occurrence of the object within the selected section (part) of the image, for which the virtual sensor is responsible. In borderline situations, where the probability is not sufficient, the annotator is invited to check the source image data manually. If done correctly, such manual re-training rises exponentially the accuracy of the virtual sensor. Ideally, pretraining of the model on which the virtual sensor is based is made during initial setup. Upon initial pre-training of the model, the accuracy on probability of occurrence of the object in the virtual sensor shall be at least 95 %, preferably at least 98 % and optimally over 99 %. After deploying the virtual sensor into production, accuracy values below 80 % trigger the escalation and the annotator is ihvited to make manual verification and / or correction.
[0029] To sum up, the virtual sensor is relying on a computer program that takes a simplified representation of 3D space on input and returns the occurrence probability of a specific object in a certain position and certain time on the output, same as an equivalent HW sensor would produce the corresponding data in physical space.
[0030] The sensor is repeatedly applied to processed images and provides a time series of probabilities of a presence or motion of an object at a specific position, which can be used for consolidation i.e. strengthening signal across several virtual sensors and detecting the presence or motion at high accuracy.
[0031] By using the present method, a digital twin of the observed area can be obtained. To obtain best results, multiple cameras shall be deployed in the area to avoid blind areas where those are expected. Such a digital twin involves any types of objects for which the virtual sensors are trained to recognize them. This eliminates dealing with abundant data and objects which are of no importance for the tasks to be accomplished, which cannot be achieved by selecting robust algorithms such as the abovementioned YOLO.
[0032] In one embodiment of this invention, a small neural network with five convolutional layers having input size 32 x 32 has been used. The training comprises backtracking (use of backpropagation algorithm) of up to 40 iterations (epoch).
[0033] In one embodiment of the invention, the components include multiple cameras and computer processors, which perform one er more of the operations described above. The method also requires the involvement of a human annotator.
[0034] Advantageously, the method comprises use of at least one computer processor connected to the camera in its proximity, preferably not extending the distance of 1 km, more preferably being placed within the distance of 1m from the camera, wherein at least one processor is close to a remote human operating the data correction tools. The term ’’remote” is defined as maximum distance being equal to speed of light multiplied by the required latency for the controlled output data. For post processing of the old video recording stored on the cloud server or on any computer device, the computer processor does not need to be in the proximity of the camera, since the distance of the camera is irrelevant in such a case.
[0035] Further prerequisite to be met is latency for human-in-the-loop Al, wherein the information regarding object’s exact position is required with certainty only after a time period sufficient for the human operator to check the data quality (at least a few seconds), and a mistake in the data is admissible for the ‘real-time’ data. The detection accuracy of the virtual sensor is set to reach 98 - 99,9 % range. However, over time, the accuracy may be affected by small alterations of the object, resulting in decreasing the range of accuracy of the signal by approximately 3 - 5 %, depending on the alteration gravity. This results in the need for re-training of the virtual sensor to increase the accuracy to required level.
[0036] Advantageously, the physical setup of the tracking system is adopted to cover all positions of the tracked object so that the object is visible during its movement by at least one camera to ensure conserving its identification data (ID) in the system. It is not necessary to secure the visibility of the object by the same camera, but it is sufficient to obtain information on the movement from different viewpoints provided by different cameras.[Q037] It is advantageous that the camera frame rate, its resolution and contrast is high enough for recognition and identification the object including its movement within the observed area. The recognition and identification of the object and its movement must be identified reliably by a human after a suitable transformation of the image obtained by the camera (e.g. contrast filter on pixels, noise reduction using a sequence of images etc.). In case that these conditions are nut met, the probability of error in tracking rises substantially.
[0038] The requirements on the quality of the image shall take into account the size of the object being tracked, its maximum speed and the resolution of the camera. Based on these facts, it is possible to calculate the minimum required frame rate of image recording.[003S] One aspect of the present invention related to reducing the overall computational workload lies in that the date from cameras are processed at the local processor using the virtual sensor approach, which has been briefly described above, with the details being introduced in the following part of the description and examples of embodiment.
[0040] Advantageously, several virtual sensors are employed and compared to strengthen the signal and detect outlier data, which relates to a potential error. The reliable data reached by consensus of sensors may be compared against the data from other devices, in particular from other cameras, or can also be compared to the data from relevant HW sensors, such as odometers, lidars etc. The detection of an object is provided by thresholding the SW signal and cross-checking against the physical possibility of the detected outcome (e.g. the object cannot disappear from a 10 m area within <1 s). The disagreements between sensors or sensor information and physical possibilities are sent to the remote human operator for correction.
[0041] Such approach demonstrates substantial advantages, since it reduces the need for data flows, as most data are processed at local processors, further it improves latency due to fast local processing, improves availability of the system in terms of losing the connection between the local and remote processor (if the connection is lost,the system still continues working), while allowing high reliability of the detection through ex-post checks and using external processors if needed.
[0042] Further to the aforementioned advantage of reduced computational workload, the advantage of using virtual sensors lies in the possibility to easily interpret their position in physical space. Also, they can be easily verified (using the underlying image), automatically set up, and can be set up ex post on past data.
[0043] Preferably, the virtual sensors are initially advantageously set-up using a pretrained convolutional neural network (CNN), which is set on a specific position in the simplified 3D model The pre-training process involves training a CNN on a large dataset for a related task before fine-tuning it on the specific task of interest. Use of CNN has the advantage in that features learned during pre-training on a large dataset are often transferable to ether tasks. Also, it is data efficient, since pre-training helps in overcoming data limitations for the target task by leveraging knowledge gained from a larger dataset. Moreover, pre-trained model often converges faster during fine-tuning compared to training from scratch.
[0044] Within the pre-training process, values of virtual sensors obtained over a period of time are shown to a human for determining whether the specific object is or is not in the area.
[0045] The annotated data is then used to retrain the CNN. This approach is done at the initial setup, but also continually to improve the accuracy of the system by:- correcting any potential mistakes (disagreement between sensors) and- continually improving the quality of the CNN and adjusting to any changes in the production.
[0046] Such approach is particularly advantageous by eliminating the need for SW experts or data scientists, whereby the majority of work is done by the pre-trained virtual sensor. Thus, only employees with minimum training are needed, the reliability of the system is very high right from the beginning, and the virtual sensor dynamically adjusts to changes in the production.
[0047] Data processing workflow as a part of the method may involve any of the following steps (combination of steps or a permutation of this order):« preferably, arranging of images using 3D knowledge of the area to obtain video footage (this step serves to give human annotators better overview of the data to be controlled, in particular if not arranged by machine upon defining the 3D position of each virtual sensor); o applying CNNs in order to turn parts of the individual image into scores representing likelihood of occurrence of an object in a corresponding 3D area, thus processing the video footage into a time series data using CNN or other feature assessment techniques, wherein the object of interest may also include material, material changes, and vehicles; o preferably, combining time series data from multiple sources to achieve consensus or escalate the error for human correction (this step significantly improves the accuracy, is an important element of the human in the loop); o in case of tracking of transportation means used to carry an object to be recognized and tracked, following steps are performed; o for any detected object, its position in time t+dt is updated from its position in time t assuming a certain maximum displacement for items associated with transportation means and no motion for items not associated with transportation means; o association of tracked object and transportation means; o connecting of pick-ups and pick-downs by transportation means (association of object and transportation mean) into full transport trajectories; o updating of the database of instantaneous objects and transportation means positions by simulating the transport trajectories.
[0048] Ths above-described steps are advantageous in that, at any critical step of the workflow, they provide easy human control or even fully manual processing of the step by remote human operators (annotators) through web data correction tools.
[0049] It is advantageous to use 3D modelling for spatiai consolidation of virtual sensor data, wherein physical interpretation of any sensor can be derived from 3D position of the cameras and the spatial arrangement of objects inside the monitored area. Thepositions, mutual orientations and intrinsic, properties of cameras must be set to setup the methods of combining sensor data into object positions.
[0050] it is also advantageous to combine deep learning tools and / or Artificial Intelligence (Al) to pre-train the CNN model for recognition of the object.
[0051] Further, it is advantageous to implement cross-checking across unreliable sources by combining data sources to execute one or more of the steps. This provides interface to a human correcting the database. The result of the above step is higher strength of virtual sensor signal and highlighting potential errors, which makes it easier to find and remove the error in the data on object tracking.
[0052] By using the present method, the identity of each item in the area is inferred by association of objects and transporting means (vehicles) with events and tracking the position of objects and transport vehicles in time. This provides the advantage to track objects which are visually indistinguishable, and thus eliminating the need to use labels; or tags.
[0053] Furthermore, by implementing the above steps, the projection of sensors for easy data interpretation and improved accuracy is dimensionally reduced through 1 D tracking methods, which is more straightforward to interpret by human annotator and eliminates risk of error.
[0054] As suggested previously, the implementation of the human-in-the-loop scheme by using manual corrections greatly improves the accuracy of the object tracking path and serves as a basis for further improvement of the system by automatic teaching from previous errors. Once the human annotator makes the correction, the system will learn how to deal with similar scenario in the future.
[0055] Further improvement of the method is achieved by combining tracking algorithms for improved automatic accuracy, which can be selected from: o expectation-maximization (EM) algorithm, which is defined as an Iterative method to find (local) maximum likelihood or maximum a posteriori (MAP) estimates of parameters in statistical models, where the model depends on unobserved latent variables;o Gaussian curve fitting , which approximates a curve which is specified point for point by an x-channel and a y-channel, with a Gaussian curve according to the known formula. The process of Gaussian curve fitting involves finding the values of p and G that maximize the likelihood of the observed data given the Gaussian distribution. This is often done using methods such as maximum likelihood estimation (MLE) or least squares fitting; o NEB (Nudged Elastic Band) technique, which works by dividing the transition pathway into a series of images, representing different configurations along the path. These images are then connected by springs, forming an elastic band. The band is "nudged" iteratively until it converges to the minimum energy pathway, providing insights into the energetics and mechanisms of the studied process.
[0056] Present method has been found particularly suitable for various industrial and non-industrial applications, such as tracking of trains, cars on roads, highways or parking lots, tracking of products in industrial complexes and warehouses, tracking of cranes, conveyors, etc.BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The invention will be elucidated based on an exemplary embodiment shown in the attached drawings, in which:Fig. 1 shows an example of placement of several virtual detectors into the image;Fig, 2 shows a scheme of flow diagram employing steps involved in the present method.DETAILED DESCRIPTION OF THE INVENTION
[0058] The method according to this invention is demonstrated on the following example. This example is introduced merely for the purpose of elucidating the results of real-world implementation of the method and shall not be construed in any ways as limiting the extent of the present method,
[0059] A stamp shop composed of 6 different, machines with independent operating system with typical cycle times 1 - 6 s is digitized using virtual sensors within hours:o An edge device is pieced within suitable distance from each machine, for example 2 m, and a stationary camera attached to it is oriented to observe the die moving up or down. o An annotator sees images from more than 5 full machine cycles and identifies a region of interest (Rol) on the image that changes the most significantly during the cycle and classifies more than 10 images as half die open and half die closed. The images are augmented and used to train CNN. o CNN is deployed at the Rol at fps more than 10 times higher than the cycle time to clearly identify each cycle for cycle time measurement: o Escalation to a more expensive operator (for annotation). o Escalation to Another Rol is similarly used on the same camera and the signal is compared. If a discrepancy is found between signals from the two virtual sensors, the image is sent for annotation. o If a signal from the virtual sensor corresponding to the sensor accuracy gets below $0 % reliability threshold, the image is sent for annotation. o- If the cycle time becomes for 10 consecutive cycles irregular beyond 3 times average steps to reproduce, annotation is made. o If the position of a new image in the information space (after applying the neural network before the last layer, called softmax) is outside any hypersphere defining a certain state (i.e. the distance is greater than 3 times the radius of the hypersphere; the radius is defined as the square root of the average of the squares of the distances of the image positions assigned to state x from the centre of the cluster). o if the cycle time becomes 1.5 times (or more) shorter than the shortest expected or reported by the customer, an annotation is made. o Annotated images are added to the training set and CNN retrained. The new CNN is used for the past images to test that it does not lower quality, and on the new images to realize the benefit from annotation. o If the annotation is not clear, the annotator sends it through the ticketing system to the developer, who assesses if a new Rol is needed, or more states, or the issue needs escalation to the operations expert.INDUSTRIAL UTILIZATION
[0060] The disclosed solution has good industrial applicability especially for tracking larger objects of known shape or appearance, for example, in large industrial halls, warehouses without fixed shelf systems, container docks, train docks, car parks etc. It is particularly suitable for tracking coils, metal pieces and larger construction elements. The method provides reliable data source for obtaining digital twin of objects to be tracked, representing real-time information on each stock keeping unit (SKU) in a warehouse (WH), which brings significant benefits to WH managers and thus saving human labor spent on looking for material, improving management flow, reducing the overall equipment effectiveness (OEE) of vehicles, improving throughput, quality etc.
Claims
AMENDED CLAIMS received by the International Bureau on 24 October 2025 (24.10.2025)CLAIMS1. A method of observing objects using computer vision with reduced computational workload, comprising use of computer vision to turn one or more video streams obtained from a camera into a verifiable signal with an interpretation similar to an industrial loT sensor, and the method further comprises use of: a computer program having instructions adapted to provide the occurrence probability of a specific object in a certain position and certain time on the output, at least one camera having resolution and frame rate that allow for human recognition and identification of the object and its exemplary states, plurality of virtual sensors, wherein each virtual sensor is assigned to a region of interest representing a section of the image and is further assigned to a specific camera used to obtain such image, a computer processor to process the data obtained from at least one camera, at least one human annotator involved in the process of training and / or retraining of the virtual sensor, wherein the virtual sensor is repeatedly applied to processed images obtained from camera, and provides a time series of probabilities of a presence or motion of an object at a specific position, which are used for consolidation i.e. strengthening signal across several virtual sensors and detecting the presence or motion at high accuracy, and the virtual sensor is adapted to measure probability and detect object appearance and disappearance by comparing images of the same part of the observed area in the region of interest, wherein the virtual sensor is pre-trained by convolutional neural network (CNN) in order to turn parts of the individual image into scores representing likelihood of occurrence of an object in a corresponding 3D area, thus processing the video footage into a time series data using CNN or other feature assessment techniques, wherein the object of interest may include material, material changes, and vehicles, characterized in that the virtual sensor is configured to maintain a defined minimum detection accuracy threshold during operation, wherein outputs below said threshold trigger escalation for human processing, verification and retraining, and that the method employs a hierarchical system ofmultiple executors to execute such processing, verification and retraining, including at least one human annotator, wherein processing, verification and retraining tasks are allocated among such executors based on their reliability, availability, and latency, thereby minimizing computational cost while preserving target accuracy, and wherein the processing, verification, and retraining steps are performed within the timescale of the observed process.
2. The method according to claims 1 , wherein the manual correction of the data is used for re-training of the CNN to increase the accuracy of the virtual sensor.
3. The method according to claim 1 or 2, wherein pre-training the CNN model for recognition of the object via virtual sensor is made using deep learning tools and / or Artificial Intelligence (Al).
4. The method according to any of claims 1 to 3, wherein the image data obtained from camera are processed by computer processor(s) situated in the proximity of the camera not exceeding 1 km of distance, preferably tens of meters, wherein the camera and the processor form edge device.
5. The method according to any of claims 1 to 4, wherein all cameras are stationary.
6. Use of the method according to any of claims 1 to 5 for tracking of trains, cranes, cars on roads, highways or parking lots, tracking of products in industrial complexes and warehouses.