Set prediction for visual inspection
The set prediction technique addresses labor-intensive and inaccurate infrastructure inspection issues by using a specialized model architecture to generate precise defect predictions, improving efficiency and accuracy in defect identification and visualization.
Patent Information
- Application Number
- PCT/US2024/020774
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-20
- Publication Date
- 2025-09-25
AI Technical Summary
Conventional inspection methods for infrastructure, such as sewer lines and bridges, are labor-intensive and prone to inaccuracies due to the time-consuming nature of manual video analysis, and existing machine learning systems face challenges in accurately linking defect detections across frames, leading to redundant reports and classification errors.
A set prediction technique using a specialized model architecture that generates precise predictions for infrastructure defects, incorporating a feature extractor, transformer-based feature encoder, and predictor to identify keyframes with high confidence scores, reducing annotation costs and improving defect identification efficiency.
The method significantly reduces annotation efforts and enhances inspection efficiency by providing accurate, concise defect identification and visualization, minimizing redundant notifications and misclassifications.
Smart Images

Figure US2024020774_25092025_PF_FP_ABST
Abstract
Description
SET PREDICTION FOR VISUAL INSPECTIONBACKGROUNDField
[0001] The present disclosure is generally directed to inspection systems and methods for identifying events in image data, and more specifically, to infrastructure defect detection using machine learning methods.
[0002] Related Art
[0003] Regular inspection of infrastructure, such as sewer lines, tunnels, and bridges, is vital for ensuring their safety and reliability. During inspections, inspectors assess the condition of these structures and document any observed damages, their severity, and precise locations. This information aids in forecasting potential future collapses and strategizing repair priorities for issues that pose significant risks. Such inspections are critical in ensuring uninterrupted utility’ services and reducing the risk of severe injuries.
[0004] Visual inspection is a widely used method for infrastructure inspections that involves documentation of visually observed damages. The process typically involves the analysis of video recordings of the target infrastructure by certified inspectors. This method allows inspectors to visually assess the surfaces of the infrastructure, which improves the accuracy of damage evaluations and enables infrastructure owners to take prompt action to prevent severe damage. Anticipating future structural failures aids in ensuring the safety of the infrastructure and its users.
[0005] However, the detailed video analysis required by inspectors is both time-consuming and labor-intensive. Further, it carries the risk of overlooking existing damages, thus potentially compromising the accuracy of the evaluations.
[0006] Video analytics-based systems have been developed for automation in industrial settings. These systems use machine learning algorithms to analyze videos and provide inspection information such as the locations, sizes, and other details of defects. This information can be used to identify frames with defects and other issues, thereby helping inspectors quickly find damages and bypass non-problematic frames. As a result, machine learning-based video analysis can significantly reduce inspection time and labor costs.
[0007] One existing approach involves a vision-based defect detection method that feeds surface images into machine-learning models to detect defects in input images. If the modeldetects any defects, their locations are displayed inside images using visualization methods such as bounding boxes around defect locations.
[0008] Yet, to accurately evaluate infrastructure conditions, it is essential to report each defect, necessitating a method for link detection results across video frames. One way to achieve this is by using Tracking-by-detection, which is a common method to associate detected defects between frames. In this technique, defect detectors are first applied to each frame of an input video, and output detection results are obtained for each frame. Subsequently, detection results such as bounding boxes are associated between frames based on the location of the box and its appearance. Through this tracking, one track result is obtained for each defect.
[0009] The process of generating training data for these models requires extensive manual labor. For example, recent deep-neural network-based methods require tens of thousands of annotated images to achieve accurate prediction results, meaning that tens of thousands of images need to be annotated with bounding boxes and their corresponding defect class labels. This can result in substantial annotation costs. The problem is further exacerbated when defect locations are not well-defined by bounding boxes, which can cause confusion and prolong the annotation process. Additionally, if trackers need to be trained to obtain highly accurate tracking results, ground truth bounding boxes must be annotated with frame association information so that trackers can extract crucial information to associate bounding boxes during training. This annotation requirement further increases the labor involved.
[0010] Therefore, it is desirable to have systems and methods that overcome the limitations of existing methods.SUMMARY
[0011] Methods and systems herein allow for the accurate identification of infrastructure defects and similar events through analysis of inspection videos and similar collections of images. Set prediction techniques cause a model to generate a singular, precise prediction for each identified defect, thereby eliminating redundant predictions for any single defect. In embodiments, this may be achieved by the creation of a specialized model architecture that is configured to produce a concise set of defect predictions. Furthermore, a training method enables the model to accurately predict such a set of defects.
[0012] Some aspects of the disclosure, comprise obtaining a sequence of frames that are associated with one or more events, and using a feature extractor to transform each frame in the sequence into a feature embedding. The feature embeddings are provided to a feature encoder that, for each identified event, activates a specific feature embedding, resulting in a collection of activated feature embeddings. Activated embeddings are then supplied to apredictor that generates an output that includes prediction results. Each prediction result comprises a confidence score that indicates the likelihood of each event's occurrence. Based on these confidence scores, a keyframe, e.g., a frame that best displays the event determined by the highest confidence score for that event, may be selected for each activated feature embedding. The keyframe provides a visual representation of the identified events as the output.
[0013] In some aspects of the disclosure, each activated feature embedding is treated as a one-dimensional vector, with the sequence of frames mapping out a timeline of events. The feature extractor may be trained by a model designed to extract features that contribute to the output of prediction results. These results may comprise an event’s classification, location, severity, or other relevant characteristics. The feature extractor might utilize a 3D CNN to extract features and process the sequence for temporal feature embeddings or a 2D CNN for spatial analysis.
[0014] In some aspects of the disclosure, a transformer-based feature encoder may be equipped with a self-attention mechanism, which facilitates the exchange of contextual information between feature embeddings across frames, identifying embeddings common to the sequence. This encoder can capture temporal progressions and is configured to disregard embeddings unrelated to the events. The feature encoder may comprise an FNN, utilizing either activation functions or linear layers to activate the specific embeddings. The predictor may comprise a set of FNNs, each tasked with generating a part of the prediction results.
[0015] In some aspects of the disclosure, in a training phase, the outputs are matched with ground truth labels to ensure continuity and accuracy in event detection across frames. This phase includes training the model with broadly annotated frames and utilizing an objective function to iteratively adjust the model’s parameters for improved accuracy and reliability in event identification. The matching process, essential for linking events across frames, may employ the Hungarian algorithm to determine the most reliable confidence scores.
[0016] Aspects of the present disclosure can involve a system, which can involve means for obtaining a sequence of frames that are associated with one or more events and transforming each frame in the sequence into a feature embedding. The feature embeddings are provided to means for activating, for each identified event, a specific feature embedding, resulting in a collection of activated feature embeddings. Activated embeddings are then supplied to a means for generating an output that includes prediction results. Each prediction result comprises a confidence score that indicates the likelihood of each event’s occurrence. Based on these confidence scores, a keyframe, e.g., a frame that best displays the event determined by thehighest confidence score for that event, may be selected for each activated feature embedding. The keyframe provides a visual representation of the identified events in the output, as described.
[0017] As a result, methods and systems herein substantially reduce annotation costs by simplifying the task for annotation workers to merely identifying frames that depict target events. Monitoring or inspection operations become more efficient as the occurrence of redundant notifications and inaccuracies in event classification is significantly reduced. The system’s capability to highlight keyframes drastically simplifies monitoring and evaluations by inspection staff, enabling swift and more accurate identification of events.BRIEF DESCRIPTION OF DRAWINGS
[0018] FIG. 1 illustrates an exemplary diagram of a model architecture and the flow for set prediction, in accordance with an example implementation.
[0019] FIG. 2 is a flowchart for a defect set prediction process, in accordance with an example implementation.
[0020] FIG. 3 is a more generalized flowchart for an event set prediction process, in accordance with an example implementation.
[0021] FIG. 4 illustrates an exemplary feature extractor, in accordance with an example implementation.
[0022] FIG. 5 illustrates an exemplary feature encoder, in accordance with an example implementation.
[0023] FIG. 6 illustrates an exemplar}' predictor, in accordance with an example implementation.
[0024] FIG. 7 illustrates an exemplary model architecture and the flow for a defect set prediction training process, in accordance with an example implementation.
[0025] FIG. 8 is a flowchart for a defect set prediction training process, in accordance with an example implementation.
[0026] FIG. 9 illustrates another exemplary' feature encoder, in accordance with an example implementation, and FIG. 10 illustrates another exemplary predictor, in accordance with an example implementation.
[0027] FIG. 1 1 illustrates yet another exemplary encoder, in accordance with an example implementation.
[0028] FIG. 12 illustrates an example computing environment with an example computer device suitable for use in some example implementations.DETAILED DESCRIPTION
[0029] The following detailed description provides details of the figures and example implementations of the present application. Reference numerals and descriptions of redundant elements between figures are omitted for clarity. Terms used throughout the description are provided as examples and are not intended to be limiting. For example, the use of the term “automatic” may involve fully automatic or semi-automatic implementations involving user or administrator control over certain aspects of the implementation, depending on the desired implementation of one of ordinary skill in the art practicing implementations of the present application. Selection can be conducted by a user through a user interface or other input means, or can be implemented through a desired algorithm. Example implementations as described herein can be utilized either singularly or in combination and the functionality of the example implementations can be implemented through any means according to the desired implementations.
[0030] Conventional consolidation techniques such as tracking, commonly utilized to combine multiple detection results into one defect report, frequently face challenges in achieving accurate results due to limited visual features or ambiguous definitions that do not allow for accurate categorization. This ambiguity oftentimes leads to the generation of multiple redundant reports for the same defect by machine learning models. This, in turn, necessitates intervention by inspectors to manually eliminate the excess reports produced by such methods, which ultimately reduces overall inspection efficiency. Thus, despite advancements, the efficiency of inspections remains hampered by the limitations of conventional defect recognition methods.
[0031] Further, in existing sy stems, detection and tracking are performed separately, which impedes the ability to recognize defects by using temporal context. This makes it challenging to differentiate defects that are captured from distant locations, thus leading to classification errors. As an example, a fracture defect may be misidentified as a simple crack when captured from a distance, and the predicted class name may change as the camera approaches the fracture defect. Such classification inconsistencies make it challenging to consolidate the detection results and impose additional burdens on inspectors to manually remove misclassified defects, further contributing to an inefficient inspection process.
[0032] Therefore, it is desirable to have systems and methods that can accurately and efficiently identify specific events from image data to monitor and identify specific events based on spatial and temporal analysis without misidentification and unnecessary duplication, thereby, significantly reducing manual efforts required for data annotation and examination.
[0033] In this document, the term “set"’ is understood in the mathematical sense, where each identified event or defect is represented as a unique element within that set.
[0034] FIG. 1 illustrates an exemplary diagram of a model architecture and the flow for set prediction, in accordance with an example implementation. In embodiments, architecture 150 comprises camera 100, mobile unit 101, pipe 102, frames 103, feature extractor 104, feature embeddings 105, feature encoder 106, activated feature embedding 107, predictor 108, prediction results 109, and graphical user interface (GUI) 110.
[0035] Camera 100 may be any type of sensor that can generate image frames. As depicted in FIG. 1, camera 100 is either permanently installed or mounted on mobile unit 101, such as a rover or drone, and configured to record video footage of any type of surface defect. The recorded videos can be stored in the camera's memory devices and then transferred to a visual inspection platform for subsequent processing. Alternatively, videos can be directly streamed to a video platform via wireless networks for immediate processing.
[0036] In operation, mobile unit 101 may access pipe 102, which may be a sewer line or any other infrastructure to be inspected. The video footage comprises frames 103 that are input to feature extractor 104, e.g., after being converted into a sequence of frames. Feature extractor 104 is configured to process frames 103 and convert them into one-dimensional feature embeddings 105. Feature extractor 104 is further trained to extract essential features from each frame 103 to predict defect information, such as defect classes, defect locations within frames, and defect severities.
[0037] Feature extraction can be applied to each frame 103 separately, e.g., using a 2- dimensional convolutional neural network (CNN), or to an entire frame sequence by using a 3- dimensional CNN. As a person of skill in the art will appreciate, using 2D CNNs allow s leveraging longer temporal context, as 2D CNNs have less computational complexity’ than 3D CNNs, and can thus handle relatively more frames. Typically no direct temporal information is used in the feature extraction phase of 2D CNNs. Conversely, 3D CNNs allow users to leverage temporal information, typically at the cost of the length of the temporal context becoming shorter.
[0038] Once the feature embeddings 105 are extracted, they are directed to a feature encoder 106, which may comprise a Transformer network that allows extracted feature embeddings 105 to communicate and exchange information extracted from each frame 103. During this communication, temporal context is leveraged to increase defect identification accuracy. This ensures that specific feature embeddings are activated for identified defectsacross frames 103, with the remainder being deactivated, thereby enabling the model to generate a single predictive result 109 per defect.
[0039] These activated feature embeddings 107 are then input into predictor 108, which may comprise either a straightforward linear layer or more complex feed-forward neural networks (FNNs) that transform activated feature embeddings 107 into definitive prediction results 109. In such set prediction embodiments, a high score (e.g., a confidence score) is assigned to a specific defect, even if several frames 103 capture that same defect by using scores to serve as primary indicators of defects for defect detection.
[0040] The recognition outcomes can be visualized on graphical user interface (GUI) 110, and those frames that are associated with activated feature embeddings can be highlighted as keyframes. Such keyframes are distinguished by their clear depiction of target defects attributed to the model’s refined training method. The visualization not only expedites the process of understanding defects by inspectors but also improves the efficiency of the inspection workflow.
[0041] FIG. 2 is a flowchart for a defect set prediction process, in accordance with an example implementation. In embodiments, defect set prediction process 200 may start, for example, after obtaining image data from a camera that generates video footage depicting surface defects on the surface of some infrastructure (e.g., pipe 102 depicted in FIG. 1). At step 202, the image data may be obtained in form of frames that are extracted from the video footage.
[0042] At step 204, any number of such frames are input to a feature extractor that processes the frames by converting each frame into a feature embedding vector to obtain feature embeddings. Such a conversion process ensures that every frame is analyzed individually, allowing for a granular examination of potential defects.
[0043] At step 206, the obtained feature embeddings are then fed into a feature encoder to undergo further processing to obtain refined or enhanced feature embeddings. The refinement process may leverage deep learning techniques to highlight critical defect-related features from each frame.
[0044] At step 206, the refined feature embeddings are then fed into a predictive model, such as predictor 108 shown in FIG. 1, which evaluates feature embeddings to generate prediction results, e.g., until the video ends, which is determined at step 210. As discussed in greater detail further below, prediction results not only assign class scores but may comprise defect class scores and other information, such as locations of defects within frames and anassessment of their severity. In this model, only the feature embedding that is deemed the most relevant among those with high scores is identified as selected.
[0045] Once a detected defect is designated in this manner, and the video ends, the implicated frame is displayed in GUI window 110, at step 212. The frame with the high score may be visually accentuated as a keyframe, e.g., surrounded by adjacent frames such as to enable a comprehensive view of the defect and facilitate efficient and accurate inspection.
[0046] FIG. 3 is a more generalized flowchart for an event set prediction process, in accordance with an example implementation. In embodiments, event set prediction process 300 may start at step 302, when, in response to acquiring a sequence of frames associated with one or more events of interest, a feature extractor is used to convert each frame within the obtained sequence, transforming it into a feature embedding. This process may be performed for each of a set of frames to capture essential data and characteristics.
[0047] At step, 304 the feature embeddings may be fed into a feature encoder, which is configured to identity7and activate a specific feature embedding for each event, resulting in a set of activated feature embeddings that represent the identified events.
[0048] At step 306, the set of activated feature embeddings may be provided to a predictor. The predictor is tasked with generating prediction results, which may comprise a confidence score for each event. The confidence score serves as an indicator of the likelihood that the event is present.
[0049] At step 308, the confidence scores obtained from the predictor are used to select a keyframe for each activated feature embedding. The selected keyframe is the frame within the sequence that best displays the event, as determined by the highest confidence score.
[0050] Finally, at step 310, any number of selected keyframes may be output as a representation of the events identified from the sequence of frames. A keyframe serves as a visual summary of an event that may be further analyzed or reported.
[0051] One skilled in the art shall recognize that: (1) certain steps may optionally be performed; (2) steps may not be limited to the specific order set forth herein; (3) certain steps may be performed in different orders; and (4) certain steps may be done concurrently.
[0052] FIG. 4 illustrates an exemplary feature extractor, in accordance with an example implementation. In embodiments, the core of feature extractor 104 may comprise CNNs that contain learnable weight and bias parameters, which are updated during the training process to accurately classify defects and generate relevant outputs. Initially, the CNNs apply convolution operations to input video frames using these learnable parameters. Convolution operations are followed by activation functions (e.g., ReLU) that introduce a non-linearity7. Additionalprocessing operations may include, for example, normalization and regularization operations. The sequence of convolution and activation operations is iterated several times and converts the input video frames into feature maps that represent crucial information required for defect prediction. To transition from multi-dimensional feature maps (e.g., 3-dimensional tensors) to one-dimensional feature embeddings, a global average pooling (GAP) technique 402 may be used, which averages values in feature maps over height and width directions. This feature extraction phase is designed to extract information, such as color, edges, background contexts, and object contexts, from video frames, preparing them for further analysis by the feature encoder and processing by the predictive model.
[0053] FIG. 5 illustrates an exemplary feature encoder, in accordance with an example implementation. In embodiments, feature encoder 106 utilizes advanced Transformer encoder networks that have two main components: self-attention modules (e.g., 502) and feed-forward neural networks (FNNs) (e.g., 504). Self-attention module 502 is designed to facilitate interaction among the input feature embeddings by using attention mechanisms. The learnable parameters of the self-attention modules allow one feature embedding to be activated for one defect while suppressing others. This interaction ensures that each detected defect is represented by a unique, activated embedding. Therefore, if a frame sequence contains two different defects, two distinct embeddings will be activated, and the rest will be deactivated or ignored. FNN 504 comprises two linear layers having weight and bias parameters and an activation function between them (not shown). FNN 504 refines feature embeddings after communication in self-attention module 502. Self-attention operations and FNN operations are iteratively applied to optimize the embeddings for accurate defect prediction.
[0054] FIG. 6 illustrates an exemplary predictor, in accordance with an example implementation. In embodiments, predictor 108 comprises FNNs (e.g., 602), each designed to predict a distinct component of defect analysis. For instance, one FNN may be dedicated to determining likelihood scores associated with various defect classes based on feature embeddings, while another FNN may be used to evaluate the severity7of identified defects using the same feature embeddings for its analysis. The main function of predictor 108 is to extract and process the information contained in feature embeddings 107 and transform them into output values to facilitate comprehensive defect assessment.
[0055] FIG. 7 illustrates an exemplary model architecture and the flow for a defect set prediction training process, in accordance with an example implementation. In embodiments, training the model involves the preparation of a set of ground truth labels, which may comprise information such as defect classes, their approximate temporal occurrences, and otherinformation for accurate prediction. Once ground truth labels 706 are established, the training phase can commence. Training may start with generating initial prediction results 702, e.g., by using the previously mentioned prediction process. Then prediction results 702 and ground truth labels 706 are matched, e.g., by using a minimum-cost matching algorithm that is configured to define the prediction that most closely corresponds to each actual defect. A main goal of the algorithm is to identify the most representative prediction result that matches a ground truth label 706. This matching helps to train the model7s components, such as feature extractor 104, feature encoder 106, and predictor 108, to generate one accurate prediction result for each defect, irrespective of the number of frames that capture that event, such that a number of activated feature embeddings corresponds to the one or more events, thereby suppressing redundant identification of the same defect across frames. After the matching process, trainer 704 evaluates objective functions derived from prediction results 702 and ground truth labels 706 to update the model’s parameters such as to optimize objective functions. Updating ensures that matched prediction results have high classification scores, whereas unmatched prediction results have low classification scores. As a result, this training enables the identification of the most representative frame for each defect, which is then highlighted, in the subsequent prediction phase, with a high predictive score.
[0056] FIG. 8 is a flowchart for a defect set prediction training process, in accordance with an example implementation. In embodiments, in a training phase, images or frames, which have been obtained from a camera, may be provided to a model that uses a feature extractor to obtain feature embeddings at step 802. At step 804, the feature embeddings are provided to a feature encoder that refines them and, at step 806, provides them to a predictor to obtain prediction outcomes. A matcher, such as matcher 710 in FIG. 7, is used to match the model’s prediction outcomes with the ground truth labels at step 808. Ground truth labels may comprise defects details such as their classification, severity, and temporal location. At step 810, the matcher may calculate a matching cost between prediction results and ground truth labels and find an optimal assignment of the prediction results to the ground truth labels, e.g., by using a minimum cost matching algorithm such as the Hungarian algorithm. Based on the matching results, a trainer may then calculate objective functions, such as cross-entropy loss for defect classification. The trainer may further update the parameters of feature extractor, the feature encoder, and the predictor based on the values of the objective functions and employ update algorithms, such as stochastic gradient descent, for neural network models.
[0057] The matching process may incorporate several factors, including defect class scores. As an example, to predict defect class scores and the severity of defects, the matching cost of the z-th prediction result and the z-th ground truth label can be obtained as follows:
[0058] Wherein J~C,7is a matching cost value between the z-th prediction result and the z- th ground truth label;is the matching cost of defect class scores between them; JfV is the matching cost of severity between them;is the matching cost of temporal location distance, and a are weight hyper-parameters for the costs, which are used to adjust the importance of each cost: cLE [0,1]Ncis predicted defect class scores; Cj E {0,1}^ is ground truth defect class labels; stE [0,1] is a predicted severity score; S}E [0,1] is a ground truth severity score; d, E Z is a frame number that generates the z-th prediction results; and dj E L is a ground truth frame number of the frame that captures a target defect.
[0059] The model utilizes matched prediction results with ground truth labels to refine objective functions to ensure that accurate predictions yield high scores for the corresponding defect categories, closely mirroring ground truth data. Conversely, predictions that diverge from ground truth labels result in the diminution of their scores. In this manner, the objective function may be used to penalize the model for inaccurately connecting events across a sequence of frames and reward accurate identification or linking of events. By rewarding precise matches and penalizing discrepancies, the model thereby increases the reliability' and precision of defect classification. At step 812 it is determined whether training should end, e.g., because the objective function met a desired accuracy. If not, training may resume with step 802 by receiving another set of frames. Otherwise, at step 814, the trained parameters (e.g., weight parameters) may be stored, e.g., for use in a validation phase followed by an inference phase in which the model is presented with previously unseen frames.
[0060] FIG. 9 illustrates another exemplary feature encoder, in accordance with an example implementation, and FIG. 10 illustrates another exemplary' predictor, in accordance with an example implementation. In embodiments, feature encoder 106 may comprise both a transformer encoder network and a transformer decoder network. The decoder netw ork may comprise three primary components: cross-attention module 902. FNNs 904, and queryembeddings 906. Cross-attention modules allow 902 query embeddings to aggregate features from input feature embeddings using attention mechanisms. The learnable parameters of these modules enable query embedding 906 to be activated for a single defect while deactivating others. If there are two different defects in the frame sequence, then two different embeddings are activated while others are deactivated. FNNs 904 have two linear layers with weight and bias parameters along with an activation function between them. They further refine query embeddings 906 after aggregation in cross-attention modules 902. The cross-attention operation and FNNs 902 are applied iteratively.
[0061] When transformer decoder networks are used, it is unclear which frame features the feature embeddings have aggregated in cross-attention modules 902. as shown in FIG. 10, to identify the frame numbers whose information the feature embeddings contain, predictor 108 requires an FNN that converts input feature embeddings to frame numbers. If transformer decoder networks are used, it becomes easier to learn to obtain one activated embedding for one defect and deactivate other embeddings, thereby improving recognition performance.
[0062] FIG. 11 illustrates yet another exemplary encoder, in accordance with an example implementation. In embodiments, feature encoder 106 may comprise memory module 1102 that feeds previously activated embeddings to self-attention modules 1104. Memory module 1102 enables feature embeddings to leverage longer temporal context, thereby improving the recognition performance.
[0063] FIG. 12 illustrates an example computing environment with an example computer device suitable for use in some example implementations. Computer device 1205 in computing environment 1200 can include one or more processing units, cores, or processors 1210, memory' 1215 (e.g., RAM, ROM, and / or the like), internal storage 1220 (e.g., magnetic, optical, solid-state storage, and / or organic), and / or I / O interface 1225, any of which can be coupled on a communication mechanism or bus 1230 for communicating information or embedded in the computer device 1205. I / O interface 1225 is also configured to receive images from cameras or provide images to projectors or displays, depending on the desired implementation.
[0064] Computer device 1205 can be communicatively coupled to input / user interface 1235 and output device / interface 1240. Either one or both of input / user interface 1235 and output device / interface 1240 can be a wired or wireless interface and can be detachable. Input / user interface 1235 may include any device, component, sensor, or interface, physical or virtual, that can be used to provide input (e.g., buttons, touch-screen interface, keyboard, a pointing / cursor control, microphone, camera, braille, motion sensor, optical reader, and / or the like). Output device / interface 1240 may include a display, television, monitor, printer, speaker,braille, or the like. In some example implementations, input / user interface 1235 and output device / interface 1240 can be embedded with or physically coupled to the computer device 1205. In other example implementations, other computer devices may function as or provide the functions of input / user interface 1235 and output device / interface 1240 for a computer device 1205.
[0065] Examples of computer device 1205 may include highly mobile devices (e.g., smartphones, devices in vehicles and other machines, devices earned by humans and animals, and the like), mobile devices (e g., tablets, notebooks, laptops, personal computers, portable televisions, radios, and the like), and devices not designed for mobility (e.g., desktop computers, other computers, information kiosks, televisions with one or more processors embedded therein and / or coupled thereto, radios, and the like).
[0066] Computer device 1205 can be communicatively coupled (e.g., via I / O interface 1225) to external storage 1245 and network 1250 for communicating with any number of networked components, devices, and systems, including one or more computer devices of the same or different configuration. Computer device 1205 or any connected computer device can be functioning as, providing services of, or referred to as a server, client, thin server, general machine, special-purpose machine, or another label.
[0067] I / O interface 1225 can include wired and / or wireless interfaces using any communication or I / O protocols or standards (e.g., Ethernet, 802.1 lx, Universal System Bus, WiMax, modem, a cellular network protocol, and the like) for communicating information to and / or from at least all the connected components, devices, and network in computing environment 1200. Network 1250 can be any network or combination of networks (e.g., the Internet, local area network, wide area network, a telephonic network, a cellular network, a satellite network, and the like).
[0068] Computer device 1205 can use and / or communicate using computer-usable or computer-readable media, including transitory media and non-transitory media. Transitory media include transmission media (e.g., metal cables, fiber optics), signals, carrier waves, and the like. Non-transitory media include magnetic media (e.g.. disks and tapes), optical media (e.g.. CD ROM, digital video disks, Blu-ray disks), solid-state media (e.g., RAM, ROM. flash memory, solid-state storage), and other non-volatile storage or memory.
[0069] Computer device 1205 can be used to implement techniques, methods, applications, processes, or computer-executable instructions in some example computing environments. Computer-executable instructions can be retrieved from transitory media, and stored on and retrieved from non-transitory media. The executable instructions can originate from one ormore of any programming, scripting, and machine languages (e.g., C, C++, C#, Java, Visual Basic, Python, Perl, JavaScript, and others).
[0070] Processor(s) 1210 can execute under any operating system (OS) (not shown), in a native or virtual environment. One or more applications can be deployed that include logic unit 1260, application programming interface (API) unit 1265, input unit 1270, output unit 1275, and inter- unit communication mechanism 1295 for the different units to communicate with each other, with the OS, and with other applications (not shown). The described units and elements can be varied in design, function, configuration, or implementation and are not limited to the descriptions provided. Processor(s) 1210 can be in the form of hardw are processors such as central processing units (CPUs) or in a combination of hardw are and softw are units.
[0071] In some example implementations, when information or an execution instruction is received by API unit 1265, it may be communicated to one or more other units (e.g., logic unit 1260, input unit 1270, output unit 1275). In some instances, logic unit 1260 may be configured to control the information flow among the units and direct the services provided by API unit 1265, input unit 1270, output unit 1275, in some example implementations described above. For example, the flow of one or more processes or implementations may be controlled by logic unit 1260 alone or in conjunction with API unit 1265. The input unit 1270 may be configured to obtain input for the calculations described in the example implementations, and the output unit 1275 may be configured to provide output based on the calculations described in example implementations.
[0072] Processor(s) 1210 can be configured to execute a method or computer instructions which can involve, obtaining a sequence of frames that are associated with one or more events, and using a feature extractor to transform each frame in the sequence into a feature embedding. The feature embeddings are provided to a feature encoder that, for each identified event, activates a specific feature embedding, resulting in a collection of activated feature embeddings. Activated embeddings are then supplied to a predictor that generates an output that includes prediction results. Each prediction result comprises a confidence score that indicates the likelihood of each event’s occurrence. Based on these confidence scores, a keyframe, e.g.. a frame that best displays the event determined by the highest confidence score for that event, may be selected for each activated feature embedding. The keyframe provides a visual representation of the identified events is the output, as described, for example, with respect to the flowchart in FIG. 3.
[0073] In addition, each activated feature embedding is treated as a one-dimensional vector, with the sequence of frames mapping out a timeline of events. The feature extractormay be trained by a model designed to extract features that contribute to the output of prediction results. These results may comprise an event's classification, location, severity, or other relevant characteristics. The feature extractor might utilize a 3D CNN to extract features and process the sequence for temporal feature embeddings or a 2D CNN for spatial analysis.
[0074] A transformer-based feature encoder may be equipped with a self-attention mechanism, which facilitates the exchange of contextual information between feature embeddings across frames, identifying embeddings common to the sequence. This encoder can capture temporal progressions and is configured to disregard embeddings unrelated to the events. The feature encoder may comprise an FNN, utilizing either activation functions or linear layers to activate the specific embeddings. And the predictor may comprise a set of FNNs, each tasked with generating a part of the prediction results.
[0075] In a training phase, the outputs are matched with ground truth labels to ensure continuity and accuracy in event detection across frames. This phase includes training the model with broadly annotated frames and utilizing an objective function to iteratively adjust the model's parameters for improved accuracy and reliability in event identification. The matching process, essential for linking events across frames, may employ the Hungarian algorithm to determine the most reliable confidence scores. It is noted the annotations in such broadly annotated frames may be general information about the content without detailed labeling. An exemplary coarsely annotated frame may indicate the presence of a defect in an image without specifying its exact type.
[0076] Some portions of the detailed description are presented in terms of algorithms and symbolic representations of operations within a computer. These algorithmic descriptions and symbolic representations are the means used by those skilled in the data processing arts to convey the essence of their innovations to others skilled in the art. An algorithm is a series of defined steps leading to a desired end state or result. In example implementations, the steps carried out require physical manipulations of tangible quantities for achieving a tangible result.
[0077] Unless specifically stated otherwise, as apparent from the discussion, it is appreciated that throughout the description, discussions utilizing terms such as "‘processing,” “computing,” “calculating,” “determining,” “displaying,” or the like, can include the actions and processes of a computer system or other information processing device that manipulates and transforms data represented as physical (electronic) quantities within the computer system’s registers and memories into other data similarly represented as physical quantities within the computer system’s memories or registers or other information storage, transmission or display devices.
[0078] Example implementations may also relate to an apparatus for performing the operations herein. This apparatus may be specially constructed for the required purposes, or it may include one or more general-purpose computers selectively activated or reconfigured by one or more computer programs. Such computer programs may be stored in a computer- readable medium, such as a computer-readable storage medium or a computer-readable signal medium. A computer-readable storage medium may involve tangible mediums such as optical disks, magnetic disks, read-only memories, random access memories, solid-state devices and drives, or any other types of tangible or non-transitory media suitable for storing electronic information. A computer-readable signal medium may include mediums such as carrier waves. The algorithms and displays presented herein are not inherently related to any particular computer or other apparatus. Computer programs can involve pure software implementations that involve instructions that perform the operations of the desired implementation.
[0079] Various general-purpose systems may be used with programs and modules in accordance with the examples herein, or it may prove convenient to construct a more specialized apparatus to perform desired method steps. In addition, the example implementations are not described with reference to any particular programming language. It will be appreciated that a variety of programming languages may be used to implement the techniques of the example implementations as described herein. The instructions of the programming language(s) may be executed by one or more processing devices, e.g., central processing units (CPUs), processors, or controllers.
[0080] As is known in the art, the operations described above can be performed by hardware, software, or some combination of software and hardw are. Various aspects of the example implementations may be implemented using circuits and logic devices (hardware), while other aspects may be implemented using instructions stored on a machine-readable medium (softw are), which if executed by a processor, would cause the processor to perform a method to carry out implementations of the present application. Further, some example implementations of the present application may be performed solely in hardw are, w hereas other example implementations may be performed solely in software. Moreover, the various functions described can be performed in a single unit, or can be spread across a number of components in any number of ways. When performed by software, the methods may be executed by a processor, such as a general-purpose computer, based on instructions stored on a computer-readable medium. If desired, the instructions can be stored on the medium in a compressed and / or encrypted format.
[0081] Moreover, other implementations of the present application will be apparent to those skilled in the art from consideration of the specification and practice of the techniques of the present application. Various aspects and / or components of the described example implementations may be used singly or in any combination. It is intended that the specification and example implementations be considered as examples only, with the true scope and spirit of the present application being indicated by the following claims.
Claims
CLAIMSWhat is claimed is:
1. A method for identifying an event from a set of images, the method comprising: in response to obtaining a sequence of frames associated with one or more events, using a feature extractor to convert each frame in the sequence of frames into a feature embedding; providing the feature embeddings to a feature encoder that, for each of the one or more events, activates one feature embedding to output a set of activated feature embeddings; providing the set of activated feature embeddings to a predictor that generates an output comprising prediction results that comprise, for each event, a confidence score indicative of that event; using the confidence score to select, for each activated feature embedding, a keyframe among the sequence of frames that displays the event; and outputting the keyframe.
2. The method according to claim 1, wherein each activated feature embedding is a onedimensional vector and the sequence of frames represents a timeline.
3. The method according to claim 1 , wherein the feature extractor has been trained by a model that extracts features to output the prediction results, the prediction results comprising at least one of an event, an event class, an event location, or an event severity.
4. The method according to claim 3, wherein the feature extractor comprises at least one of a 3-dimensional convolutional neural network (CNN) configured to extract the features and process the sequence of frames to extract temporal feature embeddings or a tw o-dimensional CNN.
5. The method according to claim 1, wherein the feature encoder is a transformer-based feature encoder that comprises a self-attention mechanism and enables exchanging contextual information between feature embeddings in subsequent frames to identify feature embeddings that are common across the sequence of frames.
6. The method according to claim 5, wherein the transformer-based feature encoder obtains a temporal progression.
7. The method according to claim 5, wherein the self-attention mechanism is configured to deactivate feature embeddings that do not comprise the event.
8. The method according to claim 1 , wherein the feature encoder comprises a feed-forward neural network (FNN) comprising at least one of an activation function that activates the one feature embedding or a linear layer.
9. The method according to claim 1, wherein the predictor comprises a set of FNNs that is each is configured to generate one of the prediction results.
10. The method according to claim 4, further comprising, in a training phase, performing steps compnsing: matching the output with a ground truth label to link events across the sequence of frames; training the model using a dataset of broadly annotated frames; and determining and using an objective function to iteratively adjust model parameters.1 1 . The method according to claim 10, wherein the objective function penalizes the model for inaccurately connecting the events across the sequence of frames and rewards accurate identification or linking of the events.
12. The method according to claim 10, wherein the matching uses the Hungarian algorithm to identify a highest score among the confidence scores.
13. The method according to claim 10, wherein the events are comprised in an inspection video and identify an infrastructure event and the ground truth label comprises at least one of a defect class, a defect location, or a defect severity associated with an infrastructure defect.
14. A non-transitory computer-readable medium for storing instructions for executing a process, the instructions comprising:responsive to obtaining a sequence of frames associated with one or more events, using a feature extractor to convert each frame in the sequence of frames into a feature embedding; providing the feature embeddings to a feature encoder that, for each of the one or more events, activates one feature embedding to output a set of activated feature embeddings; providing the set of activated feature embeddings to a predictor that generates an output comprising prediction results that comprise, for each event, a confidence score indicative of that event; using the confidence score to select, for each activated feature embedding, a keyframe among the sequence of frames that displays the event; and outputting the keyframe.
15. A system for identifying defects in infrastructure inspection videos, the system comprising: a camera configured to gather a sequence of frames associated with one or more events; a feature extractor configured to convert each frame in the sequence of frames into a feature embedding; a feature encoder that, for each of the one or more events, activates one feature embedding to output a set of activated feature embeddings; a predictor that, in response to receiving the set of activated feature embeddings, generates an output comprising prediction results that comprise, for each event, a confidence score indicative of that event, the predictor further configured to use the confidence score to select, for each activated feature embedding, a keyframe among the sequence of frames that displays the event; and a display configured to output the keyframe.
Citation Information
Patent Citations
Machine learning image processing techniques
US20210279866A1
Defect detection method and apparatus
US20210374928A1
Classifying elements and predicting properties in an infrastructure model through prototype networks and weakly supervised learning
US20220358360A1
Predictive modeling and control system for building equipment with generative adversarial network
US20230316066A1