Computer system and computer-implemented method for a computer vision-based rapid diagnostic test result interpretation platform
A computer vision system with deep learning and augmented reality aids in accurately capturing and interpreting rapid diagnostic test results by guiding optimal image capture and analysis, addressing inconsistencies and lighting issues.
Patent Information
- Application Number
- JP2025507683
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-10
- Filing Date
- 2023-08-10
- Publication Date
- 2025-10-01
AI Technical Summary
Existing rapid diagnostic tests face challenges in accurately capturing and analyzing results due to issues like inconsistent image acquisition, reliance on specific device features, and sensitivity to lighting conditions, leading to erroneous results.
A computer vision-based system using deep learning modules and neural networks to guide optimal image capture and analyze lateral flow-based diagnostic tests, incorporating augmented reality for user guidance and real-time result interpretation.
Enhances the accuracy and reliability of diagnostic test result interpretation by optimizing image capture and analysis, reducing errors caused by positioning and lighting variations.
Smart Images

Figure 2025532468000001_ABST
Abstract
Description
[Technical Field]
[0001] Generally, the present disclosure relates to computer-implemented methods and computer systems for a rapid diagnostic test result interpretation platform using computer vision. [Background technology]
[0002] The use of lateral flow-based in vitro rapid diagnostic tests (IV-RDTs) in disease diagnosis has expanded significantly, highlighted by the need for large-scale population self-screening during the COVID-19 pandemic. Summary of the Invention
[0003] In some aspects, the techniques described herein include receiving, by a processor, a first image stream of an inspection device from a camera, the inspection device including an inspection area where the inspection device displays a visual indicator; applying, by the processor, at least one computer vision technique of a deep learning module to identify a plurality of device features of the inspection device in the first image stream and classifying at least one first image in the first image stream to at least one first reference image in a corpus of reference images of the inspection device based on the plurality of device features of the inspection device; and selecting, by the processor, based on the at least one first reference image, at least one imaging instruction for the camera to capture a second image stream, the second image stream including the inspection area. receiving, by the processor, the second image stream from a camera tuned by the at least one imaging instruction; applying, by the processor, at least one computer vision technique of a deep learning module to identify a plurality of device features of the inspection area in the second image stream and classifying, based on the plurality of device features of the inspection area, at least one second image in the second image stream to at least one second reference image in a corpus of reference images of the inspection area; and identifying, by the processor, an inspection result based on a visual indicator in the at least one second image based on the at least one second reference image.
[0004] In some aspects, the techniques described herein relate to a method further including transmitting, by a processor, user identification information to a user device including a camera; maintaining, by the processor, a connection to the user device based on the user identification information; and receiving, by the processor, a first image stream and a second image stream over the connection.
[0005] In some aspects, the technology described herein relates to methods in which the plurality of device features are structural features.
[0006] In some aspects, the technology described herein relates to a method in which at least one computer vision technology of a deep learning module includes at least one modular neural network, each modular neural network including at least one convolutional neural network input.
[0007] In some aspects, the technology described herein relates to a method in which at least one convolutional neural network identifies an area of interest in at least one first image in a first image stream.
[0008] In some aspects, the technology described herein relates to a method in which at least one modular neural network identifies at least one candidate imaging instruction command, and at least one convolutional neural network selects at least one imaging instruction command from the at least one candidate imaging instruction command.
[0009] In some aspects, the techniques described herein relate to a method further including generating, by a processor, at least one imaging instruction command for a camera to capture a second image stream based on imaging instruction metadata associated with the at least one first reference image, wherein the second image stream includes the inspection area.
[0010] In some aspects, the technology described herein relates to a method in which instructing to perform at least one imaging instruction command includes transmitting, by a processor, the at least one imaging instruction command to a camera to cause the camera to automatically generate a second image stream based on the at least one imaging instruction command.
[0011] In some aspects, the techniques described herein relate to a method in which instructing to perform at least one imaging instruction includes transmitting, by a processor, to a user device including a camera, at least one imaging instruction including augmented reality instructions for the user device to overlay on a first image stream to instruct the user to reposition the camera to capture the inspection device.
[0012] In some aspects, the technology described herein relates to a method further including storing, by the processor, the test results; and transmitting, by the processor, the test results to an administrator device for display.
[0013] In some aspects, the techniques described herein include receiving a first image stream of an inspection device from a camera, the inspection device including an inspection area where the inspection device displays a visual indicator; applying at least one computer vision technique of a deep learning module to identify a plurality of device features of the inspection device in the first image stream; classifying at least one first image in the first image stream to at least one first reference image in a corpus of reference images of the inspection device based on the plurality of device features of the inspection device; and selecting at least one imaging instruction for the camera to capture a second image stream based on the at least one first reference image, the second image stream including the inspection area, the second image stream displaying a visual indicator. A system includes a processor that commands execution of at least one imaging instruction to automatically generate an image stream, receives a second image stream from a camera regulated by the at least one imaging instruction, applies at least one computer vision technique of a deep learning module to identify a plurality of device features of an inspection area in the second image stream, classifies at least one second image in the second image stream to at least one second reference image in a corpus of reference images of the inspection area based on the plurality of device features of the inspection area, and identifies an inspection result based on a visual indicator in the at least one second image based on the at least one second reference image.
[0014] In some aspects, the technology described herein relates to a system in which a processor transmits user identification information to a user device including a camera, maintains a connection to the user device based on the user identification information, and receives a first image stream and a second image stream over the connection.
[0015] In some aspects, the technology described herein relates to a system in which a plurality of device features are structural features.
[0016] In some aspects, the technology described herein relates to a system in which at least one computer vision technology of a deep learning module includes at least one modular neural network in communication with at least one convolutional neural network input.
[0017] In some aspects, the technology described herein relates to a system in which at least one convolutional neural network identifies an area of interest in at least one first image in a first image stream.
[0018] In some aspects, the technology described herein relates to a system in which at least one computer vision technique of a deep learning module identifies at least one candidate imaging instruction command, and at least one convolutional neural network selects at least one imaging instruction command from the at least one candidate imaging instruction command.
[0019] In some aspects, the technology described herein relates to a system in which the processor is further configured to generate at least one imaging instruction command for the camera to capture a second image stream based on imaging instruction metadata associated with at least one second reference image, the second image stream including the inspection area.
[0020] In some aspects, the technology described herein further relates to a system in which the processor sends at least one imaging instruction command to the camera to instruct the camera to perform the at least one imaging instruction command and to cause the camera to automatically generate a second image stream based on the at least one imaging instruction command.
[0021] In some aspects, the technology described herein further relates to a system in which the processor sends at least one imaging instruction command to a user device including a camera to instruct the user device to perform the at least one imaging instruction command, the at least one imaging instruction command including augmented reality instructions for the user device to overlay on the first image stream and instruct the user to reposition the camera to capture the inspection device.
[0022] In some aspects, the technology described herein further relates to a system where the processor stores the test results and transmits the test results to an administrator device for display.
[0023] In some aspects, the techniques described herein include: identifying, by a processor of a user device from a camera of the user device, a first image stream of an inspection device, the inspection device including an inspection area where the inspection device displays a visual indicator; applying, by the processor, at least one computer vision technique of a deep learning module, to identify a plurality of device features of the inspection device in the first image stream, transmitting the plurality of device features of the inspection device to a server including a corpus of reference images of the inspection device; receiving, from the server, at least one first reference image in the corpus of reference images, and classifying, based on the plurality of device features, the at least one first image in the first image stream into at least one first reference image; and selecting, by the processor, at least one imaging instruction for the camera to capture a second image stream, the second image stream including the inspection area. and instructing a processor to perform at least one imaging instruction to automatically generate a second image stream. The processor then identifies, by the processor, a second image stream from a camera tuned by the at least one imaging instruction. The processor then applies at least one computer vision technique of a deep learning module to identify a plurality of device features of the inspection area in the second image stream, transmitting the plurality of device features of the inspection area to a server including a corpus of reference images of the inspection area, receiving from the server at least one second reference image in the corpus of reference images of the inspection area, and classifying the at least one second image in the second image stream into the at least one second reference image based on the plurality of device features of the inspection area. The processor then identifies, by the processor, an inspection result based on a visual indicator in the at least one second image based on the at least one second reference image.
[0024] In some aspects, the techniques described herein relate to a method further including receiving, by a processor, user identification information from a server; maintaining, by the processor, a connection to the server based on the user identification information; and transmitting, by the processor, a plurality of device characteristics of the inspection device and a plurality of device characteristics of the inspection area to the server via the connection.
[0025] In some aspects, the technology described herein relates to methods in which the plurality of device features are structural features.
[0026] In some aspects, the technology described herein relates to a method in which at least one multi-agent system includes at least one deep learning model and at least one augmented reality model.
[0027] In some aspects, the technology described herein relates to a method in which at least one convolutional neural network identifies an area of interest in at least one first image in a first image stream.
[0028] In some aspects, the techniques described herein relate to a method in which at least one computer vision technique of a deep learning module identifies at least one candidate imaging instruction, and at least one convolutional neural network selects at least one imaging instruction from the at least one candidate imaging instruction.
[0029] In some aspects, the techniques described herein relate to a method further including receiving, by a processor, from a server, imaging instruction metadata associated with at least one second reference image; and generating, by the processor, at least one imaging instruction command for a camera to capture a second image stream based on the imaging instruction metadata associated with the at least one second reference image, wherein the second image stream includes an inspection area.
[0030] In some aspects, the technology described herein relates to a method in which instructing the camera to perform at least one image capture instruction includes generating, by a processor, at least one image capture instruction for the camera to automatically generate a second image stream based on the at least one image capture instruction.
[0031] In some aspects, the technology described herein relates to a method in which instructing a user device to perform at least one imaging instruction includes causing a processor to display, on a user device, at least one imaging instruction including augmented reality instructions overlaid on an image stream to instruct the user to reposition a camera to capture an inspection image.
[0032] In some aspects, the technology described herein relates to a method further including storing, by the processor, the test results; and transmitting, by the processor, the test results to an administrator device for display. [Brief explanation of the drawings]
[0033] Embodiments of the present disclosure, briefly summarized above and described in more detail below, can be understood by reference to exemplary embodiments of the present disclosure as illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings illustrate only typical embodiments of the present disclosure and, therefore, should not be considered limiting of its scope, as the present disclosure may admit of other equally effective embodiments. The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Patent and Trademark Office upon request and payment of the necessary fee.
[0034] [Figure 1] 1A-1D depict exemplary aspects of the present disclosure, in accordance with at least some principles of at least some embodiments of the present disclosure. [Figure 2]1A-1D depict exemplary aspects of the present disclosure, in accordance with at least some principles of at least some embodiments of the present disclosure. [Figure 3] 1A-1D depict exemplary aspects of the present disclosure, in accordance with at least some principles of at least some embodiments of the present disclosure. [Figure 4] 1A-1D depict exemplary aspects of the present disclosure, in accordance with at least some principles of at least some embodiments of the present disclosure. [Figure 5] 1A-1D depict exemplary aspects of the present disclosure, in accordance with at least some principles of at least some embodiments of the present disclosure. [Figure 6] 1A-1D depict exemplary aspects of the present disclosure, in accordance with at least some principles of at least some embodiments of the present disclosure. [Figure 7] 1A-1D depict exemplary aspects of the present disclosure, in accordance with at least some principles of at least some embodiments of the present disclosure. [Figure 8] 1A-1D depict exemplary aspects of the present disclosure, in accordance with at least some principles of at least some embodiments of the present disclosure. [Figure 9] 1A-1D depict exemplary aspects of the present disclosure, in accordance with at least some principles of at least some embodiments of the present disclosure. [Figure 10] 1A-1D depict exemplary aspects of the present disclosure, in accordance with at least some principles of at least some embodiments of the present disclosure. [Figure 11] 1A-1D depict exemplary aspects of the present disclosure, in accordance with at least some principles of at least some embodiments of the present disclosure. [Figure 12] 1A-1D depict exemplary aspects of the present disclosure, in accordance with at least some principles of at least some embodiments of the present disclosure. [Figure 13] 1A-1D depict exemplary aspects of the present disclosure, in accordance with at least some principles of at least some embodiments of the present disclosure. [Figure 14] 1A-1D depict exemplary aspects of the present disclosure, in accordance with at least some principles of at least some embodiments of the present disclosure. [Figure 15] 1A-1D depict exemplary aspects of the present disclosure, in accordance with at least some principles of at least some embodiments of the present disclosure. [Figure 16] 1A-1D depict exemplary aspects of the present disclosure, in accordance with at least some principles of at least some embodiments of the present disclosure.
[0035] For ease of understanding, the same reference numerals have been used, where possible, to indicate identical elements common to the figures. The figures are not drawn to scale and may use simplified representations for clarity. It is contemplated that elements and features of one embodiment may be beneficially incorporated in other embodiments without further recitation. DETAILED DESCRIPTION OF THE INVENTION
[0036] Among these advantages and technical solutions disclosed, other objects and advantages of the present disclosure may become apparent from the following description taken in conjunction with the accompanying drawings. Detailed embodiments of the present disclosure are disclosed herein; however, it should be understood that the disclosed embodiments are merely exemplary of the present disclosure, which may be embodied in various forms. Additionally, each of the examples given in connection with various embodiments of the present disclosure are intended to be illustrative and not limiting.
[0037] Throughout this specification, the following terms shall have the meanings expressly associated therewith, unless the context clearly dictates otherwise. As used herein, the phrases "in one embodiment" and "in some embodiments" do not necessarily refer to the same embodiment, although they may. Additionally, as used herein, the phrases "in another embodiment" and "in some other embodiments" do not necessarily refer to different embodiments, although they may. Thus, as described below, various embodiments of the present disclosure can be readily combined without departing from the scope or spirit of the present disclosure. Furthermore, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is contemplated that it is within the knowledge of one of ordinary skill in the art to achieve such feature, structure, or characteristic in connection with other embodiments, whether or not explicitly described herein.
[0038] The term "based on" is not exclusive and allows for the basis of additional unrecited factors unless the context clearly dictates otherwise. Additionally, throughout this specification, the meaning of the singular ("a," "an," and "the") includes plural references. The meaning of "in" includes "into" and "on."
[0039] It is understood that at least one aspect / function of the various embodiments described herein can be performed in real time and / or dynamically. As used herein, the term “real time” refers to an event / action that can occur instantaneously or near-instantaneously while another event / action is occurring. For example, “real-time processing,” “real-time computation,” and “real-time execution” all relate to the performance of a computation during the actual time that an associated physical process (e.g., a user interacting with an application on a mobile device) is occurring, and the results of the computation can be used to guide the physical process.
[0040] As used herein, the term "dynamically" means that events and / or actions can be triggered and / or occur without human intervention. In some embodiments, events and / or actions according to the present disclosure can occur in real time and / or based on a predetermined periodicity of at least one of nanoseconds, nanoseconds, microseconds, microseconds, milliseconds, milliseconds, seconds, seconds, minutes, minutes, an hour, hours, daily, daily, weekly, monthly, etc.
[0041] As used herein, the term "runtime" corresponds to any behavior that is dynamically determined during the execution of a software application or at least a portion of a software application.
[0042] In some embodiments, the specially programmed computing system of the present invention with associated devices operates in a distributed network environment, communicating over a suitable data communications network (e.g., the Internet, etc.) and utilizing at least one suitable data communications protocol (e.g., IPX / SPX, X.25, AX.25, AppleTalk®, TCP / IP (e.g., HTTP), etc.). Of note, the embodiments described herein may, of course, be implemented using any suitable hardware and / or computing software language. In this regard, those skilled in the art are familiar with the types of computer hardware that may be used, the types of computer programming techniques that may be used (e.g., object-oriented programming), and the types of computer programming languages that may be used (e.g., C++, Objective-C, Swift, Java, Javascript). The above examples are, of course, illustrative and not limiting.
[0043] As used herein, the terms "image" and "image data" are used interchangeably to identify data representing visual content, including, but not limited to, images encoded in various computer formats (e.g., ".jpg", ".bmp", etc.), streaming video based on various protocols (e.g., Real Time Streaming Protocol (RTSP), Real Time Transport Protocol (RTP), Real Time Transport Control Protocol (RTCP), etc.), recorded / generated non-streaming video in various formats (e.g., ".mov", ".mpg", ".wmv", ".avi", ".flv", etc.), and real-time visual images captured via a camera application on a mobile device.
[0044] The materials disclosed herein may be implemented in software or firmware, or a combination thereof, or as instructions stored on a machine-readable medium that can be read and executed by at least one processor. A machine-readable medium may include any medium and / or mechanism for storing or transmitting information in a form readable by a machine (e.g., a computing device). For example, a machine-readable medium may include read-only memory (ROM), random-access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, electrical, optical, acoustic, or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.), and the like.
[0045] In another form, a non-transitory article, such as a non-transitory computer-readable medium, may be used in conjunction with any of the above examples or other examples, except that it does not include the transitory signal itself. It includes elements other than the signal itself that can temporarily hold data in a "transitory" manner, such as RAM.
[0046] As used herein, the terms "computer engine" and "engine" identify at least one software component and / or a combination of at least one software component and at least one hardware component that is designed / programmed / configured to manage / control other software and / or hardware components (e.g., libraries, software development kits (SDKs), objects, etc.).
[0047] Examples of hardware elements may include devices, logic devices, components, processors, microprocessors, circuits, processors, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, application specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), memory units, logic gates, registers, semiconductor devices, chips, microchips, chipsets, etc. In some embodiments, the at least one processor may be implemented as a complex instruction set computer (CISC) or reduced instruction set computer (RISC) processor, an x86 instruction set compatible processor, a multi-core, or any other microprocessor, graphics processing unit (GPU), or central processing unit (CPU). In various implementations, the at least one processor may be a dual-core processor, a dual-core mobile processor, etc.
[0048] Examples of software may include a software component, program, application, computer program, application program, system program, machine program, operating system software, middleware, firmware, software module, routine, subroutine, function, method, procedure, software interface, application program interface (API), instruction set, computing code, computer code, code segment, computer code segment, word, value, symbol, or any combination thereof. The determination of whether an embodiment is implemented using hardware and / or software elements may vary according to any number of factors, such as desired computational speed, power level, thermal tolerance, processing cycle budget, input data rate, output data rate, memory resources, data bus speed, and other design or performance constraints.
[0049] One or more aspects of at least one embodiment may be embodied by representative instructions stored on a machine-readable medium that represent various logic within a processor, and that, when read by a machine, cause the machine to create logic for performing the techniques described herein. Such representations, known as "IP cores," may be stored on tangible machine-readable media and supplied to various customers or manufacturing facilities for loading into manufacturing machines that actually configure the logic or processor.
[0050] As used herein, the term "user" shall mean at least one user.
[0051] In some embodiments, the exemplary computing device of the present invention may be utilized for a variety of goals, including, but not limited to, computer vision, recommended positioning based on images from a mobile device camera, and applications. In some embodiments, the exemplary camera may be either a video camera or a digital still camera. In some embodiments, the exemplary camera may operate based on transmitting analog and / or digital signals to at least one storage device that may reside in at least one of a desktop computer, a laptop computer, or an output device such as, but not limited to, VR glasses, lenses, a smartwatch, and / or a mobile device screen.
[0052] In some embodiments, as detailed herein, an exemplary inventive computing device of the present disclosure may be programmed / configured to process visual feeds from various visual recording devices (e.g., without limitation, mobile device cameras, computer cameras, display screens, or any other cameras for similar purposes) to detect, recognize, and track (e.g., in real time) inspection devices appearing in the visual recordings at one time and / or over a period of time. In some embodiments, the exemplary inventive computing device may be utilized for various purposes, such as, but not limited to, computer vision, positioning-based recommendation instructions in applications, etc. In some embodiments, the exemplary camera may be either a video camera or a digital still camera. In some embodiments, the exemplary camera may operate based on transmitting analog and / or digital signals to at least one storage device, which may be present in at least one of a desktop computer, a laptop computer, or an output device, such as, but not limited to, VR glasses, lenses, a smartwatch, and / or a mobile device.
[0053] In some embodiments, the exemplary inventive process of detecting, recognizing, and tracking one or more inspection devices over time is independent of whether the visual recordings were obtained from the same or different recording devices at the same or different locations.
[0054] In some embodiments, an exemplary computing device of the present invention that detects, recognizes, and tracks one or more inspection areas of one or more inspection devices may rely on one or more centralized databases (e.g., data centers). For example, the exemplary computing device of the present invention may extract feature inputs for each inspection device that can be utilized to identify / recognize specific inspection areas or inspection results. In some embodiments, the exemplary computing device of the present invention may improve detection by training itself based on a single image or a collection of video frames or videos taken in different conditions for inspection result identification. In some embodiments, for example, in a mobile (e.g., smartphone) application configured / programmed to provide video communication capabilities with augmented reality elements, the exemplary computing device of the present invention may detect one or more inspection devices in real time based at least in part on a single frame or a series of frames without storing recognition state.
[0055] In some embodiments, the exemplary computing device of the present invention may utilize one or more techniques for identifying the inspection device and tracking the inspection area, as detailed herein, hi some embodiments, the exemplary computing device of the present invention may further utilize techniques that may enable identification of the inspection area within the image frame.
[0056] In some embodiments, the exemplary computing device of the present invention may support medical testing purposes such as, but not limited to, determining / estimating and tracking patient test results, obtaining feedback on testing devices, tracking patient status for medical and preventative purposes (e.g., disease monitoring), suitable applications in statistics, sociology, etc. In some embodiments, the exemplary computing device of the present invention may support entertainment and / or educational purposes. For example, exemplary electronic content in a mobile and / or computer-based application may be dynamically adjusted and / or triggered based at least in part on a detected testing device. In some embodiments, exemplary examples of such dynamically adjusted / triggered electronic content may be one or more of a visual mask and / or visual / audio effects (e.g., color, shape, size, etc.) that may be applied to a testing device image in an exemplary video stream. In some embodiments, the exemplary augmented content is dynamically generated by the exemplary computing device of the present invention and may consist of at least one of instructions, suggestions, facts, images, etc.
[0057] In some embodiments, the exemplary computing device of the present invention can utilize raw video or image (e.g., screenshot) input / data from any type of known camera, including both analog and digital.
[0058] In some embodiments, an exemplary computing device of the present invention may utilize deformable 3D inspection device images that can be trained to generate meta-parameters (e.g., but not limited to, coefficients defining the deviation of the inspection device from a mean shape, coefficients defining the inspection area and / or inspection results, camera position, and / or head position, etc.).
[0059] In some embodiments, the exemplary computing device of the present invention may be configured to identify test results based on a single frame as a baseline, with several frames being able to improve detection quality. In some embodiments, the exemplary computing device of the present invention may be configured to estimate more sophisticated patterns of test results, thus enabling the use of lower resolution cameras (e.g., mobile or web cameras).
[0060] In some embodiments, an exemplary inventive computing device or system may be directly connected to or operably remotely connected to an existing camera (e.g., mobile, computer-based, or otherwise). In some embodiments, an exemplary inventive computing device or system may include a specifically programmed inventive data processing module that acquires video input from one or more cameras. For example, the specifically programmed inventive data processing module may determine the source of the input video data and perform any necessary transcoding into a different format or any other appropriate adjustment so that the video input may be available for processing according to the principles of the present disclosure. In some embodiments, the input image data (e.g., input video data) may include any suitable type of source for video content, including a variety of video sources. In some embodiments, content from an input video (e.g., the video stream of FIG. 1) may include both video data and metadata. A single picture may be included in a frame. In some embodiments, the specifically programmed inventive data processing module may decode the video input in real time and separate it into frames. In some embodiments, the exemplary input video stream captured by an exemplary camera (e.g., the front camera of a mobile personal smartphone) may be segmented into frames. For example, a typical video sequence is an interleaved format of several camera shots, where a camera shot is a performance recorded consecutively with a given camera setting. Camera alignment, as used herein, can refer to the alignment of different cameras capturing video, images, or screen frames in a video or image sequence / stream. The concept of camera alignment is based on the cameras performing the reconstruction of video editing. A typical video sequence is an interleaved format of several camera shots, where a camera shot is a performance recorded consecutively with a given camera setting.By registering each camera from the incoming video frames, the original interleaved format can be separated into several sequences, each corresponding to a registered camera aligned to the original camera settings.
[0061] In some embodiments, a specifically programmed data processing module of the present invention may be configured to process each frame or series of frames using an appropriate test device detection algorithm. For example, if one or more test or control regions are detected in a frame, the specifically programmed data processing module of the present invention may extract a feature vector and store the extracted information in one or more databases. In some embodiments, an exemplary computer device of the present invention may include a specifically programmed test identification module of the present invention that can compare the extracted features with previous information stored in a database. In some embodiments, if the specifically programmed test identification module of the present invention determines a match, the new information is added to the existing data to increase the accuracy of further identification and / or improve the quality of the test result determination. In some embodiments, if corresponding data does not exist in the database, a new entry is created. In some embodiments, the resulting test result determination may be stored in a database for further analysis.
[0062] The use of lateral flow-based IV-RDTs in disease diagnosis has expanded significantly recently, highlighted by the need for large-scale population self-screening during the COVID-19 pandemic. However, a major drawback of this simple and economical solution for disease diagnosis and monitoring, particularly related to self-diagnosis, has been the need for a reliable platform for recording test results and communicating them to specialized healthcare providers / HMOs. To address this obstacle, several solutions for capturing IV-RDT results and their digital transfer have been developed. IV-RDT result capture can involve the use of a dedicated reader that scans the lateral flow membrane to identify colored or fluorescent marker lines. Test results are then communicated to the operator in a qualitative binary (positive / negative) mode.
[0063] However, identifying IV-RDT results with a camera is difficult due to the inability to detect test results. One technical issue is that image acquisition can depend on precise strip positioning relative to the mobile camera to optimize image capture. Such positioning is not monitored, and multiple images must be captured until a processable image is obtained. This can be difficult when the IV-RDT device cannot be fixed at a specific location in space (e.g., an IV-RDT that is not a cassette and therefore cannot be placed flat on a solid support).
[0064] Another technical issue is that image acquisition requires building specific templates for each type of IV-RDT that rely on external IV-RDT device features (e.g., cassette structure and specific additional patterns, coloring, background illumination, etc.) to guide the positioning of the mobile camera. Such features are not present or are limited in inspection devices that deviate from the flat cassette format.
[0065] Yet another technical problem is the difficulty of image analysis when image acquisition is not optimal: for example, the camera may be positioned too far from the object, making image analysis of blurred objects impractical.
[0066] Another technical issue is that shadows and reflections produced by different lighting positions and conditions, or those introduced by the user (e.g., blocking the light or casting shadows), can introduce false lines / features or obscure existing ones. This can lead to erroneous results obtained by the algorithm, which is expected to compromise the specificity and sensitivity of the test.
[0067] 1 illustrates one embodiment of a detection system 100 including a server 105 in communication with a client device 110 having a camera 115 and a display application 120 (e.g., a mobile application installed on the client device 110) for detecting an inspection area 130 having a visual indicator 135 that indicates an inspection result 140. The inspection area 130 may include an inspection area 130 that displays a visual indicator 135 that identifies the inspection result 140. As shown in FIG. 2, the inspection area 130 includes the inspection area 130. As shown in FIG. 3, the inspection area 130 includes the visual indicator 135.
[0068] The display application 120 can display an interface that interactively guides a user to appropriately position the camera 115 to generate a first image stream of the inspection device 125. In some embodiments, the first image stream includes a screenshot of the inspection device 125. The server 105 can include a neural network for live detection, recognition, and tracking of the inspection device 125. The server 105 can locate the inspection device 125 within the image stream and guide the user in real time on how to position the camera 115 for optimal recognition of the inspection area 130 to identify the inspection result 140. The server 105 can generate imaging instructions to adjust the camera 115 to capture a second image stream of the inspection area 130 of the inspection device 125. In some embodiments, the second image stream includes a screenshot of the inspection device 125. The neural network can analyze the second image stream and identify the inspection result 140 from the visual indicators 135. Using techniques such as backpropagation, transfer learning, and data modeling, the detection system 100 can learn, adapt, and change over time.
[0069] In some embodiments, server 105 is a standalone server that communicates with client device 110, which may be a user device, a portable device with a phone camera, or a mobile, web, or local device. In some embodiments, server 105 may include at least one processor. In some embodiments, client device 110 may include at least one processor.
[0070] The test device 125 may be a rapid diagnostic test (RDT) for in vitro diagnostics (IVD). The server 105 may visually analyze the test device 125 in real time by receiving an image stream of the test device 125 from the client device 110. The detection system 100 may be used to identify a test result 140 from a visual indicator 135 in the test area 130. In some embodiments, the test result 140 represented by the visual indicator 135 is binary (e.g., positive / negative). In some embodiments, the test result 140 represented by the visual indicator 135 is in semi-quantitative form (e.g., 1-10). In some embodiments, the test result 140 represented by the visual indicator 135 is in quantitative form (e.g., concentration, dose, level).
[0071] The server 105 may include a deep learning module 145 for classifying images from the client device 110. In some embodiments, the deep learning module 145 may be configured to utilize computer vision techniques. The deep learning module 145 may include, utilize, or be cloud-based AI computer vision (CV), computer vision (CV), artificial intelligence (AI), deep learning (DL), web applications, live video streams, or live analytics with CV. In some embodiments, the deep learning module 145 may be configured to utilize one or more exemplary AI / computer vision techniques selected from, but not limited to, decision trees, graph algorithms, boosting, support vector machines, neural networks, nearest neighbor algorithms, naive Bayes, bagging, and random forests. In some embodiments, optionally in combination with any of the embodiments described above or below, the exemplary neural network technique may be one of, but not limited to, a feedforward neural network, a radial basis function network, a recurrent neural network, a convolutional network (e.g., a U-net), or other suitable network. In some embodiments, optionally in combination with any of the embodiments described above or below, an exemplary implementation of a neural network may be performed as follows. i) Define the neural network architecture / model, ii) forwarding the input data to an exemplary neural network model; iii) incrementally training example models; iv) determining the accuracy of a certain number of time steps; v) applying the example trained model to process newly received input data; vi) Optionally, and in parallel, continue training the example trained model with a predetermined periodicity.
[0072] In some embodiments, optionally in combination with any of the above or below embodiments, the exemplary trained deep learning model may specify a neural network by at least a neural network topology, a set of activation functions, and connection weights. For example, the topology of the neural network may include the configuration of nodes in the neural network and the connections between such nodes. In some embodiments, optionally in combination with any of the above or below embodiments, the exemplary trained deep learning model may also be specified to include other parameters, including, but not limited to, a bias value / function and / or an aggregation function. For example, the activation function of a node may be a step function, a sine function, a continuous or piecewise linear function, a sigmoid function, a hyperbolic tangent function, a ReLU function, or any other type of mathematical function that represents a threshold at which the node is activated. In some embodiments, optionally in combination with any of the above or below embodiments, the exemplary aggregation function may be a mathematical function (e.g., sum, product, etc.) that combines input signals to the node. In some embodiments, optionally in combination with any of the above or below embodiments, the output of the exemplary aggregation function may be used as an input to the exemplary activation function. In some embodiments, optionally in combination with any of the embodiments described above or below, the bias may be a constant value or function that may be used by the aggregation function and / or activation function to make a node more or less likely to be activated.
[0073] In some embodiments, deep learning module 145 may include neural networks 146A-146N forming modular neural network 148 for classifying images from client device 110. In some embodiments, modular neural network 148 receives input from each of neural networks 146A-146N. Each of neural networks 146A-146N may be a convolutional neural network (CNN). In some embodiments, neural networks 146A-146N may use a combination of CNN elements to convert images / video from an image stream into inputs (e.g., nodes) for deep learning techniques, and feed the image inputs into a deep neural network (DNN) to identify inspection results 140. In some embodiments, deep learning module 145 may be a trained multi-agent system (MAS) including neural networks 146A-146N and modular neural network 148 for computer recognition task decomposition.
[0074] The server 105 may include a data store 150 for storing a corpus of reference images 152 of the inspection device 125 and a corpus of reference images 154 of the inspection area 130 .
[0075] The server 105 may include a generator 155 for generating at least one imaging instruction for the client device 110 to capture an image of the inspection device 125. The deep learning module 145 may identify the inspection result 140 generated by the inspection device 125 based on the visual indicator 135.
[0076] 4, a method 400 for a rapid diagnostic test result interpretation platform using computer vision is shown. In some embodiments, the display application 120 can display interactive instructions for a user to complete a task, such as using the test device 125 to generate a test result 140. For example, the instructions can be similar to administering a COVID-19 antigen test.
[0077] The method 400 may include the server 105 receiving a first image stream of the inspection device 125 (step 402). The deep learning module 145 may establish a connection with the client device 110 to receive the image stream from the client device 110. In some embodiments, the client device 110 may connect to the cloud-based server 105 in real time via a web socket. This socket may be established to facilitate communication between the display application 120 (e.g., a client-side model (MAR)) and the server 105 (e.g., a server-side model (CNN)) after the inspection device 125 is detected. For example, the deep learning module 145 may open a live port to an optical flow converter to optimize data flow. The deep learning module 145 may establish a request to open a live camera broadcast from the client device 110 that sends the image stream to the server 105. In some embodiments, the detection system 100 may use an API call to open a live socket to the camera 115 that broadcasts the image stream from an end user that sends a request to the detection system 100. A live socket can establish a connection for a video feed from a client device 110 to a cloud-based deep learning AI.
[0078] In some embodiments, the deep learning module 145 can send user identification information to the client device 110 including the camera 115 (e.g., a code, password, or token for the user device). For example, the deep learning module 145 can provide the user with ID for documentation. In some embodiments, the deep learning module 145 can receive user verification, such as a token for complying with privacy regulations, to capture the image stream.
[0079] In some embodiments, deep learning module 145 can establish a connection with client device 110 in response to receiving user identification information (e.g., a token). Deep learning module 145 can use the established connection via a socket so that a live stream is established from client device 110 to deep learning module 145. For example, deep learning module 145 can establish a connection with display application 120. In some embodiments, display application 120 can be a web application, a native application, an API, an SDK, a container (Docker), etc.
[0080] In some embodiments, the deep learning module 145 can receive the first and second imaging streams over a connection. For example, the deep learning module 145 can obtain a live video stream from the camera 115 of the client device 110. The deep learning module 145 can communicate with the client device 110. For example, the deep learning module 145 can use an open socket with a user identifier (e.g., a user ID) to a cloud-based server. In some embodiments, the deep learning module 145 or the display application 120 can apply optical flow methods (such as sparse feature propagation or metadata keyframe extraction) to the video streams to minimize data transactions and achieve optimal recognition times for the inspection device 125 and the inspection area 130.
[0081] The method 400 may include the server 105 identifying the inspection device 125 in the first image stream (step 404). Referring now to FIGS. 5A and 5B, images of the first image stream of the inspection device 125 are shown. In some embodiments, the deep learning module 145 may receive the first image stream of the inspection device 125 from the client device 110. The deep learning module 145 may identify the inspection device 125 in the first image stream. In some embodiments, as shown in FIGS. 5A and 5B, the deep learning module 145 may cause the display application 120 to display or highlight the inspection device 125 and the inspection region 130 in the first image stream. For example, as shown in FIG. 5B, the display application 120 may display the text “Device” after identifying the inspection device 125 in the first image stream.
[0082] Deep learning module 145 can use neural networks 146A-146N to recognize, track, or monitor inspection device 125 and its location within the image stream. In some embodiments, neural networks 146A-146N can use a combination of CNN elements to convert images / video from the first image stream into inputs (e.g., nodes) for deep learning techniques and feed the image inputs into a deep neural network (DNN) of deep learning module 145 to identify inspection device 125. For example, modular neural network 148 can receive inputs from trained neural networks 146A-146N to allow data collected from the first image stream via optical flow methods to be inserted into the modular neural network's 148's associated task. Modular neural network 148 can process the inputs through one or more of its task-driven DNNs, passing information between them. In some embodiments, the deep learning module 145 can apply optical flow methods (sparse feature propagation, metadata keyframe extraction, etc.) to the image stream to minimize data transactions and achieve optimal (live) recognition times for the inspection device 125 and inspection area 130.
[0083] Deep learning module 145 can classify images in the first image stream to identify inspection device 125. Neural networks 146A-146N can receive and analyze images of inspection device 125. In some embodiments, neural networks 146A-146N can be convolutional neural networks that are part of a modular neural network 148 that identifies inspection device 125 in at least one first image in the first image stream, as shown in FIGS. 7, 8A, 8B, 9, 10, and 11.
[0084] After recognizing the inspection device 125, the neural networks 146A-146N can track the inspection device 125, and the modular neural network 148 can provide the live position (e.g., x, y, z coordinates) of the inspection device 125 to the display application 120. In some embodiments, the display application 120 can include functionality in the deep learning module 145 to track the inspection device 125 (e.g., including an IV-RDT test) while the deep learning module 145 analyzes the image stream of the generator 155 to generate and provide imaging instructions for display by the display application 120 to guide the user to scan the test by marking the inspection device 125 and the test area 130 on a screen controlled by the display application 120, as shown in FIGS. 5A and 5B .
[0085] The display application 120 can guide the user in real time to move the inspection device 125 or camera 115, tracking and highlighting the inspection device 125 so that the inspection area 130 and visual indicators 135 can be recognized by the server 105. The deep learning module 145 can guide the user in "real time" using an interactive visual interface, text, and voice to achieve the best angle and position of objects in space and within the object to obtain the best possible view of the inspection area 130. The neural networks 146A-146N can be updated and receive input from other neural networks 146A-146N as a feedback loop to form a recurrent neural network (RNN) to retrain the neural networks 146A-146N and make corrections during the process of analyzing the image stream.
[0086] As shown in FIG. 6 , in some embodiments, the deep learning module 145 can identify multiple device features of the inspection device 125 in the first image stream. In some embodiments, the multiple device features are multiple structural features of the inspection device 125. For example, the deep learning module 145 can identify points, edges, or objects that make up the inspection device 125. For example, the deep learning module 145 can identify two features of the inspection device 125 to identify the inspection device 125 with a base level of confidence. In another example, the deep learning module 145 can identify three features of the inspection device 125 to identify an inspection device 125 with a higher confidence level. In yet another example, the deep learning module 145 can identify four features of the inspection device 125 to identify an inspection device 125 with an even higher confidence level.
[0087] In some embodiments, the corpus of reference images 152 can be images of the inspection device 125 for comparison by the neural networks 146A-146N to identify the inspection device 125 in the first image stream. In some embodiments, the corpus of reference images 154 can be images of the inspection area 130 for comparison by the neural networks 146A-146N to identify the inspection area 130 in the second image stream. In some embodiments, the deep learning module 145 can classify at least one first image in the first image stream to at least one first reference image in the corpus of reference images 152 of the inspection device 125 based on a plurality of device features of the inspection device 125.
[0088] The method 400 can include the server 105 selecting imaging instructions to coordinate image capture of the inspection device 125 (step 406). The server 105 can include a generator 155 for generating at least one imaging instruction for the client device 110 to capture an image of the inspection area 130 of the inspection device 125. The imaging instructions can be selected to image the inspection area 130 by optimizing an image stream of the inspection area 130 using the client device 110. The imaging instructions can maximize the resolution of the inspection area 130 by improving focus and lighting at the location of the inspection area 130. For example, the imaging instructions can cause the client device 110 to turn on a flash to improve lighting, implement a filter on the image, or change the focus of the camera 115. The imaging instructions can create optimal conditions for the deep learning module 145 to record a second image stream of the inspection area 130.
[0089] In some embodiments, the generator 155 can select at least one imaging instruction for the client device 110 to capture a second image stream based on the at least one first reference image. The second image stream can include the inspection area 130. For example, the imaging instruction can cause the display application 120 to display an interface that guides the user on the best angle and position of the camera 115 relative to the inspection device 125 within the space and within the object to obtain the best possible second image stream of the inspection area 130.
[0090] In some embodiments, the generator 155 can generate at least one imaging instruction for the camera 115 to capture a second image stream of the inspection area 130 based on the imaging instruction metadata associated with the at least one first reference image. In some embodiments, the deep learning module 145 can identify at least one candidate imaging instruction. For example, the generator 155 can maintain the candidate imaging instruction in the data store 150. The candidate imaging instruction can be a default instruction such as twisting the inspection device 125, turning on a light, zooming in or out, moving up or down, or requesting the user to adjust the focus of the camera 115.
[0091] In some embodiments, at least one convolutional neural network 146A-146N can select at least one imaging instruction from the at least one candidate imaging instruction. For example, generator 155 can use a deep learning (ML) algorithm to analyze the image stream and generate instructions for display by display application 120 on a user interface.
[0092] In some embodiments, generator 155 can instruct client device 110 to implement at least one imaging instruction to automatically generate the second image stream. In some embodiments, generator 155 can instruct client device 110 directly (e.g., to change the focus of camera 115). In some embodiments, to implement the at least one imaging instruction, generator 155 can send at least one imaging instruction to client device 110 to cause camera 115 to automatically generate the second image stream based on the at least one imaging instruction.
[0093] 12A, 12B, and 12C, in some embodiments, the generator 155 can instruct a user of the client device 110 (e.g., request the user to reposition the client device 110 for a different view of the inspection device 125). In some embodiments, the generator 155 can send at least one imaging instruction instruction to the client device 110, including the camera 115, including augmented reality instructions instructing the user to reposition the camera so that the display application 120 overlays the first image stream to capture the inspection device. The augmented reality instructions can indicate how the user can reposition the client device 110 to image the inspection area 130. For example, the display application 120 can display a second image stream to the user on the display of the client device 110. The augmented reality instructions can be overlaid on the inspection device 125 to indicate how the user can reposition the client device 110 to image the inspection area 130. As shown in FIGS. 12A-12C, the augmented reality instructions may request the user to rotate the inspection device 125 to bring the inspection area 130 into the field of view of the camera 115.
[0094] The method 400 may include the server 105 receiving a second image stream of the inspection area 130 of the inspection device 125 (step 408). Referring now to FIG. 13, an image of the second image stream of the inspection device 125 is shown. In some embodiments, the deep learning module 145 may receive the second image stream from the client device 110 coordinated with the at least one image capture instruction. Once the camera 115 is positioned based on the at least one image capture instruction, a deep learning algorithm of the deep learning module 145 analyzes keyframes provided by the client device 110. The deep learning module 145 may perform this analysis until a sufficient number of images have been collected to identify the inspection result 140.
[0095] The method 400 may include the server 105 identifying an inspection area 130 in the second image stream (step 410). The deep learning module 145 may classify the images in the second image stream to identify the inspection area 130 and the visual indicator 135. The neural networks 146A-146N may receive and analyze the images in the second image stream using the techniques described herein, including identifying the inspection area 130 including the visual indicator 135 and identifying the inspection result 140 itself. In some embodiments, as shown in FIGS. 7, 8A, 8B, 9, 10, and 11, the neural networks 146A-146N may be convolutional neural networks that are part of a modular neural network 148 that identifies the inspection area 130 in at least one first image in the second image stream.
[0096] The deep learning module 145 can identify the inspection area 130 in the second image stream. In some embodiments, the deep learning module 145 can identify multiple device features of the inspection area 130 in the second image stream. In some embodiments, the multiple device features of the inspection area 130 are multiple structural features of the inspection area 130. For example, the deep learning module 145 can identify points, edges, or objects that make up the inspection area 130. In some embodiments, the deep learning module 145 can identify multiple features that make up the inspection area 130. For example, the deep learning module 145 can identify two features of the inspection area 130 to identify the inspection area 130 with a base confidence. In another example, the deep learning module 145 can identify three features of the inspection area 130 to identify the inspection area 130 with a higher confidence level. In yet another example, the deep learning module 145 can identify four features of the inspection area 130 to identify the inspection area 130 with an even higher confidence.
[0097] The deep learning module 145 can use neural networks 146A-146N to recognize, track, or monitor the inspection area 130 and its position relative to the inspection device 125. In some embodiments, the deep learning module 145 can use a combination of convolutional neural networks 146A-146N to identify the inspection device 125. In some embodiments, the deep learning module 145 can classify at least one second image in the second image stream to at least one second reference image in a corpus of reference images 154 of the inspection area 130 based on multiple device features of the inspection area 130. For example, the deep learning module 145 can apply neural networks, algorithms, and image recognition methods to analyze the second image stream including the inspection area 130 to identify the visual indicator 135.
[0098] In some embodiments, the deep learning module 145 may cause the display application 120 to display or highlight the inspection device 125 and the inspection region 130 in the second image stream, as shown in Figure 13. For example, as shown in Figure 13, the display application 120 may display the text "Strip Window" after identifying the inspection device 125 in the first image stream.
[0099] In some embodiments, if the deep learning module 145 is unable to identify the inspection area 130 in the second image stream, the modular neural network 148 can generate at least one imaging instruction including camera settings to apply and adjust autofocus for the camera 115. The neural networks 146A-146N can maintain a feedback loop forming a recurrent neural network (RNN) to optimize camera sensitivity and lighting settings. The generator 155 can monitor as camera frames are collected, analyzed, and filtered to select frames that optimize the quality of the images that the deep learning module 145 is using to identify the inspection result 140. In some embodiments, the generator 155 can generate additional imaging instructions as described in step 406.
[0100] The method 400 may include the server 105 identifying an inspection result 140 in the inspection area 130 of the inspection device 125 (step 412). In some embodiments, the deep learning module 145 may identify the inspection result 140 based on at least one second reference image and based on a visual indicator 135 in the inspection area 130 in at least one of the second images. For example, the deep learning module 145 may identify a visual indicator in the inspection area 130 in the second image.
[0101] In some embodiments, deep learning module 145 may cause display application 120 to display test result 140. For example, deep learning module 145 may cause display application 120 to display test result 140 as a QR code. In another example, deep learning module 145 may cause display application 120 to display test result 140 as a number or indicator (e.g., positive or negative).
[0102] In some embodiments, the deep learning module 145 can store the test results 140. For example, the deep learning module 145 can document the test results 140 and associate the test results with an identifier, such as a user or patient identifier. In some embodiments, the deep learning module 145 can store the test results 140 in a data store 150.
[0103] In some embodiments, the detection system 100 can use mobile computer vision and cloud-based AI (with a web application UX / UI) for live diagnosis of IV-RDTs. The detection system 100 can include an SDK platform for easy interaction with medical record systems and communication with specialists and HMOs. In some embodiments, the deep learning module 145 can send the test results 140 to an administrator device for display. For example, the administrator device can be an HMO, and the deep learning module 145 can send the test results 140 to the HMO and update the user's or patient's medical file. The server 105 can return the test results 140 to the user, operator, HMO, or physician for further review, recording, and medical treatment / follow-up decisions. For example, the server can provide the test results 140 as a response to the user and / or HMO / physician. The detection system 100 can close the loop between the user of the test device 125 and the HMO / physician for recording the test results 140 and providing medical treatment / follow-up.
[0104] In some embodiments, if the deep learning module 145 is unable to identify the inspection result 140 from the visual indicator 135 within the inspection area 130 of the second image stream, the generator 155 may generate additional imaging instructions as described in step 406.
[0105] From the synergistic combination of these technologies and platforms, the detection system 100 is able to overcome numerous technical challenges and includes numerous technical solutions.
[0106] One technical solution is for the display application 120 to include a pre-trained 3D object detector and tracker MAR (Mobile Augmented Reality) model that detects and tracks the inspection device 125 by using the RAM of the client device 110, and for the GPU of the server 105 to analyze the image stream and generate imaging instructions. By splitting the processing between the server 105 and the client device 110, the embodiments described herein enable less memory usage while ensuring faster processing times, as the inspection device 125 is tracked by the nearby client device 110 while neural networks 146A-146N forming a modular neural network 148 recursively analyze the image stream to generate imaging instructions that identify the inspection result 140.
[0107] Another technical solution of the detection system 100 is to create the display application 120 as a secure web app with a user-friendly UX / UI that accesses the camera 115. Another technical solution of the detection system 100 is the display application 120 that reduces the need to install complex applications on mobile devices. Another technical solution of the detection system 100 is that the display application 120 can be easily deployed and used on any smart client device 110.
[0108] Another technical solution is that the server 105 can use PWA (Progressive Web Applications) with the development of cloud servers to improve the accuracy of deep learning models and increase the speed of ML algorithms. The detection system 100 can use a CV algorithm with a trained (task-specific) CNN to locate sub-objects (SOs), which may be the inspection area 130, and analyze them in real time to meet specific diagnostic task requirements. Another technical solution of the detection system 100 is to reduce the need for image capture by the client device 125, which can overcome issues regarding proper positioning of the camera 115 for effective image capture.
[0109] Another technical solution of the detection system 100 is to manage an SDK with different API calls and support a dynamically guided, user-friendly application / interface. Another technical solution of the detection system 100 is to develop a user-side ML model for fast and efficient data compression and minimization for improved socket to cloud-server communication. Another technical solution of the detection system 100 is to develop a deep learning module computer vision technology as a MAS (multi-agent system) with an MNN architecture that combines CNN, RNN, and NN for different tasks and manages and manipulates them to support multiple models. Another technical solution of the detection system 100 is to develop an efficient training protocol for different CNN models using data collection and annotation to enhance model training. Another technical solution of the detection system 100 is to develop a responsive application for a "real-time" user experience. Yet another technical solution of the detection system 100 is to connect to different clients (such as HMOs) with minimal perturbations. Another technical solution of the detection system 100 is quality control to optimize and maintain the specificity and sensitivity of the test. Another technical solution for the detection system 100 is to approve these tools for clinical use through different regulatory authorities.
[0110] Another technical solution of the detection system 100 is to increase inspection accuracy by using deep learning AI algorithms over applications that use only deep learning models. Another technical solution of the detection system 100 is to reduce the need to model each inspection shape, making it easier to recognize complex geometric shapes. Another technical solution of the detection system 100 is to improve the accuracy of adaptive diagnostic inspections over time by providing semi-quantitative or quantitative inspection results. Another technical solution of the detection system 100 is to improve the overall UX / UI of the application through live guidance of a VR model that communicates with the AI model. Another technical solution of the detection system 100 is to support multiple users with live streaming via a Linux cloud-based server application. Another technical solution of the detection system 100 is to simplify switching between tasks, objects, and / or SOs. Another technical solution of the detection system 100 is to establish security between the server 105 and the client device 110 through a gateway connection to a secure, different backend server / proxy.
[0111] 14, in some embodiments, a system 1400 is shown that includes a client device 110 that includes a camera 115 and an application 1410 that includes a display application 120, a deep learning module 145, and a generator 155. The system 1400 can be used to visually analyze an in-vitro rapid diagnostic test (IV-RDT) device using computer vision (CV) and mobile augmented reality (MAR) running on the client-based device. For example, the client device 110 can be hardware, such as a configured mobile phone or smart camera, that runs the application 1410, which can be a software application that performs the functions of the display application 120, the deep learning module 145, and the generator 155. In some embodiments, the application 1410 can be a web application, a native application, an API, an SDK, a container (Docker), or the like. In some embodiments, the deep learning module 145 is part of the application 1410 such that the mobile device can apply a neural network 146 to the image stream generated by the camera 115, forming a modular neural network 148. In some embodiments, client device 110 can be communicatively coupled to data store 150. For example, data store 150 can be hosted on a server or on the cloud. In some embodiments, client device 110 can include any or all of the components 115-155 described herein, including data store 150, such that client device 110 can perform the functions described herein while offline.
[0112] 15, a method 1500 for a rapid diagnostic test result interpretation platform using computer vision is shown. Method 1500 can include an application 1410 capturing a first image stream of a test device 125 (step 1502). In some embodiments, display application 120 can display interactive instructions for a user to complete a task, such as using test device 125 to generate test result 140. For example, the instructions can be similar to administering a COVID-19 antigen test.
[0113] The method 1500 may include the application 1410 identifying the inspection device 125 in the first image stream (step 1504). Referring now to FIGS. 5A and 5B, images of the first image stream of the inspection device 125 are shown. In some embodiments, the application 1410 may capture the first image stream of the inspection device 125. The application 1410 may identify the inspection device 125 in the first image stream. In some embodiments, as shown in FIGS. 5A and 5B, the display application 120 may display or highlight the inspection device 125 and the inspection area 130 in the first image stream. For example, as shown in FIG. 5B, the display application 120 may display the text "Device" after identifying the inspection device 125 in the first image stream.
[0114] In some embodiments, the application 1410 can apply a deep learning module 145 to classify images to identify the inspection area 130 and the visual indicators 135. For example, the application 1410 can apply optical flow methods (sparse feature propagation, metadata keyframe extraction, etc.) to the image stream to minimize data transactions and achieve optimal (live) recognition times for the inspection device 125 and the inspection area 130.
[0115] The application 1410 can use neural networks 146A-146N to recognize, track, or monitor the inspection device 125 and its position within the image stream. Each model can include at least one convolutional neural network or other deep learning algorithm. For example, each section of the MNN can include a convolutional neural network (CNN) trained to allow data collected from the first image stream via optical flow techniques to be inserted into the associated model task. In some embodiments, the application 1410 can identify the inspection device 125 using a combination of CNN elements.
[0116] In some embodiments, application 1410 can include neural networks 146A-146N forming a modular neural network 148 that analyzes images of inspection device 125 using techniques described herein, including identifying inspection area 130 including visual indicator 135 and identifying inspection result 140 itself. In some embodiments, deep learning module 145 includes at least one modular neural network 148. For example, application 1410 can include neural networks 146A-146N and a modular neural network 148 for computer recognition task segmentation.
[0117] In some embodiments, modular neural network 148 receives input from each of neural networks 146A-146N. In some embodiments, at least one neural network 146A-146N is part of modular neural network 148 that identifies inspection device 125 in at least one first image in the first image stream, as shown in Figures 7, 8A, 8B, 9, 10, and 11. Each of neural networks 146A-146N may be a convolutional neural network (CNN). In some embodiments, neural networks 146A-146N may identify inspection device 125 using a combination of CNN elements.
[0118] After recognizing the inspection device 125, the CNN of the application 1410 can track the inspection device 125 and provide the inspection device 125's live location (e.g., x, y, z coordinates). The application 1410 can track and highlight the inspection device 125 to guide the user in real time to move the inspection device 125 or camera 115 so that the inspection area 130 and visual indicators 135 can be recognized by the deep learning module 145. The application 1410 can guide the user in "real time" with an interactive visual interface, text, and voice to achieve the best angle and position of the object in space and within the object to obtain the best possible view of the inspection area 130. The application 1410 can receive input from other neural networks 146A-146N as a feedback loop to update and form a recurrent neural network (RNN) to retrain the neural networks 146A-146N to make corrections as the image stream is received.
[0119] 6 , in some embodiments, the application 1410 can identify multiple device characteristics of the inspection device 125 in the first image stream. For example, the application 1410 can identify features that make up the inspection device 125. In some embodiments, the multiple device characteristics are structural features. In some embodiments, the application 1410 can send the multiple device characteristics of the inspection device 125 to the data store 150. In some embodiments, the application 1410 can store the multiple device characteristics of the inspection device 125 on local storage of the client device 110.
[0120] In some embodiments, the application 1410 can access a corpus of reference images 152 and 154 stored locally on the client device 110. In some embodiments, the application 1410 can establish a connection with the data store 150 to access the corpus of reference images 152 and 154. For example, the data store 150 can open a live port to an optical flow converter to optimize data flow. In some embodiments, the application 1410 can receive user identification information from the data store 150. In some embodiments, the application 1410 can receive user verification, such as from a regulatory agency / entity, to ingest the image stream. The application 1410 can use the connection established via a socket, resulting in a live stream from the application 1410 to the data store 150. For example, the data store 150 can establish a connection with the application 1410.
[0121] The method 1500 can include the application 1410 selecting imaging instructions to coordinate the capture of images of the inspection device 125 (step 1506). The application 1410 can generate at least one imaging instruction that causes the application 1410 to capture images of the inspection area 130 of the inspection device 125. The imaging instructions can be selected to image the inspection area 130 by using the application 1410 to optimize an image stream of the inspection area 130. The imaging instructions can maximize the resolution of the inspection area 130 by improving focus and lighting at the location of the inspection area 130. For example, the imaging instructions can cause the application 1410 to turn on (e.g., light) a flash to improve lighting, implement a filter on the image, or change the focus of the camera 115. The imaging instructions can create optimal conditions for the application 1410 to record a second image stream of the inspection area 130.
[0122] In some embodiments, the application 1410 can select at least one imaging instruction that causes the application 1410 to capture a second image stream based on the at least one first reference image. The second image stream can include the inspection area 130. For example, the imaging instruction can cause the application 1410 to display an interface that guides the user on the best angle and position of the camera 115 relative to the inspection apparatus 125 in the space and within the object to obtain the best possible second image stream of the inspection area 130.
[0123] In some embodiments, the application 1410 can generate at least one imaging instruction command for the camera 115 to capture a second image stream of the inspection area 130 based on the imaging instruction metadata associated with the at least one first reference image. In some embodiments, the application 1410 can receive imaging instruction metadata associated with the at least one additional reference image from the data store 150. In some embodiments, the application 1410 can retrieve imaging instruction metadata associated with the at least one additional reference image from local storage of the system 1400.
[0124] In some embodiments, the application 1410 can identify at least one candidate imaging instruction using the deep learning module 145. For example, the application 1410 can maintain the candidate imaging instructions in local storage. In some embodiments, the application 1410 can receive the candidate imaging instructions from the data store 150. The candidate imaging instructions can be default instructions such as twisting the inspection device 125, turning on a light, zooming in or out, moving up or down, or requesting the user to adjust the focus of the camera 115.
[0125] In some embodiments, application 1410 can select at least one imaging instruction from the at least one candidate imaging instruction, for example, application 1410 can use a machine learning (ML) algorithm to analyze the image stream and generate instructions for display by display application 120 on a user interface.
[0126] In some embodiments, application 1410 can instruct camera 115 to perform at least one imaging instruction to automatically generate the second image stream. In some embodiments, application 1410 can instruct camera 115 directly (e.g., to change the focus of camera 115). In some embodiments, to perform the at least one imaging instruction, application 1410 executes at least one imaging instruction to cause camera 115 to automatically generate the second image stream based on the at least one imaging instruction.
[0127] 12A, 12B, and 12C, in some embodiments, the application 1410 can instruct the user (e.g., request the user to reposition the camera 115 for a different view of the inspection device 125). In some embodiments, the application 1410 can instruct the user to reposition the camera to capture the inspection device based on at least one imaging instruction, including augmented reality instructions, causing the application 1410 to overlay a first image stream. The augmented reality instructions can indicate how the user can reposition the camera 115 to image the inspection area 130. For example, the display application 120 can display a second image stream to the user on the display of the application 1410. The augmented reality instructions can be overlaid on the inspection device 125 to indicate how the user can reposition the camera 115 to image the inspection area 130. As shown in FIGS. 12A-12C, the augmented reality instructions can request the user to rotate the inspection device 125 to bring the inspection area 130 into the field of view of the camera 115.
[0128] The method 1500 may include the application 1410 capturing (step 1508) a second image stream of the inspection area 130 of the inspection device 125. Referring now to FIG. 13 , an image of the second image stream of the inspection device 125 is shown. In some embodiments, the application 1410 may receive the second image stream from the camera 115 adjusted with at least one image capture command. Once the camera 115 is positioned based on the at least one image capture command, a deep learning algorithm of the application 1410 analyzes keyframes provided by the camera 115. The deep learning module 145 may perform this analysis until a sufficient number of images have been collected to identify the inspection result 140.
[0129] The method 1500 may include an application 1410 that identifies an inspection area 130 in a second image stream (step 1510). A deep learning module 145 may classify images in the second image stream to identify the inspection area 130 and the visual indicator 135. Neural networks 146A-146N may receive and analyze images in the second image stream using techniques described herein, including identifying the inspection area 130, including the visual indicator 135, and identifying the inspection result 140 itself. In some embodiments, as shown in FIGS. 7, 8A, 8B, 9, 10, and 11, the neural networks 146A-146N may be convolutional neural networks that are part of a modular neural network 148 that identifies the inspection area 130 in at least one first image in the second image stream.
[0130] The application 1410 can identify the inspection area 130 in the second image stream. In some embodiments, the application 1410 can apply the deep learning module 145 to identify multiple device features of the inspection area 130 in the second image stream. In some embodiments, the application 1410 can send the multiple device features of the inspection area 130 to the data store 150. In some embodiments, the application 1410 can store the multiple device features of the inspection area 130 in local storage of the system 1400.
[0131] The application 1410 may use a neural network to recognize, track, or monitor the inspection area 130 and its position relative to the inspection device 125. In some embodiments, the application 1410 may use a combination of CNN elements to identify the inspection device 125.
[0132] In some embodiments, the application 1410 may apply the deep learning module 145 to classify at least one second image in the second image stream to at least one second reference image in the corpus of reference images 154 of the inspection area 130 based on a plurality of device features of the inspection area 130. For example, the application 1410 may apply neural networks, algorithms, and image recognition methods to analyze the second image stream including the inspection area 130 to identify the visual indicators 135.
[0133] In some embodiments, application 1410 may cause display application 120 to display or highlight inspection device 125 and inspection area 130 in the second image stream, as shown in Figure 13. For example, as shown in Figure 13, display application 120 may display the text "Strip Window" after identifying inspection device 125 in the first image stream.
[0134] In some embodiments, if application 1410 is unable to identify inspection area 130 in the second image stream, modular neural network 148 can generate at least one imaging instruction including camera settings to apply and adjust autofocus for camera 115. Neural networks 146A-146N can maintain a feedback loop forming a recurrent neural network (RNN) to optimize camera sensitivity and lighting settings. Application 1410 can monitor as camera frames are collected, analyzed, and filtered to select frames that optimize the quality of the images used by application 1410 to identify inspection result 140. In some embodiments, application 1410 can generate additional imaging instructions as described in step 1506.
[0135] The method 1500 may include an application 1410 that identifies an inspection result 140 in an inspection area 130 of the inspection device 125 (step 1512). In some embodiments, the application 1410 may identify the inspection result 140 based on at least one second reference image and based on a visual indicator 135 in the inspection area 130 in at least one of the second images. For example, the application 1410 may identify a visual indicator in the inspection area 130 in the second image.
[0136] In some embodiments, application 1410 can cause test results 140 to be displayed on display application 120. For example, application 1410 can display test results 140 as a QR code. In another example, application 1410 can display test results 140 as a number or indicator (e.g., positive or negative).
[0137] In some embodiments, the application 1410 can store the test results 140. For example, the application 1410 can document the test results 140 and associate the test results with an identifier, such as a user or patient identifier. In some embodiments, the application 1410 can store the test results 140 in a data store 150. In some embodiments, the application 1410 can send the test results 140 to an administrator device for display. For example, the administrator device can be an HMO, and the application 1410 can send the test results 140 to the HMO and update the user's or patient's medical file.
[0138] In some embodiments, if the application 1410 is unable to identify the inspection result 140 from the visual indicator 135 in the inspection area 130 of the second image stream, the application 1410 may generate additional imaging instructions as described in step 1506.
[0139] From the synergistic combination of these technologies and platforms, the system 1400 is able to overcome numerous technical challenges and includes numerous technical solutions.
[0140] One technical solution is that the server 105 can use PWA (Progressive Web Applications) with the development of cloud servers to improve the accuracy of deep learning models and increase the speed of ML algorithms. The detection system 100 can use the CV algorithm with trained (task-specific) CNNs to locate sub-objects (SOs), which may be the inspection area 130, and analyze them in real time to meet specific diagnostic task requirements.
[0141] Another technical solution of the system 1400 is to manage SDKs with different API calls and support dynamically guided user-friendly applications / interfaces. Another technical solution of the system 1400 is to develop user-side ML models for fast and efficient data compression and minimization for improved socket to cloud-server communication. Another technical solution of the system 1400 is to develop computer vision technology for deep learning modules as a multi-agent system (MAS) with an MNN architecture that combines CNNs, RNNs, and NNs for different tasks and manages and manipulates them to support multiple models. Another technical solution of the system 1400 is to develop efficient training protocols for different CNN models using data collection and annotation to enhance model training. Another technical solution of the system 1400 is to develop responsive applications for a "real-time" user experience. Yet another technical solution of the system 1400 is to connect to different clients (such as HMOs) with minimal perturbations. Another technical solution of the system 1400 is quality control to optimize and maintain test specificity and sensitivity. Another technical solution for the system 1400 is to approve these tools for clinical use through different regulatory authorities.
[0142] Another technical solution of the system 1400 is the display application 120 reducing the need to install complex applications on mobile devices. Another technical solution of the system 1400 is creating the display application 120 as a secure web app with a user-friendly UX / UI that accesses the camera 115. Another technical solution of the system 1400 is improving the overall UX / UI of the application through live guidance of a VR model that communicates with an AI model. Another technical solution of the system 1400 is being able to easily deploy and use the display application 120 on any smart client device 110.
[0143] Another technical solution of the system 1400 is to reduce the need for image capture by the client device 125 and overcome the problem of proper positioning of the camera 115 for effective image capture. Another technical solution of the system 1400 is to be able to establish security between the server 105 and the client device 110 via a gateway connection to a secure and different backend server / proxy.
[0144] Another technical solution of system 1400 is to increase inspection accuracy by using deep learning AI algorithms over applications that use only deep learning models. Another technical solution of system 1400 is to reduce the need to model each inspection shape, making it easier to recognize complex geometric shapes. Another technical solution of system 1400 is to improve the accuracy of adaptive diagnostic inspections over time through the ability to obtain semi-quantitative or even quantitative inspection results. Another technical solution of system 1400 is to support multiple users with live streaming via a Linux® cloud-based server application. Another technical solution of system 1400 is to simplify switching between tasks, objects, and / or SOs.
[0145] FIG. 16 illustrates a block diagram of a computer-based system and platform 800 according to one or more embodiments of the present disclosure. However, not all of these components are required to practice one or more embodiments, and variations in the arrangement and type of components may be made without departing from the spirit or scope of various embodiments of the present disclosure. In some embodiments, the exemplary computer devices and exemplary computing components of the exemplary computer-based system and platform 800 are capable of managing a large number of clients and concurrent transactions, as described in detail herein. In some embodiments, the exemplary computer-based system and platform 800 may be based on a scalable computer and network architecture incorporating various strategies for data evaluation, caching, retrieval, and / or database connection pooling. An example of a scalable architecture is one that is capable of operating multiple servers.
[0146] 16 , member computing devices 802, 803, and 804 (e.g., clients) of exemplary computer-based system and platform 800 may include virtually any computing device capable of sending and receiving messages to and from each other and to another computing device, such as servers 806 and 807, over a network (e.g., a cloud network), such as network 805. In some embodiments, member devices 802-804 may be personal computers, multiprocessor systems, microprocessor-based or programmable consumer electronics devices, network PCs, and the like. In some embodiments, one or more member devices within member devices 802-804 may typically include computing devices that connect using a wireless communications medium, such as a cell phone, a smartphone, a pager, a walkie-talkie, a radio frequency (RF) device, an infrared (IR) device, a CB, an integrated device combining one or more of the foregoing devices, or virtually any mobile computing device. In some embodiments, one or more member devices among member devices 802-804 may be a device capable of connecting using a wired or wireless communication medium, such as a PDA, a pocket PC, a wearable computer, a laptop, a tablet, a desktop computer, a netbook, a video game device, a pager, a smartphone, an ultra-mobile personal computer (UMPC), AR glasses / lenses, and / or any other device equipped to communicate via a wired and / or wireless communication medium (e.g., NFC, RFID, NBIOT, 3G, 4G, 5G, GSM, GPRS, WiFi, WiMax, CDMA, satellite, Bluetooth, ZigBee, etc.).
[0147] In some embodiments, one or more member devices among member devices 802-804 may be capable of running one or more applications such as an internet browser, mobile applications, voice calls, video games, video conferencing, and email, among others. In some embodiments, one or more member devices among member devices 802-804 may be configured to receive and send web pages, etc. In some embodiments, an exemplary specifically programmed browser application of the present disclosure may be configured to receive and display graphics, text, multimedia, etc. using virtually any web-based language, including, but not limited to, Standard Generalized Markup Language (SMGL) such as Hypertext Markup Language (HTML), Wireless Application Protocol (WAP), Handheld Device Markup Language (HDML) such as Wireless Markup Language (WML), WMLScript, XML, JavaScript, etc. In some embodiments, member devices among member devices 802-804 may be specifically programmed in Java, Python.Net, QT, C, C++, and / or any other suitable programming language. In some embodiments, one or more member devices within member devices 802-804 may include or be specifically programmed to execute applications that perform a variety of possible tasks, such as, but not limited to, messaging functions, browsing, searching, playing, streaming, or displaying various forms of content, including locally stored or uploaded messages, images and / or videos, and / or games.
[0148] In some embodiments, the exemplary network 805 may provide network access, data transfer, and / or other services to any computing device coupled thereto. In some embodiments, the exemplary network 805 may include and implement at least one dedicated network architecture that may be based at least in part on one or more standards set by, for example, but not limited to, the Global System for Mobile Communications (GSM) Association, the Internet Engineering Task Force (IETF), and the Worldwide Interoperability for Microwave Access (WiMAX) Forum. In some embodiments, the exemplary network 805 may implement one or more of the GSM architecture, the General Packet Radio Service (GPRS) architecture, the Universal Mobile Telecommunications System (UMTS) architecture, and the evolution of UMTS called Long Term Evolution (LTE). In some embodiments, the exemplary network 805 may include and implement, alternatively or in conjunction with, the WiMAX architecture defined by the WiMAX Forum. In some embodiments, optionally in combination with any of the embodiments described above or below, the exemplary network 805 may also include, for example, at least one of a local area network (LAN), a wide area network (WAN), the Internet, a virtual LAN (VLAN), a corporate LAN, a Layer 3 virtual private network (VPN), a corporate IP network, or any combination thereof. In some embodiments, optionally in combination with any of the embodiments described above or below, the at least one computer network communication over the exemplary network 805 may be transmitted based at least in part on one of a plurality of communication modes, such as, but not limited to, NFC, RFID, Narrowband Internet of Things (NBIOT), ZigBee, 3G, 4G, 5G, GSM, GPRS, WiFi, WiMax, CDMA, satellite, and any combination thereof.In some embodiments, the exemplary network 805 may also include mass storage, such as network-attached storage (NAS), a storage area network (SAN), a content delivery network (CDN), or other forms of computer- or machine-readable media.
[0149] In some embodiments, exemplary server 806 or exemplary server 807 may be a web server (or a series of servers) running a network operating system, examples of which may include, but are not limited to, Microsoft Windows Server, Novell Netware, or Linux. In some embodiments, exemplary server 806 or exemplary server 807 may be used and / or provided for cloud and / or network computing. Although not shown in FIG. 16 , in some embodiments, exemplary server 806 or exemplary server 807 may have connections to external systems, such as email, SMS messaging, text messaging, advertising content providers, etc. Any of the features of exemplary server 806 may also be implemented in exemplary server 807, and vice versa.
[0150] In some embodiments, one or more of the exemplary servers 806 and 807 may be specifically programmed to perform as, by way of non-limiting example, an authentication server, a search server, an email server, a social networking service server, an SMS server, an IM server, an MMS server, an exchange server, a photo sharing service server, an advertisement serving server, a financial / banking related service server, a travel service server, or any similarly suitable service-based server for users of the member computing devices 801-804.
[0151] In some embodiments, optionally in combination with any of the above or below described embodiments, for example, one or more of exemplary computing member devices 802-804, exemplary server 806, and / or exemplary server 807 may include specially programmed software modules that may send, process, and receive information using scripting languages, remote procedure calls, email, tweets, short message service (SMS), multimedia message service (MMS), instant messaging (IM), Internet Relay Chat (IRC), mIRC, Jabber, application programming interfaces, simple object access protocol (SOAP) methods, common object request broker architecture (CORBA), HTTP (hypertext transfer protocol), REST (representational state transfer), or any combination thereof.
[0152] The description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of exemplary embodiments provides those skilled in the art with an enabling description of implementing one or more exemplary embodiments. It will be understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope of the embodiments of the present disclosure. Example embodiments are described below with reference to the drawings. Elements that are identical, similar, or identically acting in various figures are identified with the same reference numerals, and repeated descriptions of these elements are partially omitted to avoid redundancy.
[0153] From the foregoing description, it will be apparent that variations and modifications can be made to the disclosed embodiments to adapt them to various applications and conditions, and such embodiments are within the scope of the following claims.
[0154] The recitation of a list of elements in any definition of a variable herein includes definitions of that variable as any single element or combination (or subcombination) of the listed elements. The recitation of an embodiment herein includes that embodiment as any single embodiment or in combination with any other embodiment or portion thereof.
Claims
1. receiving, by a processor, at least one first image of a first device; training, by the processor, a deep learning module including a plurality of neural networks; identifying one of the presence or absence of the first device in the at least one first image and a first inspection area of the first device in the at least one first image based on classifying a plurality of device features and a plurality of inspection area features identified in the at least one first image compared to a first reference image of the first device; when the first care area is identified as not being present in the at least one first image or when the classification of the plurality of care area features is below a base level confidence; generating at least one imaging instruction to capture at least one second image of at least one of the first device or the first inspection area of the first device; training at least one neural network of the plurality of neural networks based on at least one computer vision technique; receiving, by the processor, the at least one second image based on the at least one image capture instruction; inputting, by the processor, the at least one second image of the first device into the deep learning module; classifying the plurality of device features and the plurality of inspection area features of the first inspection area identified in the at least one second image by comparing them to a second reference image of a visual indicator indicative of a first inspection result in the first inspection area of the first device; when the presence of the first care area in the at least one second image is identified or when the classification of the plurality of care area features is higher than the base level confidence; retraining the deep learning module to improve identification of the first device based on the plurality of device features and the plurality of inspection area features identified in the at least one first image and the at least one second image; training the deep learning module to identify the first inspection result in the first inspection area based on a plurality of device features and a plurality of inspection area features identified in at least one first image and at least one second image; receiving, by the processor, at least one third image of a second device; inputting the at least one third image into the deep learning module to identify a second inspection result in a second inspection area of the second device in the at least one third image; A method comprising:
2. training the deep learning module for the improved identification of the first device includes training a first neural network on the plurality of device features and the plurality of inspection area features identified in the at least one first image and the at least one second image; retraining the deep learning module to identify the first device includes retraining the first neural network on the device features and the inspection area features identified in the at least one first image and the at least one second image. The method of claim 1.
3. controlling, by the processor, a camera to capture the at least one second image of the first device based on the at least one image capture instruction; applying, by the processor, a filter to the at least one second image of the first device based on the at least one capture instruction; The method of claim 1 further comprising:
4. controlling, by the processor, a camera flash based on the at least one image capture instruction to capture the at least one second image of the first device; applying, by the processor, a filter to the at least one second image of the first device based on the at least one capture instruction; The method of claim 1 further comprising:
5. controlling, by the processor, a focus of a camera to capture the at least one second image of the first device based on the at least one image capture instruction; applying, by the processor, a filter to the at least one second image of the first device based on the at least one capture instruction; The method of claim 1 further comprising:
6. training the deep learning module for the improved identification of the first inspection area includes training a second neural network on the plurality of device features and the plurality of inspection area features in the at least one first image and the at least one second image; retraining the deep learning module to identify the first inspection area includes retraining the second neural network on the device features and the inspection area features identified in the at least one first image and the at least one second image. The method of claim 2 further comprising:
7. generating, by the processor, a representation of the first device within the at least one first image; generating, by the processor, instructions in the at least one first image to move a camera to capture the at least one second image of the first device based on the at least one image capture instruction; The method of claim 1 further comprising:
8. generating, by the processor, instructions in the at least one first image to move the first device to capture the first inspection area of the first device in the at least one second image based on the at least one imaging instruction; generating, by the processor, a representation of the first examination region of the first device within the at least one second image; The method of claim 1 further comprising:
9. transmitting, by the processor, the second test results to a medical records system for displaying and updating the user's medical file associated with the second test results. The method of claim 1 further comprising:
10. Identifying the first test result includes: generating, by the processor, additional image capture instructions for capturing additional images of the first device; identifying, by the processor, within the first inspection area of the first device in the additional image, the first inspection result generated by the first device within the first inspection area of the first device; The method of claim 1 , comprising:
11. 1. A system comprising: a processor, the processor comprising: receiving at least one first image on a first device; training a deep learning module including a plurality of neural networks, identifying one of the presence or absence of the first device in the at least one first image and a first inspection area of the first device in the at least one first image based on classifying a plurality of device features and a plurality of inspection area features identified in the at least one first image compared to a first reference image of the first device; when the first care area is identified as not being present in the at least one first image or when the classification of the plurality of care area features is below a base level confidence; generating at least one imaging instruction to capture at least one second image of at least one of the first device or the first inspection area of the first device; training at least one neural network of the plurality of neural networks based on at least one computer vision technique; receiving the at least one second image based on the at least one imaging instruction command; inputting the at least one second image of the first device into the deep learning module; classifying the plurality of device features and the plurality of inspection area features of the first inspection area identified in the at least one second image by comparing them to a second reference image of a visual indicator indicative of a first inspection result in the first inspection area of the first device; when the presence of the first care area in the at least one second image is identified or when the classification of the plurality of care area features is higher than the base level confidence; retraining the deep learning module to improve identification of the first device based on the plurality of device features and the plurality of inspection area features identified in the at least one first image and the at least one second image; training the deep learning module to identify the first inspection result in the first inspection area based on the plurality of device features and the plurality of inspection area features identified in the at least one first image and the at least one second image; receiving at least one third image on a second device; inputting the at least one third image into the deep learning module to identify a second inspection result in a second inspection area of the second device in the at least one third image; To run the system.
12. The processor further comprises: training the deep learning module for the improved identification of the first device includes training a first neural network on the plurality of device features and the plurality of inspection area features identified in the at least one first image and the at least one second image; retraining the deep learning module to identify the first device includes retraining the first neural network on the device features and the inspection area features identified in the at least one first image and the at least one second image. The system of claim 11 , wherein
13. The processor further comprises: controlling a camera of the first device to capture the at least one second image based on the at least one image capture instruction; applying a filter to the at least one second image of the first device based on the at least one image capture instruction; The system of claim 11 , further comprising:
14. The processor further comprises: controlling a camera flash to capture the at least one second image of the first device based on the at least one image capture instruction; applying a filter to the at least one second image of the first device based on the at least one image capture instruction; The system of claim 11 , further comprising:
15. The processor further comprises: controlling a focus of a camera to capture the at least one second image of the first device based on the at least one image capture instruction; applying a filter to the at least one second image of the first device based on the at least one image capture instruction; The system of claim 11 , further comprising:
16. The processor further comprises: training the deep learning module for the improved identification of the first inspection area includes training a second neural network on the plurality of device features and the plurality of inspection area features in the at least one first image and the at least one second image; retraining the deep learning module to identify the first inspection area includes retraining the second neural network on the device features and the inspection area features identified in the at least one first image and the at least one second image. The system of claim 12 , wherein
17. The processor further comprises: generating a representation of the first device within the at least one first image; generating instructions in the at least one first image to move a camera of the first device to capture the at least one second image based on the at least one image capture instruction; The system of claim 11 , further comprising:
18. The processor further comprises: generating, based on the at least one imaging instruction, instructions in the at least one first image to move the first device to capture the first inspection area of the first device in the at least one second image; generating a representation of the first examination area of the first device within the at least one second image; The system of claim 11 , further comprising:
19. The processor further comprises: transmitting the second test results to a medical records system for displaying and updating the user's medical file associated with the second test results; The system of claim 11 , further comprising:
20. The processor further comprises: generating additional image capture instructions to capture additional images of the first device; identifying, within the first inspection area of the first device in the additional image, the first inspection result generated by the first device within the first inspection area of the first device; The system of claim 11 , further comprising:
21. At least one computer-readable storage medium encoded with computer-executable instructions, When executed by a computer, the computer-executable instructions cause the computer to: receiving at least one first image on a first device; training a deep learning module including a plurality of neural networks, identifying one of the presence or absence of the first device in the at least one first image and a first inspection area of the first device in the at least one first image based on classifying a plurality of device features and a plurality of inspection area features identified in the at least one first image compared to a first reference image of the first device; when the first care area is identified as not being present in the at least one first image or when the classification of the plurality of care area features is below a base level confidence; generating at least one imaging instruction to capture at least one second image of at least one of the first device or the first inspection area of the first device; training at least one neural network of the plurality of neural networks based on at least one computer vision technique; receiving the at least one second image based on the at least one imaging instruction command; inputting the at least one second image of the first device into the deep learning module; classifying the plurality of device features and the plurality of inspection area features of the first inspection area identified in the at least one second image by comparing them to a second reference image of a visual indicator indicative of a first inspection result in the first inspection area of the first device; when the presence of the first care area in the at least one second image is identified or when the classification of the plurality of care area features is higher than the base level confidence; retraining the deep learning module to improve identification of the first device based on the plurality of device features and the plurality of inspection area features identified in the at least one first image and the at least one second image; training the deep learning module to identify the first inspection result in the first inspection area based on the plurality of device features and the plurality of inspection area features identified in at least one first image and at least one second image; receiving at least one third image on a second device; inputting the at least one third image into the deep learning module to identify a second inspection result in a second inspection area of the second device in the at least one third image; performing a method including A computer-readable storage medium.