Computer system and computer-implemented method for rapid diagnostic test result interpretation platform using computer vision
Through the deep learning module, computer vision technology is applied to identify IV-RDT devices and area features, and imaging guidance commands are generated, which solves the problem of image acquisition and analysis of IV-RDT results in the prior art, and achieves efficient and accurate image capture and analysis.
Patent Information
- Application Number
- CN202380071994.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-08-10
- Filing Date
- 2023-08-10
- Publication Date
- 2025-05-23
AI Technical Summary
When using cameras to identify IV-RDT results, the prior art faces that image acquisition depends on accurate strip positioning. Especially when the IV-RDT device cannot fix the position, image acquisition and analysis are difficult, and it is necessary to build a specific template for each type of IV-RDT. The template is missing or limited when deviating from the standard format.
Through the deep learning module, computer vision technology is applied to identify the device features of the test equipment and test areas, and the images are classified based on the reference image corpus, and imaging guidance commands are generated to optimize image capture and realize automated image analysis.
It improves the accuracy and efficiency of image acquisition and analysis, reduces the dependence on the fixed position of IV-RDT devices, adapts to different types of IV-RDT devices, and enhances the identification and transmission capabilities of test results.
Smart Images

Figure CN120035839A_ABST
Abstract
Description
Technical Field
[0001] Generally speaking, the disclosure relates to computer-implemented methods and computer systems configured for a rapid diagnostic test result interpretation platform employing computer vision. Background Art
[0002] The use of lateral flow-based in vitro rapid diagnostic tests (IV-RDTs) in disease diagnosis has expanded dramatically, highlighted by the need for large population self-screening during the COVID-19 pandemic. Summary of the invention
[0003] In some aspects, the technology described herein relates to a method comprising: receiving, by a processor, a first image stream of a test device from a camera, the test device including a test area displaying a visual indicator; applying, by the processor, at least one computer vision technique of a deep learning module to: identify multiple device features of the test device in the first image stream, and classify at least one first image in the first image stream based on the multiple device features of the test device against at least one first reference image in a reference image corpus of the test device; selecting, by the processor, at least one imaging guidance command for a camera to capture a second image stream based on the at least one first reference image, wherein the second image stream includes the test area; instructing, by the processor, to implement at least one imaging guidance command to automatically generate the second image stream; receiving, by the processor, the second image stream from a camera adjusted using the at least one imaging guidance command; applying, by the processor, at least one computer vision technique of a deep learning module to: identify multiple device features of the test area in the second image stream, and classify at least one second image in the second image stream based on the multiple device features of the test area against at least one second reference image in a reference image corpus of the test area; and identifying, by the processor, a test result based on a visual indicator in the at least one second image based on the at least one second reference image.
[0004] In some aspects, the technology described herein relates to a method, further comprising: transmitting, by a processor, a user identification to a user device including a camera; maintaining, by the processor, a connection with the user device based on the user identification; and receiving, by the processor, a first image stream and a second image stream via the connection.
[0005] In some aspects, the technology described herein relates to a method in which a plurality of device features are structural features.
[0006] In some aspects, the technology described herein relates to a method in which at least one computer vision technique of a deep learning module includes at least one modular neural network, each modular neural network including at least one convolutional neural network input.
[0007] In some aspects, the techniques described herein relate to a method in which at least one convolutional neural network identifies a test region in at least one first image in a first image stream.
[0008] In some aspects, the technology described herein relates to a method in which at least one modular neural network identifies at least one candidate imaging guidance command; and wherein at least one convolutional neural network selects at least one imaging guidance command of the at least one candidate imaging guidance command.
[0009] In some aspects, the technology described herein relates to a method that also includes: generating, by a processor, at least one imaging guidance command for a camera to capture a second image stream based on imaging guidance metadata associated with at least one first reference image, wherein the second image stream includes a test area.
[0010] In some aspects, the technology described herein relates to a method in which instructing implementation of at least one imaging guidance command includes transmitting, by a processor, the at least one imaging guidance command to a camera so that the camera automatically generates a second image stream based on the at least one imaging guidance command.
[0011] In some aspects, the technology described herein relates to a method wherein instructing implementation of at least one imaging guidance command includes transmitting, by a processor, at least one imaging guidance command to a user device that includes a camera, the at least one imaging guidance command including augmented reality instructions for the user device overlaid on a first image stream to guide a user to reposition the camera to capture a test device.
[0012] In some aspects, the technology described herein relates to a method, further comprising: storing, by a processor, the test results; and transmitting, by the processor, the test results to an administrator device for display.
[0013] In some aspects, the technology described herein relates to a system comprising: a processor configured to: receive a first image stream of a test device from a camera, the test device including a test area displaying a visual indicator; apply at least one computer vision technique of a deep learning module to: identify multiple device features of the test device in the first image stream, and classify at least one first image in the first image stream based on the multiple device features of the test device against at least one first reference image in a reference image corpus of the test device; select at least one imaging guidance command for a camera to capture a second image stream based on the at least one first reference image, wherein the second image stream includes the test area; instruct implementation of the at least one imaging guidance command to automatically generate the second image stream; receive the second image stream from the camera adjusted using the at least one imaging guidance command; apply at least one computer vision technique of the deep learning module to: identify multiple device features of the test area in the second image stream, and classify at least one second image in the second image stream based on the multiple device features of the test area against at least one second reference image in a reference image corpus of the test area; and identify a test result based on visual indicators in the at least one second image based on the at least one second reference image.
[0014] In some aspects, the technology described herein relates to a system in which the processor is further configured to: transmit a user identification to a user device including a camera; maintain a connection with the user device based on the user identification; and receive a first image stream and a second image stream via the connection.
[0015] In some aspects, the technology described herein relates to a system in which a plurality of device features are structural features.
[0016] In some aspects, the techniques described herein relate to a system wherein at least one computer vision technique of a deep learning module includes at least one modular neural network in communication with at least one convolutional neural network input.
[0017] In some aspects, the techniques described herein relate to a system in which at least one convolutional neural network identifies a test region in at least one first image in a first image stream.
[0018] In some aspects, the techniques described herein relate to a system wherein at least one computer vision technique of a deep learning module identifies at least one candidate imaging guidance command; and wherein at least one convolutional neural network selects at least one imaging guidance command from the at least one candidate imaging guidance command.
[0019] In some aspects, the technology described herein relates to a system wherein the processor is further configured to generate at least one imaging guidance command for a camera to capture a second image stream based on imaging guidance metadata associated with at least one second reference image, wherein the second image stream includes a test area.
[0020] In some aspects, the technology described herein relates to a system, wherein to instruct implementation of at least one imaging guidance command, the processor is further configured to: transmit the at least one imaging guidance command to the camera so that the camera automatically generates a second image stream based on the at least one imaging guidance command.
[0021] In some aspects, the technology described herein relates to a system wherein, to instruct implementation of at least one imaging guidance command, the processor is further configured to: transmit at least one imaging guidance command to a user device including a camera, the at least one imaging guidance command including augmented reality instructions for the user device overlaid on a first image stream to guide the user to reposition the camera to capture a test device.
[0022] In some aspects, the technology described herein relates to a system wherein the processor is further configured to: store the test results; and transmit the test results to an administrator device for display.
[0023] In some aspects, the technology described herein relates to a method comprising: identifying, by a processor of a user device from a camera of the user device, a first image stream of a test device, the test device including a test area displaying a visual indicator; applying, by the processor, at least one computer vision technique of a deep learning module to: identify multiple device features of the test device in the first image stream, transmit the multiple device features of the test device to a server including a corpus of reference images of the test device, receive from the server at least one first reference image in the corpus of reference images, and classify at least one first image in the first image stream against the at least one first reference image based on the multiple device features; and selecting, by the processor, at least one imaging guidance command for the camera to capture a second image stream based on the at least one first reference image. The image stream includes a test area; the processor instructs the implementation of at least one imaging guidance command to automatically generate a second image stream; the processor identifies the second image stream from a camera adjusted using the at least one imaging guidance command; the processor applies at least one computer vision technology of a deep learning module to: identify multiple device features of the test area in the second image stream, transmit the multiple device features of the test area to a server including a reference image corpus of the test area, receive at least one second reference image in the reference image corpus of the test area from the server, and classify at least one second image in the second image stream based on the multiple device features of the test area against the at least one second reference image; and the processor identifies a test result based on a visual indicator in the at least one second image based on the at least one second reference image.
[0024] In some aspects, the technology described herein relates to a method that also includes: receiving, by a processor, a user identification from a server; maintaining, by the processor, a connection with the server based on the user identification; and transmitting, by the processor, to the server via the connection, multiple device characteristics of a test device and multiple device characteristics of a test area.
[0025] In some aspects, the technology described herein relates to a method in which a plurality of device features are structural features.
[0026] In some aspects, the technology described herein relates to a method in which at least one multi-agent system includes at least one deep learning model and at least one augmented reality model.
[0027] In some aspects, the techniques described herein relate to a method in which at least one convolutional neural network identifies a test region in at least one first image in a first image stream.
[0028] In some aspects, the technology described herein relates to a method in which at least one computer vision technique of a deep learning module identifies at least one candidate imaging guidance command; and wherein at least one convolutional neural network selects at least one imaging guidance command from the at least one candidate imaging guidance command.
[0029] In some aspects, the technology described herein relates to a method that also includes: receiving, by a processor, imaging guidance metadata associated with at least one second reference image from a server; and generating, by the processor, at least one imaging guidance command for a camera to capture a second image stream based on the imaging guidance metadata associated with the at least one second reference image, wherein the second image stream includes a test area.
[0030] In some aspects, the technology described herein relates to a method in which instructing implementation of at least one imaging guidance command includes generating, by a processor, the at least one imaging guidance command to cause a camera to automatically generate a second image stream based on the at least one imaging guidance command.
[0031] In some aspects, the technology described herein relates to a method in which instructing implementation of at least one imaging guidance command includes: causing a processor to cause a user device to display at least one imaging guidance command including augmented reality instructions overlaid on an image stream to guide a user to reposition a camera to capture a test device.
[0032] In some aspects, the technology described herein relates to a method, further comprising: storing, by a processor, the test results; and transmitting, by the processor, the test results to an administrator device for display. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The embodiments of the disclosure briefly summarized above and discussed in more detail below may be understood by reference to the illustrative embodiments of the disclosure depicted in the accompanying drawings. It should be noted, however, that the drawings illustrate only typical embodiments of the disclosure and, therefore, should not be construed as limiting the scope of the disclosure, as the disclosure may admit of other equally effective embodiments. This patent or application file contains at least one color drawing. Copies of this patent or patent application publication with color drawings will be provided by the Patent Office upon request and payment of the necessary fee.
[0034] Figures 1 to 16 Some exemplary aspects of the disclosure are presented in accordance with at least some principles of at least some embodiments of the disclosure.
[0035] To facilitate understanding, the same reference numerals are used, where possible, to represent the same elements common to the drawings. The drawings are not drawn to scale and may be simplified representations for clarity. It is contemplated that elements and features of one embodiment may be beneficially incorporated into other embodiments without further description. DETAILED DESCRIPTION
[0036] In addition to those benefits and technical solutions already disclosed, other purposes and advantages of the disclosure will become apparent from the following description in conjunction with the accompanying drawings. Detailed embodiments of the disclosure are disclosed herein; however, it should be understood that the disclosed embodiments are merely illustrative of the disclosure that can be implemented in various forms. In addition, each example given in conjunction with the various embodiments of the disclosure is intended to be illustrative rather than restrictive.
[0037] Throughout the specification, the following terms take the meanings clearly associated with this article, unless the context clearly stipulates otherwise. As used herein, the phrases "in one embodiment" and "in some embodiments" do not necessarily refer to the same embodiment, although it may refer to the same embodiment. In addition, as used herein, the phrases "in another embodiment" and "in some other embodiments" do not necessarily refer to different embodiments, although it may refer to different embodiments. Therefore, as described below, without departing from the scope or spirit of the open text, the various embodiments of the open text can be easily combined. In addition, when describing a particular feature, structure or characteristic in conjunction with an implementation, it is within the knowledge of those skilled in the art to realize such features, structures or characteristics in conjunction with other implementations (regardless of whether this paper is clearly described).
[0038] The term "based on" is not exclusive and allows for being based on additional factors not described unless the context clearly dictates otherwise. In addition, throughout the specification, the meanings of "a", "an" and "the" include plural references. The meaning of "in" includes "in" and "on".
[0039] It should be understood that at least one aspect / functionality of the various embodiments described herein can be performed in real time and / or dynamically. As used herein, the term "real time" refers to an event / action that can occur instantly or almost instantly when another event / action occurs. For example, "real-time processing", "real-time computing" and "real-time execution" all refer to performing calculations during the actual time when a related physical process (e.g., a user interacts with an application on a mobile device) occurs, so that the results of the calculations can be used to guide the physical process.
[0040] As used herein, the term "dynamically" means that events and / or actions can be triggered and / or occur without any human intervention. In some embodiments, events and / or actions according to the disclosure can occur in real time and / or based on a predetermined periodicity of at least one of the following: nanoseconds, several nanoseconds, microseconds, several microseconds, milliseconds, several milliseconds, seconds, several seconds, minutes, minutes, hours, hours, days, weeks, months, etc.
[0041] As used herein, the term "runtime" corresponds to any behavior that is dynamically determined during execution of a software application or at least a portion of a software application.
[0042] In some embodiments, the specially programmed computing system of the present invention with associated devices is configured to operate in a distributed network environment, communicating through a suitable data communications network (e.g., the Internet, etc.) and utilizing at least one suitable data communications protocol (e.g., IPX / SPX, X.25, AX.25, AppleTalk, etc.). TM , TCP / IP (e.g., HTTP), etc.). Of course, it should be noted that any suitable hardware and / or computing software language may be used to implement the embodiments described herein. In this regard, those of ordinary skill in the art are familiar with the types of computer hardware that may be used, the types of computer programming techniques that may be used (e.g., object-oriented programming), and the types of computer programming languages that may be used (e.g., C++, Objective-C, Swift, Java, Javascript). Of course, the foregoing examples are illustrative and non-restrictive.
[0043] As used herein, the terms "image" and "image data" are used interchangeably to identify data representing visual content, including but not limited to images encoded in various computer formats (e.g., ".jpg"; ".bmp", etc.), streaming video based on various protocols (e.g., Real-Time Streaming Protocol (RTSP), Real-time Transport Protocol (RTP), Real-time Transport Control Protocol (RTCP), etc.), recorded / generated non-streaming video in various formats (e.g., ".mov", ".mpg", ".wmv", ".avi", ".flv", etc.), and real-time visual images acquired through a camera application on a mobile device.
[0044] The materials disclosed herein may be implemented in software or firmware or a combination thereof, or as instructions stored on a machine-readable medium, which may be read and executed by at least one processor. A machine-readable medium may include any medium and / or mechanism for storing or transmitting information in a form readable by a machine (e.g., a computing device). For example, a machine-readable medium may include a read-only memory (ROM); a random access memory (RAM); a magnetic disk storage medium; an optical storage medium; a flash memory device; an electrical, optical, acoustic, or other form of propagated signal (e.g., a carrier wave, an infrared signal, a digital signal, etc.), and the like.
[0045] In another form, a non-transitory artifact (such as a non-transitory computer readable medium) can be used with any of the examples mentioned above or other examples, except that it does not include the transient signal itself. It does include those elements other than the signal itself that can temporarily store data in a "transitory" manner, such as RAM, etc.
[0046] As used herein, the terms "computer engine" and "engine" identify at least one software component and / or a combination of at least one software component and at least one hardware component that is designed / programmed / configured to manage / control other software and / or hardware components (such as a library, software development kit (SDK), object, etc.).
[0047] Examples of hardware elements may include processors, microprocessors, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, application specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), logic gates, registers, semiconductor devices, chips, microchips, chipsets, etc. In some embodiments, at least one processor may be implemented as a complex instruction set computer (CISC) or a reduced instruction set computer (RISC) processor; an x86 instruction set compatible processor, a multi-core or any other microprocessor, a graphics processing unit (GPU), or a central processing unit (CPU). In various implementations, at least one processor may be a dual-core processor, a dual-core mobile processor, etc.
[0048] Examples of software may include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, processes, software interfaces, application program interfaces (APIs), instruction sets, computing codes, computer codes, code segments, computer code segments, words, values, symbols, or any combination thereof. Determining whether an embodiment is implemented using hardware elements and / or software elements may vary depending on any number of factors, such as desired computing rates, power levels, thermal tolerances, processing cycle budgets, input data rates, output data rates, memory resources, data bus speeds, and other design or performance constraints.
[0049] One or more aspects of at least one embodiment may be implemented by representative instructions stored on a machine-readable medium, which represents various logic within the processor and, when read by a machine, causes the machine to fabricate logic to perform the techniques described herein. Such representations, known as "IP cores," may be stored on a tangible, machine-readable medium and supplied to various customers or manufacturing facilities to load into the fabrication machines that actually make the logic or processor.
[0050] As used herein, the term "user" shall have the meaning of at least one user.
[0051] In some embodiments, the exemplary computing device of the present invention can be configured for various purposes, such as but not limited to computer vision, mobile device cameras, and image-based recommendation positioning of applications, etc. In some embodiments, the exemplary camera can be a video camera or a digital still camera. In some embodiments, the exemplary camera can be configured to work based on sending analog and / or digital signals to at least one storage device that can reside in at least one of the following: a desktop computer, a laptop computer, or an output device, such as but not limited to VR glasses, lenses, smart watches, and / or a screen of a mobile device.
[0052] In some embodiments, as detailed herein, the exemplary computing device of the present invention of the disclosure can be programmed / configured to process visual feeds from various visual recording devices (such as but not limited to mobile device cameras, computer cameras, display screens, or any other camera for similar purposes) to detect, identify, and track (e.g., in real time) test devices that appear in the visual recording at one time and / or over a period of time. In some embodiments, the exemplary computing device of the present invention can be configured for various purposes, such as but not limited to computer vision, location-based recommended commands in applications, etc. In some embodiments, the exemplary camera can be a video camera or a digital still camera. In some embodiments, the exemplary camera can be configured to work based on sending analog and / or digital signals to at least one storage device that can reside in at least one of the following: a desktop computer, a laptop computer, or an output device, such as but not limited to VR glasses, lenses, smart watches, and / or mobile devices.
[0053] In some embodiments, the exemplary inventive process of detecting, identifying, and tracking one or more test devices over time is agnostic as to whether visual recordings are being obtained from the same or different recording devices at the same or different locations.
[0054] In some embodiments, an exemplary computing device of the present invention configured to detect, identify, and track one or more test areas of one or more test devices may rely on one or more centralized databases (e.g., a data center). For example, an exemplary computing device of the present invention may be configured to extract feature inputs for each test device that can be used to identify / recognize a specific test area or test result. In some embodiments, an exemplary computing device of the present invention may be configured to improve detection by training itself based on a single (multiple) image for test result identification or based on a collection of video frames or videos taken under different conditions. In some embodiments, for example, in a mobile (e.g., smart phone) application configured / programmed to provide video communication capabilities with augmented reality elements, an exemplary computing device of the present invention may be configured to detect one or more test devices in real time based at least in part on a single frame or a series of frames without saving recognition state.
[0055] In some embodiments, an exemplary inventive computing device may be configured to utilize one or more techniques for test device identification and test area tracking, as described in detail herein. In some embodiments, an exemplary inventive computing device may be configured to further utilize techniques that may allow identification of a test area in an image frame.
[0056] In some embodiments, the exemplary computing device of the present invention can be configured to assist medical testing purposes, such as but not limited to: determining / estimating and tracking the test results of patients; obtaining feedback on test equipment; tracking the patient's condition for medical and preventive purposes (e.g., monitoring diseases); suitable applications in statistics; suitable applications in sociology, etc. In some embodiments, the exemplary computing device of the present invention can be configured to assist entertainment or / and educational purposes. For example, exemplary electronic content in mobile applications and / or computer-based applications can be dynamically adjusted and / or triggered based at least in part on the detected test equipment. In some embodiments, an illustrative example of such dynamically adjusted / triggered electronic content can be one or more of a visual mask and / or a visual / audio effect (such as color, shape, size, etc.), which can be applied to the test equipment image in the exemplary video stream. In some embodiments, exemplary enhanced content can be dynamically generated by the exemplary computing device of the present invention and consists of at least one of a command, a suggestion, a fact, an image, etc.
[0057] In some embodiments, an exemplary inventive computing device may be configured to utilize raw video or image (eg, screen capture) input / data from any type of known camera, including both analog and digital cameras.
[0058] In some embodiments, an exemplary computing device of the present invention may be configured to utilize a deformable three-dimensional test device image that can be trained to generate meta-parameters (such as, but not limited to, coefficients defining the deviation of the test device from a mean shape; coefficients defining test areas and / or test results, camera positions and / or head positions, etc.).
[0059] In some embodiments, an exemplary inventive computing device may be configured to identify test results based on a single frame as a baseline; at the same time, using several frames may improve the quality of detection. In some embodiments, an exemplary inventive computing device may be configured to estimate more refined patterns in the test results, thus enabling the use of a lower resolution camera (e.g., a mobile or web camera, etc.).
[0060] In some embodiments, an exemplary computing device or system of the present invention may be configured to connect directly to an existing camera (e.g., mobile, computer-based, etc.) or to be operably and remotely connected. In some embodiments, an exemplary computing device or system of the present invention may be configured to include a specially programmed data processing module of the present invention, which may be configured to obtain video input from one or more cameras. For example, in accordance with the principles of the disclosed text, a specially programmed data processing module of the present invention may be configured to determine the source of the input video data, the need for code conversion to a different format, or to perform any other appropriate adjustments so that the video input can be used for processing. In some embodiments, the input image data (e.g., input video data) may include any appropriate type of source for the video content, and may include a variety of video sources. In some embodiments, the video data from the input video (e.g., Figure 1 The content of a video stream (of a video stream) may include both video data and metadata. A single picture may also be included in a frame. In some embodiments, a specially programmed data processing module of the present invention may be configured to decode the video input in real time and separate it into frames. In some embodiments, an exemplary input video stream captured by an exemplary camera (e.g., the front camera of a mobile personal smart phone) may be divided into frames. For example, a typical movie sequence is an interlaced format of multiple camera lenses, and camera shooting is a performance recorded continuously using a given camera setting. As used herein, camera registration may refer to the registration of different cameras that capture video, images, or screen frames in a video or image sequence / stream. The concept of camera registration is based on camera shooting in the reconstruction of video editing. A typical video sequence is an interlaced format of multiple camera lenses, and camera shooting is a performance recorded continuously using a given camera setting. By registering each camera from the incoming video frame, the original interlaced format can be divided into multiple sequences, each of which corresponds to a registered camera aligned with the original camera setting.
[0061] In some embodiments, the specially programmed data processing module of the present invention can be configured to process each frame or a series of frames using a suitable test device detection algorithm. For example, if one or more test or control areas are detected in a frame, the specially programmed data processing module of the present invention can be configured to extract feature vectors and store the extracted information in one or more databases. In some embodiments, an exemplary computing device of the present invention can be configured to include a specially programmed test identification module of the present invention, and the specially programmed test identification module of the present invention can be configured to compare the extracted features with the previous information stored in the database. In some embodiments, if the specially programmed test identification module of the present invention determines a match, new information is added to the existing data to increase the accuracy of further identification and / or improve the quality of the test result determination. In some embodiments, if the corresponding data is missing in the database, a new entry is created. In some embodiments, the obtained test results can be determined to be stored in a database for further analysis.
[0062] The use of lateral flow-based IV-RDT in disease diagnosis has expanded greatly recently, with the need for large populations to self-screen during the COVID-19 pandemic further accentuating this. However, a major drawback of this simple and economical solution for disease diagnosis and monitoring, especially involving self-testing, is the need for a reliable platform for recording and transmitting test results to healthcare providers / HMOs. In an attempt to address this obstacle, several solutions for IV-RDT result capture and their digital transmission have been developed. Capturing IV-RDT results can involve the use of a dedicated reader that scans the lateral flow membrane to identify the colored / fluorescent marker lines. Thereafter, the test results are transmitted to the operator in a qualitative binary (positive / negative) mode.
[0063] However, it is challenging to identify IV-RDT results using a camera because it cannot detect test results. One technical issue is that image acquisition can rely on accurate strip positioning relative to a moving camera to optimize image capture. Such positioning can be unsupervised and require capturing multiple images until a processable image is obtained. This can be challenging in situations where the IV-RDT device cannot be fixed to a certain position in space (e.g., the IV-RDT is not a box and therefore cannot be laid flat on a solid support).
[0064] Another technical issue is that image acquisition requires building specific templates for each type of IV-RDT, which rely on external IV-RDT device features (e.g., box structure and specific added patterns, coloring, background lighting, etc.) to guide mobile camera positioning. In test devices that deviate from the flat box format, such specificity is missing or limited.
[0065] Yet another technical problem is that image analysis is difficult when image acquisition is not ideal. For example, the camera may be positioned too far from the object, and image analysis of blurred objects is not feasible.
[0066] Another technical problem is that shadows and reflections produced by different lighting positions and conditions or introduced by the user (e.g., blocking lighting or casting shadows) may cause erroneous lines / features or hide existing lines / features. This may cause erroneous results to be obtained by the algorithm, which is expected to compromise the specificity and sensitivity of the test.
[0067] Figure 1 An embodiment of a detection system 100 is depicted that includes a server 105 in communication with a client device 110 having a camera 115 and a display application 120 (e.g., a mobile application installed on the client device 110) for detecting a test device 125 to detect a test area 130 having a visual indicator 135 indicating a test result 140. The test device 125 may include a test area 130 that displays a visual indicator 135 identifying the test result 140. Figure 2 As shown, the test equipment 125 includes a test area 130. Figure 3 As shown, the test device 125 includes a visual indicator 135 .
[0068] The display application 120 may display an interface for interactively guiding the user to correctly position the camera 115 to generate a first image stream of the test device 125. In some embodiments, the first image stream includes a screenshot of the test device 125. The server 105 may include a neural network for real-time detection, identification, and tracking of the test device 125. The server 105 may locate the test device 125 in the image stream and guide the user in real time how to position the camera 115 to best identify the test area 130 to identify the test results 140. The server 105 may generate an imaging guidance command to adjust the camera 115 to capture a second image stream of the test area 130 of the test device 125. In some embodiments, the second image stream includes a screenshot of the test device 125. The neural network may analyze the second image stream and identify the test results 140 from the visual indicator 135. Using techniques such as back propagation, transfer learning, and data modeling, the detection system 100 can learn, adopt, and change over time.
[0069] In some embodiments, the server 105 is a stand-alone server that communicates with a client device 110, which may be a user device, a mobile device with a phone camera, or a mobile, web, or local device. In some embodiments, the server 105 may include at least one processor. In some embodiments, the client device 110 may include at least one processor.
[0070] The test device 125 may be a rapid diagnostic test (RDT) for in vitro diagnostics (IVD). The server 105 may visually analyze the test device 125 in real time by receiving an image stream of the test device 125 from the client device 110. The detection system 100 may be used to identify a test result 140 from a visual indicator 135 in the test area 130. In some embodiments, the test result 140 represented by the visual indicator 135 is binary (e.g., positive / negative). In some embodiments, the test result 140 represented by the visual indicator 135 is in a semi-quantitative format (e.g., 1-10). In some embodiments, the test result 140 represented by the visual indicator 135 is in a quantitative format (e.g., concentration, dose, level).
[0071] The server 105 may include a deep learning module 145 for classifying images from the client device 110. In some embodiments, the deep learning module 145 may be configured to utilize computer vision technology. The deep learning module 145 may include, utilize, or may be cloud-based AI computer vision (CV), artificial intelligence (AI), deep learning (DL), web applications, live video streams, or live analysis by CV. In some embodiments, the deep learning module 145 may be configured to utilize one or more exemplary AI / computer vision technologies selected from but not limited to the following: decision trees, graph algorithms, enhancements, support vector machines, neural networks, nearest neighbor algorithms, naive Bayes, bagging, random forests. In some embodiments, and optionally, in combination with any of the embodiments described above or below, the exemplary neural network technology may be, but not limited to, one of the following: a feedforward neural network, a radial basis function network, a recurrent neural network, a convolutional network (e.g., U-net), or other suitable networks. In some embodiments, and optionally, in combination with any of the embodiments described above or below, an exemplary implementation of a neural network may be performed as follows:
[0072] i) define the neural network architecture / model,
[0073] ii) passing input data to the exemplary neural network model,
[0074] iii) incrementally training an exemplary model,
[0075] iv) determine the accuracy for a specific number of time steps,
[0076] v) applying the exemplary trained model to process newly received input data,
[0077] vi) Optionally and in parallel, continue training the exemplary trained model at a predetermined periodicity.
[0078] In some embodiments, and optionally, in combination with any of the embodiments described above or below, an exemplary trained deep learning model may specify a neural network by at least a neural network topology, a series of activation functions, and connection weights. For example, the topology of a neural network may include configurations of nodes of the neural network and connections between such nodes. In some embodiments, and optionally, in combination with any of the embodiments described above or below, an exemplary trained deep learning model may also be specified to include other parameters, including but not limited to bias values / functions and / or aggregation functions. For example, the activation function of a node may be a step function, a sine function, a continuous or piecewise linear function, a sigmoid function, a hyperbolic tangent function, a ReLU function, or other types of mathematical functions representing thresholds for activating a node. In some embodiments, and optionally, in combination with any of the embodiments described above or below, an exemplary aggregation function may be a mathematical function that combines (e.g., sums, products, etc.) input signals of a node. In some embodiments, and optionally, in combination with any of the embodiments described above or below, the output of an exemplary aggregation function may be used as an input to an exemplary activation function. In some embodiments, and optionally in combination with any of the embodiments described above or below, the bias can be a constant value or a function that can be used by the aggregation function and / or the activation function to make a node more or less likely to be activated.
[0079] In some embodiments, the deep learning module 145 may include neural networks 146A to 146N that form a modular neural network 148 for classifying images from the client device 110. In some embodiments, the modular neural network 148 receives input from each of the neural networks 146A to 146N. Each of the neural networks 146A to 146N may be a convolutional neural network (CNN). In some embodiments, the neural networks 146A to 146N may use a combination of CNN elements to convert images / videos from image streams into inputs (e.g., nodes) for deep learning techniques to feed the image inputs into a deep neural network (DNN) to identify the test results 140. In some embodiments, the deep learning module 145 may be a trained MAS (multi-agent system) including the neural networks 146A to 146N and the modular neural network 148 for computer recognition task division.
[0080] The server 105 may include a data repository 150 for storing a reference image corpus 152 of the test device 125 and a reference image corpus 154 of the test area 130 .
[0081] The server 105 may include a generator 155 for generating at least one imaging guidance command to cause the client device 110 to capture an image of the test device 125 The deep learning module 145 may identify the test result 140 generated by the test device 125 based on the visual indicator 135 .
[0082] Reference now Figure 4 , a method 400 for a rapid diagnostic test result interpretation platform employing computer vision is shown. In some embodiments, the display application 120 can display interactive instructions for task completion (such as for a user to use the test device 125 to generate a test result 140). For example, the instructions can be similar to taking a COVID-19 antigen test.
[0083] The method 400 may include the server 105 receiving a first image stream of the test device 125 (step 402). The deep learning module 145 may establish a connection with the client device 110 to receive the image stream from the client device 110. In some embodiments, the client device 110 may connect to the cloud-based server 105 in real time via a web socket. Once the test device 125 is detected, a socket may be established to facilitate communication between the display application 120 (e.g., the client-side model (MAR)) and the server 105 (e.g., the server-side model (CNN)). For example, the deep learning module 145 may open a live port to the optical flow converter to optimize the data flow. The deep learning module 145 may establish a request to open a live camera broadcast from the client device 110, which sends the image stream to the server 105. In some embodiments, the detection system 100 may use an API call to open a live socket to the camera 115, which broadcasts the image stream from the end user who sent the request to the detection system 100. The live socket may establish a connection for a video feed from the client device 110 to the cloud-based deep learning AI.
[0084] In some embodiments, the deep learning module 145 can transmit a user identification (e.g., a code, password, or token of the user's device) to the client device 110 including the camera 115. For example, the deep learning module 145 can provide the user with an ID for documentation. In some embodiments, the deep learning module 145 can receive user verification for capturing the image stream, such as a token that complies with privacy regulations.
[0085] In some embodiments, the deep learning module 145 can establish a connection with the client device 110 in response to receiving a user identification (e.g., a token). The deep learning module 145 can use the connection established through the socket so that a live stream is established from the client device 110 to the deep learning module 145. For example, the deep learning module 145 can establish a connection with the display application 120. In some embodiments, the display application 120 can be a web application, a native application, an API, an SDK, a container (Docker), etc.
[0086] In some embodiments, the deep learning module 145 can receive the first imaging stream and the second imaging stream via the connection. For example, the deep learning module 145 can obtain a live video stream from the camera 115 of the client device 110. The deep learning module 145 can communicate with the client device 110. For example, the deep learning module 145 can connect to a cloud-based server using an open socket with a user identifier (e.g., a user ID). In some embodiments, the deep learning module 145 or the display application 120 can apply an optical flow method (such as sparse feature propagation or metadata keyframe extraction) to the video stream to minimize data transactions and achieve optimal recognition time for the test device 125 and the test area 130.
[0087] The method 400 may include the server 105 identifying the test device 125 in the first image stream (step 404). Figure 5A and Figure 5B , showing an image of a first image stream of test device 125. In some embodiments, deep learning module 145 may receive a first image stream of test device 125 from client device 110. Deep learning module 145 may identify test device 125 in the first image stream. In some embodiments, such as Figure 5A and Figure 5B As shown, the deep learning module 145 can cause the display application 120 to display or highlight the test device 125 and the test area 130 in the first image stream. Figure 5B As shown, the display application 120 may display the text “Device” after identifying the test device 125 in the first image stream.
[0088] The deep learning module 145 may use the neural networks 146A to 146N to identify, track, or monitor the test device 125 and its location in the image stream. In some embodiments, the neural networks 146A to 146N may use a combination of CNN elements to convert images / videos from the first image stream into inputs (e.g., nodes) for deep learning techniques to feed the image inputs into the deep neural network (DNN) of the deep learning module 145 to identify the test device 125. For example, the modular neural network 148 may receive inputs from the trained neural networks 146A to 146N to allow data collected from the first image stream by optical flow methods to be inserted into the relevant tasks of the modular neural network 148. The modular neural network 148 may process the inputs through one or more of its task-driven DNNs that pass information between them. In some embodiments, the deep learning module 145 may apply optical flow methods (such as sparse feature propagation, metadata keyframe extraction) to the image stream to minimize data transactions and achieve optimal (live) recognition time of the test device 125 and the test area 130.
[0089] Deep learning module 145 may classify images in the first image stream to identify test device 125. Neural networks 146A through 146N may receive and analyze images of test device 125. In some embodiments, Figure 7 , Fig. 8A , Figure 8B , Fig. 9 , Fig.10 and Fig.11 As shown, neural networks 146A- 146N may be convolutional neural networks that are part of modular neural network 148 that identifies test device 125 in at least one first image in the first image stream.
[0090] After identifying the test device 125, the neural networks 146A-146N can track the test device 125, and the modular neural network 148 can provide the real-time location (e.g., x, y, z coordinates) of the test device 125 to the display application 120. In some embodiments, the display application 120 can include the functionality of the deep learning module 145 to track the test device 125 (e.g., including IV-RDT tests) while the deep learning module 145 analyzes the image stream for the generator 155 to generate and provide imaging guidance commands for display by the display application 120 to guide the user to scan the test by marking the test device 125 and the test area 130 on the screen controlled by the display application 120, such as Figure 5A and Figure 5B shown.
[0091] The display application 120 may track and highlight the test device 125 to guide the user in real time to move the test device 125 or the camera 115 so that the test area 130 and the visual indicator 135 may be identified by the server 105. The deep learning module 145 may guide the user in "real time" using an interactive visual interface, text, and voice to achieve the best angle and position of objects in space and within objects to obtain the best possible view of the test area 130. The neural networks 146A-146N may be updated and receive input from other neural networks 146A-146N as a feedback loop forming a recurrent neural network (RNN) to retrain the neural networks 146A-146N to make corrections during the process of analyzing the image stream.
[0092] like Figure 6 As shown, in some embodiments, the deep learning module 145 can identify multiple device features of the test device 125 in the first image stream. In some embodiments, the multiple device features are multiple structural features of the test device 125. For example, the deep learning module 145 can identify points, edges, or objects that constitute the test device 125. For example, the deep learning module 145 can identify two features of the test device 125, thereby identifying the test device 125 with a baseline level of confidence. In another example, the deep learning module 145 can identify three features of the test device 125, thereby identifying the test device 125 with a higher confidence level. In yet another example, the deep learning module 145 can identify four features of the test device 125, thereby identifying the test device 125 with an even higher confidence level.
[0093] In some embodiments, reference image corpus 152 may be images of test device 125 for comparison by neural networks 146A-146N to identify test device 125 in the first image stream. In some embodiments, reference image corpus 154 may be images of test area 130 for comparison by neural networks 146A-146N to identify test area 130 in the second image stream. In some embodiments, deep learning module 145 may classify at least one first image in the first image stream against at least one first reference image in reference image corpus 152 of test device 125 based on a plurality of device features of test device 125.
[0094] Method 400 may include the server 105 selecting an imaging guidance command to adjust the image capture of the test device 125 (step 406). The server 105 may include a generator 155 for generating at least one imaging guidance command to cause the client device 110 to capture an image of the test area 130 of the test device 125. The imaging guidance command may be selected to optimize the image stream of the test area 130 by imaging the test area 130 using the client device 110. The imaging guidance command may maximize the resolution of the test area 130 by improving the focus and illumination at the location of the test area 130. For example, the imaging guidance command may cause the client device 110 to turn on the flash to improve illumination, filter the image, or change the focus of the camera 115. The imaging guidance command may create optimal conditions for the deep learning module 145 to record a second image stream of the test area 130.
[0095] In some embodiments, the generator 155 may select at least one imaging guidance command for the client device 110 to capture a second image stream based on at least one first reference image. The second image stream may include the test area 130. For example, the imaging guidance command may cause the display application 120 to display interfaces that guide the user to achieve the optimal angle and position of the camera 115 in space and within the object relative to the test device 125 to obtain the best possible second image stream of the test area 130.
[0096] In some embodiments, the generator 155 may generate at least one imaging guidance command for the camera 115 to capture a second image stream of the test area 130 based on imaging guidance metadata associated with at least one first reference image. In some embodiments, the deep learning module 145 may identify at least one candidate imaging guidance command. For example, the generator 155 may maintain candidate imaging guidance commands in the data repository 150. The candidate imaging guidance commands may be default commands such as requesting the user to twist the test device 125, turn on the light, zoom in or out, move up or down, or adjust the focus of the camera 115.
[0097] In some embodiments, at least one of the convolutional neural networks 146A to 146N may select at least one imaging guidance command from among the at least one candidate imaging guidance commands. For example, the generator 155 may use an ML (deep learning) algorithm to analyze the image stream and generate commands for display by the display application 120 on the user interface.
[0098] In some embodiments, the generator 155 may instruct to implement at least one imaging guidance command to automatically generate a second image stream. In some embodiments, the generator 155 may directly instruct the client device 110 (e.g., change the focus of the camera 115). In some embodiments, to implement at least one imaging guidance command, the generator 155 may transmit at least one imaging guidance command to the client device 110 such that the camera 115 automatically generates a second image stream based on the at least one imaging guidance command.
[0099] Now refer to Fig. 12A , Fig. 12B and Fig. 12C , in some embodiments, the generator 155 may instruct the user of the client device 110 (e.g., request the user to reposition the client device 110 to achieve a different view of the test device 125). In some embodiments, the generator 155 may transmit at least one imaging guidance command to the client device 110 including a camera 115, the at least one imaging guidance command including augmented reality instructions overlaid on the first image stream for the display application 120 to guide the user to reposition the camera to capture the test device. The augmented reality instructions may indicate how the user may reposition the client device 110 to image the test area 130. For example, the display application 120 may display the second image stream to the user on the display of the client device 110. The augmented reality instructions may be overlaid on the test device 125 to indicate how the user may reposition the client device 110 to image the test area 130. As FIG. 12A to FIG. 12C shown, the augmented reality instructions may request the user to rotate the test device 125 to place the test area 130 within the field of view of the camera 115.
[0100] Method 400 may include the server 105 receiving a second image stream of the test area 130 of the test device 125 (step 408). Now refer to Fig.13 , an image of the second image stream of the test device 125 is shown. In some embodiments, the deep learning module 145 may receive the second image stream from the client device 110 adjusted using at least one imaging guidance command. Once the camera 115 is positioned based on the at least one imaging guidance command, the deep learning algorithm of the deep learning module 145 analyzes the key frames provided by the client device 110. The deep learning module 145 may perform the analysis until a sufficient number of images have been collected to identify the test result 140.
[0101] Method 400 may include server 105 identifying test area 130 in the second image stream (step 410). Deep learning module 145 may classify images in the second image stream to identify test area 130 and visual indicator 135. Neural networks 146A to 146N may receive and analyze images in the second image stream using techniques described herein, including identifying test area 130 including visual indicator 135 and identifying test result 140 itself. In some embodiments, such as Figure 7 , Fig. 8A , Figure 8B , Fig. 9 , Fig.10 and Fig.11 As shown, neural networks 146A- 146N may be convolutional neural networks that are part of modular neural network 148 that identifies test region 130 in at least one first image in the second image stream.
[0102] The deep learning module 145 may identify the test area 130 in the second image stream. In some embodiments, the deep learning module 145 may identify a plurality of device features of the test area 130 in the second image stream. In some embodiments, the plurality of device features of the test area 130 are a plurality of structural features of the test area 130. For example, the deep learning module 145 may identify points, edges, or objects that constitute the test area 130. In some embodiments, the deep learning module 145 may identify a plurality of features that constitute the test area 130. For example, the deep learning module 145 may identify two features of the test area 130, thereby identifying the test area 130 with a baseline confidence level. In another example, the deep learning module 145 may identify three features of the test area 130, thereby identifying the test area 130 with a higher confidence level. In yet another example, the deep learning module 145 may identify four features of the test area 130, thereby identifying the test area 130 with an even higher confidence level.
[0103] The deep learning module 145 may use the neural networks 146A-146N to identify, track, or monitor the test area 130 and its location relative to the test device 125. In some embodiments, the deep learning module 145 may use a combination of the convolutional neural networks 146A-146N to identify the test device 125. In some embodiments, the deep learning module 145 may classify at least one second image in the second image stream against at least one second reference image in the reference image corpus 154 of the test area 130 based on a plurality of device features of the test area 130. For example, the deep learning module 145 may apply neural networks, algorithms, and image recognition methods to analyze the second image stream including the test area 130 to identify the visual indicator 135.
[0104] In some embodiments, such as Fig.13 As shown, the deep learning module 145 can cause the display application 120 to display or highlight the test device 125 and the test area 130 in the second image stream. Fig.13 As shown, the display application 120 may display the text “bar window” after identifying the test device 125 in the first image stream.
[0105] In some embodiments, if the deep learning module 145 cannot identify the test area 130 in the second image stream, the modular neural network 148 can generate at least one imaging guidance command including camera settings so that the automated focus of the camera 115 can be applied and adjusted. The neural networks 146A to 146N can maintain a feedback loop forming a recurrent neural network (RNN) to optimize camera sensitivity and lighting settings. The generator 155 can monitor as the camera frames are collected, analyzed, and filtered to select frames that optimize image quality, which the deep learning module 145 uses to identify the test results 140. In some embodiments, the generator 155 can generate additional imaging guidance commands, as discussed in step 406.
[0106] The method 400 may include the server 105 identifying the test result 140 in the test area 130 of the test device 125 (step 412). In some embodiments, the deep learning module 145 may identify the test result 140 based on the at least one second reference image, based on the visual indicator 135 in the test area 130 in at least one of the second images. For example, the deep learning module 145 may identify the visual indicator in the test area 130 in the second image.
[0107] In some embodiments, the deep learning module 145 can cause the display application 120 to display the test result 140. For example, the deep learning module 145 can cause the display application 120 to display the test result 140 as a QR code. In another example, the deep learning module 145 can cause the display application 120 to display the test result 140 as a number or an indicator (e.g., positive or negative).
[0108] In some embodiments, the deep learning module 145 can store the test results 140. For example, the deep learning module 145 can record the test results 140 and associate the test results with an identifier (such as an identifier of a user or patient). In some embodiments, the deep learning module 145 can store the test results 140 in a data repository 150.
[0109] In some embodiments, the detection system 100 can use mobile computer vision and cloud-based AI (with web application UX / UI) to perform live diagnosis of IV-RDT. The detection system 100 can include an SDK platform for simple interaction with medical record systems and communication with professionals and HMOs. In some embodiments, the deep learning module 145 can transmit the test results 140 to the administrator device for display. For example, the administrator device can be an HMO, and the deep learning module 145 can transmit the test results 140 to the HMO and update the medical files of the user or patient. The server 105 can transmit the test results 140 back to the user, operator, HMO or physician for further confirmation, recording and medical treatment decision / follow-up. For example, the server can provide the test results 140 to the user and / or HMO / physician as a response. The detection system 100 can close the loop between the user and HMO / physician of the test device 125 for recording the test results 140 and providing medical treatment / follow-up.
[0110] In some embodiments, if the deep learning module 145 cannot identify the test result 140 from the visual indicator 135 in the test area 130 of the second image stream, the generator 155 may generate additional imaging guidance commands, as discussed in step 406 .
[0111] Through the synergistic combination of these technologies and platforms, the detection system 100 can overcome numerous technical challenges and include numerous technical solutions.
[0112] One technical solution is for the display application 120 to include a pre-trained 3D object detector and tracker MAR (Mobile Augmented Reality) model to detect and track the test device 125 by using the RAM of the client device 110 while the GPU of the server 105 analyzes the image stream and generates imaging guidance commands. By dividing the processing between the server 105 and the client device 110, the embodiments described herein enable the use of less memory while ensuring faster processing time because the test device 125 is tracked by the nearby client device 110 while the neural networks 146A to 146N forming the modular neural network 148 recursively analyze the image stream to generate imaging guidance commands to identify the test result 140.
[0113] Another technical solution of the detection system 100 is to create a display application 120 as a secure web application with a user-friendly UX / UI to access the camera 115. Another technical solution of the detection system 100 is the display application 120, which reduces the need to install complex applications on mobile devices. Another technical solution of the detection system 100 is that the display application 120 can be easily deployed and used on any smart client device 110.
[0114] Another technical solution is that with the development of cloud servers, the server 105 can use PWA (Progressive Web Application) to improve the accuracy of deep learning models and increase the speed of ML algorithms. The detection system 100 can use a CV algorithm with a (task) trained CNN to locate sub-objects (SO) (which can be the test area 130) and analyze them in real time to meet specific diagnostic task requirements. Another technical solution of the detection system 100 is to reduce the need for image capture by the client device 125, which can overcome the problems associated with the proper positioning of the camera 115 to achieve effective image capture.
[0115] Another technical solution for the detection system 100 is to manage SDKs with different API calls and user-friendly applications / interfaces that support dynamic guidance. Another technical solution for the detection system 100 is to develop a user-side ML model for fast and efficient data compression and minimization to achieve improved sockets for communication to cloud servers. Another technical solution for the detection system 100 is to develop computer vision technology for deep learning modules, such as MAS (multi-agent system) with MNN architecture, which combines CNN, RNN and NN for different tasks, and manages and manipulates them to support multiple models. Another technical solution for the detection system 100 is to use data collection and annotation to develop efficient training protocols for different CNN models to enhance model training. Another technical solution for the detection system 100 is to develop responsive applications for "real-time" user experience. Another technical solution for the detection system 100 is to connect to different clients (HMO, etc.) with minimal disturbance. Another technical solution for the detection system 100 is quality management to optimize and maintain test specificity and sensitivity. Another technical solution for the detection system 100 is to approve these tools for clinical use by different regulatory agencies.
[0116] Another technical solution of the detection system 100 is to increase the test accuracy by using a deep learning AI algorithm instead of only using the application of a deep learning model. Another technical solution of the detection system 100 reduces the need to model each test shape and makes it easy to identify complex geometric shapes. Another technical solution of the detection system 100 is to increase the accuracy of adaptive diagnostic tests over time, which has the ability to obtain semi-quantitative or even quantitative test results. Another technical solution of the detection system 100 is to improve the overall UX / UI of the application through the live guidance of the VR model that communicates with the AI model. Another technical solution of the detection system 100 is to support multiple users for live streaming through a Linux cloud-based server application. Another technical solution of the detection system 100 is to simplify switching between tasks, objects and / or SOs. Another technical solution of the detection system 100 is to establish security between the server 105 and the client device 110 through gateway connectivity with secure and different back-end servers / agents.
[0117] Reference now Fig.14 , shows a system 1400, which in some embodiments includes a client device 110, the client device includes a camera 115 and an application 1410, the application includes a display application 120, a deep learning module 145 and a generator 155. The system 1400 can be used to visually analyze an in vitro rapid diagnostic test (IV-RDT) device using computer vision (CV) and mobile augmented reality (MAR) executed on a client-based device. For example, the client device 110 can be hardware configured to execute the application 1410, such as a mobile phone or a smart camera, and the application can be a software application configured to execute the functionality of the display application 120, the deep learning module 145 and the generator 155. In some embodiments, the application 1410 can be a web application, a native application, an API, an SDK, a container (Docker), etc. In some embodiments, the deep learning module 145 is part of the application 1410, so that the mobile device can apply the neural network 146 forming the modular neural network 148 to the image stream generated by the camera 115. In some embodiments, the client device 110 can be communicatively connected to the data repository 150. For example, data repository 150 may be hosted on a server or cloud. In some embodiments, client device 110 may include any or all of components 115 to 155 (including data repository 150) described herein, such that client device 110 may perform the functionality described herein while offline.
[0118] Reference now Fig.15, a method 1500 for a rapid diagnostic test result interpretation platform employing computer vision is shown. The method 1500 may include an application 1410 capturing a first image stream of a test device 125 (step 1502). In some embodiments, the display application 120 may display interactive instructions for task completion (such as for a user to use the test device 125 to generate a test result 140). For example, the instructions may be similar to taking a COVID-19 antigen test.
[0119] The method 1500 may include the application 1410 identifying the test device 125 in the first image stream (step 1504). Figure 5A and Figure 5B , showing an image of a first image stream of a test device 125. In some embodiments, the application 1410 can capture a first image stream of a test device 125. The application 1410 can identify the test device 125 in the first image stream. In some embodiments, such as Figure 5A and Figure 5B As shown, the display application 120 can display or highlight the test device 125 and the test area 130 in the first image stream. Figure 5B As shown, the display application 120 may display the text “Device” after identifying the test device 125 in the first image stream.
[0120] In some embodiments, the application 1410 can apply the deep learning module 145 to classify the image to identify the test area 130 and the visual indicator 135. For example, the application 1410 can apply optical flow methods (such as sparse feature propagation, metadata keyframe extraction) to the image stream to minimize data transactions and achieve optimal (live) recognition time of the test device 125 and the test area 130.
[0121] The application 1410 can use the neural networks 146A to 146N to identify, track or monitor the test device 125 and its location in the image stream. Each model can include at least one convolutional neural network or other deep learning algorithm. For example, each segment of the MNN can include a trained CNN (convolutional neural network) to allow data collected from the first image stream by optical flow methods to be inserted into the relevant model task. In some embodiments, the application 1410 can use a combination of CNN elements to identify the test device 125.
[0122] In some embodiments, application 1410 may include neural networks 146A-146N forming modular neural network 148 to analyze images of test device 125 using the techniques described herein, including identifying test area 130 including visual indicator 135 and identifying test result 140 itself. In some embodiments, deep learning module 145 includes at least one modular neural network 148. For example, application 1410 may include neural networks 146A-146N and modular neural network 148 for computer recognition task partitioning.
[0123] In some embodiments, modular neural network 148 receives input from each of neural networks 146A through 146N. Figure 7 , Fig. 8A , Figure 8B , Fig. 9 , Fig.10 and Fig.11 As shown, at least one neural network 146A to 146N is part of a modular neural network 148 that identifies the test device 125 in at least one first image in the first image stream. Each of the neural networks 146A to 146N can be a convolutional neural network (CNN). In some embodiments, the neural networks 146A to 146N can use a combination of CNN elements to identify the test device 125.
[0124] After identifying the test device 125, the CNN of the application 1410 can track the test device 125 and provide a live position (e.g., x, y, z coordinates) of the test device 125. The application 1410 can track and highlight the test device 125 to guide the user in real time to move the test device 125 or the camera 115 so that the test area 130 and the visual indicator 135 can be identified by the deep learning module 145. The application 1410 can guide the user in "real time" using an interactive visual interface, text, and voice to achieve the best angle and position of objects in space and within the object to obtain the best possible view of the test area 130. The application 1410 can be updated and receive input from other neural networks 146A to 146N as a feedback loop to form a recurrent neural network (RNN) to retrain the neural networks 146A to 146N for corrections during the process of receiving the image stream.
[0125] like Figure 6As shown, in some embodiments, the application 1410 can identify multiple device features of the test device 125 in the first image stream. For example, the application 1410 can identify the features that constitute the test device 125. In some embodiments, the multiple device features are structural features. In some embodiments, the application 1410 can transmit the multiple device features of the test device 125 to the data repository 150. In some embodiments, the application 1410 can store the multiple device features of the test device 125 on a local storage device of the client device 110.
[0126] In some embodiments, the application 1410 can access the reference image corpus 152 and 154 stored locally on the client device 110. In some embodiments, the application 1410 can establish a connection with the data repository 150 to access the reference image corpus 152 and 154. For example, the data repository 150 can open a live port to the optical flow converter to optimize the data flow. In some embodiments, the application 1410 can receive a user identification from the data repository 150. In some embodiments, the application 1410 can receive a user verification for capturing the image stream, such as for a regulatory agency / entity. The application 1410 can use the connection established by the socket so that a live stream from the application 1410 to the data repository 150 is established. For example, the data repository 150 can establish a connection with the application 1410.
[0127] The method 1500 may include the application 1410 selecting an imaging guidance command to adjust the image capture of the test device 125 (step 1506). The application 1410 may generate at least one imaging guidance command to cause the application 1410 to capture an image of the test area 130 of the test device 125. The imaging guidance command may be selected to optimize the image stream of the test area 130 by imaging the test area 130 using the application 1410. The imaging guidance command may maximize the resolution of the test area 130 by improving the focus and lighting at the location of the test area 130. For example, the imaging guidance command may cause the application 1410 to turn on a flash to improve the lighting (e.g., turn on a light), implement filtering on the image, or change the focus of the camera 115. The imaging guidance command may create optimal conditions for the application 1410 to record a second image stream of the test area 130.
[0128] In some embodiments, the application 1410 may select, based on the at least one first reference image, at least one imaging guidance command for the application 1410 to capture a second image stream. The second image stream may include the test area 130. For example, the imaging guidance command may cause the application 1410 to display interfaces that guide a user to achieve an optimal angle and position of the camera 115 relative to the test device 125 in space and within the object to obtain the best possible second image stream of the test area 130.
[0129] In some embodiments, application 1410 may generate at least one imaging guidance command for camera 115 to capture a second image stream of test area 130 based on the imaging guidance metadata associated with the at least one first reference image. In some embodiments, application 1410 may receive imaging guidance metadata associated with at least one additional reference image from data repository 150. In some embodiments, application 1410 may retrieve imaging guidance metadata associated with at least one additional reference image from a local storage device of system 1400.
[0130] In some embodiments, the application 1410 can use the deep learning module 145 to identify at least one candidate imaging guidance command. For example, the application 1410 can maintain the candidate imaging guidance commands in a local storage device. In some embodiments, the application 1410 can receive the candidate imaging guidance commands from the data repository 150. The candidate imaging guidance commands can be default commands, such as requesting the user to twist the test device 125, turn on the light, zoom in or out, move up or down, or adjust the focus of the camera 115.
[0131] In some embodiments, the application 1410 may select at least one of the at least one candidate imaging guidance command. For example, the application 1410 may use an ML (machine learning) algorithm to analyze the image stream and generate a command for display by the display application 120 on the user interface.
[0132] In some embodiments, the application 1410 may instruct implementation of at least one imaging guidance command to automatically generate a second image stream. In some embodiments, the application 1410 may directly instruct the camera 115 (e.g., change the focus of the camera 115). In some embodiments, to implement at least one imaging guidance command, the application 1410 may execute at least one imaging guidance command so that the camera 115 automatically generates a second image stream based on the at least one imaging guidance command.
[0133] Reference now Fig. 12A , Fig. 12B and Fig. 12CIn some embodiments, application 1410 may instruct the user (e.g., requesting the user to reposition camera 115 to achieve a different view of test device 125). In some embodiments, application 1410 may be based on at least one imaging guidance command including augmented reality instructions superimposed on the first image stream so that application 1410 guides the user to reposition the camera to capture the test device. The augmented reality instructions may indicate how the user may reposition camera 115 to image test area 130. For example, display application 120 may display the second image stream to the user on a display of application 1410. Augmented reality instructions may be overlaid on test device 125 to indicate how the user may reposition camera 115 to image test area 130. FIG. 12A to FIG. 12C As shown, the augmented reality instructions may request the user to rotate the test device 125 to bring the test area 130 into the field of view of the camera 115 .
[0134] The method 1500 may include the application 1410 capturing a second image stream of the test area 130 of the test device 125 (step 1508). Fig.13 , showing an image of a second image stream of the test device 125. In some embodiments, the application 1410 can receive the second image stream from the camera 115 adjusted using the at least one imaging guidance command. Once the camera 115 is positioned based on the at least one imaging guidance command, the deep learning algorithm of the application 1410 analyzes the key frames provided by the camera 115. The deep learning module 145 can perform the analysis until a sufficient number of images have been collected to identify the test result 140.
[0135] Method 1500 may include applying 1410 to identify test areas 130 in the second image stream (step 1510). Deep learning module 145 may classify images in the second image stream to identify test areas 130 and visual indicators 135. Neural networks 146A to 146N may receive and analyze images in the second image stream using the techniques described herein, including identifying test areas 130 including visual indicators 135 and identifying test results 140 themselves. In some embodiments, such as Figure 7 , Fig. 8A , Figure 8B , Fig. 9 , Fig.10 and Fig.11 As shown, neural networks 146A- 146N may be convolutional neural networks that are part of modular neural network 148 that identifies test region 130 in at least one first image in the second image stream.
[0136] The application 1410 may identify the test area 130 in the second image stream. In some embodiments, the application 1410 may apply the deep learning module 145 to identify a plurality of device features of the test area 130 in the second image stream. In some embodiments, the application 1410 may transmit the plurality of device features of the test area 130 to the data repository 150. In some embodiments, the application 1410 may store the plurality of device features of the test area 130 in a local storage device of the system 1400.
[0137] The application 1410 can use a neural network to identify, track, or monitor the test area 130 and its location relative to the test device 125. In some embodiments, the application 1410 can identify the test device 125 using a combination of CNN elements.
[0138] In some embodiments, application 1410 may apply deep learning module 145 to classify at least one second image in the second image stream against at least one second reference image in reference image corpus 154 of test area 130 based on the plurality of device features of test area 130. For example, application 1410 may apply neural networks, algorithms, and image recognition methods to analyze the second image stream including test area 130 to identify visual indicator 135.
[0139] In some embodiments, such as Fig.13 As shown, application 1410 may cause display application 120 to display or highlight test device 125 and test area 130 in the second image stream. Fig.13 As shown, the display application 120 may display the text “bar window” after identifying the test device 125 in the first image stream.
[0140] In some embodiments, if the application 1410 cannot identify the test area 130 in the second image stream, the modular neural network 148 can generate at least one imaging guidance command including camera settings so that the automated focus of the camera 115 can be applied and adjusted. The neural networks 146A to 146N can maintain a feedback loop forming a recurrent neural network (RNN) to optimize camera sensitivity and lighting settings. The application 1410 can monitor as the camera frames are collected, analyzed, and filtered to select frames that optimize image quality, which the application 1410 uses to identify the test results 140. In some embodiments, the application 1410 can generate additional imaging guidance commands, as discussed in step 1506.
[0141] The method 1500 may include the application 1410 identifying the test result 140 in the test area 130 of the test device 125 (step 1512). In some embodiments, the application 1410 may identify the test result 140 based on the at least one second reference image, based on the visual indicator 135 in the test area 130 in at least one of the second images. For example, the application 1410 may identify the visual indicator in the test area 130 in the second image.
[0142] In some embodiments, the application 1410 may cause the display application 120 to display the test result 140. For example, the application 1410 may display the test result 140 as a QR code. In another example, the application 1410 may display the test result 140 as a number or an indicator (e.g., positive or negative).
[0143] In some embodiments, the application 1410 can store the test results 140. For example, the application 1410 can record the test results 140 and associate the test results with an identifier (such as an identifier of a user or patient). In some embodiments, the application 1410 can store the test results 140 in a data repository 150. In some embodiments, the application 1410 can transmit the test results 140 to an administrator device for display. For example, the administrator device can be an HMO, and the application 1410 can transmit the test results 140 to the HMO and update the medical file of the user or patient.
[0144] In some embodiments, if the application 1410 cannot identify the test result 140 from the visual indicator 135 in the test area 130 of the second image stream, the application 1410 may generate additional imaging guidance commands, as discussed in step 1506 .
[0145] Through the synergistic combination of these technologies and platforms, system 1400 can overcome numerous technical challenges and include numerous technical solutions.
[0146] One technical solution is that along with the development of cloud servers, the server 105 can use PWA (Progressive Web Application) to improve the accuracy of deep learning models and increase the speed of ML algorithms. The detection system 100 can use a CV algorithm with a (task-based) trained CNN to locate sub-objects (SO) (which can be the test area 130) and analyze them in real time to meet specific diagnostic task requirements.
[0147] Another technical solution of system 1400 is to manage SDK with different API calls and support user-friendly applications / interfaces for dynamic guidance. Another technical solution of system 1400 is to develop user-side ML models for fast and efficient data compression and minimization to achieve improved sockets for communication to cloud servers. Another technical solution of system 1400 is to develop computer vision technology for deep learning modules, such as MAS (multi-agent system) with MNN architecture, which combines CNN, RNN and NN for different tasks, and manages and manipulates them to support multiple models. Another technical solution of system 1400 is to use data collection and annotation to develop efficient training protocols for different CNN models to enhance model training. Another technical solution of system 1400 is to develop responsive applications for "real-time" user experience. Another technical solution of system 1400 is to connect to different clients (HMO, etc.) with minimal disturbance. Another technical solution of system 1400 is quality management to optimize and maintain test specificity and sensitivity. Another technical solution of system 1400 is to approve these tools for clinical use by different regulatory agencies.
[0148] Another technical solution of the system 1400 is a display application 120 that reduces the need to install complex applications on mobile devices. Another technical solution of the system 1400 is to create the display application 120 as a secure web application with a user-friendly UX / UI to access the camera 115. Another technical solution of the system 1400 is to improve the overall UX / UI of the application through live guidance of the VR model that communicates with the AI model. Another technical solution of the system 1400 is that the display application 120 can be easily deployed and used on any smart client device 110.
[0149] Another technical solution of the system 1400 is to reduce the need for image capture by the client device 125, which can overcome the problems associated with proper positioning of the camera 115 to achieve effective image capture. Another technical solution of the system 1400 is that security can be established between the server 105 and the client device 110 through gateway connectivity with a secure and different backend server / proxy.
[0150] Another technical solution of system 1400 is to increase test accuracy by using deep learning AI algorithms instead of only using deep learning models. Another technical solution of system 1400 reduces the need to model each test shape and makes it easy to identify complex geometric shapes. Another technical solution of system 1400 is to increase adaptive diagnostic test accuracy over time, which has the ability to obtain semi-quantitative or even quantitative test results. Another technical solution of system 1400 is to support multiple users for live streaming through a Linux cloud-based server application. Another technical solution of system 1400 is to simplify switching between tasks, objects and / or SOs.
[0151] Fig.16 Depicted is a block diagram of a computer-based system and platform 800 according to one or more embodiments of the disclosure. However, not all of these parts are necessary to practice one or more embodiments, and without departing from the spirit or scope of the various embodiments of the disclosure, the arrangement and type of parts can be changed. In some embodiments, exemplary computer-based system and the exemplary computing device and the exemplary computing components of platform 800 can be configured to manage a large number of clients and concurrent transactions, as described in detail herein. In some embodiments, exemplary computer-based system and platform 800 can be based on scalable computer and network architecture, and architecture combines various strategies for evaluating data, caching, searching and / or database connection pooling. An example of scalable architecture is the architecture that can operate multiple servers.
[0152] In some embodiments, reference Fig.16, the member computing devices 802, 803, and 804 (e.g., clients) of the exemplary computer-based system and platform 800 may include virtually any computing device capable of receiving messages from and sending messages to another computing device (such as servers 806 and 807) over a network (e.g., a cloud network) such as network 805. In some embodiments, the member devices 802 to 804 may be personal computers, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, and the like. In some embodiments, one or more of the member devices 802 to 804 may include computing devices that typically connect using a wireless communication medium, such as a cellular phone, a smart phone, a pager, a walkie-talkie, a radio frequency (RF) device, an infrared (IR) device, a CB, an integrated device combining one or more of the foregoing devices, or virtually any mobile computing device, and the like. In some embodiments, one or more member devices within member devices 802 to 804 may be devices capable of connecting using a wired or wireless communication medium, such as a PDA, a POCKET PC, a wearable computer, a laptop computer, a tablet computer, a desktop computer, a netbook, a video gaming device, a pager, a smart phone, an ultra mobile personal computer (UMPC), AR glasses / lenses, and / or any other device equipped to communicate via a wired and / or wireless communication medium (e.g., NFC, RFID, NBIOT, 3G, 4G, 5G, GSM, GPRS, WiFi, WiMax, CDMA, satellite, Bluetooth, ZigBee, etc.).
[0153] In some embodiments, one or more of the member devices 802 to 804 may run one or more applications, such as an Internet browser, mobile applications, voice calls, video games, video conferencing, and email, etc. In some embodiments, one or more of the member devices 802 to 804 may be configured to receive and send web pages, etc. In some embodiments, the exemplary specially programmed browser application of the disclosure text may be configured to receive and display graphics, text, multimedia, etc., using almost any web-based language, including but not limited to Standard Generalized Markup Language (SMGL), such as Hypertext Markup Language (HTML), Wireless Application Protocol (WAP), Handheld Device Markup Language (HDML), such as Wireless Markup Language (WML), WMLScript, XML, JavaScript, etc. In some embodiments, the member devices within the member devices 802 to 804 may be specially programmed by Java, Python.Net, QT, C, C++, and / or other suitable programming languages. In some embodiments, one or more member devices within member devices 802 to 804 may be specifically programmed to include or execute applications to perform a variety of possible tasks, such as, but not limited to, instant messaging functionality, browsing, searching, playing, streaming, or displaying various forms of content, including locally stored or uploaded messages, images and / or videos, and / or games.
[0154] In some embodiments, the exemplary network 805 can provide network access, data transmission and / or other services to any computing device connected thereto. In some embodiments, the exemplary network 805 can include and implement at least one dedicated network architecture, which can be based at least in part on one or more standards set by, for example, but not limited to, the Global System for Mobile Communications (GSM) Association, the Internet Engineering Task Force (IETF), and the Worldwide Interoperability for Microwave Access (WiMAX) Forum. In some embodiments, the exemplary network 805 can implement one or more of the GSM architecture, the General Packet Radio Service (GPRS) architecture, the Universal Mobile Telecommunications System (UMTS) architecture, and the UMTS evolution known as Long Term Evolution (LTE). In some embodiments, as an alternative or combination of one or more of the above, the exemplary network 805 can include and implement the WiMAX architecture defined by the WiMAX Forum. In some embodiments, and optionally, in combination with any of the embodiments described above or below, the exemplary network 805 can also include, for example, at least one of a local area network (LAN), a wide area network (WAN), the Internet, a virtual LAN (VLAN), an enterprise LAN, a layer 3 virtual private network (VPN), an enterprise IP network, or any combination thereof. In some embodiments, and optionally, in conjunction with any of the embodiments described above or below, at least one computer network communication on the exemplary network 805 may be transmitted at least in part based on one of a plurality of communication modes, such as, but not limited to: NFC, RFID, Narrowband Internet of Things (NBIOT), ZigBee, 3G, 4G, 5G, GSM, GPRS, WiFi, WiMax, CDMA, satellite, and any combination thereof. In some embodiments, the exemplary network 805 may also include a mass storage device, such as a network attached storage (NAS), a storage area network (SAN), a content delivery network (CDN), or other forms of computer or machine readable media.
[0155] In some embodiments, the exemplary server 806 or the exemplary server 807 may be a web server (or a series of servers) running a network operating system, examples of which may include, but are not limited to, Microsoft Windows Server, Novell NetWare, or Linux. In some embodiments, the exemplary server 806 or the exemplary server 807 may be used for and / or provide cloud and / or network computing. Although in Fig.16 807, but in some embodiments, exemplary server 806 or exemplary server 807 may have connections with external systems such as email, SMS instant messaging, text instant messaging, advertising content providers, etc. Any features of exemplary server 806 may also be implemented in exemplary server 807, and vice versa.
[0156] In some embodiments, in a non-limiting example, one or more of the exemplary servers 806 and 807 may be specifically programmed to perform as an authentication server, a search server, an email server, a social networking service server, an SMS server, an IM server, an MMS server, an exchange server, a photo sharing service server, an advertising serving server, a financial / banking related service server, a travel service server, or any similar suitable service-based server for users of member computing devices 801 to 804.
[0157] In some embodiments, and optionally, in combination with any of the embodiments described above or below, for example, one or more of the exemplary computing member devices 802 to 804, the exemplary server 806 and / or the exemplary server 807 may include specially programmed software modules that may be configured to send, process, and receive information using scripting languages, remote procedure calls, email, Twitter, Short Message Service (SMS), Multimedia Messaging Service (MMS), Instant Messaging (IM), Internet Relay Chat (IRC), mIRC, Jabber, application programming interfaces, Simple Object Access Protocol (SOAP) methods, Common Object Request Broker Architecture (CORBA), HTTP (Hypertext Transfer Protocol), REST (Representational State Transfer), or any combination thereof.
[0158] This specification only provides exemplary embodiments, and is not intended to limit the scope, applicability or configuration of the disclosure. On the contrary, the following description of exemplary embodiments will provide an enabling description for realizing one or more exemplary embodiments to those skilled in the art. It should be understood that, without departing from the spirit and scope of the embodiments disclosed in the present invention, various changes may be made to the function and arrangement of elements. Embodiment examples are described below with reference to the accompanying drawings. The elements of the same, similar or identical functions are identified with the same reference numerals in each of the accompanying drawings, and the repeated description of these elements is partially omitted to avoid redundancy.
[0159] From the foregoing description it will be apparent that the embodiments of the disclosure may be varied and modified to adapt them to various usages and conditions. Such embodiments are also within the scope of the appended claims.
[0160] The recitation of a list of elements with any definition of a variable herein includes defining the variable as any single element in the listed elements or any combination (or subcombination) of the listed elements. The recitation of an embodiment herein includes the embodiment as any single embodiment or in combination with any other embodiment or portion thereof.
Claims
1. A method, include: receiving, by a processor, at least one first image of a first device; A deep learning module comprising a plurality of neural networks is trained by the processor to: classifying a plurality of device features and a plurality of test area features identified in the at least one first image based on comparison with a first reference image of the first device, identifying the first device in the at least one first image and identifying one of the presence or absence of a first test area of the first device in the at least one first image; When it is identified that the first test area is not present in the at least one first image or when the classification of the plurality of test area features is below a baseline level of confidence: generating at least one imaging guidance command to capture at least one second image of at least one of the first device or the first test area of the first device; wherein at least one neural network of the plurality of neural networks is configured based on at least one computer vision technique; receiving, by the processor, the at least one second image based on the at least one imaging guidance command; Inputting, by the processor, the at least one second image of the first device into the deep learning module to: classifying the plurality of device features and the plurality of test area features of the first test area identified in the at least one second image by comparison to a second reference image of a visual indicator indicating a first test result in the first test area of the first device; When it is identified that the first test area exists in the at least one second image or when the classification of the plurality of test area features is above the reference level confidence: retraining the deep learning module for improved identification of the first device based on the plurality of device features and the plurality of test area features identified in the at least one first image and the at least one second image; as well as training the deep learning module for identifying the first test result in the first test area based on the plurality of device features and the plurality of test area features identified in the at least one first image and the at least one second image; receiving, by the processor, at least one third image of a second device; and The at least one third image is input by the processor into the deep learning module to identify a second test result in a second test area of the second device in the at least one third image.
2. The method according to claim 1, wherein training the deep learning module for the improved identification of the first device comprises training a first neural network for the plurality of device features and the plurality of test area features identified in the at least one first image and the at least one second image; Wherein retraining the deep learning module for the identification of the first device comprises retraining the first neural network for the plurality of device features and the plurality of test area features identified in the at least one first image and the at least one second image.
3. The method according to claim 1, further comprising: include: controlling, by the processor, a camera based on the at least one imaging guidance command to capture the at least one second image of the first device; and A filter is applied, by the processor, to the at least one second image of the first device based on the at least one imaging guidance command.
4. The method according to claim 1, further comprising: include: controlling, by the processor, a flash of a camera based on the at least one imaging directing command to capture the at least one second image of the first device; and A filter is applied, by the processor, to the at least one second image of the first device based on the at least one imaging guidance command.
5. The method according to claim 1, further comprising: include: controlling, by the processor, a focus of a camera based on the at least one imaging steering command to capture the at least one second image of the first device; and A filter is applied, by the processor, to the at least one second image of the first device based on the at least one imaging guidance command.
6. The method according to claim 2, further comprising: include: Among them training The deep learning module for the improved identification of the first test area includes training a second neural network for the plurality of device features and the plurality of test area features in the at least one first image and the at least one second image; Wherein retraining the deep learning module for the identification of the first test area includes retraining the second neural network for the multiple device features and the multiple test area features identified in the at least one first image and the at least one second image.
7. The method according to claim 1, further comprising: include: generating, by the processor, an indication of the first device in the at least one first image; and Instructions are generated, by the processor, in the at least one first image based on the at least one imaging guidance command, for moving a camera to capture the at least one second image of the first device.
8. The method according to claim 1, further comprising: include: generating, by the processor, instructions in the at least one first image for moving the first device to capture the first test area of the first device in the at least one second image based on the at least one imaging guidance command; and An indication of the first test area of the first device in the at least one second image is generated by the processor.
9. The method according to claim 1, further comprising: include: The second test result is transmitted by the processor to a medical record system for display and updating of a user's medical file associated with the second test result.
10. The method of claim 1, wherein the first test result is identified include: generating, by the processor, additional imaging guidance commands to capture additional images of the first device; and The first test result generated by the first device in the first test area of the first device in the additional image is identified by the processor in the first test area of the first device.
11. A system, include: A processor configured to: receiving at least one first image of a first device; Train a deep learning module consisting of multiple neural networks to: classifying a plurality of device features and a plurality of test area features identified in the at least one first image based on comparison with a first reference image of the first device, identifying the first device in the at least one first image and identifying one of the presence or absence of a first test area of the first device in the at least one first image; When it is identified that the first test area is not present in the at least one first image or when the classification of the plurality of test area features is below a baseline level of confidence: generating at least one imaging guidance command to capture at least one second image of at least one of the first device or the first test area of the first device; wherein at least one neural network of the plurality of neural networks is configured based on at least one computer vision technique; receiving the at least one second image based on the at least one imaging guidance command; Inputting the at least one second image of the first device into the deep learning module to: classifying the plurality of device features and the plurality of test area features of the first test area identified in the at least one second image by comparison to a second reference image of a visual indicator indicating a first test result in the first test area of the first device; When it is identified that the first test area exists in the at least one second image or when the classification of the plurality of test area features is above the reference level confidence: retraining the deep learning module for improved identification of the first device based on the plurality of device features and the plurality of test area features identified in the at least one first image and the at least one second image; as well as training the deep learning module for identifying the first test result in the first test area based on the plurality of device features and the plurality of test area features identified in the at least one first image and the at least one second image; receiving at least one third image of a second device; and The at least one third image is input into the deep learning module to identify a second test result in a second test area of the second device in the at least one third image.
12. The system of claim 11, wherein the processor is further configured to: wherein training the deep learning module for the improved identification of the first device comprises training a first neural network for the plurality of device features and the plurality of test area features identified in the at least one first image and the at least one second image; Wherein retraining the deep learning module to achieve the identification of the first device includes retraining the first neural network for the multiple device features and the multiple test area features identified in the at least one first image and the at least one second image.
13. The system of claim 11, wherein the processor is further configured to: Based on the at least one imaging guidance command, controlling a camera to capture the at least one second image of the first device; and Based on the at least one imaging guidance command, a filter is applied to the at least one second image of the first device.
14. The system of claim 11, wherein the processor is further configured to: Based on the at least one imaging directing command, controlling a flash of a camera to capture the at least one second image of the first device; and Based on the at least one imaging guidance command, a filter is applied to the at least one second image of the first device.
15. The system of claim 11, wherein the processor is further configured to: Based on the at least one imaging guidance command, controlling the focus of a camera to capture the at least one second image of the first device; and Based on the at least one imaging guidance command, a filter is applied to the at least one second image of the first device.
16. The system of claim 12, wherein the processor is further configured to: wherein training the deep learning module to achieve the improved identification of the first test area comprises training a second neural network for the plurality of device features and the plurality of test area features in the at least one first image and the at least one second image; Wherein retraining the deep learning module to achieve the identification of the first test area includes retraining the second neural network for the multiple device features and the multiple test area features identified in the at least one first image and the at least one second image.
17. The system of claim 11, wherein the processor is further configured to: generating an indication of the first device in the at least one first image; and Based on the at least one imaging guidance command, instructions are generated in the at least one first image to move a camera to capture the at least one second image of the first device.
18. The system of claim 11, wherein the processor is further configured to: generating, based on the at least one imaging guidance command, instructions in the at least one first image for moving the first device to capture the first test area of the first device in the at least one second image; and An indication of the first test area of the first device in the at least one second image is generated.
19. The system of claim 11, wherein the processor is further configured to: The second test result is transmitted to a medical record system for display and updating of a user's medical file associated with the second test result.
20. The system of claim 11, wherein the processor is further configured to: generating additional imaging guidance commands to capture additional images of the first device; and The first test result generated by the first device in the first test area of the first device is identified in the first test area of the first device in the additional image.
21. At least one computer-readable storage medium encoded with computer-executable instructions that, when executed by a computer, cause the computer to perform the following method: receiving at least one first image of a first device; Train a deep learning module consisting of multiple neural networks to: classifying a plurality of device features and a plurality of test area features identified in the at least one first image based on comparison with a first reference image of the first device, identifying the first device in the at least one first image and identifying one of the presence or absence of a first test area of the first device in the at least one first image; When it is identified that the first test area is not present in the at least one first image or when the classification of the plurality of test area features is below a baseline level of confidence: generating at least one imaging guidance command to capture at least one second image of at least one of the first device or the first test area of the first device; wherein at least one neural network of the plurality of neural networks is configured based on at least one computer vision technique; Based on the at least one imaging guidance command, receiving the at least one second image; inputting the at least one second image of the first device into the deep learning module to: classifying the plurality of device features and the plurality of test area features of the first test area identified in the at least one second image by comparison to a second reference image of a visual indicator indicating a first test result in the first test area of the first device; When it is identified that the first test area exists in the at least one second image or when the classification of the plurality of test area features is above the reference level confidence: retraining the deep learning module for improved identification of the first device based on the plurality of device features and the plurality of test area features identified in the at least one first image and the at least one second image; as well as training the deep learning module for identifying the first test result in the first test area based on the plurality of device features and the plurality of test area features identified in the at least one first image and the at least one second image; receiving at least one third image of a second device; and The at least one third image is input into the deep learning module to identify a second test result in a second test area of the second device in the at least one third image.