Apparatus and method for detecting associations between different types of datasets
The apparatus and method improve the accuracy of dataset association by training a classifier to identify and display relationships between different types of datasets, addressing errors from existing techniques.
Patent Information
- Application Number
- JP2025500142
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-06-30
- Filing Date
- 2023-06-30
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2043-06-30
AI Technical Summary
Current techniques for creating relationships between different types of datasets, such as image and text data, are flawed and prone to errors due to factors like image quality and data inconsistencies, leading to inaccurate information extraction and association.
An apparatus and method using a processor to identify and generate associations between datasets by training a second association classifier with data entries, including a first set of associations, and displaying the resulting associations using a display device.
Enhances the accuracy and reliability of information extraction and association across multiple media types by leveraging a trained classifier to detect and display meaningful relationships between diverse datasets.
Smart Images

Figure 2025524574000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of priority of U.S. Provisional Application No. 63 / 357,978, filed on July 1, 2022, entitled "SYSTEMS AND METHODS FOR DETECTING ASSOCIATIONS AMONG DATASETS OF DIFFERENT TYPES", and claims the benefit of priority of U.S. Non - Provisional / Provisional Application No. 18 / 217,378, filed on June 30, 2023, entitled "APPARATUS AND A METHOD FOR DETECTING ASSOCIATIONS AMONG DATASETS OF DIFFERENT TYPES". Each of U.S. Non - Provisional / Provisional Application No. 18 / 217,378 and U.S. Provisional Patent Application No. 63 / 357,978 is hereby incorporated by reference in its entirety.
[0002] The present invention generally relates to the field of detecting associations between different types of datasets. In particular, the present invention is directed to an apparatus and a method for detecting associations between different types of datasets.
Background Art
[0003] In various fields such as healthcare, pathology, logistics, and document management, the need for accurate and reliable information extraction and association across multiple media is significant. However, current techniques used to create relationships between data across two or more media are flawed and can introduce errors due to factors such as image quality, font variations, or data inconsistencies.
Summary of the Invention
[0004] In one aspect, an exemplary apparatus for detecting associations between different types of datasets includes at least a processor and a memory communicatively coupled to the at least one processor, the memory configured to receive a plurality of datasets including a first dataset and a second dataset, identify a first set of associations between a first subset of the first dataset and a first subset of the second dataset, use a second association classifier to generate, in response to the first set of associations, a second set of associations between a second subset of the first dataset and a second subset of the second dataset, where generating the second set of associations includes training the second association classifier using second association training data, the second association training data including a plurality of data entries including the first set of associations, and generating the second set of associations in response to the first set of associations using the trained second association classifier, and cause the at least one processor to display the second set of associations using a display device.
[0005] In another aspect, an exemplary method for detecting associations between different types of datasets includes receiving, using at least a processor, a plurality of datasets including a first dataset and a second dataset; identifying, using at least a processor, a first set of associations between a first subset of the first dataset and a first subset of the second dataset; generating, using at least a processor and using a second association classifier, a second set of associations between a second subset of the first dataset and a second subset of the second dataset in response to the first set of associations, generating the second set of associations including training the second association classifier using second association training data, the second association training data including a plurality of data entries including the first set of associations, and generating the second set of associations in response to the first set of associations using the trained second association classifier; and displaying, using a display device, the second set of associations.
[0006] In another aspect, another exemplary apparatus for detecting associations between different types of datasets includes at least a processor and a memory communicatively coupled to the at least one processor, the memory being configured to receive a plurality of datasets including a first dataset and a second dataset, the first and second datasets including data elements of different types, to identify an initial set of associations between the first dataset and the second dataset, each association including one or more first data elements from the first dataset and one or more second data elements from the second dataset, to train a neural network model to detect additional associations between the first dataset and the second dataset, the initial set of associations being used as training data for training the neural network model, and to configure the at least one processor to use the trained neural network model to detect one or more additional associations between the first dataset and the second dataset.
[0007] In another aspect, another exemplary method for detecting associations between different types of datasets is to receive, using at least a processor, a plurality of datasets including a first dataset and a second dataset, wherein the first and second datasets include data elements of different types; to identify, using at least a processor, an initial set of associations between the first dataset and the second dataset, wherein each association includes one or more first data elements from the first dataset and one or more second data elements from the second dataset; to train, using at least a processor, a neural network model to detect additional associations between the first dataset and the second dataset, wherein the initial set of associations is used as training data for training the neural network model; and to detect, using at least a processor and the trained neural network model, one or more additional associations between the first dataset and the second dataset.
[0008] Details of one or more variations of the subject matter described in this specification are set forth in the accompanying drawings and the following description. Other features and advantages of the subject matter described in this specification will be apparent from the description and drawings, and from the claims.
[0009] For purposes of illustrating the invention, the drawings show aspects of one or more embodiments of the invention. It should be understood, however, that the invention is not limited to the precise arrangements and instrumentalities shown in the drawings.
Brief Description of the Drawings
[0010]
Fig. 1
Fig. 2
Fig. 3
Fig. 4
Fig. 5
Fig. 6
Fig. 7
Fig. 8
Fig. 9
Fig. 10
DETAILED DESCRIPTION OF THE INVENTION
[0011] The drawings are not necessarily to scale and may be shown by imaginary lines, diagrams, and partial views. In certain cases, details that are not necessary for understanding the embodiments or details that make it difficult to perceive other details may be omitted. Like reference numerals in the various drawings indicate like elements.
[0012] Broadly, aspects of the present disclosure are directed to an apparatus and method for detecting associations between different types of datasets. The apparatus includes at least a processor and a memory communicatively coupled to the at least one processor. The memory instructs the processor to receive a plurality of datasets from a user. The memory instructs the processor to identify a first set of associations between the plurality of datasets. The memory instructs the processor to generate a second set of associations in response to the first set of associations using a second association classifier. Generating the second set of associations includes training the second association classifier using second association training data, the second association training data including a plurality of data entries including a first set of associations as an input that correlates to a second set of associations as an output. The memory instructs the processor to display the second set of associations using a display device. Exemplary embodiments showing aspects of the present disclosure are described below in the context of several specific examples.
[0013] Referring now to FIG. 1, an exemplary embodiment of an apparatus 100 for detecting associations between different types of datasets is shown. Apparatus 100 includes a processor 104. Processor 104 can include any computing device as described in the present disclosure, including, but not limited to, a microcontroller, a microprocessor, a digital signal processor (DSP), and / or a system-on-chip (SoC) as described in the present disclosure. The computing device can be included in, and / or communicate with, a mobile device, including mobile devices such as cellular phones or smartphones. Processor 104 can include a single computing device operating independently, or can include two or more computing devices operating in cooperation, in parallel, sequentially, etc., where the two or more computing devices can be included in a single computing device or together in two or more computing devices. Processor 104 can interface or communicate with one or more additional devices, as will be described in more detail later, via a network interface device. The network interface device can be utilized to connect processor 104 to one or more of various networks and one or more devices. Examples of network interface devices include, but are not limited to, network interface cards (e.g., mobile network interface cards, LAN cards), modems, and any combination thereof. Examples of networks include wide area networks (e.g., the Internet, corporate networks), local area networks (e.g., networks associated with offices, buildings, campuses, or other relatively small geographic spaces), telephone networks, data networks associated with telephone / voice providers (e.g., mobile communication provider data and / or voice networks), direct connections between two computing devices, and any combination thereof, but are not limited thereto. The network can use wired and / or wireless communication modes. Generally, any network topology can be used.Information (e.g., data, software, etc.) can be communicated between a computer and / or computing device. Processor 104 can include, but is not limited to, for example, a computing device or cluster of computing devices at a first location and a second computing device or cluster of computing devices at a second location. Processor 104 can include one or more computing devices dedicated for data storage, security, distribution of traffic for load balancing, etc. Processor 104 can distribute one or more computing tasks, as described hereinafter, across multiple computing devices that can operate in parallel, in series, redundantly, or in any other manner used for task or memory distribution among computing devices. Processor 104 can be implemented using a “shared nothing” architecture in which data is cached at the worker, which in one embodiment can enable scalability of apparatus 100 and / or computing devices.
[0014] Continuing to refer to FIG. 1, the processor 104 can be designed and / or configured to execute any method, method step, or sequence of method steps in any embodiment described in the present disclosure in any order and any number of repetitions. For example, the processor 104 can be configured to repeatedly execute a single step or sequence until a desired or commanded result is achieved, and the repetition of the step or sequence of steps is executed by repeatedly and / or recursively using the output of a previous repetition as an input to a subsequent repetition, aggregating the input and / or output of the repetition to generate an aggregated result, reducing or decrementing one or more variables such as global variables, and / or dividing a larger processing task into a set of smaller processing tasks that are repeatedly addressed. The processor 104 can execute any step or sequence of steps as described in the present disclosure in parallel, such as using two or more parallel threads, processor cores, etc. to execute a step two or more times simultaneously and / or substantially simultaneously, and the division of tasks between the parallel threads and / or processes can be executed according to any protocol suitable for dividing tasks between repetitions. Those skilled in the art will recognize, upon considering the entire disclosure, various ways in which steps, sequences of steps, processing tasks, and / or data can be subdivided, shared, or otherwise processed using iteration, recursion, and / or parallel processing.
[0015] Continuing to refer to FIG. 1, the apparatus 100 includes a memory. The memory is communicatively coupled to the processor 104. The memory may include instructions that configure the processor 104 to perform the tasks disclosed in the present disclosure. As used in the present disclosure, "communicatively coupled" means being connected by a connection, attachment, or link therebetween that enables the receipt and / or transmission of information between two or more related elements. For example, without limitation, this connection can be between two or more components, circuits, devices, systems, apparatuses, etc., that is wired or wireless, direct, or indirect, and enables the receipt and / or transmission of data and / or signals therebetween. The data and / or signals therebetween can include, but are not limited to, electrical, electromagnetic, magnetic, video, audio, wireless, and microwave data and / or signals, combinations thereof, etc. The communication connection can be achieved, for example, without limitation, directly or through one or more intervening devices or components via wired or wireless electronic communication, digital communication, or analog communication. Further, the communication connection can include electrically coupling or connecting at least one output of one device, component, or circuit to at least one input of another device, component, or circuit. For example, without limitation, via a bus or other facility for mutual communication between elements of a computing device. The communication connection can also include, for example, without limitation, an indirect connection via a wireless connection, wireless communication, low power wide area network, optical communication, magnetic coupling, capacitive coupling, or optical coupling. In some cases, the term "communicatively coupled" can be used in the present disclosure instead of communicatively connected.
[0016] Continuing to refer to FIG. 1, the processor 104 is configured to receive a plurality of data sets 108 from a user. As used in this disclosure, a "data set" is a collection of data. The data set 108 is a structured collection of data that is compiled and presented in a specific format for analysis and interpretation. This consists of individual data points or observations, each representing specific information. The data set 108 can be generated via various means including manual data entry, data collection from sensors or devices, scraping of data from websites, or extraction of data from existing databases. A data set can include a plurality of individual data points, often referred to as records, instances, or observations. Each data point represents a distinct unit of information such as a customer, transaction, measurement, or any other relevant entity.
[0017] Continuing to refer to FIG. 1, the plurality of data sets 108 may include a plurality of metadata. As used in this disclosure, "metadata" refers to descriptive information or attributes that provide context, structure, and meaning to data. Metadata is essentially data about data. Metadata helps in understanding and managing various aspects of data such as its origin, content, format, quality, and use. This plays an important role in effectively organizing, searching, and interpreting data. Metadata can include descriptive metadata, structural metadata, administrative metadata, technical metadata, provenance metadata, usage metadata, and the like. Metadata can be organized and managed via a metadata schema, standard, or framework. These provide guidelines and specifications for capturing, storing, and exchanging metadata in a consistent and structured manner. Common metadata standards include Dublin Core, Metadata Object Description Schema (MODS), and United States Federal Geographic Data Committee (FGDC) metadata standards. In some cases, metadata may be associated with text data or image data. Metadata may also be associated with a pathology slide. Metadata can provide additional descriptive information or attributes linked to the image data or text data associated with the pathology slide. The metadata associated with the plurality of data sets 108 may include patient information. Patient information may include data such as the patient's name, unique patient identifier (ID), age, gender, and any other relevant demographic information. Patient information helps in identifying the slide and associating it with the correct individual's medical record. Metadata may also include case-specific details, which may include information about a particular case or clinical scenario related to the slide. Case-specific details may include information about the case number, attending physician, clinical history, related symptoms, or any other relevant details that help in understanding the context of the slide. In some cases, metadata may include information about the specific specimen type of the slide. This may include the type of tissue or sample that the slide represents. Metadata may include notes, comments, or observations made by a pathologist or other healthcare provider.These annotations can highlight specific features, anomalies, or aspects of note in the slide that are important for interpretation or follow-up analysis. The date and time when the slide was prepared, analyzed, or labeled can be associated as metadata. This information helps to track and maintain a chronological record of slide-related activities. This can be breast tissue, a lung biopsy, a skin lesion, or any other anatomical or pathological specimen. In some embodiments, the metadata can include information regarding staining or preparation techniques, pathological diagnosis, and the like.
[0018] Continuing to refer to FIG. 1, the plurality of data sets 108 includes a first data set 112. As used in this disclosure, a "first data set" is a data set 108 that includes a plurality of data. In one embodiment, the first data set 112 can be a data set 108 that includes a plurality of text data. As used in this disclosure, "text data" is a collection of data consisting of text-based information. Examples of text data can include documents, captions, sentences, paragraphs, free text fields, transcripts, prognosis labels, and the like. In some embodiments, the text data within the first data set can be related to a pathology slide. As used in this disclosure, a "pathology slide" is a slide glass that includes a portion of a biological substance biopsied from a patient. The pathology slide can include tissue biopsied from the patient, and this biopsied tissue is sliced into very thin layers and placed on the slide glass. The first data set 112 can include written descriptions of various aspects of the pathology slide. The first data set 112 can include documents surrounding the pathology slide. These documents can include information regarding the examination, analysis, storage, and disposal of the pathology slide by medical personnel. The first data set 112 includes information that describes and provides the context of the pathology slide, which can enable researchers, clinicians, or data analysts to effectively understand and analyze the slide. By way of non-limiting example, the first data set 112 can include slide identifiers, slide descriptions, clinical information, pathology reports, annotations, notes, laboratory findings, and the like. The first data set 112 can include a brief description or summary of the pathology slide, providing an overview of its content, specimen type, and related characteristics. This description can include details such as tissue type, staining techniques used, and any specific features or abnormalities present on the slide. In other embodiments, the first data set 112 can include a pathology report generated by a medical professional or data scientist after analyzing the slide. This report can provide detailed findings, observations, interpretations, and diagnoses based on the examination of the slide. This can include descriptions of tissue structure, cell morphology, tumor grade classification, stage classification, and other related pathological features.
[0019] Continuing to refer to FIG. 1, the plurality of data sets 108 includes a second data set 116. As used in this disclosure, a "second data set" is a data set 108 that includes a plurality of data. In one embodiment, the second data set 116 can be a data set 108 that includes a plurality of image data. As used in this disclosure, "image data" is a collection of data consisting of data associated with a plurality of images. The second data set 116 can be a collection of a plurality of images compiled and presented in a specific format for analysis, training of a machine learning model, or any other image-related task. This can consist of a variety of images captured from various sources such as digital cameras, satellites, medical imaging devices, microscopes, or other image acquisition methods. The image data can be stored in a digital format such as JPEG, PNG, or TIFF. The image data can include both color images and / or grayscale images. The second data set 116 can include a plurality of medical images. Medical images can include X-rays, CT scans, MRIs, ultrasounds, PET scans, electrocardiogram scans, etc. In some cases, the image data can include either 2D and / or 3D medical images. In one embodiment, the second data set 116 can include a plurality of images associated with a pathology slide. The image data associated with a pathology slide can include a visual representation of the slide captured by an imaging technique. The image data can represent cell structure, tissue morphology, and any pathological features observed on the slide. This can include a series of images that capture the pathology slide. In one embodiment, this can include a plurality of images of the pathology slide at various levels of magnification and various resolutions. The image data associated with a pathology slide can be captured at various magnification levels ranging from low to high magnification. Lower magnification images can provide a wider view of the tissue and can help identify the overall structure, while higher magnification images can provide a more detailed examination of individual cells and cell structures. The image data can include images obtained by digital scanning technology or digital microscopy technology. In some cases, each image can represent a specific pathology slide and include visual information regarding cell structure, tissue morphology, and any abnormalities or features present on the slide.The second dataset 116 may also include metadata that provides additional information regarding each pathological slide image. This metadata can include details such as the slide specimen type (e.g., tissue, cell, biopsy), the staining technique used, the magnification level, the imaging modality, and any relevant contextual information regarding the slide. In some embodiments, optionally, the image data may include annotations or overlays that highlight specific regions or features of interest. These annotations can be added manually by a pathologist or generated by an automated algorithm to assist in the identification and analysis of specific pathological findings.
[0020] Continuing to refer to FIG. 1, the plurality of data sets 108 may include a plurality of media data. As used in this disclosure, "media data" is an element of data associated with a plurality of media. Media can include audio recordings, video recordings, images, digital media, graphs, interrelated data structures, and the like. In some embodiments, the media data may include video data. Video data can capture a sequence of frames that often depicts temporal changes or dynamic processes. Video data can be utilized in applications such as surveillance, action recognition, and motion analysis. Each video within the data set can include a plurality of frames, and the associated metadata can provide details such as duration, frame rate, or time stamps. In some cases, the media data may include audio data. Audio data can include audio files, which are another type of media data found in the data set. The audio files can include recorded sounds, voices, or sounds associated with pathology slides. Audio data is frequently used in speech recognition, speech classification, and acoustic analysis. The metadata of the audio data can include attributes such as duration, sample rate, or audio format. In some cases, the media data may include document data. Document data can include text documents such as PDF, Word files, or plain text files. These documents can include research papers, reports, or any other form of written information. Document data is often used in natural language processing tasks, information retrieval, or text classification. The metadata of the document data can include information such as document title, author, publication date, or word count. In some cases, the media data may be associated with pathology slides. Media data associated with pathology slides can refer to additional information and records related to the slide, such as images, videos, or metadata. These media elements can complement the slide itself and help provide further insights into the underlying pathology.
[0021] Continuing to refer to FIG. 1, the media data may include graphs or other interrelated data structures. In some cases, the contents of the graphs and data structures may be associated with the pathology slides. The media data may include a tissue graph, where individual cells or regions within the tissue are represented as nodes and the connections or edges between them indicate spatial relationships or interactions. This graph structure enables the exploration of cell networks, the patterns of cell organization, and the identification of abnormal cell clusters. Further, the media data may include a diagnostic pathway graph that maps the progression of a diagnosis based on observations and test results. This represents the decision-making process followed by a pathologist and indicates potential branching paths and alternative diagnoses. Data structures such as databases or file systems are used to store and manage the various data associated with the pathology slides. These structures may include patient information, slide metadata, diagnostic annotations, digital images, and other media data. These facilitate the efficient search, compilation, and retrieval of relevant information for clinical purposes, research, and education.
[0022] Continuing to refer to FIG. 1, the processor 104 is configured to identify a first set 120 of associations between a first dataset 112 and a second dataset 116. As used in this disclosure, a "first set of associations" refers to a relationship or connection established between two or more data types. This can include a relationship or connection established between image data and text data. This includes linking or integrating the image data of the second dataset 116 with the descriptive or explanatory text data of the first dataset 112 to provide additional context, enhance understanding, and convey relevant information. In one embodiment, the first set 120 of associations can include identifying an object or set of objects within the image data of the second dataset 116 and using the text data of the first dataset 116 to identify that object. In a non-limiting example, the processor 104 can identify a group of abnormal objects within a plurality of image data associated with the second dataset 116. The processor 104 can identify an object based on a combination of metadata and image data from the second dataset 116, where the metadata identifies the tissue biopsied from the slide and the image data includes an image of the abnormality. The processor 104 can then pair that abnormality with the text data of the first dataset 112, where the text data includes a written description of the abnormality. The pairing of the text data of the first dataset 112 and the image data of the second dataset 116 can be described as the first set 120 of associations. In one embodiment, each association of the first set 120 of associations can include a feature within the image data of the second dataset 116 that correlates to a character string from the first dataset 112, where the character string can linguistically describe this feature. In a non-limiting example, the first set of associations can include a relationship or connection established between several different data types. This can include a relationship between two or more of audio data, image data, text data, video data, graphs, data structures, and the like.
[0023] Continuing to refer to FIG. 1, the processor 104 may be configured to generate a first set 120 of associations using a bootstrap process. As used in this disclosure, a "bootstrap process" is a resampling technique used to estimate the sampling distribution of a statistic or to evaluate the uncertainty associated with a sample. The bootstrap process may include generating multiple resamples of the original dataset by random sampling with replacement. Each resample is the same size as the original dataset, but some observations may appear multiple times and other observations may be excluded. This process enables the creation of a pseudo-population from which statistical estimates can be derived. Once the resamples are obtained, the desired statistic is calculated for each resample. This statistic can be an average, median, standard deviation, correlation coefficient, or any other relevant measure. By repeating this resampling process many times (often thousands of times), a distribution of the statistic, known as the bootstrap distribution, is obtained. In this case, the bootstrap process of the processor 104 can start by generating multiple resamples of the dataset. For example, each resample can consist of paired samples of text data and image data, which can be any other data referred to herein. These resamples are created by randomly selecting instances from the first and second datasets with replacement, ensuring that both text data and image data are held together in each resample. For each resample, the text data and image data pair are analyzed together to explore the relationship or association between them. Various techniques can be applied based on a particular task or objective. The strength of the relationship or association between the text data and the image data can be evaluated by measuring a performance metric or a statistical measure. For a classification task, accuracy, precision, recall, or an F1 score can be calculated. Alternatively, a correlation coefficient, mutual information, or other statistical measure can be used to quantify the relationship between the two data types.If the strength of the relationship or association between text data and image data exceeds a predetermined threshold, a first association can be created between the text data and the image data. The bootstrap process may be repeated multiple times, generating a different resample each time. This repetition enables the estimation of variability and uncertainty in the relationship or association metric. By analyzing the results across the resampled datasets, confidence intervals can be constructed, hypothesis tests can be performed, or stability assessments can be conducted to evaluate the significance and robustness of the relationship.
[0024] Continuing to refer to FIG. 1, the first set of associations 120 may include one or more direct correlations between the first data set 112 and the second data set 116. As used in this disclosure, "direct correlation" is a direct correspondence or alignment between two or more data types. This may include a direct correspondence or alignment between the visual content of image data and the accompanying text data. The direct correlation between the first data set 112 and the second data set 116 may include a clear and explicit correspondence between the image data and the accompanying text data. In a non-limiting example, the direct connection may include pairing an image with a caption that accurately describes an object, scene, or visual feature depicted in the image. The text data is directly related to and closely aligned with the visual information within the image. In the context of a pathology slide, this connection is intended to provide a comprehensive description of the visual pathological features observed in the second data set 116 via the relevant text data of the first data set. Non-limiting examples of direct correlations may include annotations of pathology slides. The image data of the second data set 116 may be annotated with text data from the first data set 112 that explicitly describes the features, structures, or abnormalities observed within the image. The annotation functions as a direct link between the image data and the associated text data and may provide a specific description aligned with the visual content. The direct correlation may include pairing a prognosis label with all or a portion of the image. The prognosis label may describe abnormalities, inflammation, coloring, size, tissue structure, etc. The prognosis label functions as a direct representation of the visual findings in the image data.
[0025] Continuing to refer to FIG. 1, the processor 104 can generate a first set 120 of associations by extracting visual features from the second dataset 116. The processor 104 can extract visual features from the image data using machine vision. As used in this disclosure, "visual features" are one or more objects of interest located within the image data. Visual features within the image data can include the presence or absence of cell structures or tissue morphology. The second dataset 116 can include microscopic images of tissue samples, and the visual features can include various elements that provide information regarding cell composition and cell constitution. Non-limiting examples of visual features can include cell nuclei, tissue structures, cell arrangements, cell differentiation, inflammatory cells, cell abnormalities, staining patterns, and the like. In some cases, the morphology and characteristics of the cell nucleus can serve as visual features. This can include any abnormalities such as size, shape, chromatin pattern, presence of nucleoli, and nuclear enlargement or irregularity. In other cases, the arrangement and organization of tissue components can be identified as visual features. This can include the presence of glands, tubules, ducts, blood vessels, or other anatomical structures within the tissue sample. The spatial distribution and relationships between cells can also be identified as visual features. The spatial distribution and relationships can indicate a particular pathological state. Features such as cell crowding, overlapping, irregular patterns, or loss of normal tissue structure can be observed and analyzed. The degree of cell differentiation or maturation can be a visual feature evaluated by a pathologist. This can include evaluating the similarity between a cell and its normal counterpart and identifying any abnormal or undifferentiated cells. The presence and distribution of inflammatory cells such as lymphocytes, neutrophils, or macrophages within the tissue sample can be visual features indicating an immune response or inflammation. Visual features can include the presence of abnormal cells such as cancer cells or cells having abnormal shapes, sizes, or staining patterns. These abnormalities can indicate a neoplastic state or other pathological processes. Different staining techniques are used in pathology to highlight specific components or structures. Staining patterns such as eosinophilia, basophilia, or immunohistochemical staining can serve as visual features for identifying specific cell characteristics or pathological markers.
[0026] Continuing to refer to FIG. 1, the processor 104 may identify visual features within the second data set 116 using a machine vision system. The machine vision system may identify characteristic points or regions within an image that can be used as visual features. The machine vision system may use images from the second data set 116 to make determinations regarding scenes, spaces, and / or objects within the image data. For example, in some cases, the machine vision system can be used for world modeling or alignment of objects within a space. In some cases, alignment may include, but is not limited to, image processing such as object recognition, feature detection, edge / corner detection, etc. Non-limiting examples of feature detection may include Scale-Invariant Feature Transform (SIFT), Canny edge detection, Shi Tomasi corner detection, etc. In some cases, alignment may include one or more transforms for orienting a camera frame (or image or video stream) with respect to a three-dimensional coordinate system, and exemplary transforms include, but are not limited to, homography transforms and affine transforms. In one embodiment, alignment of the first frame to the coordinate system can be verified and / or corrected using object identification and / or computer vision as described above. However, for example, but not limited to, an initial alignment to two dimensions represented as alignment to, for example, the x and y coordinates may be performed using a two-dimensional projection of a three-dimensional point onto the first frame. A third dimension of alignment representing depth and / or the z-axis can be detected by comparing the two frames. For example, if the first frame includes a pair of frames captured using a pair of cameras (e.g., a stereo camera, also referred to as a stereo camera in the present disclosure), image recognition and / or edge detection software can be used to detect a pair of stereograms of an object's image, and the two stereograms can be compared to derive the z-axis value of a point on the object, for example, using interpolation to enable derivation of additional z-axis points inside and / or around the object. This can be repeated using multiple objects within the field of view, including, but not limited to, environmental features of interest identified by an object classifier and / or indicated by an operator.In one embodiment, the x and y axes can be selected so as to span a common plane for two cameras used for stereoscopic image capture and / or the xy plane of the first frame, and as a result, as described above, for the affine transformation of the coordinates of an object, the x and y translation components and φ can be pre-entered into the translation matrix and the rotation matrix. The initial x and y coordinates and / or estimations in the transformation matrix may alternatively or additionally be performed between the first frame and the second frame, as described above. As described above, for a plurality of points on the object and / or each point of the edge and / or plurality of edges of the object, the initial estimated value of the z coordinate can be input to the x and y coordinates of the first stereoscopic frame based on assumptions about the object, such as the assumption that the ground is substantially parallel to the xy plane as selected above. Then, the Z coordinate aligned using the image capture and / or object identification process as described above, and / or the x, y, and z coordinates, can be compared with the coordinates predicted using the initial guess in the transformation matrix, and an error function can be calculated by comparing the two sets of points, and new x, y, and / or z coordinates can be iteratively estimated and compared until the error function falls below a threshold level. In some cases, the machine vision system can use a classifier, such as any of the classifiers described throughout the present disclosure.
[0027] Continuing to refer to FIG. 1, the processor 104 can match the visual features extracted from the second data set 116 to similar visual features. Feature matching involves comparing visual features extracted from different image data to identify corresponding or similar features. This aims to establish correspondences between features in different image data, which can be useful for tasks such as image alignment, object recognition, or image search. Once the extracted visual features are identified, the processor 104 can generate a description of the extracted visual features. The description can capture the local appearance or characteristics around each visual feature. The description can include information such as which type of tissue is located within the slide, the description of any visual feature, etc. The description can include information regarding the shape, texture, color, inflammation, staining pattern, or gradient of the region surrounding the visual feature. The processor 104 can pair the description of the extracted visual features with the descriptions of other previously identified visual features. In some cases, the processor 104 can compare a measure of similarity or correspondence between visual features in different pathological slides. This information can be used to quantify the similarity or dissimilarity between slides based on shared or distinct visual features. This enables comparing slides based on common morphological patterns, cell characteristics, or other visual attributes. In some embodiments, visual feature pairing can be used to find similar or corresponding regions of interest (ROIs) in different pathological slides. For example, in tumor detection or tracking, matching features in different slides can identify the same tumor region over time or across different patient samples.
[0028] Continuing to refer to FIG. 1, the processor 104 may be configured to generate a plurality of named entities in response to the first dataset 112 using a named entity recognition process. As used in the present disclosure, a "named entity" is a particular type of word or phrase that represents a real-world object having a unique identity. A named entity can be a person, place, idea, concept, or thing that indicates a particular individual, organization, location, date, time, product, event, quantity, disease, tissue sample, and other entities that can be uniquely identified. These entities play an important role in understanding the context and extracting meaningful information from text. A named entity can provide context information and function as a reference point for understanding the meaning and relationships within the text. Recognizing and extracting named entities from text data is a fundamental task in various other applications where it is important to understand natural language processing (NLP), information extraction, text mining, and the semantics of text to identify key elements. The named entities generated from the text data associated with the first dataset 112 can include specific terms or entities that provide information about the slide, its characteristics, or the observed pathological findings. Non-limiting examples of named entities can include diseases, condition names, tissue or organ names, cell types, cell structures, staining methods, staining techniques, gene names, protein names, diagnostic terms, medical abbreviations, and the like.
[0029] Continuing to refer to FIG. 1, the processor 104 can be configured to generate a plurality of named entities using a named entity recognition (NER) system. As used in this disclosure, a "named entity recognition (NER) system" is software that identifies a plurality of named entities from text. The NER system can be configured to identify a plurality of named entities from a first dataset 112. The input to the NER system can include a plurality of datasets 108, the first dataset 112, metadata, text data, and the like. The output of the named entity recognition system can include a plurality of named entities. A named entity can typically include a structured representation of the identified named entity in the form of an annotation or tag attached to the original text.
[0030] Continuing to refer to FIG. 1, the NER system can generate a plurality of named entities using a natural language processing model. As used in this disclosure, a "natural language processing (NLP) model" is a computational model designed to process and understand human language. This harnesses techniques from machine learning, linguistics, and computer science to enable a computer to understand, interpret, and generate natural language text. The NLP model may preprocess the text data, where the input text may include the first dataset 112, or any other data referred to herein. Preprocessing the input text may include tasks such as tokenization (splitting the text into individual words or subword units), normalizing the text (lowercasing, removing punctuation, etc.), and encoding the text into a numerical representation suitable for the model. The NLP model may include a transformer architecture, where a transformer is a deep learning model that uses an attention mechanism to capture the relationships between words or subword units within a text sequence. These consist of multiple layers of self-attention and feed-forward neural networks. The NLP model can weight the importance of different words or subword words within a text sequence while considering the context. This enables the model to capture the dependencies and relationships between words while considering both local and global context. This process can be used to identify a plurality of named entities. The language processing model is automatically generated by the processor 104 and / or the named entity recognition system, generates an association between one or more important terms extracted from the first dataset 112, and may include a program that detects associations between such important terms, including but not limited to, mathematical associations. The associations between language elements (language elements include important terms extracted for the purposes of this specification), the relationships of such categories to other such terms may include, but are not limited to, mathematical associations including statistical correlations between any language element and any other language element and / or multiple language elements.Statistical correlations and / or mathematical associations can include, for example, probabilistic formulas or relationships that indicate the likelihood that a given extracted key term indicates a given category of semantic meaning. As a further example, statistical correlations and / or mathematical associations can include probabilistic formulas or relationships that indicate positive and / or negative associations between at least the extracted key terms and / or given semantic relationships, where the positive or negative indicators can include an indication of whether a given document exhibits a categorical semantic relationship. Whether a phrase, sentence, word, or other text element within the first dataset 112 constitutes a positive or negative indicator can, in one embodiment, be determined by, for example, a mathematical association between detected key terms, a comparison with phrases and / or words indicating positive and / or negative indicators stored in the memory of the processor 104, and the like.
[0031] Continuing to refer to FIG. 1, the processor 104 can generate a first set 120 of associations by pairing one or more visual features with a plurality of named entities. In some cases, the first set 120 of associations can include a co-expression of the named entities and the one or more visual features. The visual features of the second dataset 116 and the named entities of the first dataset 112 may be combined to create a co-expression of the prognostic slide. In some cases, the co-expression can be used as training data for a machine learning model, as described later in this specification. In some cases, the pairing of one or more visual features with a plurality of named entities can include that the annotations overlaid on the image function as a direct correlation between the visual data and the text data. These annotations can highlight specific regions or structures of interest within the pathology slide, and the corresponding text provides a detailed description or diagnosis associated with those regions. The annotations can indicate the presence of a tumor, grade classification or stage classification information, or specific abnormalities observed. In other cases, the pairing of one or more visual features with a plurality of named entities can include a pathology report that directly describes the visual features observed in the image. This report can provide detailed findings, interpretations, and diagnoses based on the examination of the pathology slide. This aligns directly with the visual information describing the tissue structure, cell morphology, tumor characteristics, and any other relevant pathological observations.
[0032] Continuing to refer to FIG. 1, the processor 104 can generate a first set 120 of associations using a first association classifier 124. As used in this disclosure, a "first association classifier" is a classifier configured to generate a first set 120 of associations. The first association classifier 124 may be consistent with the classifier described later in FIG. 2. Inputs to the first association classifier 124 may include a plurality of data sets 108, a first data set 112, a second data set 116, examples of the first set 120 of associations, and the like. Outputs from the first association classifier 124 may include a first set 120 of associations adjusted to match the first data set 112 and the second data set 116. The first association training data may include a plurality of data entries including a plurality of inputs correlated with a plurality of outputs for training the processor by a machine learning process. In one embodiment, the first association training data may include a plurality of first data sets 112 and a plurality of second data sets 116 correlated with examples of the first set 120 of associations. In another embodiment, the first association training data may include a plurality of named entities correlated with a plurality of visual features. The first association training data may be received from the database 300. The first association training data may include information regarding a plurality of data sets 108, a first data set 112, a second data set 116, examples of the first set 120 of associations, and the like. In one embodiment, the first association training data may be iteratively updated according to the input and output results of a past first association classifier or any other classifier mentioned throughout this disclosure. Classifiers may use, but are not limited to, linear classifiers such as logistic regression classifiers and / or naive Bayes classifiers, nearest neighbor classifiers such as k-nearest neighbor classifiers, support vector machines, least squares support vector machines, Fisher's linear discriminant, quadratic classifiers, decision trees, boosted trees, random forest classifiers, and the like.
[0033] Continuing to refer to FIG. 1, the processor 104 can assign a pathology identifier 128 to each visual feature or named entity. As used in this disclosure, a "pathology identifier" is a unique identification code or label assigned to both the image data and text data associated with a pathology slide. This identifier functions as a common reference or link between the image data and the corresponding text information, enabling their association and retrieval. The pathology identifier 128 can represent the content of both the image data and the text data. In a non-limiting example, the pathology identifier 128 can identify both an image and text containing information about the structure of the tissue within the pathology slide. The pathology identifier 128 can be a numeric code, alphanumeric code, barcode, QR code (registered trademark), etc. By using the same identifier for both modalities, a direct or indirect association is established, ensuring that the image data and text data represent the same pathology slide. In some cases, the pathology identifier 128 can be used to represent a first set 120 of associations or a second set of associations. The pathology identifier 128 may be assigned to the first set 120 of associations or the second set 132 of associations using the processor 104. This identifier functions as a common reference point for establishing a connection between the visual content captured in the image data and the corresponding text information such as pathology reports or clinical notes and images associated with the pathology slide. This identifier facilitates seamless integration and retrieval of information, enabling efficient compilation, analysis, and retrieval of pathology data. Researchers, pathologists, and healthcare providers can use the pathology identifier 128 to search for and access the image data and text data together, enabling a comprehensive understanding of the content of the pathology slide and facilitating accurate diagnosis, research, and collaboration.
[0034] Continuing to refer to FIG. 1, the processor 104 is configured to identify a second set 132 of associations between the first data set 112 and the second data set 116 in response to a first set 120 of associations. As used in this disclosure, "second set of associations" refers to a relationship or connection established between two or more types of data. This may include a relationship or connection established between image data and text data. The second set 128 of associations may be generated in a similar manner as the first set of associations. The second set 132 of associations may include linking or integrating the image data of the second data set 116 with the descriptive or explanatory text data of the first data set 112 to provide additional context, enhance understanding, and convey relevant information. The second set 132 of associations may be similar to the first set 120 of associations. However, the second set 132 of associations may include multiple indirect correlations between the first data set 112 and the second data set 116. As used in this disclosure, "indirect correlation" refers to a relationship in which the first data set 112 provides text data related to the image data of the second data set 112 or supplementary information for its image data, but does not explicitly describe the visual features observed in the image. Indirect correlations may provide context information related to the image content. This may include patient demographics, medical history, chief symptoms, diagnosis, treatment details, or other relevant clinical information associated with the pathology slide. Although not directly describing the visual appearance, this information helps to provide context and insight into the underlying pathology. In one embodiment, the indirect correlation may include diagnostic findings or interpretations made by a pathologist or healthcare provider based on the examination of the pathology slide. These findings may not explicitly describe the visual appearance and provide insight into the presence of abnormalities, disease states, tumor grade or stage classification, or other diagnostic observables related to the image data. The indirect correlation between the first data set 112 and the second data set 116 may enable a broader understanding of the clinical context, the patient's medical history, and the diagnostic interpretation associated with the visual information captured in the image.Text data may not explicitly describe visual features while providing supplementary information that enhances the interpretation and analysis of pathological slides.
[0035] Continuing to refer to FIG. 1, the processor 104 is configured to identify a second set 132 of associations between a first dataset 112 and a second dataset 116 as a first set 120 of functions of associations. The processor 104 may generate a second set 132a of associations as a function of a plurality of pathology identifiers 128. In one embodiment, the processor 104 pairs the data of the first dataset 112 and the second dataset 116 based on the degree of similarity by the pathology identifier 128. The processor 104 can generate a similarity score between the query image features and the features of other text descriptions in the dataset, and vice versa. Thereby, the processor 104 can identify the most similar and relevant image-text pairs. In some embodiments, the processor can perform image-text similarity calculations or text-image similarity calculations. Using these calculations, it is possible to evaluate the degree of similarity between the content of the image data and the content text data. As used in the present disclosure, a "similarity score" is a score that reflects the degree of similarity between the content of the image data and the content of the text data. The similarity score can be generated by preprocessing both the image data and the text data. The preprocessing may include resizing, normalizing, and extracting relevant visual features using computer vision techniques such as convolutional neural networks (CNNs). Alternatively, preprocessing the text data can be done by tokenizing, removing stop words, and converting it to a numerical representation using techniques such as word embedding or language models. Next, the processor 104 can associate each visual feature and named entity with a unique pathology identifier 128. This identifier functions as a link between the image and text data associated with the same pathology slide. Next, the processor 104 generates a similarity score between the visual features of the second dataset 116 and the text features of the first dataset 112. Various similarity metrics can be used, such as cosine similarity, Euclidean distance, or other distance measurements that capture the similarity between feature vectors.Next, the processor 104 can sort the pairs of visual features and named entities in descending order based on their similarity scores. Thereby, the most similar image-text pairs are identified. The processor 104 can further filter the associations by considering only the pairs that share the same pathology identifier. Thereby, it is ensured that the generated associations are based on the content similarity of the images and text data associated with the same pathology slide.
[0036] Continuing to refer to FIG. 1, the processor 104 can generate a second set 132 of associations using a second association classifier 136. As used in this disclosure, a "second association classifier" is a classifier configured to generate a second set 132 of associations. The second association classifier 136 may be consistent with the classifier described later in FIG. 2. Inputs to the second association classifier 136 can include a plurality of data sets 108, a first data set 112, a second data set 116, a first set 120 of associations, visual features, named entities, pathology identifiers, examples of the second set 132 of associations, and the like. Outputs from the second association classifier 136 can include a second set 136 of associations adjusted to match the first data set 112 and the second data set 116. Outputs from the second association classifier 136 can further include similarity scores. The second association training data can include a plurality of data entries including a plurality of inputs correlated with a plurality of outputs for training the processor by a machine learning process. In one embodiment, the second association training data can include a first set 120 of associations correlated with examples of the second set 136 of associations. In another embodiment, the second association training data can include a plurality of visual features correlated with a plurality of named entities. The second association training data can be received from the database 300. The second association training data can include information regarding a plurality of data sets 108, a first data set 112, a second data set 116, a first set 120 of associations, visual features, named entities, pathology identifiers, examples of the second set 132 of associations, and the like. In one embodiment, the second association training data can be iteratively updated according to the input and output results of a past first association classifier or any other classifier mentioned throughout this disclosure. Classifiers can use, but are not limited to, linear classifiers such as, but not limited to, logistic regression classifiers and / or naive Bayes classifiers, nearest neighbor classifiers such as k-nearest neighbor classifiers, support vector machines, least squares support vector machines, Fisher's linear discriminant, quadratic classifiers, decision trees, boosted trees, random forest classifiers, and the like.
[0037] Continuing to refer to FIG. 1, the processor 104 can generate a second set 132 of associations using comparative fuzzy inference. As used in this disclosure, "comparative fuzzy inference" is a method of interpreting values within an input vector (i.e., visual features and named entities) and assigning values to an output vector based on a set of rules. The set of fuzzy rules can include a set of linguistic variables that describe how the system should make decisions regarding the classification of inputs or the control of outputs. Fuzzy inference rules operate on fuzzy sets and provide a framework for mapping input variables to output variables via linguistic rules. Fuzzy inference rules can operate using linguistic variables that represent imprecise or vague concepts rather than exact numerical values. Linguistic variables are defined by membership functions that describe the degree of membership or truth for different linguistic terms or categories. In a non-limiting example, the linguistic variables associated with the second set 132 of associations can have linguistic terms such as "highly relevant," "moderately relevant," and / or "not relevant," each having its corresponding membership function. Fuzzy inference rules typically follow a conditional "IF-THEN" structure. This consists of an antecedent (the IF part) and a consequent (the THEN part). The antecedent specifies the conditions or criteria based on which the rule is applied, and the consequent determines the output or conclusion of the rule. In one embodiment, the second set 132 of associations can be determined by comparing the degree of match between a first fuzzy set and a second fuzzy set and / or between single values within or among them, which is sufficient for the purposes of the matching process.
[0038] Referring further to FIG. 1, a second set 132 of associations can be determined in response to an intersection between two fuzzy sets, where each fuzzy set can represent a visual feature and a named entity, respectively. Comparing the visual feature and the named entity can include using a fuzzy set inference system as described hereinafter in this specification, or any scoring method as described throughout the present disclosure. For example, but not limited to, the processor 104 can use a fuzzy logic model to determine a first set 120 of associations or a second set 136 of associations according to the fuzzy set comparison techniques as described in the present disclosure. In some embodiments, each piece of information associated with a visual feature can be compared with a named entity, and the second set 132 of associations can be represented using linguistic variables over a range of potential numerical values, where the values of the linguistic variables can be represented as fuzzy sets over that range, and a "good" or "ideal" fuzzy set can correspond to a range of values that can be characterized as being ideal, while other fuzzy sets can correspond to ranges and / or values that can be characterized as normal, bad, or not otherwise ideal ranges and / or values. In an embodiment, these variables can be used to compare the visual feature and the named entity to determine a second set 132 of associations specific to the visual feature. The fuzzy inference system combines such linguistic variable values according to one or more fuzzy inference rules including any type of fuzzy inference system and / or rules as described in the present disclosure to determine the degree of membership in one or more output linguistic variables having values representing ideal overall performance, normal or moderate overall performance, and / or low or inadequate overall performance, and such mapping can be "defuzzified" as described in more detail hereinafter to provide an overall output and / or evaluation.
[0039] Referring further to FIG. 1, the processor may be configured to generate a machine learning model, such as a second association classifier 136, using a Naive Bayes classification algorithm. The Naive Bayes classification algorithm generates a classifier by assigning class labels to problem instances represented as vectors of element values. The class labels are drawn from a finite set. The Naive Bayes classification algorithm may include generating a family of algorithms that assume that the value of a particular element is independent of the value of any other element given the class variable. The Naive Bayes classification algorithm may be based on Bayes' theorem, expressed as P(A / B) = P(B / A)P(A) ÷ P(B), where P(A / B) is the probability of hypothesis A given data B, also known as the posterior probability, P(B / A) is the probability of data B assuming hypothesis A is true, P(A) is the probability that hypothesis A is true regardless of the data, also known as the prior probability of A, and P(B) is the probability of the data regardless of the hypothesis. The Naive Bayes algorithm can be generated by first converting the training data into a frequency distribution table. The processor 104 can then calculate a likelihood table by calculating the probabilities of different data entries and classification labels. The processor 104 can calculate the posterior probability of each class using the Naive Bayes equation. The class with the highest posterior probability is the result of the prediction. The Naive Bayes classification algorithm may include a Gaussian model that follows a normal distribution. The Naive Bayes classification algorithm may include a multinomial model used for discrete counts. The Naive Bayes classification algorithm may include a Bernoulli model that can be used when the vector is binary.
[0040] Referring further to FIG. 1, the processor 104 may be configured to generate a machine learning model, such as a second association classifier 136, using a k-nearest neighbor (KNN) algorithm. As used in this disclosure, the "k-nearest neighbor algorithm" utilizes feature similarity to analyze how similar out-of-sample features are to training data and classify input data into one or more clusters and / or categories of features as represented by the training data, which may be performed by representing both the training data and the input data in vector form, identifying classifications within the training data using one or more measures of vector similarity, and determining the classification of the input data. The k-nearest neighbor algorithm may include specifying a k-value, or a number that instructs the classifier to select the k most similar entry training data for a given sample, determining the most common classifier of the entries within the database, and classifying known samples, which may be performed recursively and / or iteratively to generate a classifier that may be used to classify the input data as a further sample. For example, an initial set of samples may be performed to cover initial heuristics and / or a "first guess" in the output and / or relationships, which may be seeded using expert input received according to any process, including but not limited to, as described herein. By way of non-limiting example, the initial heuristic may include a ranking of associations between the input and elements of the training data. The heuristic may include selecting some of the top-ranked associations and / or training data elements.
[0041] Continuing to refer to FIG. 1, generating the k-nearest neighbor algorithm can generate a first vector output including data entry clusters, generate a second vector output including input data, and calculate the distance between the first vector output and the second vector output using any suitable norm such as cosine similarity, Euclidean distance measurement, etc. Each vector output can be represented as an n-tuple of values, where n is at least two values. Each value of the n-tuple of values can represent a measurement or other quantitative value associated with a given category or attribute of the data, examples of which are provided in more detail below, and the vectors can be represented in n-dimensional space using axes for each category of values within the n-tuple of values such that the vectors have a geometric direction characterizing the relative amounts of the attributes within the n-tuple as compared to each other. Two vectors can be considered equivalent if their directions and / or the relative amounts of the values within each vector as compared to each other are the same. Thus, as a non-limiting example, a vector represented as [5,10,15] can be treated as equivalent to a vector represented as [1,2,3] for the purposes of the present disclosure. Vectors can be more similar if their directions are more similar and more different if their directions are more diverse, but vector similarity can alternatively or additionally be determined using the average of the similarities between similar attributes, or any other measure of similarity suitable for an n-tuple of values, or an aggregation of numerical similarity measurements for the purpose of a loss function as described in more detail below. Any vector as described herein can be scaled such that each vector represents each attribute along an equivalent scale of values. Each vector can be "normalized" or divided by a "length" attribute such as a length attribute l derived using the Pythagorean norm
Number
[0042] Continuing with reference to FIG. 1 , the processor 104 may be configured to generate the third dataset 140 using the second set of associations 132. As used in this disclosure, a “third dataset” is a dataset 108 that includes a plurality of text data and / or a plurality of image data. The third dataset 140 is generated based on the second set of associations 132. In one embodiment, the text data and / or image data may be used in conjunction with the second set of associations 132 to identify additional text data or image data that have an indirect or direct correlation to the initial text data and / or image data. The additional text data and / or image data may be identified from the first dataset 112, the second dataset 116, or the database 300. The processor 104 may identify the third dataset using a cross-modal search process. As used in this disclosure, a “cross-modal search process” refers to a process of searching for data from one modality (e.g., image) based on a query from another modality (e.g., text), or vice versa. The cross-modal search process involves finding relevant instances in one modality that are semantically related to a query in a different modality. The cross-modality search process may be performed in response to a pathology identifier 128. The processor 104 may assign a pathology identifier 128 to the initial text data and / or image data. The pathology identifier 128 associated with the initial text data and / or image data may be grouped with other related pathology identifiers 128 to identify image data and / or text data in the third dataset 140. In a non-limiting example, the processor 104 may receive image data depicting inflamed tissue along with the second set of associations 132. The processor uses the cross-modal search process to generate the third dataset 140, which includes a plurality of text data indirectly correlated to the image data.
[0043] Continuing to refer to FIG. 1, the processor 104 can generate a third data set 140 using a data set classifier. As used in this disclosure, a "data set classifier" is a classifier configured to generate a third data set 140. The data set classifier may be consistent with the classifier described later in FIG. 2. Inputs to the data set classifier may include a plurality of data sets 108, a first data set 112, a second data set 116, a first set of associations 120, visual features, named entities, pathology identifiers, a second set of associations 132, examples of the third data set 140, and the like. Outputs from the data set classifier may include a third data set 140 that includes text data and / or image data. The data set training data may include a plurality of data entries that include a plurality of inputs that correlate to a plurality of outputs for training the processor by a machine learning process. In one embodiment, the data set training data may include a second set of associations 136 that correlates to examples of the third data set 140. In another embodiment, the data set training data may include a plurality of visual features that correlate to a plurality of named entities. The data set training data may be received from the database 300. The data set training data may include information regarding a plurality of data sets 108, a first data set 112, a second data set 116, a first set of associations 120, visual features, named entities, pathology identifiers, a second set of associations 132, examples of the third data set 140, and the like. In one embodiment, the data set training data may be iteratively updated according to the input and output results of a past first association classifier or any other classifier mentioned throughout this disclosure. Classifiers may use, but are not limited to, linear classifiers such as, but not limited to, logistic regression classifiers and / or naive Bayes classifiers, nearest neighbor classifiers such as k-nearest neighbor classifiers, support vector machines, least squares support vector machines, Fisher's linear discriminant, quadratic classifiers, decision trees, boosted trees, random forest classifiers, and the like.
[0044] Referring further to FIG. 1, the processor 104 may be configured to display a second set 132 of associations using the display device 144. As used in this disclosure, a "display device" is a device used to display content. The display device 144 may include a user interface. A "user interface," as used herein, is a means by which a user and a computer system interact with each other, for example, through the use of input devices and software. The user interface may include a graphical user interface (GUI), a command line interface (CLI), a menu-driven user interface, a touch user interface, a voice user interface (VUI), a form-based user interface, any combination thereof, and the like. The user interface may include a smartphone, a smart tablet, a desktop, or a laptop operated by a user. In one embodiment, the user interface may include a graphical user interface. A "graphical user interface (GUI)" is a graphical form of a user interface that enables a user to interact with an electronic device, as used herein. In some embodiments, the GUI may include icons, menus, other visual indicators, or representations (graphics), audio indicators such as primary notations, as well as display information and associated user controls. A menu may include a list of options and may enable a user to select one of them. A menu bar, such as a pull-down menu, may be displayed horizontally across the screen. When any option in this menu is clicked, a pull-down menu may appear. A menu may include a context menu that appears only when a user performs a specific action. This example is pressing the right mouse button. When this is done, a menu may appear under the cursor. Files, programs, web pages, etc. can be represented using small pictures within the graphical user interface. For example, a link to a decentralized platform as described in this disclosure may be incorporated using an icon.Using icons can be a quick way to open documents, run programs, etc., because they provide immediate access when clicked. The information included in the user interface may be directly affected using graphical control elements such as widgets. As used herein, a "widget" is a user control element that enables a user to control and change the appearance of elements within the user interface. In this context, a widget may refer to a general GUI element such as a checkbox, button, or scrollbar to an instance of that element, or a customized set of such elements used for a particular function or application (such as a dialog box for a user to customize the appearance of a computer screen). User interface controls may include software components with which a user interacts directly via manipulation to read or edit information displayed via the user interface. Widgets may be used to display lists of related items, navigate the system using links, tabs, and manipulate data using checkboxes, radio boxes, etc.
[0045] Continuing to refer to FIG. 1, an apparatus 100 for detecting associations between different types of datasets may include at least a processor 104 and a memory communicatively coupled to the at least one processor. The memory may include instructions that configure the processor 104 to receive a plurality of datasets 108. The plurality of datasets 108 may include a first dataset 112 and a second dataset 116. The instructions may instruct the processor to identify a first set 120 of associations between a first subset of the first dataset 112 and a first subset of the second dataset 116. In some cases, generating the first set of associations may include identifying a plurality of named entities within the first dataset 112, each of the plurality of named entities being associated with at least a data element within the second dataset 116. For example, the first set of associations may be made via explicit naming of the configuration data within one or more of the first dataset 112 and the second dataset 116. Alternatively or additionally, the first set of associations may be received from expert input or user input.
[0046] Continuing to refer to FIG. 1, the instructions may further instruct the processor 104 to generate a second set 132 of associations between a second subset of the first dataset 112 and a second subset of the second dataset 116. In some cases, one or more of the second subset of the first dataset 112 and the second subset of the second dataset 116 are larger (e.g., have more data or data elements) than one or more of the first subset of the first dataset 112 and the first subset of the second dataset 116. Similarly, in some cases, the second subset may include some or all of the first subset of the first dataset 112 and / or the second dataset. In some cases, the generation of the second set 132 of associations may be performed in response to the first set 120 of associations, for example, by using a second association classifier 136. In some cases, generating the second set 132 of associations may include training the second association classifier 136 using second association training data 300. The second association classifier 136 may include any classifier described in this disclosure, for example, with reference to FIG. 2. The second association training data may include any training data described in this disclosure, for example, with reference to FIG. 2. In some versions, the second association training data may include a plurality of data entries including the first set 120 of associations. In some cases, generating the second set 132 of associations may be performed in response to the first set of associations using the trained second association classifier 136. Finally, the instructions may instruct the processor 104 to display the second set of associations using a display device. In some cases, displaying the second set of associations may include displaying only a portion of the second set 132 of associations and / or displaying data that requires the second set 132 of associations but does not explicitly represent the second set 132 of associations.
[0047] Referring further to FIG. 1, in some embodiments, the second association training data 300 may include a first subset of the first data set 112 as an input that correlates to a first subset of the second data set 116 as an output, i.e., the second association training data 300 may include a first set of associations 120.
[0048] Referring further to FIG. 1, in some embodiments, generating the second set of associations 132 may further include training a generative machine learning process using the first data set 112, synthesizing first synthetic data according to the first data set using generative machine learning, and generating the second set of associations 132 according to the first synthetic data and the first set of associations. The generative machine learning process may include any generative machine learning process described in the present disclosure with reference to, for example, FIG. 2. The synthetic data may include any generated data described in the present disclosure with reference to, for example, FIG. 2. In some cases, the second association training data may include first synthetic data as an input that correlates to one or more subsets of the first or second data set as an output.
[0049] Referring further to FIG. 1, in some embodiments, one or more of the first data set and the second data set may include text or text data. In some cases, generating the second set of associations 132 may further include associating text data within one or more of the first data set 112 and the second data set 116 using a natural language processing model, and generating the second set of associations 132 according to the associated text data. The natural language processing model may include any language processing model or process described in the present disclosure with reference to, for example, FIG. 2.
[0050] Referring further to FIG. 1, in some embodiments, generating the second set 132 of associations may further include calculating distances between data elements within one or more of the first data set 112 and the second data set 116, and generating the second set 132 of associations according to the distances between the data elements. The distance may include any distance described in the present disclosure, such as a vector distance, including the disclosure in FIG. 2.
[0051] Referring further to FIG. 1, in some embodiments, one or more of the first data set 112 and the second data set 116 include metadata, and generating the second set 132 of associations may further include associating metadata within one or more of the first data set and the second data set, and generating the second set 132 of associations according to the associated metadata. The metadata may include any metadata or context data described in the present disclosure, for example, with reference to FIG. 2.
[0052] Referring further to FIG. 1, in some embodiments, one or more of the first data set 112 and the second data set 116 may include image data, and generating the second set 132 of associations may include using a machine vision system to identify one or more visual features within one or more of the first data set and the second data set.
[0053] Referring now to FIG. 2, an exemplary embodiment of a machine learning module 200 capable of executing one or more machine learning processes as described in the present disclosure is shown. The machine learning module may use machine learning processes to perform determination, classification, and / or analysis steps, methods, processes, etc. as described in the present disclosure. A "machine learning process," as used in the present disclosure, is a process that automatically uses training data 204 to generate an algorithm instantiated with hardware or software logic, data structures, and / or functions executed by a computing device / module to generate an output 208 when provided with data provided as an input 212, which is in contrast to a non-machine learning software program in which the commands to be executed are pre-determined by a user and written in a programming language.
[0054] Referring further to FIG. 2, "training data", as used herein, is data that includes correlations that can be used by a machine learning process to model relationships between two or more categories of data elements. For example, without limitation, training data 204 can include a plurality of data entries, also known as "training examples", each entry representing a set of data elements that were recorded, received, and / or generated together, and the data elements can be correlated, for example, by their co-presence in a given data entry, by their proximity in a given data entry, etc. The plurality of data entries within training data 204 can manifest one or more trends of correlations between categories of data elements, for example, without limitation, a higher value of a first data element belonging to a first category of data elements tends to correlate with a higher value of a second data element belonging to a second category of data elements, indicating a potential proportional or other mathematical relationship linking the values belonging to the two categories. The plurality of categories of data elements can be related in training data 204 according to various correlations, and the correlations can indicate causal and / or predictive links between categories of data elements, which can be modeled as relationships, such as mathematical relationships, by a machine learning process, as will be described in more detail below. Training data 204 can be formatted and / or organized by categories of data elements, for example, by associating data elements with one or more descriptors corresponding to the categories of data elements. As a non-limiting example, training data 204 can include data input in a standardized format by a person or process such that an entry of a given data element within a given field in a form can be mapped to one or more descriptors of a category.Elements within the training data 204 can be linked to the category descriptors by tags, tokens, or other data elements, for example, but not limited to, providing the training data 204 in a fixed-length format, a format that links the positions of the data to categories such as the Comma-Separated Value (CSV) format, and / or a self-describing format such as Extensible Markup Language (XML), JavaScript® Object Notation (JSON), enabling a process or device to detect the categories of the data.
[0055] Alternatively or additionally, still referring to FIG. 2, the training data 204 may include one or more unclassified elements, i.e., the training data 204 may not be formatted or may not include descriptors for some elements of the data. Machine learning algorithms and / or other processes can sort the training data 204 according to one or more classifications, for example, using natural language processing algorithms, tokenization, detection of correlation values in raw data, etc., and the categories can be generated using correlation and / or other processing algorithms. As a non-limiting example, in a corpus of text, phrases that make up the number “n” of multi-word expressions such as nouns modified by other nouns can be identified according to the statistically significant morbidity rate of n-grams that contain such words in a particular order, and such n-grams are classified as elements of a language such as “words” that will be tracked in the same way as single words, and new categories can be generated as a result of statistical analysis. Similarly, in a data entry containing some text data, a person's name can be identified by referring to a list, dictionary, or other list of terms, enabling ad-hoc classification by a machine learning algorithm and / or an automated association of the data and descriptors within the data entry or to a given format. The ability to automatically classify data entries may make it possible to apply the same training data 204 to two or more separate machine learning algorithms, as will be explained in more detail below. The training data 204 used by the machine learning module 200 can correlate any input data as described in this disclosure to any output data as described in this disclosure. Non-limiting and exemplary examples include a first set 120 of associations, direct correlations, examples of a second set 132 of associations, or any data referred to throughout this disclosure.
[0056] Referring further to FIG. 2, as will be described in further detail below, one or more supervised and / or unsupervised machine learning processes and / or models can be used to filter, sort, and / or select training data, such models can include, but are not limited to, a training data classifier 216. The training data classifier 216, as used in the present disclosure, represents and / or uses a machine learning model, such as a data structure representing a mathematical model, neural net, or program generated by a machine learning algorithm known as a “classification algorithm,” as defined below, which sorts inputs into categories or bins of data and outputs the categories or bins of data and / or labels associated therewith. The classifier can be configured to output at least one data that labels or otherwise identifies, for example, data sets that are clustered together and found to be close under a distance metric, as described below. The distance metric can include, but is not limited to, any norm such as the Pythagorean norm. The machine learning module 200 can generate a classifier using a classification algorithm as defined as the process by which a computing device and / or any module and / or component operating therein derives a classifier from the training data 204. Classification can be performed using, but not limited to, linear classifiers such as logistic regression classifiers and / or naive Bayes classifiers, nearest neighbor classifiers such as k-nearest neighbor classifiers, support vector machines, least squares support vector machines, Fisher's linear discriminant, quadratic classifiers, decision trees, boosted trees, random forest classifiers, learning vector quantization, and / or neural network-based classifiers. By way of non-limiting example, the training data classifier 216 can classify elements of the training data into a first set 120 of associations.
[0057] Referring further to FIG. 2, training examples for use as training data can be selected from a population of potential examples according to a cohort related to an analytical problem or classification task to be solved. Alternatively or additionally, the training data can be selected to span possible situations or sets of inputs for the machine learning model and / or process encountered during deployment. For example, without limitation, for each category of input data to a machine learning process or model that may exist within a range of values within a set of phenomena such as images, user data, process data, physical data, etc., a computing device, processor, and / or machine learning model can select training examples representing each possible value on such a range and / or representative samples of values on such a range. The selection of representative samples can include, for example, selecting training examples at a rate that matches the statistically determined and / or predicted distribution of such values according to relative frequency such that values that occur more frequently in the population of data so analyzed are represented by more training examples than values that occur less frequently. Alternatively or additionally, the set of training examples can be compared to a set of representative values in a database and / or presented to a user such that the process can automatically or via user input detect one or more values not included in the set of training examples. A computing device, processor, and / or module can automatically generate missing training examples, which can be done by receiving and / or searching for missing input and / or output values and correlating the missing input and / or output values with corresponding output and / or input values collocated in the data record with the retrieved values provided by a user and / or other devices, etc.
[0058] Referring further to FIG. 2, a computer, processor, and / or module may be configured to sanitize training data. "Sanitizing" the training data, as used in the present disclosure, is a process by which training examples that impede the convergence of a machine learning model and / or the processing to useful results are removed. For example, but not limited to, a training example may include input and / or output values that are outliers from typically encountered values such that a machine learning algorithm using the training example is adapted to a low likelihood amount as input and / or output, e.g., values exceeding a threshold number of standard deviations from a mean value, average, or expected value may be excluded. Alternatively or additionally, one or more training examples may be identified as having low-quality data, where "low-quality" is defined as having a signal-to-noise ratio below a threshold.
[0059] As a non-limiting example, further referring to FIG. 2, images used to train an image classifier or other machine learning model and / or process that takes an image as input or generates an image as output may be rejected if the image quality falls below a threshold. For example, but not limited to, a computing device, processor, and / or module may perform blur detection, and excluding one or more blur detections may be performed, as a non-limiting example, by performing an approximation such as a Fourier transform of the image, or a fast Fourier transform (FFT), and analyzing the distribution of low and high frequencies in the resulting frequency domain depiction of the image, and the number of high frequency values below a threshold level may indicate blurriness. As a further non-limiting example, blur detection may be performed by convolving an image or a channel of an image with a Laplacian kernel, which may generate a numerical score that reflects some rapid changes in the intensity shown in the image such that a high score indicates sharpness and a low score indicates blurriness. Blur detection may be performed using a gradient-based operator that measures an operator based on the gradient or first derivative of the image, based on the hypothesis that rapid changes indicate sharp edges within the image and thus a lower degree of blurriness. Blur detection may be performed using a wavelet-based operator that utilizes the ability of the coefficients of a discrete wavelet transform to describe the frequency and spatial content of the image. Blur detection may be performed using a statistic-based operator that utilizes some image statistics as texture descriptors to calculate a focus level. Blur detection may be performed by using discrete cosine transform (DCT) coefficients to calculate the focus level of the image from its frequency components.
[0060] Continuing to refer to FIG. 2, a computing device, processor, and / or module may be configured to precondition one or more training examples. For example, without limitation, if a machine learning model and / or process has one or more inputs and / or outputs that require, transmit, or receive a particular number of bits, samples, or other units of data, elements of the one or more training examples that are to be used as or compared to the inputs and / or outputs can be modified to have such a number of units of data. For example, a computing device, processor, and / or module may convert a smaller number of units, such as within a low pixel count image, to a desired number of units, for example by upsampling and interpolation. As a non-limiting example, a low pixel count image may have 100 pixels, but the desired number of pixels may be 128. A processor may be able to interpolate the low pixel count image to convert 100 pixels to 128 pixels. It should be noted that those skilled in the art can understand various methods for interpolating a smaller number of data units, such as samples, pixels, bits, etc., to a desired number of such units by reading the present disclosure. In some cases, a set of interpolation rules may be trained by a neural network or other machine learning model that is trained to predict interpolated pixel values using a set of very detailed inputs and / or outputs and a corresponding set of inputs and / or outputs downsampled to a smaller number of units and training data. As a non-limiting example, sample inputs and / or outputs, such as a sample picture having sample extended data units (e.g., pixels added between original pixels), may be input into a neural network or machine learning model and output a pseudo replica sample picture having dummy values assigned to the pixels between the original pixels based on a set of interpolation rules.As a non-limiting example, in the context of an image classifier, a machine learning model can have a set of interpolation rules trained by a set of very detailed images and a set of images downsampled to a smaller number of pixels, and a neural network or other machine learning model trained to use those examples to predict interpolation pixel values in a face picture context. As a result, an input having sample-expanded data units (added between the original data units with dummy values) can be run through the trained neural network and / or model, which can input values to replace the dummy values. Alternatively or additionally, a processor, computing device, and / or module can utilize a sample expander method, a low-pass filter, or both. As used in this disclosure, a "low-pass filter" is a filter that passes signals of frequencies lower than a selected cut-off frequency and attenuates signals of frequencies higher than the cut-off frequency. The exact frequency response of the filter depends on the filter design. A computing device, processor, and / or module can use averaging such as luma or chroma averaging within an image to fill in data units between the original data units.
[0061] In some embodiments, continuing to refer to FIG. 2, a computing device, processor, and / or module may downsample elements of a training example to a desired fewer number of data elements. As a non-limiting example, a high pixel count image may have 256 pixels, but the desired number of pixels may be 128. The processor may downsample the high pixel count image to convert 256 pixels to 128 pixels. In some embodiments, the processor may be configured to perform downsampling on data. Downsampling, also known as decimation, may include removing every Nth entry, everything other than every Nth entry, etc. in a sequence of samples, which is a process known as "compression" and may be performed, for example, by an N-sample compressor implemented using hardware or software. An anti-aliasing filter and / or an anti-imaging filter, and / or a low-pass filter may be used to remove side effects of the compression.
[0062] Referring further to FIG. 2, the machine learning module 200 can be configured to execute a delayed learning process 220 and / or protocol, which alternatively may be referred to as a “lazy loading” or “call-on-demand” process and / or protocol, and when receiving an input that is to be converted into an output, combines this input and the training set to derive an algorithm that will be used to generate the output on demand, which can be the process by which machine learning is performed. For example, an initial set of simulations can be executed to cover initial heuristics and / or “first guesses” in the output and / or relationships. By way of non-limiting example, the initial heuristics can include rankings of associations between the input and elements of the training data 204. The heuristics can include selecting some of the top-ranked associations and / or elements of the training data 204. Delayed learning can implement any suitable delayed learning algorithm including, but not limited to, the k-nearest neighbor algorithm, the lazy naive Bayes algorithm, etc., and one of ordinary skill in the art, upon considering the entire disclosure, will recognize various delayed learning algorithms that can be applied to generate the output as described in the present disclosure, including the delayed learning application of machine learning algorithms as will be described in more detail below, without limitation.
[0063] Alternatively or additionally, still referring to FIG. 2, a machine learning model 224 can be generated using a machine learning process as described in the present disclosure. A "machine learning model", as used in the present disclosure, is a mathematical representation and / or algorithmic representation of the relationship between inputs and outputs, and / or an instance thereof, that is generated using any machine learning process, including but not limited to any of the processes described above, and stored in memory, where the inputs are presented to a previously created machine learning model 224 that generates an output based on the derived relationship. For example, but not limited to, a linear regression model generated using a linear regression algorithm can calculate a linear combination of input data using coefficients derived during the machine learning process to compute output data. As a further non-limiting example, the machine learning model 224 can be generated by creating an artificial neural network, such as a convolutional neural network, that includes an input layer of nodes, one or more intermediate layers, and an output layer of nodes. The connections between the nodes can be created via a process of "training" the network where elements from a set of training data 204 are applied to the input nodes, and then appropriate training algorithms (such as the Levenberg-Marquardt method, conjugate gradient, annealing method, or other algorithms) are used to adjust the connections and weights between the nodes of adjacent layers of the neural network to generate a desired value at the output nodes. This process is sometimes referred to as deep learning.
[0064] Referring further to FIG. 2, the machine learning algorithm may include at least a supervised machine learning process 228. The at least supervised machine learning process 228, as defined herein, receives a training set that relates some inputs to some outputs and attempts to generate one or more data structures that represent and / or instantiate one or more mathematical relationships that relate the inputs to the outputs, where each of the one or more mathematical relationships is optimal for some criteria specified for the algorithm using some scoring functions. For example, the supervised learning algorithm may include a first set 120 of associations as described above as an input, a second set 120 of associations as an output, and a scoring function that represents a desired form of relationship that will be detected between the input and the output, and the scoring function may, for example, maximize the probability that a given input and / or combination of element inputs is associated with a given output and minimize the probability that a given input is not associated with a given output. The scoring function can be represented as a risk function that represents the "expected loss" of the algorithm that relates the input to the output, and the loss is calculated as an error function that represents the degree to which the prediction generated by the relationship is inaccurate when compared to a given input-output pair provided in the training data 204. Those skilled in the art will recognize various possible variations of the at least supervised machine learning process 228 that can be used to determine the relationship between the input and the output upon reviewing the entire disclosure. The supervised machine learning process may include a classification algorithm as defined above.
[0065] Referring further to FIG. 2, training a supervised machine learning process can include, but is not limited to, iteratively updating coefficients, biases, and / or weights based on an error function, an expected loss, and / or a risk function. For example, the output generated by a supervised machine learning model using an input example in a training example can be compared to the output example from the training example, and an error function can be generated based on the comparison that includes any error function suitable for use in any machine learning algorithm described in this disclosure, such as the sum of the squares of the differences between one or more sets of comparison values. Such an error function can be used to update one or more weights, biases, coefficients, or other parameters of the machine learning model via any suitable process, including, but not limited to, a gradient descent process, a least squares process, and / or other processes described in this disclosure. This can be done iteratively and / or recursively to gradually adjust such weights, biases, coefficients, or other parameters. The update can be performed using one or more backpropagation algorithms in a neural network. The iterative and / or recursive update of weights, biases, coefficients, or other parameters as described above can be performed until the currently available training data is exhausted and / or until a convergence test is passed, where a "convergence test" is a test for a condition selected as indicating that the model and / or its weights, biases, coefficients, or other parameters have reached a certain level of accuracy. In a convergence test, for example, the difference between two or more consecutive errors or error function values can be compared, where a difference below a threshold amount can be interpreted as indicating convergence. Alternatively or additionally, one or more errors and / or error function values evaluated in a training iteration can be compared to a threshold.
[0066] Referring further to FIG. 2, a computing device, processor, and / or module may be configured to execute the methods, method steps, sequences of method steps, and / or algorithms described with reference to this figure in any order and any degree of repetition. For example, a computing device, processor, and / or module may be configured to repeatedly execute a single step, sequence, and / or algorithm until a desired or commanded result is achieved, and the repetition of a step or sequence of steps may be executed by repeatedly and / or recursively using the output of a previous repetition as an input to a subsequent repetition, aggregating the input and / or output of the repetition to generate an aggregated result, reducing or decrementing one or more variables such as global variables, and / or dividing a larger processing task into a set of smaller processing tasks to be repeatedly addressed. A computing device, processor, and / or module may execute any step, sequence of steps, or algorithm in parallel, such as using two or more parallel threads, processor cores, etc. to execute a step two or more times simultaneously and / or substantially simultaneously, and the division of tasks between the parallel threads and / or processes may be executed according to any protocol suitable for the division of tasks between repetitions. Those skilled in the art will recognize various ways in which steps, sequences of steps, processing tasks, and / or data can be subdivided, shared, or otherwise processed using iteration, recursion, and / or parallel processing upon considering the entire disclosure.
[0067] Referring further to FIG. 2, the machine learning process may include at least an unsupervised machine learning process 232. An unsupervised machine learning process, as used herein, is a process of deriving inferences within a dataset without regard to labels, and as a result, an unsupervised machine learning process can freely discover any structure, relationship, and / or correlation provided in the data. An unsupervised process may not require a response variable, and an unsupervised process can be used to find interesting patterns and / or inferences between variables and determine, for example, the degree of correlation between two or more variables.
[0068] Referring further to FIG. 2, the machine learning module 200 can be designed and configured to create a machine learning model 224 using techniques for developing a linear regression model. The linear regression model can include ordinary least squares regression that aims to minimize the sum of the squares of the differences between the predicted results and the actual results according to an appropriate norm (e.g., vector space distance norm) for measuring such differences, and the coefficients of the resulting linear equations can be modified to improve the minimization. The linear regression model can include ridge regression, where the function to be minimized includes the least squares function and a term that multiplies the square of each coefficient by a scalar amount to impose a penalty on large coefficients. The linear regression model can include a least absolute shrinkage and selection operator (LASSO) model, where ridge regression is combined with multiplying the least squares term by a coefficient divided by twice the number of samples. The linear regression model can include a multi-task LASSO model where the norm applied to the least squares term of the LASSO model is the Frobenius norm corresponding to the square root of the sum of the squares of all terms. The linear regression model can include an elastic net model, a multi-task elastic net model, a least angle regression model, a LARS LASSO model, an orthogonal matching pursuit model, a Bayesian regression model, a logistic regression model, a stochastic gradient descent model, a perceptron model, a passive-aggressive algorithm, a robust regression model, a Huber regression model, or any other suitable model that can be envisioned by one of ordinary skill in the art upon consideration of the entire disclosure. In one embodiment, the linear regression model can be generalized to a polynomial regression model, thereby obtaining a polynomial (e.g., quadratic, cubic, or higher order equation) that provides the best prediction output / actual output fit. As will be apparent to one of ordinary skill in the art upon consideration of the entire disclosure, methods similar to those described above can be applied to minimize the error function.
[0069] Continuing to refer to FIG. 2, the machine learning algorithm may include, but is not limited to, linear discriminant analysis. The machine learning algorithm may include quadratic discriminant analysis. The machine learning algorithm may include kernel ridge regression. The machine learning algorithm may include a support vector machine that includes, but is not limited to, a regression process based on support vector classification. The machine learning algorithm may include a stochastic gradient descent algorithm that includes classification and regression algorithms based on stochastic gradient descent. The machine learning algorithm may include a nearest neighbor algorithm. The machine learning algorithm may include various forms of latent space regularization such as variational regularization. The machine learning algorithm may include a Gaussian process such as Gaussian process regression. The machine learning algorithm may include a cross-decomposition algorithm that includes partial least squares and / or canonical correlation analysis. The machine learning algorithm may include a naive Bayes method. The machine learning algorithm may include a decision tree-based algorithm such as a decision tree classification or regression algorithm. The machine learning algorithm may include an ensemble method such as a bagging meta-estimator, a random forest of trees, AdaBoost, gradient tree boosting, and / or a voting classifier method. The machine learning algorithm may include a neural network algorithm that includes a convolutional neural network process.
[0070] Referring further to FIG. 2, the machine learning model and / or process can be deployed or instantiated by being incorporated into a program, apparatus, system, and / or module. For example, but not limited to, a machine learning model, neural network, and / or some or all of their parameters can be stored and / or deployed in any memory or circuit. Parameters such as coefficients, weights, and / or biases can be stored as an array of wires set to logic “1” and “0” voltage levels within a logic circuit and / or as circuit-based constants such as binary inputs and / or outputs to represent numbers by any suitable coding system including, for example, two's complement, or can be stored in any volatile and / or non-volatile memory. Similarly, mathematical operations and the input and / or output of data to and / or from models, neural network layers, etc. can be instantiated in hardware circuits and / or in the form of instructions within firmware, machine code such as binary operation code instructions, assembly language, or any high-level programming language. Any technique for the hardware and / or software instantiation of memory, instructions, data structures, and / or algorithms can be used to instantiate the machine learning process and / or model, which includes, but is not limited to, the manufacture and / or configuration of non-reconfigurable hardware elements, circuits, and / or modules such as ASICs, the manufacture and / or configuration of reconfigurable hardware elements, circuits, and / or modules such as FPGAs, the manufacture and / or of non-reconfigurable and / or configuration non-rewritable memory elements, circuits, and / or modules such as non-rewritable ROMs, the manufacture and / or configuration of reconfigurable and / or rewritable memory elements, circuits, and / or modules such as rewritable ROMs or other memory technologies described in the present disclosure, and / or any combination of the manufacture and / or configuration of any computing device and / or its components described in the present disclosure.Such deployed and / or instantiated machine learning models and / or algorithms can receive input from any other processes, modules, and / or components described in this disclosure and generate output to any other processes, modules, and / or components described in this disclosure.
[0071] Continuing to refer to FIG. 2, after initial deployment and / or instantiation, any process of training, retraining, deploying, and / or instantiating any machine learning model and / or algorithm can be performed and / or repeated to modify, refine, and / or improve the machine learning model and / or algorithm. Such retraining, deployment, and / or instantiation can be processed as periodic or regular processes such as retraining, deployment, and / or instantiation during a regular elapsed period, after some measure of quantity such as the number of bytes of processed data or other metric, the number of uses or executions of the processes described in this disclosure, and / or according to a software, firmware, or other update schedule. Alternatively or additionally, retraining, deployment, and / or instantiation can be event-based and can be triggered, without limitation, by user input indicating sub-optimal or otherwise problematic performance and / or by an automated field test and / or auditing process, which can compare the machine learning model and / or algorithm, and / or the output of its error and / or error function, to any threshold, convergence test, etc., and / or can compare the output of the processes described herein to similar thresholds, convergence tests, etc. Event-based retraining, deployment, and / or instantiation can alternatively or additionally be triggered by the receipt and / or generation of one or more new training examples, and some new training examples can be compared to a preconfigured threshold, where exceeding the preconfigured threshold can trigger retraining, deployment, and / or instantiation.
[0072] Referring further to FIG. 2, retraining and / or additional training can be performed using any currently or previously deployed version of a machine learning model and / or algorithm as a starting point and using any process for training as described above. Training data for retraining can be collected, pre-conditioned, sorted, classified, sanitized, or otherwise processed according to any process described in this disclosure. The training data includes, but is not limited to, training examples that include inputs and correlated outputs used, received, and / or generated from any version of any system, module, machine learning model or algorithm, device, and / or method described in this disclosure, such examples can be modified and / or labeled according to user feedback or other processes to indicate a desired result, and / or the actual or measured results from a process modeled and / or predicted by a system, module, machine learning model or algorithm, device, and / or method can be used as the "desired" result to be compared with the output for a training process as described above.
[0073] Redployment can be performed using any reconfiguration and / or rewriting of reconfigurable and / or rewritable circuitry and / or memory elements, or alternatively, redployment can be performed by generating new hardware and / or software components, circuitry, instructions, etc., which can be added to and / or replace existing hardware and / or software components, circuitry, instructions, etc.
[0074] Continuing to refer to FIG. 2, the machine learning process may include a generative machine learning process. As used in this disclosure, a "generative machine learning process" is a process that uses a prompt (i.e., an input) to automatically generate an output that is consistent with training data, which is in contrast to a non-machine learning software program where the output is pre-determined by a user and written in a programming language. Typically, a generative machine learning process determines patterns and structures from training data and uses these patterns and structures to synthesize new data with similar characteristics in response to an input.
[0075] Continuing to refer to FIG. 2, generative machine learning processes can synthesize different types or domains of data, including but not limited to text, code, images, molecules, audio (e.g., music), video, and robotic actions (e.g., electromechanical system operations). Exemplary generative machine learning systems trained on words or word tokens operating in the text domain include GPT-3, LaMDA, LLaMA, BLOOM, GPT-4, etc. Exemplary machine learning processes trained on programming language text (i.e., code) include, but are not limited to, OpenAI Codex. Exemplary machine learning processes trained on a set of images (e.g., with text captions) include Imagen, DALL-E, Midjourney, Adobe Firefly, Stable Diffusion, etc., and image generative machine learning processes can, in some cases, be trained for text-to-image generation and / or neural style transfer. Exemplary generative machine learning processes trained on molecular data include, but are not limited to, AlphaFold, which can be used for protein structure prediction and drug discovery. Generative machine learning processes trained on audio training data include MusicLM, which can be trained on audio waveforms of music correlated with text annotations, and music generative machine learning processes can, in some cases, generate new music samples based on text descriptions. Exemplary generative machine learning processes trained on video include, but are not limited to, RunwayML and Make-A-Video by Meta Platforms. Finally, exemplary generative machine learning processes trained using robotic action data include, but are not limited to, UniPi by Google Research.
[0076] Continuing to refer to FIG. 2, in some cases, the generative machine learning process may include a generative adversarial network. As used in this disclosure, a "generative adversarial network (GAN)" is a machine learning process that includes at least two adversarial networks configured to synthesize data according to defined rules (e.g., the rules of a game). In some cases, the generative adversarial network may include a generator network and a discriminator network, where the generator network generates candidate data and the discriminator network evaluates the candidate data. An exemplary GAN can be described according to the following game: Each probability space (Ω, μ ref ) defines a GAN game. There are two adversarial networks, namely a generator network and a discriminator network. The set of generator network strategies is P(Ω), i.e., the set of all probability measures on Ω, μ G . The set of discriminator network strategies is the set of Markov kernels μ D : Ω → P[0,1], where P[0,1] is the set of probability measures on [0,1]. The GAN game can be a zero-sum game with the following objective function.
[0077]
Equation
[0078] Referring further to FIG. 2, in some embodiments, the generative machine learning process may include, but is not limited to, adversarial generative networks (GANs) such as CycleGAN. CycleGAN may use a pair of mutually dependent neural network generators that depend on each other's outputs to be used as inputs. The CycleGAN process enables forward and backward domain conversions to occur simultaneously. The CycleGAN process may include a set of calculations. CycleGAN may be different from paired training data, where paired training data are training examples where the correspondence between x i and y i already exists.
Number
Number
Number
Number
Number
Number
[0079] This process is iterable in the reverse direction and is motivated by reducing the cycle consistency loss as follows.
Number
Number
[0080] Referring further to FIG. 2, in another non-limiting embodiment, the generative machine learning model may include a diffusion process and / or a diffusion model. An exemplary diffusion model may include an energy-guided stochastic differential equation (EGSDE) process. EGSDE relies on a score-based diffusion model to transform data from one domain to another. In contrast to the CycleGAN process described above, EGSDE uses a score-based diffusion model (SBDM) to perturb an initial dataset into Gaussian noise and then reverses the process to transform the noise back into the data distribution. This diffusion model employs a pre-trained energy function based on data from the initial source domain and data from the final target domain to guide the inference process of a pre-trained stochastic differential equation (SDE). As used herein, an "energy function" is defined as an approximate expression of a transfer function for transforming data from a first dataset from the source domain into usable data within the target domain (e.g., the domain of a second dataset). The energy function consists of two terms. The first guiding term is a realistic expert, which prioritizes focusing on the energy function discarding source-domain-specific features. The second guiding term is a faithful expert, which prioritizes focusing on the energy function preserving domain-independent features. By combining these two functions, a target-domain dataset is obtained that does not depend on the source-domain data protocol but retains the substantial features shown by the data. In combination with the pre-trained energy function, the EGSDE method uses three experts (the energy function, the realistic expert, and the faithful expert) such that all contribute to the generation of the best-fit output data. A mathematical explanation of this process will be detailed for time-series data. Let q(y0) be
Number
[0081] This can then be discretized using the Euler-Maruyama solver. By formally adopting a step size of h, the iteration rule from s to t = s - h is as follows. [Number]
[0082] Furthermore, referring to Figure 2, and further the EGSDE, the set of unpaired images from the source domain [Number] and the set of unpaired images from the target domain [Number] Referring to using these as training data, the goal is to transfer the original time series data from the source domain to the target domain. This goal is achieved by designing the distribution p(y0|x0) on the target domain γ conditioned on the time series x0 ∈ χ to be transferred. The transformed time series data should be realistic for the target domain by changing domain-specific features and faithful to the source time series by preserving domain-independent features. Then, Iterative Latent Variable Refinement (ILVR) can use a diffusion model on the target domain for realism. ILVR is for y TStarting from ~N(0, I), y T is sampled from the diffusion model described immediately above. To promote faithfulness, y T is refined further by adding the residual between the sample y T perturbed source image x T through a non-trainable low-pass filter: y T ← y T + Φ(x t ) - Φ(y t ), x t ~ q t|0 (x t |x0), where Φ(·) is the low-pass filter and q t|0 (·|·) is the perturbation kernel determined by the forward SDE. To use the pre-trained energy function most accurately across both domains, the effective conditional distribution p(y0|x0) is defined by composing the pre-trained SDE and the pre-trained energy function under the following mild regularity conditions:
Equation
Equation
Equation
Equation
[0083] Continuing to refer to FIG. 2 and further referring to the use of EGSDE, the energy function is derived by balancing the need to retain domain-independent features of the initial time series data 108 while appropriately modifying domain-specific features. Based on this balance, the energy function is the sum of two logarithmic potential functions:
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0084] A more sophisticated E that not only acts as a simple low-pass filter i may be used to untangle the learning methods of different domains.
[0085] Referring further to FIG. 2 and further to the use of EGSDE, solving the energy - induced reverse - time SDE can be achieved using a pre - trained score - based model s(y,t) and an energy function E(y,x,t), and an example 116 generated from the conditional distribution p(y0|x0) can be created. A numerical solver can be used to approximate the trajectory from the SDE to achieve a fair comparison. As a non - limiting embodiment, an Euler - Maruyama solver that employs a step size h is used herein, where the iteration rule from s to t = s - h is as follows.
Number
[0086] The Monte Carlo method is used to estimate the expected value of a single sample with respect to efficiency. In a non - limiting embodiment, a variance - preserving energy - induced SDE can be used to modify the noise prediction network and incorporate it into the sampling procedure of the denoising diffusion probability model.
[0087] Referring further to FIG. 2 and further to the use of EGSDE, the product of experts is used in the discretized sampling process. The conditional distribution at time t is defined as follows:
Number
Number
Number
Number
Number
Number
Number
[0088] Referring further to FIG. 2, the one or more processes or algorithms described above can be executed by at least dedicated hardware unit 232. For the purposes of this figure, a “dedicated hardware unit” is, without limitation, a hardware component, circuit, etc. that is specifically designated or selected to execute one or more specific tasks and / or processes described with reference to this figure, such as preconditioning and / or sanitizing training data and / or training a machine learning algorithm and / or model, that is separate from the main control circuit and / or processor that executes the method steps described in this disclosure. Dedicated hardware unit 232 can include a hardware unit capable of performing iterative or intensive calculations, such as matrix-based calculations for updating or adjusting the parameters, weights, coefficients, and / or biases of a machine learning model and / or neural network, by efficiently using, without limitation, pipelining, parallel processing, etc., and such a hardware unit can be optimized for such processes by including, for example, dedicated circuits for matrix and / or signal processing operations that include multiple arithmetic and / or logic circuit units, such as multipliers and / or adders, that can operate simultaneously and / or in parallel. Such a dedicated hardware unit 232 can include, without limitation, a graphics processing unit (GPU), a dedicated signal processing module, an FPGA, or other reconfigurable hardware configured to instantiate parallel processing units for one or more specific tasks. A computing device, processor, apparatus, or module can be configured to instruct one or more dedicated hardware units 232 to perform one or more operations described herein, such as evaluating model and / or algorithm outputs, making one-time or iterative updates to parameters, coefficients, weights, and / or biases, and / or performing any other operations, such as vector and / or matrix operations as described in this disclosure.
[0089] Referring now to FIG. 3, an exemplary association database 300 is shown in block diagram form. In one embodiment, any past or current version of any of the data disclosed herein may be stored within an association database 300 including, but not limited to, a plurality of data sets 108, a first data set 112, a second data set 116, a first set of associations 120, visual features, named entities, pathology identifiers, a second set of associations 132, a third data set 140, and the like. The processor 104 can be communicatively coupled to the association database 300. For example, in some cases, the database 300 may be local to the processor 104. Alternatively or additionally, in some cases, the database 300 may be remote from the processor 104 and communicable with the processor 104 via one or more networks. The network may include, but is not limited to, a cloud network, a mesh network, and the like. As an example, a "cloud-based" system, as the term is used herein, refers to a system that includes software and / or data stored, managed, and / or processed on a network of remote servers hosted in the "cloud" via, for example, the Internet, rather than on a local server or personal computer. A "mesh network," as used in this disclosure, is a local network topology in which the infrastructure processor 104 connects directly, dynamically, and non-hierarchically to as many other computing devices as possible. A "network topology," as used in this disclosure, is the arrangement of the elements of a communication network. The association database 300 can be implemented as, but is not limited to, a relational database, a key-value search database such as a NOSQL database, or any other format or structure for use as a database that would be recognized as suitable by one of ordinary skill in the art upon consideration of the entire disclosure. The association database 300 may alternatively or additionally be implemented using a distributed data storage protocol and / or data structure such as a distributed hash table.As described above, the association database 300 may include a plurality of data entries and / or records. Data entries within the database may be flagged with one or more additional elements of information or linked to such additional elements, which may be reflected in linked tables such as tables associated by one or more indexes within the data entry cell and / or relational database. Those skilled in the art, upon considering the entire disclosure, will recognize the various ways in which data entries within the database can store, retrieve, organize, and / or reflect data and / or records as used herein, as well as categories and / or sets of data consistent with the present disclosure.
[0090] Referring now to FIG. 4, an exemplary embodiment of a neural network 400 is shown. A neural network 400, also known as an artificial neural network, is a network of "nodes", or a data structure having one or more inputs, one or more outputs, and a function that determines an output based on the inputs. Such nodes can be organized into a network such as, but not limited to, a convolutional neural network including an input layer of nodes 404, one or more intermediate layers 408, and an output layer of nodes 412. Connections between nodes can be created via a process of "training" the network where elements from a training data set are applied to the input nodes, and then appropriate training algorithms (such as the Levenberg-Marquardt method, conjugate gradient, annealing methods, or other algorithms) are used to adjust the connections and weights between nodes in adjacent layers of the neural network to produce a desired value at the output nodes. This process is sometimes referred to as deep learning. The connections may be made only from input nodes towards output nodes within a "feedforward" network, or the output of one layer may be fed back to inputs in the same or different layers within a "recurrent network". As a further non-limiting example, the neural network may include a convolutional neural network including an input layer of nodes, one or more intermediate layers, and an output layer of nodes. A "convolutional neural network", as used in this disclosure, is a neural network in which at least one hidden layer is a convolutional layer that convolves an input, known as a "kernel", with a subset of the input using one or more additional layers such as a pooling layer, a fully connected layer, etc. to the input of that layer.
[0091] Referring now to FIG. 5, an exemplary embodiment of a node of a neural network is shown. The node can receive numerical values from inputs to the neural network including the node, and / or from other nodes, including, but not limited to, a plurality of inputs x i The node can include i weights w multiplied by respective inputs x ican be used to perform a weighted sum of inputs. Additionally or alternatively, a bias b may be added to the weighted sum of the inputs such that an offset is added to each unit within a neural network layer that is independent of the input to the layer. The weighted sum can then be input into a function φ, thereby generating one or more outputs y. Input x i weights w applied to the input i can indicate whether the input is "excitatory" (e.g., by corresponding weights having large numerical values) such that the input has a strong influence on one or more outputs y, and / or whether the input is "inhibitory" (e.g., by corresponding weights having small numerical values) such that the input has a weak influence on another output y. Weights w i can be determined by training the neural network using training data, which can be performed using any suitable process as described above.
[0092] Referring now to FIG. 6, an exemplary embodiment of a fuzzy set comparison 600 is shown. In a non-limiting embodiment, the fuzzy set comparison. In a non-limiting embodiment, the fuzzy set comparison 600 may be consistent with the fuzzy set comparison of FIG. 1. In another non-limiting example, the fuzzy set comparison 600 may be consistent with name / version matching as described herein. For example, but not limited to, parameters, weights, and / or coefficients of the membership function may be adjusted using any machine learning method for name / version matching as described herein. In another non-limiting embodiment, the fuzzy set may represent visual features and named entities from FIG. 1.
[0093] Alternatively or additionally, referring further to FIG. 6, the fuzzy set comparison 600 may be generated in response to determining a data compatibility threshold. The compatibility threshold may be determined by a computing device. In some embodiments, the computing device may use a logic comparison program, such as, but not limited to, a fuzzy logic model, to determine the compatibility threshold and / or the version authenticator. Each such compatibility threshold may be represented as a value of a posting variable representing the compatibility threshold, or in other words, as a fuzzy set corresponding to a degree of compatibility and / or acceptability that is calculated using a statistical method, a machine learning method, or other methods that one of ordinary skill in the art may conceive of upon considering the entire disclosure. In some embodiments, determining the compatibility threshold and / or the version authenticator may include using a linear regression model. The linear regression model may include a machine learning model. The linear regression model may map statistics, such as, but not limited to, the frequency of the same range of version numbers, to the compatibility threshold and / or the version authenticator. In some embodiments, determining the compatibility threshold for any posting may include using a classification model. The classification model may be configured to input the collected data and the cluster data into the centroid based on, among other things, the frequency of occurrence of the range of version numbers, the language indicators of compatibility and / or acceptability, etc. The centroid may include the scores assigned to them so that scores can be assigned to each of the fitness thresholds. In some embodiments, the classification model may include a K-means clustering model. In some embodiments, the classification model may include a particle swarm optimization model. In some embodiments, determining the compatibility threshold may include using a fuzzy inference engine. The fuzzy inference engine may be configured to map one or more compatibility thresholds using fuzzy logic. In some embodiments, multiple computing devices may be arranged in a compatibility configuration by a logic comparison program. "Compatibility configuration", as used in this disclosure, is any grouping of objects and / or data based on skill level and / or output score.The membership function coefficients and / or constants as described above can be adjusted according to a classification and / or clustering algorithm. For example, but not limited to, the clustering algorithm can determine a Gaussian distribution or other distribution of questions regarding centroids corresponding to a given compatibility threshold and / or version authenticator, and using iteration or other methods, find a membership function of any membership function type as described above that minimizes the mean error from a statistically determined distribution, such as a triangular or Gaussian membership function regarding a centroid representing the center of the distribution that best matches the distribution. The error function to be minimized, and / or the method of minimization, can be carried out according to, but not limited to, the error functions as described in the present disclosure and / or an error function minimization process and / or method.
[0094] Referring further to FIG. 6, the inference engine can be implemented according to a plurality of scanned user labels 120 and inputs of a plurality of named entities. For example, the acceptance variable can represent a first measurable value regarding the classification of a plurality of visual features for a named entity. Continuing with the example, the output variable can represent a second set 132 of associations. In one embodiment, the plurality of visual features and / or named entities can be represented by their own fuzzy sets. In other embodiments, the evaluation factors can be represented according to the intersection of two fuzzy sets, as shown in FIG. 6. The inference engine can combine any rules such as its semantic versioning, semantic language, version ranges, etc. The degree to which a given input function membership matches a given rule can be determined by a triangular norm or "T-norm" of an output function having rules of commutativity (T(a,b)=T(b,a)), monotonicity: (when a≦c and b≦d, T(a,b)≦T(c,d)), (associativity: T(a,T(b,c))=T(T(a,b),c)), and the requirement that the number 1 functions as the identity element, such as min(a,b), the product of a and b, the drastic product of a and b, the Hamacher product of a and b, etc. The combination of rules (the combination of "and" or "or" for rule membership determination) can be performed using any T-conorm represented by an inversion T symbol or "⊥", such as max(a,b), the probabilistic sum of a and b (a + b - a*b), the bounded sum, and / or the drastic T-conorm, with commutativity: ⊥(a,b)=⊥(b,a), monotonicity: when a≦c and b≦d, ⊥(a,b)≦⊥(c,d), associativity: ⊥(a,⊥(b,c))=⊥(⊥(a,b),c), and any T-conorm that satisfies the property of the identity element 0. Alternatively or additionally, the T-conorm can be approximated by a sum, such as in a "product-sum" inference engine where the T-norm is a product and the T-conorm is a sum. The final output score or other fuzzy inference output can be determined from the output membership function as described above using any suitable defuzzification process, including but not limited to the average of the maximum defuzzification, the centroid of the area / centroid defuzzification, the center average defuzzification, the bisector of the area defuzzification, etc.Alternatively or additionally, the output rule may be replaced by a function according to a Takagi-Sugeno-King (TSK) fuzzy model.
[0095] The first fuzzy set 604 can be represented according to a first membership function 608 that represents, but is not limited to, the probability that an input falling within a first range 612 of values is a member of the first fuzzy set 604, where the first membership function 608 has values over a probability range such as, but not limited to, the interval [0,1], and the area under the first membership function 608 can represent a set of values within the first fuzzy set 604. In this exemplary depiction, for clarity, the first range 612 of values is shown as a range on a single numerical line or axis, but the first range 612 of values may be defined in two or more dimensions, representing, for example, a Cartesian product between a plurality of ranges, curves, axes, spaces, dimensions, etc. The first membership function 608 can include any suitable function that maps the first range 612 to a probability interval, including, but not limited to, a triangular function defined by two linear elements such as line segments or planes that intersect at the top or below of the probability interval. As a non-limiting example, a triangular membership function can be defined as follows:
Number
Number
Number
Number
Number
[0096] Upon considering the entire disclosure, those skilled in the art will recognize various alternative or additional membership functions that can be used consistently with the present disclosure.
[0097] The first fuzzy set 604 can represent any value or combination of values as described above, including any plurality of visual features and named entities. A second fuzzy set 616 that can represent any value represented by the first fuzzy set 604 can be defined by a second membership function 620 on a second range 624, where the second range 624 may be the same as and / or overlap with the first range 612 and / or be combined with the first range via a Cartesian product or the like to generate a mapping that allows for an overlap in the evaluation of the first fuzzy set 604 and the second fuzzy set 616. If the first fuzzy set 604 and the second fuzzy set 616 have an overlapping region 636, the first membership function 608 and the second membership function 620 may intersect at a point 632 that represents the probability of agreement between the first fuzzy set 604 and the second fuzzy set 616, as defined on a probability interval. Alternatively or additionally, a single value of the first and / or second fuzzy set may be located at a point 636 on the first range 612 and / or the second range 624, where the probability of membership can be taken by evaluating the first membership function 608 and / or the second membership function 620 at that range point. The probability at 628 and / or 632 can be compared to a threshold 640 to determine whether a positive match is indicated. The threshold 640 can, in a non-limiting example, represent the degree of agreement between the first fuzzy set 604 and the second fuzzy set 616 and / or a single value having each other or either set therein, which is sufficient for the purpose of the matching process. For example, a second set 132 of associations can show a sufficient degree of overlap with a fuzzy set representing a plurality of visual features and named entities for the combinations that occur as described above. Each threshold can be established by one or more user inputs. Alternatively or additionally, each threshold can be adjusted by machine learning and / or statistical processes, for example, but not limited to, as described in more detail below.
[0098] In one embodiment, the degree of match between fuzzy sets can be used to rank one resource against another. For example, if both a plurality of visual features and named entities have fuzzy sets, a second set 132 of associations may be generated by having a degree of overlap that exceeds a prediction threshold, and the processor 104 may further rank two resources by ranking the resource with the higher degree of match higher than the resource with the lower degree of match. If multiple fuzzy matches are performed, the degree of match for each respective fuzzy set can be calculated and aggregated, e.g., by addition, averaging, etc., to determine an overall degree of match, which can be used to rank the resources, and the selection between two or more matching resources can be performed by selection of the highest-ranked resource, and / or multiple notifications can be presented to the user in the order of ranking.
[0099] Referring now to FIG. 7, a flowchart of an exemplary method 800 for detecting associations between different types of datasets is shown. In step 705, method 700 includes receiving, using at least a processor, a plurality of datasets from a user. The plurality of datasets includes a first dataset that includes a plurality of text data and a second dataset that includes a plurality of image data. This can be implemented as described with reference to FIGS. 1-6. In one embodiment, the plurality of datasets includes a plurality of data associated with one or more prognostic slides.
[0100] Referring further to FIG. 7, in step 710, method 700 includes identifying a first set of associations between a first dataset and a second dataset using at least a processor. This can be implemented as described with reference to FIGS. 1-6. In one embodiment, the first set of associations includes multiple direct correlations between multiple datasets. In another embodiment, the first set of associations includes multiple pathology identifiers. Generating the first set of associations may include identifying one or more visual features within the second dataset using at least a processor. Generating the first set of associations may include generating a plurality of named entities associated with the first dataset using at least a processor. Generating the first set of associations may further include generating a plurality of named entities associated with the first dataset using at least a processor.
[0101] Referring further to FIG. 7, in step 715, method 700 includes generating a second set of associations using at least a processor and using a second association classifier in response to a first set of associations. Generating the second set of associations includes training the second association classifier using second association training data, where the second association training data includes a first set of associations as input that correlates to a second set of associations as output, and the second set of associations includes a plurality of data entries generated in response to the first set of associations of the trained second association classifier. In some versions, the second association training data may include a first subset of a first data set as input that correlates to a first subset of a second data set as output. This can be implemented as described with reference to FIGS. 1-6. In one embodiment, the second set of associations includes a plurality of abstract correlations between a plurality of data sets. In another embodiment, the method further includes generating a third data set using at least a processor in response to the second set of associations. In some cases, the method includes generating the second set of associations using at least a processor and using a fuzzy inference set.
[0102] Referring further to FIG. 7, in step 720, method 700 includes displaying the second set of associations using a display device. This can be implemented as described with reference to FIGS. 1-6.
[0103] Referring further to FIG. 7, in some embodiments, generating the second set of associations further includes, in step 715, training a generative machine learning process using the first data set, synthesizing first synthetic data according to the first data set using generative machine learning, and generating a second set of associations according to the first synthetic data and the first set of associations. The generative machine learning process may include any generative machine learning process described in the present disclosure, including referring to FIGS. 1-6 above. The first synthetic data may include any data generated by a machine learning process, such as the generative machine learning process described in the present disclosure, including referring to FIGS. 1-6 above. In some cases, the second association training data may include first synthetic data as input that correlates to a first subset of the second data set as output.
[0104] Referring further to FIG. 7, in some embodiments, the first data set includes text. The text may include any text or text data described in the present disclosure, for example by referring to FIGS. 1-6. In some cases, generating the second set of associations may further include, in step 715, associating text data within the first data set using a natural language processing model, and generating a second set of associations according to the associated text data within the first data set and the first set of associations.
[0105] Referring further to FIG. 7, in some embodiments, generating the second set of associations may further include, in step 715, calculating distances between data elements within the first data set, and generating a second set of associations according to the distances between data elements within the first data set and the first set of associations. The distance may include any distance described in the present disclosure, such as vector distance, for example by referring to FIGS. 1-6.
[0106] Referring further to FIG. 7, in some embodiments, the first data set may include metadata. The metadata may include any metadata or context data described in this disclosure with reference to FIGS. 1-6, for example. In some cases, generating the second set of associations may further include, in step 715, associating the metadata within the first data set and generating a second set of associations according to the associated metadata within the first data set and the first set of associations.
[0107] Referring further to FIG. 7, in some embodiments, the second data set may include image data. The image data may include any representative data such as graphics described in this disclosure with reference to FIGS. 1-6, for example. In some cases, generating the second set of associations may include, in step 715, using a machine vision system to identify one or more visual features within the second data set. The machine vision system may include any machine vision system described in this disclosure with reference to FIGS. 1-6, for example. The visual features may include any visual features described in this disclosure with reference to FIGS. 1-6, for example.
[0108] Referring further to FIG. 7, in some embodiments, generating the first set of associations may include, in step 710, identifying a plurality of named entities within the first data set, where each named entity of the plurality of named entities is associated with at least a data element within the second data set.
[0109] Referring now to FIG. 8, FIG. 8 is a schematic diagram of an apparatus 800 for detecting associations between different types of datasets. For example, apparatus 800 may be configured to detect an association (e.g., a text-image pairing) between a first dataset (e.g., text data 804) and a second dataset (e.g., image data 808). That is, for a given element in the first dataset (e.g., a word, phrase, sentence, or set of sentences in text data 804), apparatus 800 may detect an associated element in the second dataset (e.g., an image among image data 808). These may be implemented as disclosed with respect to FIGS. 1-7.
[0110] Continuing to refer to FIG. 8, in some embodiments, large general-purpose datasets consisting of image-text pairs have become widely available in recent years. For example, such datasets may include pairings of images and their captions. However, in more specialized fields such as medical image data, the use of these general-purpose datasets is limited. For example, general-purpose datasets often consist of naturally occurring images. In fact, neural network models such as the Contrastive Language-Image Pretraining (CLIP) model trained using naturally occurring images provide low performance. To reach a performance level suitable for actual use, it is desirable to train neural network models using image-text pairs from the same specialized field. For example, it is desirable to use image-text pairs from the medical field to train neural network models for medical use. By doing so, it becomes possible to improve the performance of the trained model and the computational efficiency (e.g., shortening the GPU time for training).
[0111] Continuing to refer to FIG. 8, in some embodiments, an exemplary source of image-text pairs in the medical field is patient notes. Generally, patient notes can include unstructured text and image data, along with an association between the text and the image, which can be direct or indirect. An example of a direct association is an identifier that exists in both the text and the image (e.g., the identifier may be embedded in the image itself or included in the accompanying metadata). An example of an indirect association is a textual description (e.g., a word, phrase, or sentence) of a pathological condition corresponding to an image that visually represents that condition. In many cases, a direct association between text and image can be easily analyzed using known techniques; for example, an algorithm can be programmed to detect text and images with matching identifiers. However, such techniques may not be suitable for detecting indirect associations.
[0112] Continuing to refer to FIG. 8, in some embodiments, according to some embodiments, an indirect association between an image and text data is detected using a neural network model to learn a representation (e.g., a vector representation) of the text and the image data. The learned representation is used to map the image and text data into a joint latent space (e.g., a vector space of the learned representation). Then, pairings are identified within the latent space. One of the challenges of this approach is that the mapping of text and image data into the latent space is a learned relationship rather than a function, regardless of whether the mapping domain is considered to be the image or the text. Thus, the techniques described below can address these challenges.
[0113] Continuing to refer to FIG. 8, in some embodiments, for simplicity and clarity, the following examples focus on the pairing of image and text data, but those skilled in the art will understand that this technique can be easily adapted to detect associations between a wide variety of other types of data. For example, the present technology can be applied to multimedia data (e.g., audio and video), time series data (e.g., electrocardiogram data), structured data (e.g., data tables, graphs, databases, models, etc.), medical scans (e.g., 2D or 3D medical imaging data), and the like. Further, the present technology can be used to detect associations between three or more data sets (e.g., a grouping of text, image, and time series data).
[0114] Continuing to refer to FIG. 8, in some embodiments, the text data 804 may correspond to a set of documents. For example, the text data 804 may include patient notes or other types of text-based medical records. The image data 808 may include a set of images related to the text data 804. For example, the text data 804 and the image data 808 can belong to the same set of medical records (e.g., the medical records can include both written components and image-based components).
[0115] Continuing to refer to FIG. 8, in some embodiments, the result of pairing text data 804 and image data 808 is a set 812 of text-image pairs. The text-image pairs 812 can be used for various downstream tasks such as supervised machine learning tasks, semi-supervised machine learning tasks, or self-supervised machine learning tasks. In some embodiments, the text-image pairs 812 include a set of one or more elements of the text data 804 and a corresponding set of one or more elements of the image data 808. For example, the apparatus 800 can pair an "image tuple" with a corresponding "text tuple". For example, a given image can be associated with multiple text elements such as a caption placed with the image and a description of the image that appears elsewhere in the body of the document. Similarly, an element of text can describe or compare multiple images, such as a caption applied to a group of images. Tuples can have certain advantages in specific downstream tasks such as sequence prediction tasks.
[0116] Continuing to refer to FIG. 8, in some embodiments, the apparatus 800 can be configured to improve the detection of text-image pairs 812 based on the text data 804 and the image data 808. For example, the apparatus 800 can detect the text-image pairs 812 more accurately and efficiently (e.g., with less GPU time) than existing techniques. By detecting a more complete set 812 of text-image pairs than existing techniques, the techniques used by the apparatus 800 can generate a sufficiently large set 812 of text-image pairs using less text data 804 and image data 808, thereby reducing the storage amount and bandwidth used for storing and transmitting the data. Further, by generating a larger and more accurate set 812 of text-image pairs, the apparatus 800 can similarly improve the accuracy and efficiency of downstream tasks, e.g., for machine learning tasks, faster training with fewer GPU cycles, etc.
[0117] Continuing to refer to FIG. 8, in some embodiments, the apparatus 800 can execute a bootstrapping process in which an initial set 816 of text-image pairs is generated based on the text data 804 and the image data 808. Various strategies can be used to generate the initial text-image pairs 816. For example, a pre-trained neural network model 820 can be utilized to generate the initial text-image pairs 816 using transfer learning. The pre-trained model 820 can be a model trained in a similar domain or a model that provides sufficient accuracy in other ways to generate the initial set. Additionally or alternatively, the initial image-text pairs 816 can be identified using direct associations such as explicit links within the text data 804 or identifiers present in both the text data 804 and the image data 808 (e.g., the identifier may be embedded in the image itself or included in the accompanying metadata). Analyzing the data set to identify the direct associations can be performed using existing symbolic techniques. In some embodiments, the identifier can include a timestamp, where the text-image pairs have matching timestamps.
[0118] Continuing to refer to FIG. 8, in some embodiments, the initial text-image pairs 816 correspond to less than the total number of text-image pairs within the text data 804 and the image data 808. Thus, using the "bootstrap" set of the initial text-image pairs 816, the apparatus 800 executes a further process to detect additional text-image pairs. For example, additional text-image pairs that are desirable to detect can include indirect pairs (e.g., text passages that describe aspects of the image without directly linking to the image), or pairs that are not detected using the pre-trained model 820 (e.g., the pre-trained model 820 is a general-purpose model or otherwise not optimized for the text data 804 and / or the image data 808).
[0119] Continuing to refer to FIG. 8, in some embodiments, to detect additional text-image pairs, the apparatus 800 can train a pairing model 824 using an initial image-text pair 816. For example, the pairing model 824 can jointly learn representations of the image data 808 and the text data 804 within the same representation space (e.g., latent space). Examples of such techniques are described in more detail in Radford, et al., 「Learning Transferable Visual Models From Natural Language Supervision」 (https: / / arxiv.org / pdf / 2103.00020.pdf).
[0120] Continuing to refer to FIG. 8, in some embodiments, the apparatus 800 can train a generation model 828 using the image data 808 and synthesize additional image data. That is, the generation model 828 is capable of generating synthetic images having attributes similar to the images within the image data 808. For example, the generation model 828 can correspond to an adversarial generation network (GAN) or a diffusion model. To efficiently train the generation model 828, an initial stage of training can include training another model (e.g., an autoencoder) to create a low-dimensional representation of the image data 808, and the low-dimensional representation is provided to the generation model 828. In some embodiments, the generation model 828 can generate image data unconditionally or conditionally. For example, image synthesis can be conditioned on text such as text from the text data 804. The initial text-image pairs 816 can be used as a training set for training the generation model 828 to generate images conditionally. In some embodiments, the generation model 828 can be used to edit the images within the image data 808, for example, by performing inpainting to fill in masked portions of the images. Examples of generation models, training techniques, and conditional and unconditional image generation are described in more detail in Rombach, et al., "High-Resolution Image Synthesis with Latent Diffusion Models" (https: / / arxiv.org / pdf / 2112.10752.pdf).
[0121] Continuing to refer to FIG. 8, in some embodiments, the apparatus 800 can use the text data 804 that is used to pre-train the language model 832. For example, the language model 832 can be configured to generate representations (e.g., vector representations) of elements of the text data 808. Exemplary examples of language models include autoregressive models such as the Generative Pretrained Transformer 3 (GPT-3) model, which is described in more detail in Brown, et al., "Language Models are Few-Shot Learners" (https: / / arxiv.org / pdf / 2005.14165.pdt), which is hereby incorporated by reference in its entirety, and masked language models (MLMs) such as the Bidirectional Encoder Representations from Transformers (BERT) model, which is described in more detail in Bao, et al., "BEiT: BERT Pre-Training of Image Transformers" (https: / / arxiv.org / pdf / 2106.08254.pdf).
[0122] Continuing to refer to FIG. 8, in some embodiments, using one or more of the language model 832, the pairing model 824, and the generation model 828, the apparatus 800 can proceed to detect additional pairs of text 836 and images 840. For example, the apparatus 800 can use one or more of the following techniques to detect additional pairs.
[0123] Continuing to refer to FIG. 8, in some embodiments, by using a text representation to find a matching image using the generative model 828, the apparatus 800 can conditionally generate an image based on elements of the text data 804 using the generative model 828. In some embodiments, the conditions can be based on the representation of text elements generated by the language model 832. The generated image can be used to identify zero or more other images within the image data 808 that are similar to the generated image. Similar images can be identified using various techniques such as determining vector similarity (e.g., cosine similarity) in the embedding space. In some embodiments, images within the image data 808 that are similar to the generated image and exceed a predetermined threshold are determined to be paired with the corresponding text elements of the text data 804.
[0124] Continuing to refer to FIG. 8, in some embodiments, by using the combined trained pairing model 824 to find matching text by using the image representation, the apparatus 800 can use the pairing model 824 to identify one or more candidate text elements in the text data 804 for a particular image in the image data 808. To improve the efficiency of this approach, the text data 804 can be pruned by various techniques such that a reduced number of text elements are considered. For example, the language model 832 can be used to cluster the elements of the text data 804 such that the pruned set of text data includes the centroids of the clusters. Further improvements can be obtained by varying the granularity of the text elements (e.g., words, phrases, sentences, etc.) to identify an appropriate level of granularity. In some embodiments, one or more of the text candidates can be paired with the corresponding image, for example, by pairing the closest match, one or more candidates that exceed a predetermined threshold, etc. In some embodiments, the generative model 828 can be used to conditionally generate an image based on the candidate text elements. Consistent with such embodiments, the determination of whether to pair a candidate text element with an image can be based on the similarity between the synthetic image and the actual image.
[0125] Continuing to refer to FIG. 8, in some embodiments, by using the implicit image magnification level and temporal order in the text to find candidate images, the apparatus 800 can use one or more context cues to make a pairing decision. The context cues can be used alone or in combination with other techniques such as those shown in blocks 107 and 108. For example, the magnification level of a microscopic image can be used as a context cue for pairing text and images that explicitly or implicitly refer to a consistent magnification level. For example, text that describes cell-level features or counts cell-level features suggests a high level of magnification, while a broader description of a pathological condition suggests a lower magnification level.
[0126] Continuing to refer to FIG. 8, in some embodiments, the time information of the text data 804 can be used as a queue for finding candidate images within the image data 808. For example, when the text of the patient note refers to a time or period (e.g., the text includes a specific date, a temporal phrase such as "the next day" or "next week"), these queues can be used to identify candidate images based on whether the image is associated with time information that matches the time information of the text (e.g., an image with a timestamp of a date one week after the patient note referring to a scan to be performed "next week"). In some examples, the time information in the text data 804 can indicate images taken in temporally consecutive order (e.g., the text can indicate that "the patient was sent back for a scan again"), which can be a queue for identifying candidate images having a matching sequence. For example, an initial diagnostic core biopsy can be paired with text reporting the initial histological grade of a tumor, and a subsequent resection specimen can be paired with text reporting a grade different from the previously reported grade. In this example, the two events can be separated by several days to months.
[0127] Continuing to refer to FIG. 8, in some embodiments, these context queues can be automatically picked up by the neural network model during training, but configuring the device 800 to explicitly use a given queue can provide additional advantages. In some embodiments, the device 800 can perform processes that facilitate the use of context queues, such as determining the magnification level of an image (e.g., by detecting and counting features within the image).
[0128] Continuing with reference to FIG. 8, in some embodiments, the subject matter described herein can be implemented as a digital electronic circuit, or computer software, firmware, or hardware, or combinations thereof, that include the structural means disclosed herein and their structural equivalents. The subject matter described herein can be implemented as one or more computer program products, such as one or more computer programs tangibly embodied in an information carrier (e.g., in a machine-readable storage device) or embodied in a propagated signal for execution by, or to control the operation of, a data processing apparatus (e.g., a programmable processor, a computer, or multiple computers). A computer program (also known as a program, software, software application, or code) can be written in any form of programming language, including a compiled or interpreted language, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file. A program can be stored in a portion of a file that holds other programs or data, in a single file dedicated to the program in question, or in multiple cooperating files (e.g., files that store one or more modules, subprograms, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers at one site or distributed across multiple sites and interconnected by a communication network.
[0129] Continuing to refer to FIG. 8, in some embodiments, the processes and logic flows described herein that include the method steps of the subject matter described herein can be performed by one or more programmable processors executing one or more computer programs that operate on input data to produce output, thereby performing the functions of the subject matter described herein. The processes and logic flows can also be performed by dedicated logic circuitry, such as an FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit), and the apparatus of the subject matter described herein can be implemented as such dedicated logic circuitry.
[0130] Continuing to refer to FIG. 8, in some embodiments, processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. In general, a processor receives instructions and data from a read only memory or a random access memory or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer also includes one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or is operatively coupled to receive data from, transfer data to, or both from and to such mass storage devices. Information carriers suitable for embodying computer program instructions and data include all forms of non-volatile memory, by way of example semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and optical disks (e.g., CD and DVD disks). The processor and the memory can be supplemented by, or incorporated in, dedicated logic circuitry.
[0131] Continuing to refer to FIG. 8, in some embodiments, to provide interaction with a user, the subject matter described herein can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and a keyboard and a pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and the input from the user can be received in any form including acoustic, voice, or tactile input.
[0132] Continuing to refer to FIG. 8, in some embodiments, the subject matter described herein can be implemented in a computing system that includes backend components (e.g., a data server), middleware components (e.g., an application server), or frontend components (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described herein), or any combination of such backend, middleware, and frontend components. The components of the system can be interconnected by digital data communication in any form or medium, such as by a communication network. Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), such as the Internet.
[0133] Referring now to FIG. 9, a flow diagram of an exemplary method 900 for detecting associations between different types of datasets is shown. Method 900 includes, at least using a processor, a step 905 of receiving a first dataset and a second dataset, where the first and second datasets include data elements of different types. In some embodiments, the first dataset may include a set of text data, the second dataset may include a set of image data, and the association may include pairs of text images. In some embodiments, the set of image data may include a visual representation of a pathological slide captured by an imaging technique. In some embodiments, the second dataset may include metadata associated with the pathological slide, and this metadata may include annotations. In some embodiments, the plurality of datasets may include media data, and this media data may include one or more audio data, video data, and document data associated with the pathological slide. In some embodiments, this media data may further include a tissue graph. In some embodiments, this media data may further include a diagnostic pathway graph. In some embodiments, one or more of the first data elements may include visual features, and one or more of the second data elements may include named entities. In some embodiments, method 900 may further include, at least using a processor, assigning a pathology identifier to an initial set of associations, where the initial set of associations may include visual features and named entities. 1. In some embodiments, method 900 may further include, at least using a processor, using a machine vision system to identify visual features. These may be implemented as disclosed with respect to FIGS. 1-8.
[0134] Continuing to refer to FIG. 9, method 900 includes step 905 of identifying an initial set of associations between a first data set and a second data set using at least a processor, each association including one or more first data elements from the first data set and one or more second data elements from the second data set. In some embodiments, method 900 includes identifying an initial set of associations using a bootstrap process using at least a processor, the bootstrap process generating a plurality of resamples of the first data set and the second data set, each resample possibly including paired samples of the first data set and the second data set, and further including determining an initial set of associations between the first data set and the second data set according to the plurality of resamples. In some embodiments, method 900 further includes determining, using at least a processor, the strength of an initial set of associations between the first data set and the second data set according to the plurality of resamples, and determining, using at least a processor, an initial set of associations between the first data set and the second data set according to the strength and a predetermined threshold. In some embodiments, method 900 further includes pairing, using at least a processor, a prognosis label with the second data set, and identifying, using at least a processor, a first data set associated with the prognosis label. In some embodiments, method 900 further includes generating an initial set of associations using a first association classifier trained with training data including a plurality of first data sets and a plurality of second data sets correlated with the initial set of associations.In some embodiments, method 900 further comprises using at least a processor to compare first visual features of a first second dataset with second visual features of a second second dataset, wherein the second second dataset may include an initial set of associations with a second first dataset; using at least a processor to match the first visual features to the second visual features in response to this comparison; and using at least a processor to determine an initial set of associations of the first second dataset in response to this matching. In some embodiments, method 900 further comprises using at least a processor to identify named entities of a first dataset using a named entity recognition process. In some embodiments, method 900 further comprises using at least a processor to preprocess the first dataset and the second dataset; and using at least a processor to generate a similarity score between the preprocessed first dataset and the preprocessed second dataset. These may be implemented as disclosed with respect to FIGS. 1-8.
[0135] Continuing to refer to FIG. 9, method 900 includes, at 905, using at least a processor to train a neural network model to detect additional associations between a first dataset and a second data, wherein the initial set of associations is used as training data for training the neural network model. These may be implemented as disclosed with respect to FIGS. 1-8.
[0136] Continuing to refer to FIG. 9, method 900 includes step 905 of detecting one or more additional associations between a first dataset and a second dataset using at least a processor and a trained neural network model. In some embodiments, method 800 may further include training a generation model to conditionally generate an image based on text, where an initial set of associations is used as training data for training this generation model. In some embodiments, method 800 may further include generating one or more images corresponding to elements from a set of text data using the generation model and comparing the one or more generated images with one or more images within a set of image data. In some embodiments, method 900 may further include generating additional associations using at least a processor according to the pathology identifiers of the initial set of associations. These may be implemented as disclosed with respect to FIGS. 1-8.
[0137] Note that any one or more of the aspects and embodiments described herein can be advantageously implemented using one or more machines (e.g., one or more computing devices utilized as user computing devices for electronic documents, one or more server devices such as document servers, etc.) programmed according to the teachings herein, as will be apparent to those skilled in the computer art. As will be apparent to those skilled in the software art, appropriate software coding can be readily created by a skilled programmer based on the teachings of this disclosure. The above aspects and implementations using software and / or software modules may also include appropriate hardware for supporting the implementation of machine-executable instructions of the software and / or software modules.
[0138] Such software may be a computer program product that uses a machine-readable storage medium. The machine-readable storage medium can store and / or encode a sequence of instructions for execution by a machine (e.g., a computing device), and can be any medium that causes a machine to execute any one of the methodologies and / or embodiments described herein. Examples of machine-readable storage media include, but are not limited to, magnetic disks, optical disks (e.g., CD, CD-R, DVD, DVD-R, etc.), magneto-optical disks, read-only memory "ROM" devices, random access memory "RAM" devices, magnetic cards, optical cards, solid-state memory devices, EPROM, EEPROM, and any combination thereof. As used herein, a machine-readable medium is intended to include a single medium as well as a collection of physically distinct media, such as a compact disc combined with computer memory or a collection of one or more hard disk drives. As used herein, a machine-readable storage medium does not include signal transmission in a transient form.
[0139] Such software may also include information (e.g., data) carried as a data signal on a data carrier such as a carrier wave. For example, the machine-executable information may be included as a data-carrying signal embodied in a data carrier, where the signal encodes a sequence of instructions for execution by a machine (e.g., a computing device) or a portion thereof, and any associated information (e.g., data structures and data) that causes a machine to execute any one of the methodologies and / or embodiments described herein.
[0140] Examples of computing devices include, but are not limited to, e - book reading devices, computer workstations, desktop computers, server computers, handheld devices (such as tablet computers, smart phones, etc.), web appliances, network routers, network switches, network bridges, any machine capable of executing a sequence of instructions specifying the actions to be taken by that machine, and any combination thereof. In one example, a computing device may include, and / or be included in, a kiosk.
[0141] FIG. 10 shows a diagrammatic representation of an embodiment of a computing device in an exemplary form of a computer device 1000 in which a set of instructions for causing any one or more of the aspects and / or methodologies of the present disclosure to be executed within a control system can be executed. Also, it is contemplated that a specially configured set of instructions for causing any one or more of the aspects and / or methodologies of the present disclosure to be executed on one or more of a plurality of devices can be implemented using a plurality of computing devices. The computer device 1000 includes a processor 1004 and a memory 1008 that communicate with each other and with other components via a bus 1012. The bus 1012 may include any of several types of bus structures including, but not limited to, a memory bus, a memory controller, a peripheral bus, a local bus, and any combination thereof, using any of a variety of bus architectures.
[0142] Processor 1004 can include any suitable processor, such as a processor incorporating logic circuitry for performing arithmetic and logical operations, such as, but not limited to, an arithmetic logic unit (ALU), which can be coordinated by a state machine and directed by operational inputs from memory and / or sensors, and the processor 1004 can be configured according to, by way of non-limiting example, a von Neumann architecture and / or a Harvard architecture. Processor 1004 can include, incorporate, and / or be incorporated in, but is not limited to, a microcontroller, a microprocessor, a digital signal processor (DSP), a field programmable gate array (FPGA), a complex programmable logic device (CPLD), a graphics processing unit (GPU), a general purpose GPU, a tensor processing unit (TPU), an analog or mixed signal processor, a trusted platform module (TPM), a floating point unit (FPU), and / or a system on chip (SoC).
[0143] Memory 1008 can include various components (e.g., machine-readable media), including, but not limited to, random access memory components, read-only components, and any combination thereof. In one example, a basic input / output system 1016 (BIOS) including basic routines that help transfer information between elements within computer device 1000 during startup, etc., can be stored in memory 1008. Memory 1008 can also include instructions (e.g., software) 1020 that embody any one or more of the aspects and / or methodologies of the present disclosure (e.g., stored on one or more machine-readable media). In another example, memory 1008 can further include any number of program modules, including, but not limited to, an operating system, one or more application programs, other program modules, program data, and any combination thereof.
[0144] The computer device 1000 may also include a storage device 1024. Examples of storage devices (e.g., storage device 1024) include, but are not limited to, hard disk drives, magnetic disk drives, optical disk drives in combination with optical media, solid state memory devices, and any combination thereof. The storage device 1024 may be connected to the bus 1012 by an appropriate interface (not shown). Exemplary interfaces include, but are not limited to, SCSI, Advanced Technology Attachment (ATA), Serial ATA, Universal Serial Bus (USB), IEEE 1394 (FIREWIRE (registered trademark)), and any combination thereof. In one example, the storage device 1024 (or one or more of its components) can be removably interfaced with the computer device 1000 (e.g., via an external port connector (not shown)). In particular, the storage device 1024 and the associated machine-readable medium 1028 can provide non-volatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for the computer device 1000. In one example, the software 1020 may be present, in whole or in part, within the machine-readable medium 1028. In another example, the software 1020 may be present, in whole or in part, within the processor 1004.
[0145] The computer device 1000 may also include an input device 1032. In one example, a user of the computer device 1000 can input commands and / or other information into the computer device 1000 via the input device 1032. Examples of the input device 1032 include, but are not limited to, alphanumeric input devices (e.g., keyboards), pointing devices, joysticks, game pads, audio input devices (e.g., microphones, voice response systems, etc.), cursor control devices (e.g., mice), touch pads, optical scanners, video capture devices (e.g., still cameras, video cameras), touch screens, and any combination thereof. The input device 1032 can interface with the bus 1012 via any of a variety of interfaces (not shown) including, but not limited to, serial interfaces, parallel interfaces, game ports, USB interfaces, FIREWIRE (registered trademark) interfaces, direct interfaces to the bus 1012, and any combination thereof. The input device 1032 can include a touch screen interface that may be part of the display 1036 or separate therefrom, as further described below. The input device 1032 can be used as a user selection device for selecting one or more graphical representations within the graphical interface, as described above.
[0146] The user can also input commands and / or other information into the computer device 1000 via a storage device 1024 (e.g., removable disk drive, flash drive, etc.) and / or a network interface device 1040. A network interface device such as the network interface device 1040 can be used to connect the computer device 1000 to one or more of various networks such as the network 1044, and one or more remote devices 1048 connected thereto. Examples of network interface devices include, but are not limited to, network interface cards (e.g., mobile network interface cards, LAN cards), modems, and any combination thereof. Examples of networks include wide area networks (e.g., the Internet, corporate networks), local area networks (e.g., networks associated with offices, buildings, campuses, or other relatively small geographical spaces), telephone networks, data networks associated with telephone / voice providers (e.g., mobile communication provider data and / or voice networks), direct connections between two computing devices, and any combination thereof, but are not limited thereto. Networks such as the network 1044 can use wired and / or wireless communication modes. Generally, any network topology can be used. Information (e.g., data, software 1020, etc.) can be communicated between the computer device 1000 via the network interface device 1040.
[0147] The computer device 1000 may further include a video display adapter 1052 for communicating a displayable image to a display device such as the display device 1036. Examples of display devices include, but are not limited to, liquid crystal displays (LCDs), cathode ray tubes (CRTs), plasma displays, light emitting diode (LED) displays, and any combination thereof. The display adapter 1052 and the display device 1036 may be utilized in combination with the processor 1004 to provide a graphical representation of aspects of the present disclosure. In addition to the display device, the computer device 1000 may include one or more other peripheral output devices including, but not limited to, audio speakers, printers, and any combination thereof. Such peripheral output devices can be connected to the bus 1012 via a peripheral interface 1056. Examples of peripheral interfaces include, but are not limited to, serial ports, USB connections, FIREWIRE® connections, parallel connections, and any combination thereof.
[0148] The foregoing was a detailed description of exemplary embodiments of the present invention. Various modifications and additions can be made without departing from the spirit and scope of the present invention. Each feature of the various embodiments described above can be combined, as necessary, with features of other described embodiments to provide combinations of numerous features in related new embodiments. Further, while the foregoing describes several separate embodiments, what is described herein are merely examples of the application of the principles of the present invention. Further, the specific methods herein can be illustrated and / or described as being performed in a particular order, but the order is highly variable within the ordinary skill in the art for achieving the methods, systems, and software according to the present disclosure. Accordingly, this description is intended to be construed as illustrative only and not as limiting the scope of the present invention.
[0149] Exemplary embodiments are disclosed above and shown in the accompanying drawings. Those skilled in the art will understand that various changes, omissions, and additions can be made to what is specifically disclosed herein without departing from the spirit and scope of the present invention.
Claims
1. An apparatus for detecting associations between different types of datasets, the apparatus comprising: at least a processor; a memory communicatively connected to the at least one processor, the memory configured to: receive a plurality of datasets including a first dataset and a second dataset; identify a first set of associations between a first subset of the first dataset and a first subset of the second dataset; generate a second set of associations between a second subset of the first dataset and a second subset of the second dataset in response to the first set of associations using a second association classifier, wherein generating the second set of associations comprises: training the second association classifier using second association training data, the second association training data including a plurality of data entries including the first set of associations; and generating the second set of associations in response to the first set of associations using the trained second association classifier; including generating; displaying the second set of associations using a display device; and instructions for configuring the at least one processor to perform the foregoing.
2. The apparatus of claim 1, wherein the second association training data includes the first subset of the first dataset as an input correlated with the first subset of the second dataset as an output.
3. Generating the second set of associations further comprises: training a generative machine learning process using the first dataset; synthesizing first synthetic data in response to the first dataset using the generative machine learning; and generating the second set of associations in response to the first synthetic data and the first set of associations. The apparatus of claim 1.
4. The apparatus of claim 3, wherein the second association training data includes the first synthetic data as an input correlated with the first subset of the second dataset as an output.
5. The first dataset includes text, and generating the second set of associations comprises: Using a natural language processing model to associate text data within the first dataset, generating a second set of associations according to the associated text data within the first dataset and the first set of associations The apparatus according to claim 1, further comprising.
6. Generating the second set of associations includes calculating distances between data elements within the first dataset, generating a second set of associations according to the distances between data elements within the first dataset and the first association The apparatus according to claim 1, further comprising.
7. The first dataset includes metadata, and generating the second set of associations includes associating metadata within the first dataset, generating a second set of associations according to the associated metadata within the first dataset and the first association The apparatus according to claim 1, further comprising.
8. The apparatus according to claim 1, wherein the first set of associations includes a plurality of pathology identifiers.
9. The second dataset includes image data, and generating the second set of associations includes using a machine vision system to identify one or more visual features within the second dataset. The apparatus according to claim 1.
10. Generating the first set of associations includes identifying a plurality of named entities within a first dataset, and each named entity of the plurality of named entities is associated with at least a data element within a second dataset. The apparatus according to claim 1.
11. A method for detecting associations between different types of datasets, the method comprising: receiving a plurality of datasets including a first dataset and a second dataset using at least a processor, identifying a first set of associations between a first subset of the first dataset and a first subset of the second dataset using the at least a processor Using the at least one processor, using a second association classifier, generating a second set of associations between a second subset of the first data set and a second subset of the second data set according to the first set of associations, wherein generating the second set of associations comprises training the second association classifier using second association training data, wherein the second association training data includes a plurality of data entries including the first set of associations, and using the trained second association classifier to generate the second set of associations according to the first set of associations including generating; displaying the second set of associations using a display device including a method.
12. The method according to claim 11, wherein the second association training data includes the first subset of the first data set as an input correlated with the first subset of the second data set as an output.
13. Generating the second set of associations comprises training a generative machine learning process using the first data set, and using the generative machine learning to synthesize first synthetic data according to the first data set, and generating the second set of associations according to the first synthetic data and the first set of associations The method according to claim 11, further comprising.
14. The method according to claim 13, wherein the second association training data includes the first synthetic data as an input correlated with the first subset of the second data set as an output.
15. The first data set includes text, and generating the second set of associations comprises associating text data in the first data set using a natural language processing model, and generating the second set of associations according to the associated text data in the first data set and the first set of associations The method according to claim 11, further comprising.
16. Generating the second set of associations comprises calculating distances between data elements in the first data set generating a second set of the associations according to the distances between data elements within the first dataset and the first set of the associations The method according to claim 11, further comprising this.
17. wherein the first dataset includes metadata, and generating the second set of the associations includes associating the metadata within the first dataset, and generating the second set of the associations according to the associated metadata within the first dataset and the first set of the associations The method according to claim 11, further comprising this.
18. The method according to claim 11, wherein the first set of the associations includes a plurality of pathology identifiers.
19. wherein the second dataset includes image data, and generating the second set of the associations includes identifying one or more visual features within the second dataset using a machine vision system. The method according to claim 11.
20. generating the first set of the associations includes identifying a plurality of named entities within a first dataset, and each named entity of the plurality of named entities is associated with at least a data element within a second dataset. The method according to claim 11.
21. An apparatus for detecting an association between different types of datasets, the apparatus comprising: at least a processor; and a memory communicably connected to the at least a processor, the memory being receiving a plurality of datasets including a first dataset and a second dataset, wherein the first and second datasets include data elements of different types; identifying an initial set of associations between the first dataset and the second dataset, each association including one or more first data elements from the first dataset and one or more second data elements from the second dataset; training a neural network model to detect additional associations between the first dataset and the second dataset, wherein the initial set of associations is used as training data for training the neural network model. Using the trained neural network model, detecting one or more additional associations between the first dataset and the second dataset An apparatus comprising instructions for configuring the at least one processor to perform the above **Claim 22** The first dataset includes a set of text data The second dataset includes a set of image data The association includes text-image pairs The apparatus according to claim 21 **Claim 23** The apparatus according to claim 22, wherein the set of image data includes a visual representation of a pathological slide captured by an imaging technique **Claim 24** The memory further includes instructions for further configuring the at least one processor to train a generation model to conditionally generate an image based on text, and the initial set of associations is used as training data for training the generation model. The apparatus according to claim 22 **Claim 25** Detecting the one or more additional associations includes Using the generation model to generate one or more images corresponding to elements from the set of text data Comparing the one or more generated images with one or more images within the set of image data The apparatus according to claim 24, further comprising the above **Claim 26** The apparatus according to claim 21, wherein the second dataset includes metadata, the metadata is associated with a pathological slide, and the metadata includes annotations **Claim 27** The apparatus according to claim 21, wherein the plurality of datasets includes media data, and the media data includes one or more audio data, video data, and document data associated with a pathological slide **Claim 28** The apparatus according to claim 27, wherein the media data further includes a tissue graph **Claim 29** The apparatus according to claim 27, wherein the media data further includes a diagnostic pathway graph **Claim 30** The memory further includes instructions for further configuring the at least one processor to identify an initial set of the associations using a bootstrap process, and the bootstrap process includes Generating a plurality of resamples of the first dataset and the second dataset, each resample including a paired sample of the first dataset and the second dataset Determining an initial set of associations between the first dataset and the second dataset according to the plurality of resamples The apparatus according to claim 21, comprising:
31. The memory is configured to Determine the strength of an initial set of associations between the first dataset and the second dataset according to the plurality of resamples; and Determine the initial set of associations between the first dataset and the second dataset according to the strength and a predetermined threshold value; The apparatus according to claim 21, further comprising instructions for further configuring the at least one processor to perform the above operations.
32. The memory is configured to Pair a prognosis label with the second dataset; and Identify the first dataset associated with the prognosis label The apparatus according to claim 21, further comprising instructions for further configuring the at least one processor to perform the above operations.
33. The apparatus according to claim 21, wherein the memory further comprises instructions for further configuring the at least one processor to generate the initial set of associations using a first association classifier trained with training data including a plurality of first datasets and a plurality of second datasets correlated with an initial set of associations.
34. The one or more first data elements include visual features; The one or more second data elements include named entities; The apparatus according to claim 21.
35. The apparatus according to claim 34, wherein the memory further comprises instructions for further configuring the at least one processor to assign a pathology identifier to the initial set of associations, and the initial set of associations includes the visual features and the named entities. A method for detecting an association between different types of datasets, the method comprising: Receiving, using at least one processor, a plurality of datasets including a first dataset and a second dataset, wherein the first and second datasets include data elements of different types Identifying an initial set of associations between the first dataset and the second dataset using the at least one processor, each association including one or more first data elements from the first dataset and one or more second data elements from the second dataset; Training a neural network model using the at least one processor to detect additional associations between the first dataset and the second data, wherein the initial set of associations is used as training data for training the neural network model; Detecting one or more additional associations between the first dataset and the second dataset using the at least one processor and the trained neural network model; A method comprising.
36. The first dataset includes a set of text data, The second dataset includes a set of image data, The association includes a text-image pair; The method according to claim 41.
37. The method according to claim 42, wherein the set of image data includes a visual representation of a pathological slide captured by an imaging technique.
38. Training a generation model using the at least one processor to conditionally generate an image based on text, wherein the initial set of associations is used as training data for training the generation model; The method according to claim 43, further comprising.
39. Generating one or more images corresponding to elements from the set of text data using the at least one processor and the generation model; Comparing the one or more generated images with one or more images within the set of image data using the at least one processor; The method according to claim 44, further comprising.
40. The method according to claim 41, wherein the second dataset includes metadata, the metadata is associated with a pathological slide, and the metadata includes annotations.
41. The method according to claim 41, wherein the plurality of data sets includes media data, and the media data includes one or more audio data, video data, and document data associated with a pathological slide.
42. The method according to claim 47, wherein the media data further includes an organizational graph.
43. The method according to claim 47, wherein the media data further includes a diagnostic pathway graph.
44. Identifying an initial set of the associations using a bootstrap process using the at least one processor, the bootstrap process comprising: Generating a plurality of resamples of the first data set and the second data set, each resample including paired samples of the first data set and the second data set; Determining an initial set of the associations between the first data set and the second data set according to the plurality of resamples Including, identifying The method according to claim 41, further comprising:
45. Determining, using the at least one processor, a strength of an initial set of the associations between the first data set and the second data set according to the plurality of resamples; Determining, using the at least one processor, an initial set of the associations between the first data set and the second data set according to the strength and a predetermined threshold The method according to claim 41, further comprising:
46. Pairing a prognostic label with the second data set using the at least one processor; Identifying, using the at least one processor, the first data set associated with the prognostic label The method according to claim 41, further comprising:
47. Generating the initial set of the associations using a first association classifier trained with training data including a plurality of first data sets and a plurality of second data sets correlated with an initial set of associations, using the at least one processor The method according to claim 41, further comprising:
48. The one or more first data elements include visual features, The one or more second data elements include named entities, The method according to claim 41.
49. Using the at least one processor to assign a pathology identifier to the initial set of associations, the initial set of associations including the visual features and the named entity, the assigning The method according to claim 54, further comprising.
50. Using the at least one processor to generate the additional association according to the pathology identifier of the initial set of associations The method according to claim 54, further comprising.
51. Using the at least one processor to identify the visual features using a machine vision system The method according to claim 54, further comprising.
52. Using the at least one processor to compare a first visual feature of a first second dataset with a second visual feature of a second second dataset, the second second dataset including an initial set of associations with a second first dataset, the comparing; Using the at least one processor to match the first visual feature to the second visual feature according to the comparison; Using the at least one processor to determine an initial set of associations of the first second dataset according to the matching The method according to claim 57, further comprising.
53. Using the at least one processor to identify the named entity of the first dataset using a named entity recognition process The method according to claim 54, further comprising.
54. Using the at least one processor to preprocess the first dataset and the second dataset; Using the at least one processor to generate a similarity score between the preprocessed first dataset and the preprocessed second dataset The method according to claim 41, further comprising.
Citation Information
Patent Citations
Cross-modal retrieval method and system based on pseudo label learning and semantic consistency
CN109784405A
Medical information processing apparatus, learning data generation program, and learning data generation method
JP2021111283A
Generation program, generation method, and generation device
JP2021194261A
Method and apparatus for training an automated dental charting system
JP2022520197A
Illustrative Medical Imaging for Functional Prognosis Estimation
US20210110914A1