Method and system for providing response on basis of image analysis through multi-resolution feature analysis

By modeling causal dependencies between low-magnification and high-magnification features in tissue pathology images using a cross-attention mechanism, the method improves prediction accuracy and efficiency, addressing the limitations of existing deep learning models in multi-resolution analysis.

WO2026111249A1PCT designated stage Publication Date: 2026-05-28LG MANAGEMENT DEV INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
LG MANAGEMENT DEV INST CO LTD
Filing Date
2025-11-04
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Existing deep learning models for tissue pathology images fail to effectively utilize multi-resolution characteristics and causal dependencies between low-magnification and high-magnification features, leading to suboptimal performance and increased analysis costs, particularly in cancer diagnosis.

Method used

A method and system that models causal dependencies between low-magnification and high-magnification features in a latent space using a cross-attention mechanism, incorporating noise terms to capture uncertainty, and integrates these features for improved prediction performance and reliability.

Benefits of technology

Enhances prediction accuracy and computational efficiency by capturing diagnostic clues from low-magnification images and efficiently selecting patches from high-magnification images, reducing GPU memory usage and inference time while improving interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025017921_28052026_PF_FP_ABST
    Figure KR2025017921_28052026_PF_FP_ABST
Patent Text Reader

Abstract

An embodiment provides a method comprising the steps of: receiving an activation signal for accessing data in at least one memory; accessing a data structure in the at least one memory according to the reception of the activation signal, wherein the data structure includes a diagnostic image; loading the diagnostic image from the at least one memory; generating, by at least one processor, at least one response with respect to the diagnostic image by using at least one artificial intelligence model using the diagnostic image as an input, wherein the at least one artificial intelligence model is pre-trained to perform image analysis on the basis of a feature representation for a causal dependency relationship between a low-magnification feature and a high-magnification feature of the diagnostic image; and inputting (ingesting) the at least one response to at least one subsequent processing component.
Need to check novelty before this filing date? Find Prior Art

Description

Method and System for Providing Responses Based on Image Analysis through Multi-Resolution Feature Analysis

[0001] The present disclosure relates to a method and system for providing a response based on image analysis, and more specifically, to a method and system for providing a response based on image analysis through multi-resolution feature interpretation that can improve the final prediction performance and accuracy of an image analysis model by going beyond learning simple associations between low-magnification features and high-magnification features in multi-resolution analysis of an image, and by modeling causal dependencies between low-magnification features and high-magnification features in a latent space and utilizing this for learning.

[0002] Understanding and interpreting high-resolution images is a critical process in various fields requiring problem diagnosis and analysis. However, due to their complexity and vast amount of information, high-resolution images necessitate skilled analysis by experts and substantial computational resources. Consequently, diagnostic and analytical standards that rely on such expert knowledge may continuously change with the emergence of new technologies and shifting requirements.

[0003] With the recent advancement of digitalization, various computational methodologies have been introduced to automate image analysis. In particular, deep learning has proven effective in understanding complex and unstructured images. However, existing deep learning models tend to operate independently without interaction between experts and algorithms due to a lack of interpretability and limitations in the ability to infer from derived results.

[0004] These problems are particularly pronounced in the analysis of tissue pathology images. Understanding and interpreting tissue pathology images is a critical process in cancer diagnosis; however, due to the complexity of these images, image examination requires expertise and significant resources. Furthermore, extensive training is required for pathologists to fully grasp the medical context.

[0005] As a method to solve this problem, a method is known in which the whole slide image (WSI) of the tissue pathology image is divided into multiple non-overlapping patches, features of each patch are extracted using a pre-trained feature extractor such as ResNet or Vision Transformer, and then a feature aggregator combines the extracted features to finally perform slide-level prediction.

[0006] However, this method has the limitation of not effectively utilizing the multi-resolution characteristics of WSI. Pathologists make diagnostic decisions by identifying the overall tissue architecture at low magnification of tissue pathology images and examining cellular structures at high magnification. Existing models fail to effectively mimic this hierarchical analysis and face a multi-resolution dilemma in which high-frequency components serving as diagnostic clues are missed at low magnification, while the number of patches increases exponentially at high magnification, leading to massively increased analysis costs.

[0007] Furthermore, existing models do not explicitly consider the causal dependencies between low-magnification and high-magnification features in the latent space when integrating them for multi-resolution analysis. While the causal relationship in the image space—where a low-magnification image is obtained by downsampling a high-magnification image—is clear, this relationship becomes ambiguous in the latent feature space after passing through the feature extractor. Consequently, the feature aggregator learns only simple associations between low-magnification and high-magnification features, leading to a failure to improve or even a decline in performance for multi-resolution analysis when labeled data is limited.

[0008] According to various embodiments of the present disclosure, the present invention aims to provide a method and system for providing an image analysis-based response through multi-resolution analysis that can provide a response of enhanced accuracy by modeling the causal dependency between low-magnification features and high-magnification features of an image in a latent space, which is a mathematical space where compressed information is projected through a feature extractor of an image analysis model, generating adjusted low-magnification feature embeddings based thereon, and utilizing them for training an image analysis model.

[0009] According to various embodiments of the present disclosure, the present invention aims to provide an image analysis-based response provision method and system capable of effectively extending an existing single-resolution analysis system to multiple resolutions and improving prediction performance and reliability by using a cross-attention mechanism in the process of integrating features of multiple-scale images and introducing unexplained factors of the latent space downsampling process as noise terms to explicitly model the causal relationships between multiple scales.

[0010] According to various embodiments of the present disclosure, the present invention aims to provide an image analysis-based response provision method and system that can effectively mitigate the multi-resolution dilemma and dramatically improve prediction performance and reliability by capturing diagnostic clues from low-magnification images through spectral analysis of images and efficiently selecting patches from high-magnification images based thereon.

[0011] According to various embodiments of the present disclosure, the present invention aims to provide an image analysis-based response providing method and system that can significantly improve the computational efficiency of high-magnification image analysis by effectively encoding geometric positional relationships and morphological similarities between patches by considering both positional information and frequency characteristic information between a plurality of patches when calculating patch relationship information used to efficiently select patches of a high-magnification image, and based thereon.

[0012] However, the technical problems that the various embodiments of the present disclosure aim to solve are not limited to the technical problems described above, and other technical problems may exist.

[0013] One embodiment is,

[0014] A method executed by a computer comprises the steps of: receiving an activation signal for accessing at least one in-memory data; accessing at least one in-memory data structure upon receiving the activation signal, wherein the data structure includes a diagnostic image and loading the diagnostic image from at least one memory; generating at least one response to the diagnostic image using at least one artificial intelligence model that takes the diagnostic image as input by at least one processor, wherein at least one artificial intelligence model is pre-trained to perform image analysis based on a feature representation of a causal dependency relationship between low-magnification features and high-magnification features of the diagnostic image; and inputting the at least one response to at least one subsequent processing component.

[0015] In another aspect, the method may further include the step of the at least one subsequent processing component manifesting the at least one response through at least one user interface.

[0016] In another aspect, the method may further include the step of providing, through a user interface, an image that highlights a portion of the diagnostic image with high importance, which is used in a calculation to determine the at least one response, based on information regarding the causal dependency relationship between the low-magnification feature and the high-magnification feature.

[0017] In another aspect, the method may further include the step of providing, through a user interface, a heatmap or highlight visualization image that highlights areas corresponding to some high-magnification regions with relatively high scores in the diagnostic image, based on the score calculated during the causal dependency modeling process.

[0018] In another aspect, the diagnostic image includes a tissue pathology image, and the at least one response may include at least one of a prediction result regarding whether the diagnostic image includes a specific disease, a prediction result regarding the grade or stage of progression of the disease, and a prediction result regarding the location of a lesion for the disease.

[0019] In another aspect, the at least one artificial intelligence model may be pre-trained by a learning method comprising the steps of: loading a training image from the at least one memory; generating a plurality of first patch feature embeddings corresponding to low-magnification features of the training image by the at least one processor; generating a plurality of second patch feature embeddings corresponding to high-magnification features of the training image by the at least one processor; generating a plurality of low-magnification dependent feature embeddings through a predetermined operation modeling the causal dependency relationship between the plurality of first patch feature embeddings and the plurality of second patch feature embeddings; and updating the parameters of the at least one artificial intelligence model such that a loss based on result data output by the at least one artificial intelligence model and correct answer data for the training image is minimized based on the predetermined operation results for the plurality of second patch feature embeddings and the plurality of low-magnification dependent feature embeddings.

[0020] In another aspect, the step of generating the plurality of low-magnification dependent feature embeddings may include the step of generating the plurality of low-magnification dependent feature embeddings through a cross-attention operation based on the plurality of first patch feature embeddings and the plurality of second patch feature embeddings.

[0021] In another aspect, the cross-attention operation may be performed based on a query based on the plurality of first patch feature embeddings and a key and value based on the plurality of second patch feature embeddings.

[0022] In another aspect, the learning method may further include the step of adjusting the plurality of low-magnification dependent feature embeddings by summing the weighted sum of the plurality of high-magnification patch level noise embeddings corresponding to the plurality of patches of the high-magnification image of the learning image and the total image level noise embedding for the high-magnification image of the learning image for each of the plurality of low-magnification dependent feature embeddings.

[0023] In another aspect, the plurality of high-magnification patch-level noise embeddings (U j ) is a predetermined probability distribution (N(μ) according to the following (formula) j , σ j 2 The learnable mean (μ) of )) j ) and standard deviation (σ j It can be sampled by ).

[0024] ceremony:

[0025] (step, is a random variable sampled from a standard normal distribution)

[0026] In another aspect, the weighted sum of the plurality of high-magnification patch-level noise embeddings can be calculated based on weights generated during a cross-attention operation process based on the plurality of first patch feature embeddings and the plurality of second patch feature embeddings.

[0027] In another aspect, the learning method further comprises the steps of: generating a plurality of patch relationship information embeddings including relationship information between the plurality of patches from a plurality of patches of the learning image; selecting some of the plurality of first patch feature embeddings based on the plurality of patch relationship information embeddings; and selecting some of the plurality of second patch feature embeddings based on the plurality of patch relationship information embeddings, wherein the step of generating the plurality of low-magnification dependent feature embeddings may include generating the plurality of low-magnification dependent feature embeddings through a predetermined operation that models the causal dependency relationship between some of the selected plurality of first patch feature embeddings and some of the selected plurality of second patch feature embeddings.

[0028] In another aspect, the step of generating the plurality of patch relationship information embeddings is based on distance information for a plurality of patches of a low-magnification image of the training image, and a distance-based adjacency matrix (A dis A step of generating ) based on similarity information between frequency features for a plurality of patches of a low-magnification image for the training image, and a frequency feature-based adjacency matrix (A HF A step of generating ) and the distance-based adjacency matrix (A dis ) and the above frequency feature-based adjacency matrix (A HF It may include the step of generating a plurality of patch relationship information embeddings including location information and characteristic relationship information between the plurality of patches by utilizing ).

[0029] In another aspect, the step of generating the plurality of patch relationship information embeddings may include: performing a spectral analysis on the distance-based adjacency matrix to obtain a plurality of position feature vectors encoding geometric position information between a plurality of patches of the low-magnification image; and performing a neural network operation to learn a graph structure on a portion of the obtained plurality of position feature vectors selected according to a predetermined condition and on the frequency feature-based adjacency matrix to generate the plurality of patch relationship information embeddings.

[0030] In another aspect, the step of obtaining the plurality of position feature vectors may include the step of calculating a plurality of eigenvectors for the graph Laplacian of the distance-based adjacency matrix according to the following (equation) as the plurality of position feature vectors.

[0031] (ceremony):

[0032] (where Δ is the graph Laplacian, I n is the identity matrix, D dis is a degree matrix, A dis is the distance-based adjacency matrix, U is the eigenvector matrix, and Λ is the eigenvalue matrix)

[0033] In another aspect, the step of generating the plurality of first patch feature embeddings includes the step of generating the plurality of first patch feature embeddings by integrating the plurality of first visual feature embeddings and the plurality of first frequency feature embeddings for the low-magnification image of the training image, and the step of generating the plurality of second patch feature embeddings may include the step of generating the plurality of second patch feature embeddings by integrating the plurality of second visual feature embeddings and the plurality of second frequency feature embeddings for the high-magnification image of the training image.

[0034] In another aspect, the step of selecting some of the plurality of first patch feature embeddings may include the step of calculating a plurality of weights for a plurality of patches of the low-magnification image based on a plurality of patch integrated embeddings that integrate the plurality of first visual feature embeddings, the plurality of first frequency feature embeddings, and the plurality of patch relationship information embeddings; and the step of selecting k first patch feature embeddings corresponding to the top k weights (k is a positive integer) among the plurality of weights from among the plurality of first patch feature embeddings.

[0035] In another aspect, the step of selecting some of the plurality of second patch feature embeddings comprises: a step of calculating a plurality of weights for a plurality of patches of the high-magnification image based on a plurality of patch integration embeddings that integrate the plurality of second visual feature embeddings, the plurality of second frequency feature embeddings, and the plurality of patch relationship information embeddings; and among the plurality of second patch feature embeddings, the upper k×M of the plurality of weights 2 k×M corresponding to weights (M is the relative magnification of the above high magnification to the low magnification) 2 It may include a step of selecting two second patch feature embeddings.

[0036] One embodiment is,

[0037] The method may include at least one memory and at least one processor that reads at least one instruction stored in the at least one memory and executes an image analysis-based response providing method, wherein the at least one instruction may include: receiving an activation signal for accessing data in the at least one memory; accessing a data structure in the at least one memory upon receiving the activation signal; wherein the data structure includes a diagnostic image and the diagnostic image is loaded from the at least one memory; wherein the at least one processor generates at least one response to the diagnostic image using at least one artificial intelligence model that takes the diagnostic image as input, wherein the at least one artificial intelligence model is pre-trained to perform image analysis based on a feature representation of a causal dependency relationship between low-magnification features and high-magnification features of the diagnostic image; and the instruction that ingests the at least one response to at least one subsequent processing component.

[0038] In another aspect, the system may further include a Field Programmable Gate Array (FPGA) implementation for a predetermined artificial neural network comprising a plurality of neurons arranged in an array including at least one register, at least one programmable logic, and at least one input interface, a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons, and at least one routing network that controls the data flow between the plurality of neurons, wherein each of the plurality of neurons is connected to at least one other neuron through the routing network to establish a transmission path for the weights.

[0039] In another aspect, the system may further include a plurality of neurons organized into an array comprising at least one register, at least one microprocessor, and at least one input, and a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons, wherein each of the plurality of neurons may further include an Application Specific Integrated Circuit (ASIC) for a predetermined artificial neural network connected to at least one other neuron through any one of the plurality of synapse circuits.

[0040] According to various embodiments of the present disclosure, by using a cross-attention mechanism based on a plurality of patch feature embeddings from a low-magnification image and a plurality of patch feature embeddings from a high-magnification image, and by introducing unexplained factors of the downsampling process from high magnification to low magnification in the latent space as noise terms, a method for training an image analysis model capable of explicitly learning causal relationships between multiple magnifications, and a method and system for providing an image analysis-based response utilizing the same can be provided. This contributes to effectively extending existing single-resolution analysis systems to multiple resolutions and improving prediction performance and reliability.

[0041] According to various embodiments of the present disclosure, diagnostic clues of a low-magnification image are captured through spectral analysis of a plurality of patches of a low-magnification image, and based on this, the multi-resolution dilemma can be effectively mitigated by efficiently selecting some of the plurality of patches of a high-magnification image that have a significant impact on the prediction result.

[0042] According to various embodiments of the present disclosure, by spectral analysis of multiple patches of low-magnification images at relatively low magnification levels of high-resolution images, important clues for diagnosis can be captured in low-magnification images, and based on this, Top-K patches can be selected to focus only on important areas instead of searching all high-magnification patches, thereby significantly reducing GPU memory usage and inference time, and dramatically improving the performance and computational efficiency of the model.

[0043] According to various embodiments of the present disclosure, the interpretability and reliability of the decision process of an image analysis model can be enhanced. By generating patch relationship information including positional and characteristic relationship information between a plurality of patches to structurally encode the positional and morphological similarities between patches, and by selecting important patches using a Gated Attention mechanism based on the patch relationship information, information can be provided that allows experts to easily understand the basis on which the model reached its conclusion. This facilitates interaction between the model and experts and can increase the reliability of image analysis.

[0044] However, the effects obtainable through the various embodiments of the present disclosure are not limited to those mentioned above, and other unmentioned effects can be clearly understood from the description below.

[0045] FIG. 1 illustrates an example of a block diagram of a computing system implementing an image analysis-based response provision service according to one embodiment.

[0046] FIG. 2 briefly illustrates the structure of a neuromorphic circuit that may be included in a processor according to one embodiment.

[0047] FIG. 3 is a block diagram of a computing device implementing an image analysis-based response provision service according to one embodiment.

[0048] FIG. 4 is a block diagram of a computing device implementing an image analysis-based diagnostic provision service according to another embodiment.

[0049] FIG. 5 is a block diagram of a computing device implementing an image analysis-based response provision service according to another embodiment.

[0050] FIG. 6 is a conceptual diagram of a framework of a multi-resolution integrated model reflecting the causal relationship between multiple magnifications according to one embodiment.

[0051] FIG. 7 is a detailed configuration diagram of a framework of a multi-resolution integrated model reflecting the causal relationship between multiple magnifications according to one embodiment.

[0052] FIG. 8 is a block diagram of a computing device implementing an image analysis-based response provision service according to another embodiment.

[0053] FIG. 9 illustrates an exemplary configuration of a spectrum analysis-based multi-resolution analysis architecture for a high-resolution image according to one embodiment.

[0054] FIG. 10 is intended to explain a method for extracting location information and characteristic relationship information between a plurality of patches of a high-resolution image according to one embodiment.

[0055] FIG. 11 is a flowchart of an image analysis-based response provision method according to one embodiment.

[0056] FIG. 12 is a flowchart of a learning method for an image analysis model according to one embodiment.

[0057] As various modifications can be made to the various embodiments of the present disclosure, specific embodiments are illustrated in the drawings and described in detail in the detailed description. The effects and features of the various embodiments of the present disclosure, and the methods for achieving them, will become clear by referring to the embodiments described in detail below together with the drawings. However, the various embodiments of the present disclosure are not limited to the embodiments disclosed below but can be implemented in various forms. In the following embodiments, terms such as "first," "second," etc., are used not in a limiting sense but for the purpose of distinguishing one component from another. Also, singular expressions include plural expressions unless the context clearly indicates otherwise. Furthermore, terms such as "include" or "have" mean that the features or components described in the specification exist, and do not preclude the possibility that one or more other features or components may be added. Additionally, in the drawings, the size of components may be exaggerated or reduced for convenience of explanation. For example, the size and thickness of each component shown in the drawings are arbitrarily depicted for convenience of explanation, so the various embodiments of the present disclosure are not necessarily limited to those depicted.

[0058] Hereinafter, various embodiments of the present disclosure will be described in detail with reference to the attached drawings. When describing with reference to the drawings, identical or corresponding components are given the same reference numerals, and redundant descriptions thereof will be omitted.

[0059]

[0060] - Image analysis-based response provision system (1000)

[0061] A system (1000) according to one embodiment can perform multi-resolution analysis on an image and provide a highly accurate response based on the results.

[0062] When learning only associations by simply combining low-magnification and high-magnification features of diagnostic images, a limitation arises in that the causal dependencies existing between resolutions in the latent space cannot be explicitly utilized.

[0063] The system (1000) can provide a response of improved accuracy for diagnostic images by using a new feature integration methodology that introduces causal modeling between the images to overcome these limitations.

[0064] The system (1000) performs a process of adjusting low-magnification feature embeddings. First, it prepares low-magnification feature embeddings and high-magnification feature embeddings for a diagnostic image through a feature extractor. Subsequently, the system (1000) can model causal dependency relationships between high-magnification features and low-magnification features by generating low-magnification dependent features through a cross-attention mechanism that utilizes low-magnification features as queries and high-magnification features as key / value pairs.

[0065] Furthermore, the system (1000) can introduce parameters of causal noise terms to capture unexplained variability or uncertainty occurring during the causal dependency modeling process, and add this to the cross-attention result to generate an adjusted low-magnification feature embedding.

[0066] Finally, the system (1000) can input the adjusted low-magnification feature embeddings and the original high-magnification feature embeddings into the final feature aggregator of the Multiple instance learning (MIL) framework to generate a final response.

[0067] Through this process, the system (1000) achieves sophisticated feature integration that goes beyond the limitations of simple association learning between multiple resolutions and even considers the uncertainty of causal relationships between resolutions, and as a result, can provide a highly reliable response to the user by dramatically improving the final prediction performance and accuracy of the image analysis model.

[0068] FIG. 1 illustrates an example of a block diagram of a computing system (1000) implementing an image analysis-based response provision service according to one embodiment.

[0069] Referring to FIG. 1, a computing system (1000) implementing an image analysis-based response providing service according to one embodiment includes a user computing device (110), a server computing system (130), and a training computing system (150), and the devices can communicate through a network (170).

[0070] A method for providing an image analysis-based diagnostic response according to one embodiment may be implemented and provided locally by a user computing device (110), implemented and provided in the form of a web service by a server computing system (130) communicating with the user computing device (110), or implemented and provided by the user computing device (110) and the server computing system (130) in conjunction with each other.

[0071] In this embodiment, the user computing device (110) and / or the server computing system (130) can train a machine learning model (120 and / or 140) through interaction with a training computing system (150) that is communicatedly connected via a network (170). The training computing system (150) may be separate from the server computing system (130) or may be part of the server computing system (130).

[0072] And at this time, the artificial intelligence model can be 1) trained directly locally by a user computing device (110), 2) trained by the server computing system (130) and the user computing device (110) interacting with each other through a network (170), and 3) trained by a separate training computing system (150) using various training and learning techniques. It may also be implemented by transmitting the artificial intelligence model trained by the training computing system (150) to the user computing device (110) and / or the server computing system (130) through the network (170) to provide / update it.

[0073] In some embodiments, the training computing system (150) may be part of the server computing system (130) or part of the user computing device (110).

[0074]

[0075] - User Computing Device (110: User Computing Device)

[0076] The user computing device (110) may include all other types of computing devices, such as a smartphone, a mobile phone, a digital broadcasting device, a PDA (personal digital assistants), a PMP (portable multimedia player), a desktop, a wearable device, an embedded computing device, a tablet PC, an augmented reality (VR) device, and / or a virtual reality (AR) device.

[0077] The user computing device (110) may include at least one processor (111) and memory (112). Here, the processor (111) may be composed of at least one of a central processing unit (CPU), a graphics processing unit (GPU), ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions, or a plurality of electrically connected processors.

[0078] In particular, according to the embodiment, this processor (111) may be configured based on a Field Programmable Gate Array (FPGA) implementation and / or an Application Specific Integrated Circuit (ASIC), which is a hardware technology for implementing a certain digital circuit.

[0079] Here, a field programmable gate array (FPGA) can refer to a flexible digital circuit that is programmable according to user needs.

[0080] As an example, a field programmable gate array implementation may include a register that temporarily stores data and controls the flow and timing of signals to maintain intermediate results or state information of operations to support synchronized operation of the FPGA, programmable logic that programs operations within the FPGA to perform specific functions or operations as logic circuits configurable according to user needs, and an input interface that receives signals from external devices or sensors and transmits them to internal circuits as a channel for receiving data from outside the FPGA.

[0081] Through the combination of the above components, a field-programmable gate array implementation can provide flexible and various types of digital circuits.

[0082] Meanwhile, an Application-Specific Integrated Circuit (ASIC) can refer to a custom integrated circuit that is fixedly designed to perform a specific use or function.

[0083] As an example, the application-dedicated integrated circuit may include a register, which is a small memory device for temporarily storing and managing data and supports the rapid processing of ASIC operations by storing intermediate calculation results or state information; a microprocessor, which is a central processing unit that performs control and operations within the ASIC and coordinates the operation of the entire system by performing various operations or generating control signals when necessary; and an input block, which is an interface for receiving data from the outside, which receives data to be processed by the ASIC and transmits it internally, and receives various input data through connections with sensors or external devices.

[0084] Through the combination of the components mentioned above, an application-specific integrated circuit can perform specific purpose tasks in an optimized manner.

[0085] For example, ASICs can have a structure of a neuromorphic circuit in the form of an array containing multiple neuron circuits.

[0086] FIG. 2 briefly illustrates the structure of a neuromorphic circuit (300) that may be included in a processor (111, 131, 151) according to one embodiment.

[0087] Referring to FIG. 2, for example, a neuromorphic circuit (300) may include a plurality of presynaptic neuron circuits (310), a plurality of presynaptic lines (311) extending laterally from the plurality of presynaptic neuron circuits (310), a plurality of postsynaptic neuron circuits (320), a plurality of postsynaptic lines (321) extending longitudinally from the plurality of postsynaptic neuron circuits (320), and a plurality of synaptic circuits (330) provided at the intersection of the plurality of presynaptic lines (311) and the plurality of postsynaptic lines (321).

[0088] A plurality of free synaptic neuron circuits (310) can transmit signals input from the outside in the form of electrical signals to a plurality of synaptic circuits (330) through a plurality of free synaptic lines (311).

[0089] Additionally, a plurality of post-synaptic neuron circuits (320) can receive electrical signals from a plurality of synaptic circuits (330) through a plurality of post-synaptic lines (321).

[0090] Furthermore, multiple post-synaptic neuron circuits (320) may transmit electrical signals to multiple synaptic circuits (330) through multiple post-synaptic lines (321).

[0091] A plurality of synapse circuits (330) can store weights included in layers constituting a neural network system implemented by a neuromorphic circuit (300) and perform a predetermined operation based on the weights and input data.

[0092] For example, each of the plurality of synaptic circuits (330) may include a resistive memory cell having a variable resistance. In this case, the resistance value of the plurality of synaptic circuits (330) changes by a voltage applied through the plurality of presynaptic neuron circuits (310) or the plurality of postsynaptic neuron circuits (320), and can store weight data according to this resistance change.

[0093] The neuromorphic circuit (300) is formed by mimicking the structure of neurons and synapses, which are essential elements of the human brain. When a deep neural network (DNN) is realized using the neuromorphic circuit (300), the data processing speed can be improved and power consumption can be reduced compared to when the existing von Neumann structure is utilized.

[0094] The memory (112) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, and combinations thereof, and may include web storage of a server that performs memory storage functions on the internet. This memory (112) may store data (113) and instructions (114) necessary for the at least one processor (111) to perform functional operations such as training an artificial intelligence model or performing image analysis through an artificial intelligence model.

[0095] In one embodiment, the user computing device (110) can store at least one machine learning model (120).

[0096] For example, the machine learning model (120) may be various machine learning models, such as multiple neural networks (e.g., deep neural networks) for performing an image analysis-based diagnostic response provision method, or other types of machine learning models including non-linear models and / or linear models, and may be composed of a combination thereof.

[0097] For example, machine learning models may include linear regression, decision trees, random forests, gradient-boosting pre-trained language models or / and deep learning models. And neural networks may include at least one of feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks or / and other forms of neural networks.

[0098] Additionally, according to various embodiments, the user computing device (110) may store a model to be used in each process and a prompt template that serves as the basis for input to the model in order to perform at least part of the process for the high-resolution image analysis-based diagnostic response provision method through a large language model (LLM).

[0099] In one embodiment, a user computing device (110) may receive at least one machine learning model (120) from a server computing system (130) through a network (170), store it in memory (112), and then execute the stored machine learning model (120) through a processor (111) to perform an operation for providing an image analysis-based response.

[0100] In another embodiment, the server computing system (130) includes at least one machine learning model (140) and performs operations through the machine learning model (140), and can provide a high-resolution image analysis-based diagnostic response provision service to the user by communicating with the user computing device (110) and related data.

[0101] For example, a user computing device (110) can perform a high-resolution image analysis-based diagnostic response provision service in which a server computing system (130) provides an output for the user's input using a machine learning model (140) via the web.

[0102] Additionally, the artificial intelligence model can be implemented in such a way that at least some of the machine learning models (120 and / or 140) are executed on a user computing device (110) and the rest are executed on a server computing system (130).

[0103] Additionally, the user computing device (110) may include at least one input component (121) for detecting user input. For example, the user input component (121) may include a touch sensor (e.g., a touch screen and / or a touch pad, etc.) for detecting a touch of a user input medium (e.g., a finger or a stylus), an image sensor for detecting user motion input, a microphone for detecting user voice input, a button, a mouse and / or a keyboard, etc. Additionally, the user input component (121) may include an interface and an external controller when receiving input to an external controller (e.g., a mouse and / or a keyboard, etc.) through an interface.

[0104]

[0105] -Server Computing System (130: Server Computing System)

[0106] The server computing system (130) can perform a series of processes to provide an image analysis-based response provision service.

[0107] Specifically, in an embodiment, the server computing system (130) can provide an image analysis-based response provision service by exchanging data necessary to enable an image analysis-based response provision service process to be driven on an external device, such as a user computing device (110).

[0108] More specifically, in an embodiment, the server computing system (130) can provide an environment in which an application for providing an image analysis-based response provision service on a user computing device (110) can operate.

[0109] To this end, the server computing system (130) may include an application program, data and / or instructions, etc. for the application to operate, and may transmit and receive various data based thereon with the external device.

[0110] A server computing system (130) may include at least one processor (131) and memory (132). Here, the processor (131) may be composed of at least one of a central processing unit (CPU), a graphics processing unit (GPU), ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, microcontrollers, microprocessors, and / or electrical units for performing other functions, or a plurality of electrically connected processors.

[0111] For example, ASICs may have a structure of a neuromorphic circuit in the form of an array containing multiple neuron circuits (see Fig. 2).

[0112] And the memory (132) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory device, magnetic disk, etc. and combinations thereof. This memory (132) may store data (133) and instructions (134) necessary for the processor (131) to perform functional operations, such as training an artificial intelligence model or executing an image analysis-based response provision method through the artificial intelligence model.

[0113] In one embodiment, the server computing system (130) may be implemented to include at least one computing device. For example, the server computing system (130) may be implemented to operate a plurality of computing devices according to a sequential computing architecture, a parallel computing architecture, or a combination thereof. Additionally, the server computing system (130) may include a plurality of computing devices connected to a network (170).

[0114] Additionally, the server computing system (130) may store at least one machine learning model (140). For example, the server computing system (130) may include a neural network and / or other multi-layer non-linear model as the machine learning model (140). Exemplary neural networks may include a feed-forward neural network, a deep neural network, a recurrent neural network, and a convolutional neural network.

[0115] In an embodiment, the server computing system (130) may further include a data store computing system (hereinafter, data store) which is a storage for continuously storing and managing raw data that forms the basis of an image analysis-based response provision service.

[0116] Such data stores may include various forms of data storage, ranging from file systems to cloud storage. For example, a data store may include at least one database among a relational database that uses a structured query language (SQL) to define and manipulate data, a NoSQL database designed for flexibility and scalability to process unstructured and semi-structured data, a data warehouse optimized for querying and analysis by centralizing large volumes of data from multiple sources as a system used for reporting and data analysis, a data warehouse that stores large volumes of raw data in basic formats such as structured data, semi-structured data, and unstructured data, and a local storage device or Network Attached Storage (NAS) that stores data in files in a format generally accessible by a computer operating system.

[0117]

[0118] - Training Computing System (150: Training Computing System)

[0119] The training computing system (150) may include at least one processor (151) and memory (152). Here, the processor (151) may be composed of at least one of a central processing unit (CPU), a graphics processing unit (GPU), ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions, or a plurality of electrically connected processors.

[0120] For example, ASICs may have a structure of a neuromorphic circuit in the form of an array containing multiple neuron circuits (see Fig. 2).

[0121] And the memory (152) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory device, magnetic disk, etc. and combinations thereof. This memory (152) may store data (153) and instructions (154) necessary for the processor (151) to perform learning of an artificial intelligence model, etc.

[0122] For example, the training computing system (150) may include a model trainer (160) that trains a machine learning model (120 and / or 140) stored in a user computing device (110) and / or a server computing system (130) using various training or learning techniques, such as back propagation of error.

[0123] For example, such a model trainer (160) can perform updates to one or more parameters of a machine learning model (120 and / or 140) for a high-resolution image analysis-based diagnostic response provision service based on a defined loss function in a backpropagation manner.

[0124] In some embodiments, performing backpropagation of the error may include performing truncated backpropagation through time. The model trainer (160) may perform a number of generalization techniques (e.g., weight devaluation, dropout and / or knowledge distillation, etc.) to improve the generalization ability of the machine learning model (120 and / or 140) being trained.

[0125] For example, a model trainer (160) can train a machine learning model (120 and / or 140) based on a series of training data (161). Here, the training data (161) may include data of different forms, such as, for example, images, audio samples and / or text.

[0126] Examples of image types that can be used may include video frames, LiDAR point clouds, X-ray images, computed tomography scans, hyperspectral images, and / or various other forms of images.

[0127] These training data (161) may be provided by a user computing device (110) and / or a server computing system (130). When the training computing device trains a machine learning model (120 and / or 140) on specific data of the user computing device (110), the machine learning model (120 and / or 140) may be characterized as a personalized model.

[0128] And the model trainer (160) includes computer logic that is utilized to provide the desired function.

[0129] Additionally, the model trainer (160) may be implemented as hardware, firmware, and / or software that controls a general-purpose processor. In one embodiment, the model trainer (160) may include a program file stored in a storage device, be loaded into memory (152), and be executed by one or more processors (151). In another embodiment, the model trainer (160) includes one or more sets of computer-executable data (153) and instructions (154) stored in a tangible computer-readable storage medium, such as a RAM hard disk or an optical or magnetic medium.

[0130] Network (170) includes, but is not limited to, 3GPP (3rd Generation Partnership Project) network, LTE (Long Term Evolution) network, WIMAX (World Interoperability for Microwave Access) network, Internet, LAN (Local Area Network), Wireless LAN (Wireless Local Area Network), WAN (Wide Area Network), PAN (Personal Area Network), Bluetooth network, satellite broadcasting network, analog broadcasting network and / or DMB (Digital Multimedia Broadcasting) network.

[0131] Generally, communication through the network (170) can be performed using any type of wired and / or wireless connection through various communication protocols (e.g., TCP / IP, HTTP, SMTP and / or FTP, etc.), encodings or formats (e.g., HTML and / or XML, etc.), and / or protection schemes (e.g., VPN, Secure HTTP and / or SSL, etc.).

[0132] FIG. 3 is a block diagram of a computing device (100) implementing an image analysis-based response provision service according to one embodiment.

[0133] Referring to FIG. 3, the computing device (100) included in the user computing device (110), server computing system (130), and training computing system (150) includes a plurality of applications (e.g., applications 1 to N). Each application may include a machine learning library and one or more machine learning models. For example, the applications may include an image processing application (e.g., detection, classification, and / or segmentation, etc.), a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and / or a chat-bot application.

[0134] In an embodiment, the computing device (100) may include a model trainer (160) for training an artificial intelligence model, and by storing and operating the trained artificial intelligence model, it may provide output data according to a predetermined input data.

[0135] Each application of the computing device (100) can communicate with a number of other components of the computing device (100), such as, for example, at least one sensor, a context manager, a device state component, and / or additional components. In one embodiment, each application can communicate with each device component using an API (e.g., a public API). In one embodiment, the API used by each application may be specific to that application.

[0136] FIG. 4 is a block diagram of a computing device (200) implementing an image analysis-based response provision service according to another embodiment.

[0137] Referring to FIG. 4, the computing device (200) includes a plurality of applications (e.g., Application 1 to Application N). Each application can communicate with a central intelligence layer. For example, applications may include an image processing application, a text messaging application, an email application, a dictation application, a virtual keyboard application and / or a browser application. In one embodiment, each application can communicate with the central intelligence layer (and a model stored therein) using an API (e.g., a common API across all applications).

[0138] The central intelligence layer may include a number of machine learning models. For example, as illustrated in FIG. 4, at least some of the machine learning models may be provided for each application and managed by the central intelligence layer. In other embodiments, two or more applications may share a single machine learning model. For example, in some embodiments, the central intelligence layer may provide a single model for all applications. In some embodiments, the central intelligence layer may be included within the operating system of the computing device (200) or otherwise implemented.

[0139] The central intelligence layer can communicate with the central device data layer. The central device data layer may be a centralized data store for the computing device (200). As illustrated in FIG. 4, the central device data layer can communicate with a number of other components of the computing device (200), such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some embodiments, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0140] The technology described herein may refer to servers, databases, software applications, and other computer-based systems, as well as actions taken and information transmitted to or from said systems. It will be recognized that the inherent flexibility of computer-based systems allows for a wide range of possible configurations, combinations, division of tasks, and functionality between and from components. For example, the processes described herein may be implemented using a single device or component or multiple devices or components operating in combination. Databases and applications may be implemented in a single system or in a distributed system across multiple systems. Distributed components may operate sequentially or in parallel.

[0141] FIG. 5 is a block diagram of a computing device (400) implementing an image analysis-based response provision service according to another embodiment. FIG. 6 is a conceptual diagram of a framework of a multi-resolution integrated model reflecting causal relationships between multiple magnifications according to one embodiment. FIG. 7 is a detailed configuration diagram of a framework of a multi-resolution integrated model reflecting causal relationships between multiple magnifications according to one embodiment. FIG. 8 is a block diagram of a computing device (401) implementing an image analysis-based response provision service according to another embodiment. FIG. 9 illustrates an exemplary configuration of a spectrum analysis-based multi-resolution analysis architecture for a high-resolution image according to one embodiment. FIG. 10 is intended to explain a method for extracting location information and characteristic relationship information between a plurality of patches of a high-resolution image according to one embodiment.

[0142] Referring to FIG. 5, the computing device (400) included in the user computing device (110), server computing system (130) and training computing system (150) may include a low-magnification patching module (1), a high-magnification patching module (2), a first visual feature extraction module (3), a second visual feature extraction module (4), a causal relationship modeling module (5), and a response generation module (6).

[0143] The low-magnification patching module (1) receives a high-resolution image (HRI) and can process it into a low-magnification image such as 10x or 5x. The low-magnification patching module (1) can efficiently obtain data for the low-magnification image that is necessary to grasp the overall context and structure among the vast amount of information in the high-resolution image (HRI).

[0144] The low-magnification patching module (1) can divide the low-magnification image generated in this way into multiple patches using a method such as a sliding window, for example. However, it is not limited to this, and various methods other than the method using a sliding window can be adopted for dividing the low-magnification image into multiple patches.

[0145] The high-magnification patching module (2) receives a high-resolution image (HRI) and can process it into a high-magnification image such as 20x or 10x. The high-magnification patching module (2) can efficiently obtain high-magnification image data necessary to identify fine morphological clues at the cellular level from among the vast amount of information in the high-resolution image (HRI).

[0146] The high-magnification patching module (2) can divide the high-magnification image generated in this way into multiple patches using a method such as a sliding window. However, it is not limited to this, and various methods other than the method using a sliding window can be adopted for dividing the high-magnification image into multiple patches.

[0147] The first patch feature extraction module (3) can receive multiple patches from the low-magnification patching module (1) and extract original features. Original features refer to the visual features of the corresponding image, and the first patch feature extraction module (3) can capture overall visual information of the image and provide basic feature representations for subsequent analysis.

[0148] The first patch feature extraction module (3) can extract visual features using a pre-trained vision encoder (e.g., ResNet), and this vision encoder can compress the visual pattern of the input patch and convert it into a latent representation. In this process, the parameters of the vision encoder can be used in a frozen state to reduce the computational burden caused by the vast number of patches in the image.

[0149] For example, referring to FIGS. 7 and 9, the first patch feature extraction module (3) receives a plurality of patches for a low-magnification image and, from this, a plurality of first patch feature embeddings (Original features for the plurality of patches) corresponding to the original features of the plurality of patches Can generate ).

[0150] The second patch feature extraction module (4) can receive multiple patches from the high-magnification patching module (2) and extract original features. Original features refer to the visual features of the corresponding image, and the second patch feature extraction module (4) can capture overall visual information of the image and provide basic feature representations for subsequent analysis.

[0151] The second patch feature extraction module (4) can extract visual features using a pre-trained vision encoder (e.g., ResNet), and this vision encoder can compress the visual pattern of the input patch and convert it into a latent representation. In this process, the parameters of the vision encoder can be used in a frozen state to reduce the computational burden caused by the vast number of patches in the image.

[0152] For example, referring to FIGS. 7 and 9, the second patch feature extraction module (4) receives a plurality of patches for a high-magnification image and, from this, a plurality of second patch feature embeddings (Original features for the plurality of patches) corresponding to the original features of the plurality of patches Can generate ).

[0153] The causal dependency modeling module (5) can apply the MAC (Multi-resolution aggregation with causal perspective) method according to the framework of a multi-resolution integrated model that reflects causal dependencies between multiple magnifications in order to dramatically improve the final prediction performance and accuracy by explicitly modeling the causal dependency between low-magnification features and high-magnification features and generating an adjusted low-magnification feature embedding that reflects causal uncertainty.

[0154] The causal dependency modeling module (5) has a plurality of first patch feature embeddings ( ) and multiple second patch feature embeddings( By performing an operation that explicitly models the causal dependency relationship between low-magnification features and high-magnification features in the projected latent space, it is possible to generate a new feature representation for the features of a low-magnification image that have been causally influenced by the features of a high-magnification image, namely, multiple low-magnification dependent feature embeddings.

[0155] For example, the causal dependency modeling module (5) models the causal dependency between multiple scales using a plurality of first patch feature embeddings ( ) and multiple second patch feature embeddings( Multiple low-scale dependent feature embeddings ( Can generate ).

[0156] This operation is performed according to the following equations (1) to (4), a plurality of first patch feature embeddings ( Query based on ) (q i ) and multiple second patch feature embeddings( Key based on ) (k ij ) and Value(v ij It can be performed based on ).

[0157] Through this process, multiple low-magnification dependent feature embeddings causally influenced by the features of high-magnification images ( ) can be generated.

[0158]

[0159] Equation (1):

[0160] Formula (2):

[0161] Equation (3):

[0162] Equation (4): = , = , = ,

[0163] (step, , , , are each learnable parameters)

[0164]

[0165] Here, the causal influence of high-magnification image features refers to the effect that high-magnification features have on low-magnification features in the latent space, such as downsampling a high-magnification image to generate a low-magnification image in the image space. The MAC method can overcome the limitations of conventional simple association learning by explicitly modeling these structural and hierarchical relationships.

[0166] Meanwhile, as described above, the causal dependency modeling module (5) has a plurality of low-magnification dependent feature embeddings ( In generating ), a plurality of first patch feature embeddings ( ) and multiple second patch feature embeddings( Rather than utilizing all of ), as described below with reference to FIGS. 8 to 10, the aforementioned series of operations is performed on some of the multiple first patch feature embeddings selected based on multiple patch relationship information embeddings and some of the multiple second patch feature embeddings selected to obtain multiple low-magnification dependent feature embeddings ( It can also generate ).

[0167] Furthermore, the causal dependency modeling module (5) identifies unexplained factors of the downsampling process that generates a low-magnification image based on a high-magnification image as noise embeddings (U, Causal modeling can be reinforced by introducing ).

[0168] Here, 'unexplained factors' refer to the fundamental uncertainty or non-deterministic variability arising from the fact that, despite the clear downsampling relationship in the image space, there is no function in the latent space that perfectly explains Z using only patch feature embeddings after passing through the feature extraction module.

[0169] For example, these factors may include variability that is difficult for a model to predict or explain, such as subtle visual variations at the patch level or local artifacts generated during the image preparation process, and noise embeddings (U, ) capable of capturing such variability By introducing ) multiple low-magnification dependent feature embeddings( Causal modeling can be reinforced by adjusting ).

[0170] For example, the causal dependency modeling module (5) has a plurality of low-magnification dependent feature embeddings ( ) weighted sum of multiple patch-level noise embeddings (U) corresponding to multiple patches of a high-magnification image for each, and the total image-level noise embedding for the high-magnification image ( Summing ) to obtain multiple integrated low-magnification dependent feature embeddings( It can generate ). These multiple integrated low-magnification dependent feature embeddings ( ) is a multiple adjusted low-magnification dependent feature embedding ( It can be referred to as ).

[0171] Here, the weights used in the weighted sum of the patch-level noise embeddings (U) are the low-magnification dependent feature embeddings ( It can be the same as the weight generated during the cross-attention operation process to calculate ).

[0172]

[0173] Equation (5):

[0174]

[0175] (where M is the relative magnification of the high magnification to the low magnification)

[0176]

[0177] In addition, multiple high-magnification patch-level noise embeddings (U j ) is a predetermined probability distribution (N(μ) according to the following equation (6). j , σ j2 The learnable mean (μ) of )) j ) and standard deviation (σ j Sampled by ), and full image-level noise embeddings for high-magnification images ( ) is a predetermined probability distribution (N( , The learnable average of )) ) and standard deviation( It can be sampled by ) (see Fig. 7).

[0178]

[0179] Equation (6):

[0180] Equation (7): ,

[0181] (step, and are random variables sampled from a standard normal distribution, respectively)

[0182]

[0183] In this way, a plurality of high-magnification patch-level noise embeddings (U j By defining this based on a predetermined probability distribution, the range of various values ​​and the potential for noise that represents variability difficult for image analysis models to predict or explain, such as subtle visual variations at the patch level or local artifacts generated during the image preparation process, can be modeled.

[0184] The response generation module (6) has a plurality of second patch feature embeddings ( ) and multiple low-magnification dependent feature embeddings( At least one response to a diagnostic image can be generated based on a predetermined operation result for ).

[0185] For example, the response generation module (6) can act as a feature aggregator for the multi-instance learning (MIL) framework, and in this case, the low-scale dependency feature embeddings generated by the causal dependency modeling module (5) ) and multiple second patch feature embeddings( It can integrate ).

[0186] Additionally, for example, the response generation module (6) has a plurality of integrated low-magnification dependent feature embeddings generated by the causal dependency modeling module (5). ) is and multiple second patch feature embeddings( It can integrate ).

[0187] Here, the integration operation may include various methods for outputting result data by combining multiple resolution features, and for example, may include the following four methods. However, it is not limited to this, and the integration operation may include various operations other than the four methods below.

[0188] First, the above integration operation may include an instance-level concatenation operation. In this case, the response generation module (6) comprises a plurality of low-magnification dependent feature embeddings ( ) and multiple second patch feature embeddings( Concatenates ) into one large feature set, and inputs the concatenated feature set into a feature aggregator to obtain the final result data ( Can generate ).

[0189] Here, the feature integrator is a component that generates bag-level features by performing a predetermined operation on an input feature set, and may include various types of operation modules that perform mean pooling, max pooling, weight-based pooling, attention-based aggregation, or multilayer perceptron (MLP) operations.

[0190] Second, the above integration operation may include a Bag-level Concatenation operation. In this case, the response generation module (6) comprises a plurality of low-magnification dependent feature embeddings ( ) and multiple second patch feature embeddings( The bag-level representations obtained by processing each of the ) through separate feature integrators (g1, g2) are input into the final prediction layer (σ) to obtain the final result data ( Can generate ).

[0191] Third, the above integration operation may include weighted pooling or attention-based aggregation operations. This involves two feature sets (multiple adjusted low-magnification dependent feature embeddings ( ) and multiple second patch feature embeddings( Instead of simply concatenating )), it is an operation that integrates them by performing a weighted sum considering the importance of each feature (e.g., attention weight).

[0192] Fourth, the above integration operation may include a transformation operation. This involves two feature sets (multiple adjusted low-magnification dependent feature embeddings ( ) and multiple second patch feature embeddings( It may include operations that fuse features by applying linear or non-linear transformations to )) and then summing or multiplying.

[0193] Based on such integrated information, the response generation module (6) can generate at least one response related to pathological diagnosis, such as cancer type prediction, cancer grade prediction (Grading), tumor tissue mutation burden (TMB) prediction, or microsatellite instability (MSI) classification. Since this response reflects the results of causal dependency learning, it can have improved accuracy and reliability.

[0194] However, it is not limited to this, and the response generation module (6) can provide a final response for various tasks such as defect classification and location identification, prediction of crop growth stage and presence or absence of disease, or detection and recognition of specific objects when the diagnostic image is an image related to various fields such as defect inspection in manufacturing, crop disease diagnosis in agriculture, or image analysis in defense, rather than the medical field.

[0195] For example, diagnostic images may include various high-resolution images related to diverse fields, such as crop images, aerial satellite images, manufactured product images, and object images.

[0196] In this case, the response generation module (6) can generate at least one of the following responses: a prediction result regarding the growth stage or presence or absence of disease of the crop included in the crop image, a prediction result regarding the progression stage of target identification or surface change in the aerial satellite image, a prediction result regarding defect classification or defect location identification in the manufactured product image, and a prediction result regarding object identification in the object capture image.

[0197] Additionally, the response generation module (6) can generate an image that highlights a part of the diagnostic image used in the calculation to determine at least one response, based on information regarding the causal dependency relationship between low-magnification features and high-magnification features, in order to provide interpretability of the final response.

[0198] In this case, by utilizing the attention score calculated during the causal dependency relationship modeling process, an image can be generated that emphasizes a high-importance portion of the diagnostic image used in the operation to determine at least one response.

[0199] This attention score is a value that quantifies the causal influence of high-magnification features on low-magnification features of a diagnostic image, and high-magnification areas with a high attention score indicate that these areas served as the most decisive clues when the image analysis model derived the final response. The response generation module (6) can generate a visualization image with a heatmap or highlights on these high-importance areas and provide it through a user interface. Such visual evidence is a key factor in increasing the confidence of pathologists or experts in the model's prediction results and can significantly improve diagnostic efficiency by allowing easy verification of whether the model focused on a reasonable medical context.

[0200] Meanwhile, in efficiently analyzing the features of patches for an image to generate at least one response, a system (1000) according to one embodiment may utilize a hierarchical analysis method that utilizes both the features of a low-magnification image and a high-magnification image to alleviate the multiple resolution dilemma.

[0201] In this case, the system (1000) receives a diagnostic image, divides it into low-magnification and high-magnification patches, and can extract both original features and high-frequency features from each patch. Subsequently, the system (1000) generates a learnable geometric position encoding (LGPE) containing information on geometric positional relationships and morphological similarities between the patches, and based on this, can efficiently select diagnostically important Top-K patches from the low-magnification image.

[0202] Additionally, the system (1000) can finally aggregate the features of high-magnification patches corresponding to the selected low-magnification patches and output a predicted value for the diagnostic image. Through this process, the system (1000) can provide a diagnostic response with high accuracy and reliability while drastically reducing the computational complexity required to analyze all high-magnification patches.

[0203] Referring to FIG. 8, a computing device (401) included in a user computing device (110), a server computing system (130), and a training computing system (150) according to one embodiment may include a low-magnification patching module (10), a high-magnification patching module (20), a first frequency feature extraction module (30), a first visual feature extraction module (40), a second frequency feature extraction module (31), a second visual feature extraction module (41), a patch relationship information embedding generation module (50), a low-magnification patch feature embedding generation module (60), a high-magnification patch feature embedding generation module (61), a low-magnification patch feature embedding selection module (70), a high-magnification patch feature embedding selection module (71), and a prediction module (80).

[0204] The low-magnification patching module (10) receives a high-resolution image (HRI) and can process it into a low-magnification image, such as 10x or 5x. The low-magnification patching module (10) can efficiently obtain data for the low-magnification image that is necessary to grasp the overall context and structure among the vast amount of information in the high-resolution image (HRI).

[0205] The low-magnification patching module (10) can divide the low-magnification image generated in this way into multiple patches using a method such as a sliding window, for example. However, it is not limited to this, and various methods other than the method using a sliding window can be adopted for dividing the low-magnification image into multiple patches.

[0206] The low-magnification patching module (10) generates patch embeddings for each of the multiple divided patches, and these embeddings can subsequently be provided as inputs to the first visual feature extraction module (40) and the first frequency feature extraction module (30). Through this process, low-magnification rich contextual information can be transmitted to the next analysis step.

[0207] The high-magnification patching module (20) receives a high-resolution image (HRI) and can process it into a high-magnification image such as 20x or 10x. The high-magnification patching module (20) can efficiently obtain high-magnification image data necessary to identify fine morphological clues at the cellular level from among the vast amount of information in the high-resolution image (HRI).

[0208] The high-magnification patching module (20) can divide the high-magnification image generated in this way into multiple patches using a method such as a sliding window. However, it is not limited to this, and various methods other than the method using a sliding window can be adopted for dividing the high-magnification image into multiple patches.

[0209] The high-magnification patching module (20) generates patch embeddings for each of the divided multiple patches, and these embeddings can subsequently be provided as inputs to the second visual feature extraction module (41) and the second frequency feature extraction module (31). Through this process, high-magnification fine information can be transmitted to the next analysis step.

[0210] The first frequency feature extraction module (30) can receive a plurality of patches divided from the low-magnification patching module (10) and extract frequency feature embeddings. Through this, the first frequency feature extraction module (30) can capture fine and diagnostically important morphological clues of a high-resolution image.

[0211] The first frequency feature extraction module (30) first performs a Fourier transform on an input patch to obtain frequency domain data, and performs thresholding on the data to remove low-frequency components from the frequency domain data and separate only the high-frequency components (HF features) necessary for diagnosis.

[0212] The high-frequency components of the patch capture fine details and sharp transitions, such as cellular structures and tissue boundaries, which are crucial for pathological diagnosis. This contributes to enhancing the accuracy and reliability of the model by capturing diagnostic clues that are easily missed in low-magnification images.

[0213] Finally, the first frequency feature extraction module (30) performs an inverse Fourier transform on the separated high-frequency components, and the data generated from the inverse Fourier transform result is used with a pre-trained vision encoder (f v By inputting into ), the frequency feature embedding (z h Can generate ).

[0214] For example, referring to FIG. 9, the first frequency feature extraction module (30) receives a plurality of patches for a low-magnification image and, from this, a plurality of first frequency feature embeddings of high-frequency components corresponding to the plurality of patches ( Can generate ) (i is a positive integer from 1 to N). Multiple first frequency feature embeddings ( ) can be provided as input to the patch feature embedding generation module (60) in a subsequent step.

[0215] The process of receiving multiple patches of the first frequency feature extraction module (30) and extracting frequency feature embeddings can be performed based on the following equations (8) and (9).

[0216]

[0217] Equation (8): ,

[0218] (step, Patch embedding, : Fourier transform function, : Fourier transform patch embeddings, : High-frequency Fourier transform patch embedding, : Critical processing function, : Hyperparameter radius)

[0219] Equation (9):

[0220] (step, : Vision encoding function, : Inverse Fourier transform function, : High-frequency Fourier transform patch embedding, : High-frequency patch embedding)

[0221]

[0222] The first visual feature extraction module (40) can receive a plurality of patches divided by the low-magnification patching module (10) and extract original features. Original features refer to the visual features of the corresponding image, and the first visual feature extraction module (40) can capture overall visual information of the image and provide basic feature representations for subsequent analysis.

[0223] The first visual feature extraction module (40) can extract visual features using a pre-trained vision encoder (e.g., ResNet), and this vision encoder can compress the visual pattern of the input patch and convert it into a latent representation. In this process, the parameters of the vision encoder can be used in a frozen state to reduce the computational burden caused by the vast number of patches in the high-resolution image.

[0224] For example, referring to FIG. 9, the first visual feature extraction module (40) receives a plurality of patches for a low-magnification image and, from this, a plurality of first visual feature embeddings ( It can generate ) multiple first visual feature embeddings ( ) can be provided as input to the patch feature embedding generation module (60) in a subsequent step.

[0225] The second frequency feature extraction module (31) can receive multiple patches divided from the high-magnification patching module (20) and extract frequency feature embeddings.

[0226] The second frequency feature extraction module (31), like the first frequency feature extraction module (30), can separate only high-frequency components through Fourier transform and threshold processing on the patch, and generate a final frequency feature embedding through inverse Fourier transform and a vision encoder.

[0227] For example, referring to FIG. 9, the second frequency feature extraction module (31) receives a plurality of patches for a high-magnification image and, from this, a plurality of second frequency feature embeddings of high-frequency components corresponding to the plurality of patches ( Can generate ) (j starts from 1 It is a positive integer up to, and (is a positive integer). Multiple second frequency feature embeddings ( ) can be provided as input to the patch feature embedding generation module (60) in a subsequent step. The detailed description of the function of the second frequency feature extraction module (31) follows the description of the first frequency feature extraction module (30).

[0228] The second visual feature extraction module (41) can receive multiple patches divided by the high-magnification patching module (20) and extract original features. Original features refer to the visual features of the corresponding image, and the second visual feature extraction module (41) can capture overall visual information of the image and provide basic feature representations for subsequent analysis.

[0229] For example, referring to FIG. 9, the second visual feature extraction module (41) receives a plurality of patches for a high-magnification image and, from this, a plurality of second visual feature embeddings (Original features for the plurality of patches) corresponding to the original features of the plurality of patches It can generate ) multiple second visual feature embeddings ( ) can be provided as input to the patch feature embedding generation module (60) in a subsequent step. The detailed description of the function of the second visual feature extraction module (41) follows the description of the first visual feature extraction module (40).

[0230] Meanwhile, in FIG. 9, the first frequency feature extraction module (30) and the second frequency feature extraction module (31) are shown as independent modules, but are not limited thereto, and a single integrated frequency feature extraction module can be implemented to extract frequency feature embeddings for both the patch from the low-magnification patching module (10) and the patch from the high-magnification patching module (20).

[0231] This integrated frequency feature extraction module can recognize scale information (e.g., 10x, 20x) of input patches to extract optimal frequency features for each patch, and can improve system efficiency by increasing the resource utilization of the module.

[0232] The patch relationship information embedding generation module (50) can encode geometric positional relationships and morphological similarities between patches based on a plurality of patches divided in the low-magnification patching module (10).

[0233] For example, referring to FIG. 10, the patch relationship information embedding generation module (50) generates an adjacency matrix (A) based on distance information between patches based on a plurality of patches from a low-magnification image. dis Adjacency matrix based on the similarity of ) and frequency features (A HF Each of ) can be generated, and using them, multiple patch relationship information embeddings corresponding to learnable geometric position encoding (LGPE) ( Can generate ).

[0234] Here, a patch relationship information embedding generation module (50) generates a plurality of patch relationship information embeddings based on a low-magnification image ( The reason for generating ) is to prevent the computational complexity and memory usage from increasing significantly due to the vast amount of multiple patches when based on high-magnification images.

[0235] In detail, multiple patch relationship information embeddings ( As described below, the process of calculating ) involves matrix operations (e.g., eigenvalue decomposition for graph Laplacian) in which the amount of computation increases rapidly as the number of patches (N) increases. Therefore, in low-magnification images with a much smaller number of patches, multiple patch relationship information embeddings ( By calculating and reusing it by aligning it to high-magnification patches, efficiency can be maximized by eliminating the need to calculate position information and characteristic relationship position encoding for multiple patches of a high-magnification image.

[0236] The patch relationship information embedding generation module (50) learns the structural relationships between low-magnification image patches and generates a learnable geometric position encoding (LGPE) necessary to effectively evaluate the importance of the patches in a subsequent step and solve the multi-resolution dilemma.

[0237] The patch relationship information embedding generation module (50) has a plurality of patch relationship information embeddings ( Two types of adjacency matrices can be utilized to generate ).

[0238] First, a distance-based adjacency matrix (A dis ) models global connectivity between patches based on the Euclidean distance between patches.

[0239] For example, the patch relationship information embedding generation module (50) considers the center points of multiple patches as nodes and constructs a graph by calculating the distance between these nodes. Distance-based adjacency matrix (A dis ) represents the connectivity between these nodes, and It has a value of 1 when the j-th patch is among the closest patches corresponding to the top-k of the i-th patch, and the distance between them does not exceed a predetermined distance threshold. Through this process, information regarding how patches are physically distributed and how close they are to each other is obtained from the distance-based adjacency matrix (A dis It can be encoded in ).

[0240] In addition, a distance-based adjacency matrix (A dis ) is used to produce multiple eigenvectors (U) that encode the geometric positions of the patches by performing eigen-decomposition on the graph Laplacian (Δ). These eigenvectors (U) possess slide rotation invariance, which is an important characteristic for high-resolution image analysis.

[0241] For example, according to the following equation (3), a distance-based adjacency matrix (A dis ) and the degree matrix (D disBy utilizing ) to define the graph Laplacian (Δ) and performing eigen-decomposition on it, a plurality of eigenvectors (U) encoding the geometric positions of the patches can be produced. The plurality of eigenvectors (U) can be referred to as a plurality of position feature vectors.

[0242]

[0243] Equation (10):

[0244] (where Δ is the graph Laplacian, I n is the identity matrix, D dis is a degree matrix, A dis is the distance-based adjacency matrix, U is the eigenvector matrix, and Λ is the eigenvalue matrix)

[0245]

[0246] Here, the degree matrix (D dis ) is a matrix representing the degree of connectivity of each patch, and is an adjacency matrix (A dis It can be created by placing the sums of the row or column elements of a matrix along the diagonal. That is, a matrix of order (D dis ) is a symmetric matrix in which only the diagonal elements have values ​​and the rest are 0, where the diagonal element values ​​correspond to the number of adjacent patches within a threshold distance among the top-k adjacent patches for the i-th patch.

[0247] The distance-based adjacency matrix (A) as shown above dis The eigenvector matrix (U) obtained as a result of eigenvector decomposition for ) encodes geometric position information between patches, and these eigenvectors are multiple patch relationship information embeddings ( These eigenvectors are used as key inputs to generate a distance-based adjacency matrix (A) generated based on the Euclidean distance between patches. disIt is generated based on ) and has important properties such as slide rotation invariance of high-resolution images.

[0248] Second, a frequency feature-based adjacency matrix (A HF ) models local connectivity between patches based on similarity (e.g., cosine similarity) between multiple high-frequency features of multiple patches, going beyond Euclidean distance.

[0249] Through this, it is possible to capture similarities in fine morphological patterns important for diagnosis, such as fat cells and cytoplasm.

[0250] For example, the patch relationship information embedding generation module (50) calculates the similarity between multiple frequency feature embeddings for multiple patches of a low-magnification image, and based on the similarity information between multiple frequency embeddings and the distance information between multiple patches of the low-magnification image, a frequency feature-based adjacency matrix (A HF Can generate ).

[0251] This process goes beyond determining connectivity based solely on distance to model local connectivity by considering even the similarity of minute morphological patterns between patches.

[0252] Specifically, according to the following equation (11), a frequency feature-based adjacency matrix (A HF Connectivity between the i-th patch and the j-th patch among multiple patches of a low-magnification image corresponding to the element of ) ) has a value of 1 when both conditions are met.

[0253] First, the j-th patch included in multiple patches of the low-magnification image is one of the top k patches with the most similar frequency feature embeddings to the i-th patch, and second, the distance between the i-th patch and the j-th patch must not exceed a specific distance threshold (τ). Based on the connectivity information generated through this process, a frequency feature-based adjacency matrix (AHF ) can ultimately be constructed.

[0254]

[0255] Equation (11):

[0256] , if and top-k ( )

[0257] , otherwise

[0258]

[0259] The patch relationship information embedding generation module (50) generates a plurality of patch relationship information embeddings corresponding to LGPE ( To finally calculate ), a compressed position feature matrix ( ) and frequency feature-based adjacency matrix (A HF ) can be input into a graph neural network (GNN).

[0260] GNN combines these two pieces of information to provide multiple patch relationship information embeddings that reflect both the location and morphological similarity of the patches ( It generates ) and can be provided as input to a low-magnification patch feature embedding selection module (70) and a high-magnification patch feature embedding selection module (71) that perform LPGA (LGPE-aware Gated Attention) operations, which are patch selection mechanisms, in a subsequent step.

[0261] Specifically, the patch relationship information embedding generation module (50) uses a compressed position feature matrix ( ) based on a portion selected according to a predetermined condition among a plurality of eigenvectors (U) according to the following equation (12). Can generate ).

[0262]

[0263] Equation (12):

[0264] (where U: eigenvector matrix, : Number of selected eigenvectors, W: Learnable weight matrix)

[0265]

[0266] Subsequently, the GNN uses the input compressed positional feature matrix ( ) and frequency feature-based adjacency matrix (A HF For ), according to the following equation (13), a neural network operation is performed to learn a graph structure modeling the relationship between multiple patches, thereby embedding multiple patch relationship information ( Can generate ).

[0267]

[0268] Equation (13):

[0269]

[0270] In this case, is, for example, a compressed position feature matrix ( ) and frequency feature-based adjacency matrix (A HF By performing matrix multiplication on ) and matrix multiplying by a predetermined weight matrix (W) thereon, a plurality of patch relationship information embeddings ( It can be implemented using a Graph Convolutional Network (GCN) computation method that produces ). However, it is not limited to this, and In addition to the GCN operation method, it can be implemented in various operation methods, such as the GraphSAGE operation method that samples and aggregates features of neighbor nodes, the GAT (Graph Attention Networks) operation method that assigns different weights to neighbor nodes using an attention mechanism, or the GIN (Graph Isomorphism Network) operation method that preserves the unique characteristics of the graph structure by summing node features and neighbor features.

[0271] The low-magnification patch feature embedding generation module (60) generates a plurality of first visual feature embeddings ( ) and a plurality of first frequency feature embeddings generated in the first frequency feature extraction module (30) By integrating ) multiple first patch feature embeddings ( ) can be generated. Referring to FIG. 9, a plurality of first patch feature embeddings ( ) is the first patch feature integration matrix ( It can correspond to ).

[0272] The low-magnification patch feature embedding generation module (60) can combine two different perspectives (visual information and frequency information) of multiple patches of a low-magnification image, thereby enabling a high-resolution image analysis model trained based on this to understand the features of the patches more richly and comprehensively.

[0273] The low-magnification patch feature embedding generation module (60) comprises a plurality of first patch feature embeddings (Z c To generate ) a plurality of first visual feature embeddings ( ) and multiple first frequency feature embeddings ( A method of concatenating ) can be used. Multiple first patch feature embeddings generated through such integration ( ) includes all information regarding fine morphological clues (high-frequency components) important for diagnosis, along with the general visual pattern of the patch. Multiple first patch feature embeddings thus generated ( ) can be provided as an input to the low-magnification patch feature embedding selection module (70).

[0274] The high-magnification patch feature embedding generation module (61) generates a plurality of second visual feature embeddings (2 visual feature embeddings generated in the second visual feature extraction module (41) ) and a plurality of second frequency feature embeddings generated in the second frequency feature extraction module (31) By integrating ) multiple second patch feature embeddings ( ) can be generated. Referring to FIG. 9, a plurality of second patch feature embeddings ( ) is the second patch feature integration matrix ( It can correspond to ).

[0275] The high-magnification patch feature embedding generation module (61) combines two different perspectives (visual information and frequency information) of multiple patches of a high-magnification image, thereby enabling a high-resolution image analysis model trained based thereon to understand the features of the patches more richly and comprehensively. A more detailed description of the function of the high-magnification patch feature embedding generation module (61) follows the description of the function of the patch feature embedding generation module (60).

[0276] Meanwhile, in FIG. 8, a low-magnification patch feature embedding generation module (60) and a high-magnification patch feature embedding generation module (61) are shown as separate modules, but are not limited thereto and can be implemented so that a single module generates both magnification patch feature embeddings. In this case, the module can recognize the magnification information of the input patch and generate feature embeddings optimized for each magnification.

[0277] The low-magnification patch feature embedding selection module (70) is a core module that implements a hierarchical patch selection method in the process of processing training data for training a high-resolution image analysis model according to various embodiments of the present disclosure.

[0278] The low-magnification patch feature embedding selection module (70) is a plurality of first patch feature embeddings ( generated by the low-magnification patch feature embedding generation module (60) Regarding ), by utilizing the LPGA (LGPE-aware Gated Attention) mechanism, it is possible to efficiently select parts corresponding to patches that correspond to a top portion in terms of diagnostic importance. Here, LPGA is a patch integration embedding for each patch in a low-magnification image ( Multiple first patch feature embeddings () using ) as input It refers to an operation that calculates a predetermined weight to select some of the ).

[0279] First, patch integration embedding for each patch of the low-magnification image ( ) is the visual feature embedding of each patch of the low-magnification image( ), frequency feature embedding( Embedding patch relationship information in ) It can be generated by integrating up to ). In this case, visual feature embeddings for each patch of the low-magnification image ( ), frequency feature embedding( ), and patch relationship information embedding( Patch integration embedding for each patch of a low-magnification image () by concatenation ) can be generated.

[0280] Subsequently, the low-magnification patch feature embedding selection module (70) selects patch integration embeddings corresponding to each patch of the low-magnification image according to the following formula (14). Multiple attention weights (αi) can be calculated based on ). These weights indicate how significantly the corresponding patch contributes to the final prediction.

[0281]

[0282] Equation (14):

[0283]

[0284] (where V, U: learnable weight matrix, η(·): sigmoid function)

[0285]

[0286] Subsequently, the low-magnification patch feature embedding selection module (70) selects a plurality of first patch feature embeddings ( Among the calculated multiple weights, k embeddings corresponding to the top k weights (k is a positive integer) can be selected. The k first patch feature embeddings thus selected ( ) can be provided as input to the subsequent step prediction module (80). For example, k is a value that can be selected differently depending on the type of high-resolution image and can be set to various numbers such as 16, 20, 300, etc.

[0287] The high-magnification patch feature embedding selection module (71) is a plurality of second patch feature embeddings ( generated by the high-magnification patch feature embedding generation module (61) Regarding ), by utilizing the LPGA mechanism, it is possible to efficiently select a portion corresponding to patches that correspond to a top portion in terms of diagnostic importance. Here, LPGA is a patch integration embedding for each patch of a high-magnification image ( Multiple second patch feature embeddings () using ) as input It refers to an operation that calculates a predetermined weight to select some of the ).

[0288] First, patch integration embedding for each patch of the high-magnification image ( ) is the visual feature embedding of each patch of a high-magnification image( ), frequency feature embedding( Embedding patch relationship information in ) It can be generated by integrating up to ). In this case, visual feature embeddings for each patch of the high-magnification image ( ), frequency feature embedding( ), and patch relationship information embedding( By concatenating ) the integrated embedding of each patch of the high-magnification image ( ) can be generated.

[0289] In this way, each patch integration embedding of the high-magnification image ( When ) is generated, patch relationship information embeddings calculated based on low-magnification images ( ) can be utilized as is. Accordingly, for multiple patches of a vast amount of high-magnification images, patch relationship information embedding ( It is possible to apply consistent geometric information to heterogeneous patches while eliminating the need for computation to generate them. This can significantly improve the computational efficiency of high-resolution analysis models and contribute to enhancing prediction performance by maintaining the overall contextual information to which multiple patches belong.

[0290] Subsequently, the high-magnification patch feature embedding selection module (71) selects patch integration embeddings corresponding to each patch of the high-magnification image according to the above equation (7). Multiple attention weights (αij) can be calculated based on ). These weights indicate how significantly the corresponding patch contributes to the final prediction.

[0291] Subsequently, the high-magnification patch feature embedding selection module (71) selects a plurality of second patch feature embeddings ( Among the multiple calculated weights, the upper k×M 2 k×M corresponding to weights (M is the relative magnification of the high magnification to the low magnification). 2 k × M embeddings can be selected. The k×M selected in this way 2 Dog's second patch feature embedding ( ) can be provided as input to the subsequent step prediction module (80).

[0292] In this way, the high-magnification patch feature embedding selection module (71) selects a plurality of second patch feature embeddings ( Among ), k×M 2 When selecting k embeddings, the k first patch feature embeddings selected by the low-magnification patch feature embedding selection module (70) ( The area occupied by patches corresponding to ) on the low-magnification image and the k×M selected by the high-magnification patch feature embedding selection module (71) 2 Dog's second patch feature embedding ( The areas occupied by the patches corresponding to ) on the high-magnification image may match.

[0293] In other words, a high-resolution image analysis model trained on data generated through this process can effectively integrate information from both magnifications by analyzing areas deemed diagnostically important in low-magnification images in the same way in high-magnification images.

[0294] The prediction module (80) can be trained to output a prediction value at the level of the entire high-resolution image by combining features selected from the low-magnification image and the high-magnification image, respectively.

[0295] The prediction module (80) acts as a feature aggregator that integrates features extracted from multiple patches in a multi-instance learning framework based on a method of dividing a high-resolution image into multiple patches according to an embodiment of the present disclosure, and can be trained to provide a final diagnostic response for the entire high-resolution image.

[0296] First, the prediction module (80) selects k first patch feature embeddings ( selected by the low-magnification patch feature embedding selection module (70) A weighted sum for ) can be calculated. This weighted sum is the patch integration embedding ( corresponding to each patch of the low-magnification image It can be calculated based on k attention weights (αi) selected from among a plurality of attention weights (αi) calculated based on ). By applying the g(·) function (linear head) to this weighted sum, k first patch feature embeddings for low-magnification images ( Bag-level representation integrating ) It can generate.

[0297] Additionally, the prediction module (80) is a k×M selected by the high-magnification patch feature embedding selection module (71). 2 Dog's second patch feature embedding ( A weighted sum for ) can be calculated. This weighted sum is the patch integration embedding ( corresponding to each patch of the high-magnification image k×M selected from multiple attention weights (αij) calculated based on ) 2 It can be calculated based on attention weights (αij). By applying the g(·) function (linear head) to this weighted sum, k×M for high-magnification images 2 Dog's second patch feature embedding ( Bag-level representation integrating ) It can generate.

[0298] Finally, the prediction module (80) uses these two hundred-level representations ( , Performing a predetermined operation including the f(·) function on ) to obtain a final full-level prediction value of the high-resolution image ( Can output ).

[0299] This f(·) function can be designed to combine two back-level representations and subject them to a linear transformation for final classification. Such a prediction module (80) produces a final predicted value ( It can be trained to minimize the error between ) and the correct data.

[0300] Thus, the prediction module (80) includes selected k first patch feature embeddings ( Weighted sum for ) and selected k×M 2 Dog's second patch feature embedding ( The final predicted value (using the weighted sum for ) By being configured to output ), it is possible to comprehensively reflect both broad contextual information of low-magnification images based on patch importance and fine clue information of high-magnification images, thereby improving the predictive performance of multi-resolution image analysis and effectively mitigating the multi-resolution dilemma.

[0301] However, it is not limited to this, and the prediction module (80) includes selected k first patch feature embeddings ( ) and selected k×M 2 Dog's second patch feature embedding ( The final predicted value ( It may also be configured to output ).

[0302]

[0303] - Image analysis-based response provision method (S100)

[0304] A method for providing a response based on image analysis (S100) according to one embodiment can overcome the limitation of the absence of causal dependency modeling between multiple resolutions in a latent space when analyzing an image at multiple resolutions, and can improve the final prediction performance of the image analysis model and the accuracy of the response.

[0305] When integrating low-magnification information and high-magnification information obtained by analyzing an image, the method (S100) does not use a method of simply listing the two pieces of information to learn the association, but adopts a method of explicitly modeling the natural causal relationship in which a low-magnification image is created from a high-magnification image in the computer's feature space.

[0306] In this modeling process, high-magnification information can be adjusted to exert a decisive influence on low-magnification information, much like a magnifying glass. Furthermore, any unpredictable variability or uncertainty arising from this adjustment process can be captured in the form of noise and reflected in the final information. By utilizing enhanced low-magnification features that account for causal uncertainty between multiple resolutions, the image analysis model ultimately learns causal dependencies between resolutions beyond simple associations, thereby improving the performance and reliability of the model's responses to diagnostic images.

[0307] FIG. 11 is a flowchart of an image analysis-based response providing method (S100) according to one embodiment. FIG. 12 is a flowchart of a learning method for an image analysis model according to one embodiment.

[0308] Referring to FIG. 11, an image analysis-based response providing method (S100) according to one embodiment may include: receiving an activation signal for accessing at least one in-memory data (S101); accessing at least one in-memory data structure upon receiving the activation signal (S103); wherein the data structure includes a diagnostic image, and loading the diagnostic image from at least one memory (S105); wherein at least one processor generates at least one response to the diagnostic image using at least one artificial intelligence model that takes the diagnostic image as input (S107); wherein at least one artificial intelligence model is pre-trained to perform image analysis based on a feature representation of a causal dependency relationship between low-magnification features and high-magnification features of the diagnostic image, and inputting at least one response to at least one subsequent processing component (S109).

[0309] In one embodiment, the method (S100) may be performed by a processor (131) included in a server computing system (130). However, it is not limited thereto, and at least a part of the method (S100) may be performed by a processor (111) of a user computing device (110) or a processor (151) of a training computing system (150), and another part may be performed by a processor (131) included in a server computing system (130).

[0310] For convenience of explanation, the following description describes a processor (131) included in a server computing system (130) performing the method (S100).

[0311] In step (S101), the processor (131) may receive an activation signal for accessing at least one in-memory data. This activation signal may occur in various forms, for example, a user's request for image analysis through a user interface, an act of a user uploading a diagnostic image to the system, or an automatic analysis request from an external system. This step (S101) serves to capture an external trigger to start the image analysis method (S100).

[0312] In step (S103), the processor (131) may access at least one in-memory data structure upon receiving an activation signal. This data structure may contain a diagnostic image, and the diagnostic image may include a high-resolution image (HRI), such as a whole slide image of histopathology (WSI). The processor (131) may access this in-memory data structure to prepare to load the image in the next step.

[0313] In step (S105), the processor (131) can load a diagnostic image from at least one memory. Step (S105) means loading a diagnostic image contained in a previously accessed data structure from at least one memory so that the processor can actually process it.

[0314] In step (S107), the processor (131) can generate at least one response to a diagnostic image using at least one pre-trained artificial intelligence model.

[0315] Here, at least one artificial intelligence model can be pre-trained to perform image analysis based on a feature representation of the causal dependency relationship between high-magnification features and low-magnification features of an image.

[0316] For example, referring to FIG. 12, an image analysis model learning method according to one embodiment may include the steps of: loading a training image from at least one memory (S1071); generating a plurality of first patch feature embeddings corresponding to low-magnification features of the training image by the at least one processor (S1073); generating a plurality of second patch feature embeddings corresponding to high-magnification features of the training image by the at least one processor (S1075); generating a plurality of low-magnification dependent feature embeddings through a predetermined operation modeling the causal dependency relationship between the plurality of first patch feature embeddings and the plurality of second patch feature embeddings (S1077); and updating the parameters of the at least one artificial intelligence model so as to minimize the loss based on result data output by the at least one artificial intelligence model and correct answer data for the training image based on the predetermined operation results for the plurality of second patch feature embeddings and the plurality of low-magnification dependent feature embeddings (S1079).

[0317] In step (S1071), the processor (131) may load a training image from at least one memory. This training image may include a high-resolution image (HRI), such as a whole slide image of histopathology (WSI).

[0318] In step (S1073), the processor (131) comprises a plurality of first patch feature embeddings corresponding to low-magnification features of a training image ( Can generate ).

[0319] For example, the processor (131) can extract original features from multiple patches of low-magnification images of the training images. Original features refer to visual features of the corresponding images, and the processor (131) can capture overall visual information of the training images and provide basic feature representations for subsequent analysis.

[0320] The processor (131) can extract visual features using a pre-trained vision encoder (e.g., ResNet), and this vision encoder can compress the visual pattern of the input patch and convert it into a latent representation.

[0321] In step (S1075), the processor (131) comprises a plurality of second patch feature embeddings corresponding to high-magnification features of the training image ( Can generate ).

[0322] For example, the processor (131) can extract original features from multiple patches of high-magnification images of the training images.

[0323] In step (S1077), the processor (131) comprises a plurality of first patch feature embeddings ( ) and multiple second patch feature embeddings( Multiple low-magnification dependent feature embeddings ( Can generate ).

[0324] Through this process, the processor (131) can overcome the limitations of learning simple associations between multiple resolutions and explicitly model causal dependencies between multiple resolutions in latent space.

[0325] For example, the processor (131) has a plurality of first patch feature embeddings ( A query based on ) and multiple second patch feature embeddings ( By performing a cross-attention operation based on key-value pairs based on ), the causal influence of high-magnification features on low-magnification features can be captured. Through this, multiple low-magnification dependent feature embeddings (which are new feature representations adjusted by integrating high-magnification information) ) can be generated.

[0326] Additionally, the processor (131) takes unexplained factors of the latent space downsampling process as noise embeddings (U, The above causal modeling can be reinforced by introducing ).

[0327] This noise captures causal uncertainty between multiple resolutions, and the processor (131) has multiple second patch feature embeddings ( Weighted sum of patch-level noise embeddings (U) for ) and full-image-level noise embeddings for high-magnification images ( ) multiple low-magnification dependent feature embeddings( Integrated low-magnification dependent feature embeddings by summing to ) Can generate ).

[0328] These noise embeddings (U, ) can be sampled based on the learnable mean and standard deviation of a given probability distribution.

[0329] Thus, the processor (131) can finally adjust the low-magnification feature embeddings by reflecting causal uncertainty beyond simple association learning between multiple resolutions.

[0330] In step (S1079), the processor (131) is the second patch feature embedding ( ) and multiple low-magnification dependent feature embeddings( or At least one parameter of an artificial intelligence model can be updated so that the loss based on the result output by the artificial intelligence model and the correct answer data for the training image is minimized based on a predetermined operation result for ).

[0331] In step (S109), the processor (131) can input at least one response to the diagnostic image generated in step (S107) into at least one subsequent processing component.

[0332] Here, the follow-up processing component can be implemented in various forms, such as a user interface module that visually displays the response to the user, a database storage module that stores the response in a medical record system, or an additional prediction module that performs further analysis based on the response.

[0333] Additionally, at least one response input to the subsequent processing component may include at least one of a prediction result regarding whether the diagnostic image contains a specific disease, a prediction result regarding the grade or stage of progression of the disease, and a prediction result regarding the location of the lesion for the disease.

[0334] Various embodiments of the present disclosure described above may be implemented in the form of program instructions that can be executed through various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the computer-readable recording medium may be those specifically designed and configured for the various embodiments of the present disclosure, or may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. Hardware devices may be modified into one or more software modules to perform processing according to the various embodiments of the present disclosure, and vice versa.

[0335] The specific embodiments described in this disclosure are exemplary and do not limit the scope of the various embodiments of this disclosure in any way. For the sake of brevity of the specification, descriptions of conventional electronic configurations, control systems, software, and other functional aspects of said systems may be omitted. Additionally, the connections of lines or connecting members between components shown in the drawings are exemplary representations of functional connections and / or physical or circuit connections, and may be replaced or additionally represented as various functional connections, physical connections, or circuit connections in actual devices. Furthermore, unless specifically stated as “essential,” “importantly,” etc., a component may not be strictly necessary for the application of the various embodiments of this disclosure.

[0336] Furthermore, although the detailed description of the present disclosure has been described with reference to preferred embodiments of the present disclosure, those skilled in the art or those with ordinary knowledge in the art will understand that various modifications and changes can be made to the various embodiments of the present disclosure without departing from the spirit and technical scope of the various embodiments of the present disclosure as set forth in the claims below. Accordingly, the technical scope of the various embodiments of the present disclosure should not be limited to the contents described in the detailed description of the specification but should be determined by the claims.

[0337] Various embodiments according to the present disclosure have industrial applicability in that they can improve the predictive performance and reliability of high-resolution images by using an image analysis model that explicitly learns the causal relationships between multiple magnifications of high-resolution images.

Claims

1. As a method executed by a computer, A step of receiving an activation signal for accessing at least one in-memory data; A step of accessing at least one in-memory data structure upon receiving the activation signal; wherein the data structure includes a diagnostic image, A step of loading the diagnostic image from at least one memory; A step in which at least one processor generates at least one response to the diagnostic image using at least one artificial intelligence model that takes the diagnostic image as input; wherein the at least one artificial intelligence model is pre-trained to perform image analysis based on a feature representation of a causal dependency relationship between low-magnification features and high-magnification features of the diagnostic image, and A method comprising the step of inputting the above at least one response into at least one subsequent processing component.

2. In Paragraph 1, A method further comprising the step of the at least one subsequent processing component manifesting the at least one response through at least one user interface.

3. In Paragraph 1, A method further comprising the step of providing, through a user interface, an image that highlights a portion of the diagnostic image with high importance used in a calculation to determine at least one response, based on information regarding the causal dependency relationship between the low-magnification feature and the high-magnification feature.

4. In Paragraph 1, A method further comprising the step of providing, through a user interface, a heatmap or highlight visualization image that highlights areas corresponding to some high-magnification regions with relatively high scores in the diagnostic image, based on a score calculated in the above-mentioned causal dependency modeling process.

5. In Paragraph 1, The above diagnostic image includes a tissue pathology image, and A method wherein the above-mentioned at least one response comprises at least one of a prediction result regarding whether the diagnostic image includes a specific disease, a prediction result regarding the grade or stage of progression of the disease, and a prediction result regarding the location of a lesion for the disease.

6. In Paragraph 1, The above diagnostic image includes at least one of a crop image, an aerial satellite image, a manufactured product image, and an object image, and A method comprising at least one of the above-mentioned response, which includes a prediction result regarding the growth stage or presence or absence of disease of a crop included in the crop image, a prediction result regarding the progression stage of target identification or surface change in the aerial satellite image, a prediction result regarding defect classification or defect location identification in the manufactured product image, and a prediction result regarding object identification in the object capture image.

7. In Paragraph 1, The above-mentioned at least one artificial intelligence model is, A step of loading a training image from at least one memory; The step of the above at least one processor generating a plurality of first patch feature embeddings corresponding to the low-magnification features of the training image; The step of the above at least one processor generating a plurality of second patch feature embeddings corresponding to high-magnification features of the training image; A step of generating a plurality of low-magnification dependent feature embeddings through a predetermined operation that models the causal dependency relationship between the plurality of first patch feature embeddings and the plurality of second patch feature embeddings; and A method pre-trained by a learning method comprising: a step of updating the parameters of at least one artificial intelligence model so as to minimize the loss based on result data output by the at least one artificial intelligence model and correct answer data for the training image, based on a predetermined operation result for the plurality of second patch feature embeddings and the plurality of low-magnification dependent feature embeddings.

8. In Paragraph 7, The step of generating the above plurality of low-magnification dependent feature embeddings is, A method comprising: a step of generating a plurality of low-magnification dependent feature embeddings through a cross-attention operation based on the plurality of first patch feature embeddings and the plurality of second patch feature embeddings.

9. In Paragraph 8, A method in which the above cross-attention operation is performed based on a query based on the plurality of first patch feature embeddings and a key and value based on the plurality of second patch feature embeddings.

10. In Paragraph 8, The above learning method is, A method further comprising the step of adjusting the plurality of low-magnification dependent feature embeddings by summing the weighted sum of the plurality of high-magnification patch level noise embeddings corresponding to the plurality of patches of the high-magnification image of the training image and the total image level noise embedding for the high-magnification image of the training image for each of the plurality of low-magnification dependent feature embeddings.

11. In Paragraph 10, The above plurality of high-magnification patch-level noise embeddings (U j ) is a predetermined probability distribution (N(μ) according to the following (formula) j , σ j 2 The learnable mean (μ) of )) j ) and standard deviation (σ j A method sampled by ). ceremony: (step, and are random variables sampled from a standard normal distribution, respectively) 12. In Paragraph 10, A method in which the weighted sum of the plurality of high-magnification patch level noise embeddings is calculated based on weights generated during a cross-attention operation process based on the plurality of first patch feature embeddings and the plurality of second patch feature embeddings.

13. In Paragraph 7, The above learning method is, A step of generating a plurality of patch relationship information embeddings including relationship information between the plurality of patches from a plurality of patches of the above-mentioned training image; A step of selecting some of the plurality of first patch feature embeddings based on the plurality of patch relationship information embeddings; and The method further includes the step of selecting some of the plurality of second patch feature embeddings based on the plurality of patch relationship information embeddings. The step of generating the above plurality of low-magnification dependent feature embeddings is, A method comprising: a step of generating a plurality of low-magnification dependent feature embeddings through a predetermined operation that models the causal dependency relationship between some of the selected plurality of first patch feature embeddings and some of the selected plurality of second patch feature embeddings.

14. In Paragraph 13, The step of generating the above plurality of patch relationship information embeddings is, A distance-based adjacency matrix (A) based on distance information for multiple patches of a low-magnification image of the above-mentioned training image dis Step of generating ); Based on similarity information between frequency features for multiple patches of a low-magnification image of the above training image, a frequency feature-based adjacency matrix (A HF Step of generating ); and The above distance-based adjacency matrix (A dis ) and the above frequency feature-based adjacency matrix (A HF A method comprising the step of generating a plurality of patch relationship information embeddings including location information and characteristic relationship information between the plurality of patches using ).

15. In Paragraph 14, The step of generating the above plurality of patch relationship information embeddings is, A step of performing spectral analysis on the distance-based adjacency matrix to obtain a plurality of position feature vectors encoding geometric position information between a plurality of patches of the low-magnification image; and A method comprising the step of generating the plurality of patch relationship information embeddings by performing a neural network operation to learn a graph structure on a portion selected according to a predetermined condition among the plurality of acquired position feature vectors and the frequency feature-based adjacency matrix.

16. In Paragraph 15, The step of acquiring the above plurality of position feature vectors is, A method comprising the step of calculating a plurality of eigenvectors for the graph Laplacian of the distance-based adjacency matrix according to the following (equation) as the plurality of position feature vectors. ceremony: (where Δ is the graph Laplacian, I n is the identity matrix, D dis is a degree matrix, A dis is the distance-based adjacency matrix, U is the eigenvector matrix, and Λ is the eigenvalue matrix) 17. In Paragraph 13, The step of generating the above plurality of first patch feature embeddings is, The method includes the step of generating the plurality of first patch feature embeddings by integrating the plurality of first visual feature embeddings and the plurality of first frequency feature embeddings for the low-magnification image of the training image; The step of generating the above plurality of second patch feature embeddings is, A method comprising the step of generating a plurality of second patch feature embeddings by integrating a plurality of second visual feature embeddings and a plurality of second frequency feature embeddings for a high-magnification image of the above-mentioned training image.

18. In Paragraph 17, The step of selecting some of the plurality of first patch feature embeddings is, A step of calculating a plurality of weights for a plurality of patches of a low-magnification image based on a plurality of patch integration embeddings that integrate the plurality of first visual feature embeddings, the plurality of first frequency feature embeddings, and the plurality of patch relationship information embeddings; and A method comprising: a step of selecting k first patch feature embeddings corresponding to the upper k weights (k is a positive integer) among the plurality of first patch feature embeddings.

19. In Paragraph 18, The step of selecting some of the plurality of second patch feature embeddings mentioned above is, A step of calculating a plurality of weights for a plurality of patches of a high-magnification image based on a plurality of patch integration embeddings that integrate the plurality of second visual feature embeddings, the plurality of second frequency feature embeddings, and the plurality of patch relationship information embeddings; and Among the plurality of second patch feature embeddings above, the upper k×M of the plurality of weights 2 k×M corresponding to weights (M is the relative magnification of the above high magnification to the low magnification) 2 A method comprising the step of selecting two second patch feature embeddings.

20. At least one memory; and At least one processor that reads at least one instruction stored in the above at least one memory and executes an image analysis-based response providing method; comprising The above at least one instruction is, A step of receiving an activation signal for accessing at least one data in memory; A step of accessing at least one in-memory data structure upon receiving the activation signal; wherein the data structure includes a diagnostic image, A step of loading the diagnostic image from at least one memory; The step of generating at least one response to a diagnostic image using at least one artificial intelligence model that takes the diagnostic image as input, wherein the at least one artificial intelligence model is pre-trained to perform image analysis based on a feature representation of a causal dependency relationship between low-magnification features and high-magnification features of the diagnostic image, and A system comprising a command to perform the step of ingesting the above at least one response to at least one subsequent processing component.

21. In Paragraph 20, A plurality of neurons comprising an array including at least one register, at least one programmable logic, and at least one input interface; a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons; and at least one routing network that controls the data flow between the plurality of neurons; comprising A system further comprising a Field Programmable Gate Array (FPGA) implementation for a predetermined artificial neural network, wherein each of the plurality of neurons is connected to at least one other neuron through the routing network to establish a transmission path for the weights.

22. In Article 20, A plurality of neurons organized into an array comprising at least one register, at least one microprocessor, and at least one input; and a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons; comprising A system comprising an Application Specific Integrated Circuit (ASIC) for a predetermined artificial neural network, wherein each of the plurality of neurons is connected to at least one other neuron through any one of the plurality of synaptic circuits.