Method, computer program, computer system (generating picture description task images using ai)

The AI-based image generation method addresses the lack of detail in conventional psychiatric and neurological images by simulating healthy and medically conditioned groups and comparing textual descriptions, enhancing diagnostic accuracy.

JP2025106193APending Publication Date: 2025-07-15INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024209836
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-03
Filing Date
2024-12-02
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Conventional image generation methods for psychiatric and neurological diagnosis lack sufficient detail and complexity, limiting the accuracy of downstream analysis and diagnosis.

Method used

A method using AI to generate images based on detailed text prompts, simulate healthy and medically conditioned groups, and compare textual descriptions to ensure rich information and quality control, enhancing diagnostic accuracy.

Benefits of technology

The method creates high-quality, detail-oriented images that improve the accuracy and reliability of psychiatric and neurological diagnosis by ensuring rich information and efficient quality control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025106193000001_ABST
    Figure 2025106193000001_ABST
Patent Text Reader

Abstract

To provide a method of generating picture description task images using AI.SOLUTION: A method includes using AI to generate a story on the basis of a user input prompt and / or a predetermined input. AI is used to generate an image on the basis of the generated story. AI is used to generate written descriptions of the generated image by simulating cohorts of healthy individuals and individuals with a predetermined condition. Diagnostic linguistic features are extracted from written descriptions of the cohorts of healthy individuals and individuals with the predetermined condition. The extracted diagnostic linguistic features of the written descriptions for the cohorts of the healthy individuals and individuals with the predetermined condition are compared. The generated image is used in a picture description task when the compared extracted features of the written descriptions of the cohorts of healthy individuals and individuals with the predetermined condition exhibit a predetermined threshold of difference.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Exemplary embodiments of the concepts of the present invention relate to the generation of image caption task images, and more particularly, to the generation of image caption task images using AI (Artificial Intelligence).

Summary of the Invention

Problems to be Solved by the Invention

[0002] The quality and richness of the information contained in the images used for image caption tasks directly affect the accuracy and effectiveness of diagnosis (e.g., neurological / psychiatric). Conventional methods of image generation often create images lacking in sufficient detail and complexity, thereby limiting the ability to provide rich and detailed descriptions of the subject. And this limitation affects the criteria for downstream psychiatric / neurological analysis and may compromise the accuracy of diagnosis.

Means for Solving the Problems

[0003] Exemplary embodiments of the concepts of the present invention relate to methods, computer program products, and systems for the generation of image caption task images using AI.

[0004] According to an exemplary embodiment of the concept of the present invention, a method for generating an image for an image description task using AI is provided. The method comprises the step of generating a story based on at least one of a user input prompt and a predetermined input using AI. The AI is used to generate an image based on the generated story. A group of healthy individuals and a group of individuals having a predetermined medical condition are simulated, and the AI is used to generate a textual description of the generated image. Diagnostic verbal features are extracted from the textual descriptions of the group of healthy individuals and the group of individuals having a predetermined medical condition. The extracted diagnostic verbal features of the textual descriptions for the group of healthy individuals and the group of individuals having a predetermined medical condition are compared. The generated image is used in an image description task when the comparison of the extracted features of the textual descriptions of the group of healthy individuals and the group of individuals having a predetermined medical condition shows a difference of at least a predetermined threshold.

[0005] According to an exemplary embodiment of the concept of the present invention, a computer program product (CPP) for generating an image for an image description task using AI is provided. The CPP includes one or more computer-readable storage media and program instructions stored on one or more non-transitory computer-readable storage media that are executable to perform the method. The method comprises the step of generating a story based on at least one of a user input prompt and a predetermined input using AI. The AI is used to generate an image based on the generated story. A group of healthy individuals and a group of individuals having a predetermined medical condition are simulated, and the AI is used to generate a textual description of the generated image. Diagnostic verbal features are extracted from the textual descriptions of the group of healthy individuals and the group of individuals having a predetermined medical condition. The extracted diagnostic verbal features of the textual descriptions for the group of healthy individuals and the group of individuals having a predetermined medical condition are compared. The generated image is used in an image description task when the comparison of the extracted features of the textual descriptions of the group of healthy individuals and the group of individuals having a predetermined medical condition shows a difference of at least a predetermined threshold.

[0006] According to an exemplary embodiment of the concept of the present invention, a computer system (CS) is provided for generating images for image captioning tasks using AI. The CS includes one or more computer processors, one or more computer-readable storage media, and program instructions stored in one or more of the computer-readable storage media for execution by at least one of the one or more processors capable of executing the method. The method includes generating a story based on at least one of a user input prompt and a predetermined input using AI. The AI is used to generate an image based on the generated story. AI is used to simulate a group of healthy individuals and a group of individuals with a predetermined medical condition to generate a textual description of the generated image. Diagnostic verbal features are extracted from the textual descriptions of the group of healthy individuals and the group of individuals with a predetermined medical condition. The extracted diagnostic verbal features of the textual descriptions for the group of healthy individuals and the group of individuals with a predetermined medical condition are compared. The generated image is used in an image captioning task when the comparison of the extracted features of the textual descriptions of the group of healthy individuals and the group of individuals with a predetermined medical condition shows a difference of at least a predetermined threshold.

Brief Description of the Drawings

[0007] The following detailed description is provided by way of example and is not intended to limit the exemplary embodiments only, and is best understood in conjunction with the accompanying drawings.

[0008]

Figure 1

[0009]

Figure 2

[0010]

Figure 3

[0011] It is understood that the drawings included are not necessarily drawn to scale / ratio. The included drawings are merely schematic examples to assist in the understanding of the concept of the present invention and are not intended to depict fixed parameters. In the drawings, like reference numerals may represent like elements.

DETAILED DESCRIPTION

[0012] Exemplary embodiments of the concept of the present invention are disclosed hereinafter. However, it is understood that the scope of the concept of the present invention is defined by the claims. The disclosed exemplary embodiments are merely examples of the claimed system, method, and computer program product. The concept of the present invention may be embodied in many different forms and should not be construed as limited to only the exemplary embodiments described herein. Rather, these exemplary embodiments included are provided to make the present disclosure complete and to facilitate understanding by those skilled in the art. In the detailed description, consideration of well-known features and techniques may be omitted to avoid unnecessarily obscuring the exemplary embodiments presented.

[0013] References in the specification to "one embodiment", "an embodiment", "an exemplary embodiment", etc., indicate that the embodiment described may include a particular feature, structure, or characteristic, but not necessarily all embodiments include that feature, structure, or characteristic. Further, such phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is submitted that it is within the knowledge of those skilled in the art to implement such feature, structure, or characteristic in connection with other embodiments, whether or not explicitly described.

[0014] To avoid obscuring the presentation of exemplary embodiments of the concepts of the present invention, in the following detailed description, some processing steps or operations known in the art may be combined for purposes of presentation and illustration and, in some instances, may not be described in detail. Further, some processing steps or operations known in the art may not be described at all. In the following detailed description, attention is directed to the distinctive features or elements of the concepts of the present invention according to various exemplary embodiments.

[0015] The proposed invention concept addresses the above-mentioned problems related to image captioning task images by systematically generating high-quality and complex images using generative AI. These images can be created based on detailed text prompts obtained from the analysis target, ensuring that each image contains rich information for the description. Further, by incorporating AI in the quality control and analysis stages, the pipeline ensures that the created images meet the required quality standards and aids in the comparison of human and AI image captions.

[0016] Through this system and method, the invention addresses the fundamental limitations of conventional image generation for psychiatric diagnosis. Thereby, it provides a robust, reliable, and efficient solution for creating rich and detail-oriented images that enhance the ability to provide detailed textual descriptions of the subject, thereby improving the accuracy of psychiatric / neurological diagnosis. The present invention is a novel system and method that employs generative artificial intelligence (AI) to create rich and detailed images for psychiatric / neurological image description tasks. It can convert the analysis target into a text prompt, which then guides the AI in generating the corresponding image. This image is then explained by the AI, and the resulting explanation is compared to the original text prompt for quality control. Finally, the explanations provided by both the human subject and the AI are compared to the original text prompt, thereby assisting in neurological / psychiatric diagnosis / analysis. This innovative pipeline design revolutionizes the way image description tasks are performed in neurological / psychiatric diagnosis. Thereby, it not only ensures the creation of high-quality images containing rich information for interpretation but also introduces an efficient AI-assisted method for quality control and analysis. The concept of the present invention significantly enhances the accuracy and reliability of psychiatric / neurological diagnosis, which is essential for informing treatment strategies and monitoring patient progress.

[0017] Figure 1 shows a schematic diagram of a computing environment 100 that includes the generation 150 of an image description task image using an AI program, according to an exemplary embodiment of the concept of the present invention.

[0018] Various aspects of the present disclosure are described by narrative text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in embodiments of a computer program product (CPP). For any flowchart, depending on the technology involved, operations may be performed in an order different from that shown in a given flowchart. For example, again depending on the technology involved, two operations shown in blocks of consecutive flowcharts may be performed in reverse order, as a single integrated step, simultaneously, or at least partially overlapping in time.

[0019] An embodiment of a computer program product (referred to herein as a "CPP embodiment" or "CPP") is a term used in the present disclosure and describes any set of one or more storage media (also referred to as "multiple media") collectively included in a set of one or more storage devices that collectively contain machine-readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP claim. A "storage device" is any tangible device capable of holding and storing instructions for use by a computer processor. By way of non-limiting example, a computer-readable storage medium may be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media are floppy disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded devices (such as punch cards or pits / lands formed on the major surfaces of discs), or any suitable combination of the foregoing. A computer-readable storage medium is not to be construed as storage in the form of a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, optical pulses passing through an optical fiber cable, electrical signals communicated through a wire, and / or other transmission media, as the term is used in the present disclosure. As will be understood by those skilled in the art, data is typically moved during normal operation of a storage device at some irregular points in time, such as during access, defragmentation, or garbage collection, but the data is not transient while it is stored, and thus the storage device is not considered to be transient for the purposes of the foregoing.

[0020] Computing environment 100 includes an example of an environment for executing at least some of the computer code involved in the execution of inventive methods such as the generation of image description task images 150 using an AI program. In addition to block 150, computing environment 100 includes, for example, computer 101, wide area network (WAN) 102, end user device (EUD) 103, remote server 104, public cloud 105, and private cloud 106. In this embodiment, computer 101 includes a processor set 110 (including processing circuit 120 and cache 121), communication fabric 111, volatile memory 112, persistent storage 113 (including operating system 122 and block 150 as shown above), a set of peripheral devices 114 (including a user interface (UI) device set 123, storage 124, and an Internet of Things (IoT) sensor set 125), and network module 115. Remote server 104 includes remote database 130. Public cloud 105 includes gateway 140, cloud orchestration module 141, a set of host physical machines 142, a set of virtual machines 143, and a set of containers 144.

[0021] Computer 101 can take the form of a desktop computer, laptop computer, tablet computer, smartphone, smartwatch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device that is currently known or will be developed in the future, capable of executing programs, accessing a network, or querying a database such as remote database 130. As is well understood in the field of computer technology, and depending on the technology, the execution of computer implementation methods can be distributed among multiple computers and / or between multiple locations. On the other hand, in this description of computing environment 100, for the sake of keeping the description as concise as possible, the detailed discussion focuses on a single computer, specifically computer 101. Although computer 101 is not shown within the cloud in FIG. 1, it can be located within the cloud. On the other hand, computer 101 does not need to exist within the cloud, except within any arbitrarily shown range.

[0022] Processor set 110 includes one or more computer processors of any type that are currently known or will be developed in the future. Processing circuit 120 can be distributed among multiple packages, for example, multiple integrated circuit chips that have been adjusted. Processing circuit 120 can implement multiple processor threads and / or multiple processor cores. Cache 121 is a memory located within the processor chip package and is typically used for high-speed access to data or code that should be available to the threads or cores executing on processor set 110. Cache memory is typically organized into multiple levels depending on its relative proximity to the processing circuit. Alternatively, some or all of the cache for the processor set can be located "off-chip". In some computing environments, processor set 110 can be designed to operate using qubits and execute quantum computing.

[0023] Computer-readable program instructions are typically loaded into computer 101 and executed by a set of processors 110 of computer 101 to perform a series of operational steps, thereby enabling a computer-implemented method. As a result, the instructions so executed instantiate the method (collectively referred to as "the method of the present invention") specified in the flowchart and / or narrative description of the computer-implemented method included herein. These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by the set of processors 110 to control and direct the execution of the method of the present invention. In computing environment 100, at least some of the instructions for performing the method of the present invention may be stored in block 150 within persistent storage 113.

[0024] Communication fabric 111 is a signal conduction path that enables various components of computer 101 to communicate with each other. Typically, this fabric is made up of switches and conductive paths, such as switches that make up a bus, bridge, physical input / output ports, etc., and conductive paths. Other types of signal communication paths, such as optical fiber communication paths and / or wireless communication paths, may be used.

[0025] Volatile memory 112 is any type of volatile memory that is currently known or will be developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory 112 is characterized by random access, although this is not required unless expressly indicated. In computer 101, volatile memory 112 is located within a single package and exists within computer 101. Alternatively or additionally, volatile memory may be distributed across multiple packages and / or located external to computer 101.

[0026] The persistent storage 113 is any form of non-volatile storage for a computer that is currently known or will be developed in the future. The non-volatility of this storage means that the stored data is maintained regardless of whether power is directly supplied to the computer 101 and / or to the persistent storage 113. The persistent storage 113 can be read-only memory (ROM), but usually at least a portion of the persistent storage enables writing, deleting, and rewriting of data. Some well-known forms of persistent storage include magnetic disks and solid-state storage devices. The operating system 122 can take several forms, such as various known proprietary operating systems that employ a kernel or open-source portable operating system interface type operating systems. The code included in block 150 typically includes at least some of the computer code involved in the execution of the method of the present invention.

[0027] The peripheral device set 114 includes a set of peripheral devices of the computer 101. Data communication connections between the peripheral devices of the computer 101 and other components may be implemented in various ways, such as a Bluetooth (registered trademark) connection, a Near Field Communication (NFC) connection, a connection by a cable (such as a Universal Serial Bus (USB) type cable), an insertion type connection (such as a Secure Digital (SD) card), a connection through a local area communication network, and even a connection through a wide area network such as the Internet. In various embodiments, the UI device set 123 may include components such as a display screen, a speaker, a microphone, wearable devices (such as goggles and smartwatches), a keyboard, a mouse, a printer, a touchpad, a game controller, and a haptic device. The storage 124 is an external storage such as an external hard drive or an insertable storage such as an SD card. The storage 124 may be persistent and / or volatile. In some embodiments, the storage 124 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where the computer 101 is required to have a large amount of storage (for example, when the computer 101 locally stores and manages a large-scale database), this storage may be provided by a peripheral storage device designed to store a very large amount of data, such as a Storage Area Network (SAN) shared by a plurality of geographically distributed computers. The IoT sensor set 125 is composed of sensors that can be used in Internet of Things applications. For example, one sensor may be a thermometer and another sensor may be a motion detector.

[0028] The network module 115 is an aggregate of computer software, hardware, and firmware that enables the computer 101 to communicate with other computers through the WAN 102. The network module 115 may include hardware such as a modem or a Wi-Fi (registered trademark) signal transceiver, software for packetizing and / or depacketizing data for communication over a communication network, and / or web browser software for communicating data over the Internet. In some embodiments, the network control function and the network transfer function of the network module 115 are executed on the same physical hardware device. In other embodiments (e.g., embodiments that utilize Software-Defined Networking (SDN)), the control function and the transfer function of the network module 115 are executed on physically separate devices such that the control function manages several different network hardware devices. The computer-readable program instructions for executing the method of the present invention can usually be downloaded to the computer 101 from an external computer or an external storage device through a network adapter card or a network interface included in the network module 115.

[0029] The WAN 102 is any wide area network (e.g., the Internet) that can communicate computer data over a non-local distance by any technology for communicating computer data that is currently known or will be developed in the future. In some embodiments, the WAN 102 may be replaced and / or supplemented by a local area network (LAN) designed to transfer data between devices located in a local area, such as a Wi-Fi network. The WAN and / or LAN typically includes computer hardware such as copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers.

[0030] An end user device (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of an enterprise operating computer 101), and can take any of the forms discussed above in relation to computer 101. The EUD 103 typically receives beneficial and useful data from the operation of computer 101. For example, in a virtual case where computer 101 is designed to provide recommendations to an end user, this recommendation will typically be communicated from the network module 115 of computer 101, via the WAN 102, to the EUD 103. In this way, the EUD 103 can display or otherwise present the recommendation to the end user. In some embodiments, the EUD 103 can be a client device such as a thin client, a thick client, a mainframe computer, and a desktop computer, etc.

[0031] A remote server 104 is any computer system that provides at least some data and / or functions to computer 101. The remote server 104 can be controlled and used by the same entity that operates computer 101. The remote server 104 represents a machine that collects and stores beneficial and useful data for use by other computers such as computer 101. For example, in a virtual case where computer 101 is designed and programmed to provide recommendations based on historical data, this historical data can be provided from the remote database 130 of the remote server 104 to computer 101.

[0032] The public cloud 105 is any computer system that provides on-demand availability of computer system resources and / or other computer functions, especially data storage (cloud storage) and computing capabilities, for use by multiple entities without direct active management by the user. Cloud computing typically exploits resource sharing to achieve coherence and economies of scale. The direct active management of the computing resources of the public cloud 105 is performed by the computer hardware and / or software of the cloud orchestration module 141. The computing resources provided by the public cloud 105 are typically implemented by virtual computing environments that run on various computers that make up the host physical machine set 142, which is the universe of physical computers within and / or available in the public cloud 105. The virtual computing environment (VCE) typically takes the form of virtual machines from the virtual machine set 143 and / or containers from the container set 144. It is understood that these VCEs can be stored as images and transferred either as images or after instantiation of the VCE among and within various physical machine hosts. The cloud orchestration module 141 manages the transfer and storage of images, deploys new instantiations of the VCE, and manages the active instantiation of the VCE deployment. The gateway 140 is an aggregate of computer software, hardware, and firmware that enables the public cloud 105 to communicate via the WAN 102.

[0033] Some further explanations of the virtualized computing environment (VCE) are provided here. The VCE can be stored as an "image". A new active instance of the VCE can be instantiated from the image. Two well-known types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to a feature of the operating system where the kernel enables the existence of multiple isolated instances of user space called containers. These isolated instances of user space typically behave as actual computers from the perspective of the programs running within them. A computer program running on a normal operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, a program running inside a container can only use the contents of the container and the devices allocated to the container, and this feature is known as containerization.

[0034] The private cloud 106 is similar to the public cloud 105, except that computing resources are available only for use by a single enterprise. The private cloud 106 is shown as being in communication with the WAN 102, but in other embodiments, the private cloud may be completely disconnected from the Internet and accessible only via a local / private network. A hybrid cloud is a composite of multiple different types of clouds (e.g., private cloud, community cloud, or public cloud types), and is often implemented by different vendors. Each of the multiple clouds remains a separate discrete entity, but the larger hybrid cloud architecture is coupled by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability among the multiple constituent clouds. In this embodiment, both the public cloud 105 and the private cloud 106 are part of a larger hybrid cloud.

[0035] FIG. 2 shows a block diagram of components included in the generation 150 of an image captioning task image using an AI program, according to an exemplary embodiment of the concepts of the present invention.

[0036] The generation 150 of image description task images using an AI program may include a generation component 210. The generation component 210 may include a user interactive interface and / or may be connected to various types of AIs (e.g., chatbots, natural language processing (NLP), text-to-speech conversion, classification models, large language models (e.g., LLM), computer vision, etc.). The generation component 210 may autonomously (e.g., via a related repository, approved and / or anonymized medical provider / entity / patient medical records, internet search, etc.) obtain image description task multimedia (e.g., audio, video, text, animation, etc.) based on a predetermined medical condition (e.g., neurological / psychiatric) and / or an image description task, including user input, analyzed user input prompts, and / or via a network. The image description task multimedia may include an image description task explanation, an image description task story, an image description task image, image description task responses from groups of healthy individuals and individuals with a predetermined medical condition (e.g., computer vision behavior analysis, text-to-speech conversion text explanations, and / or manual text explanations), image description task results / discussions in medical papers, etc.The generation component 210 can extract features such as image description task characteristics (e.g., type, morbidity rate, accuracy / validity / reliability / soundness / sensitivity, experimental design, description, disclaimer, symptoms, protocol, parameters, criteria, thresholds, etc.), predetermined disease state / differential diagnosis characteristics of the target (e.g., name, symptoms, parameters, criteria, risk factors, statistical incidence of occurrence / misdiagnosis, complications, heredity, demographics, progression, prognosis, etc.), image description task image / story characteristics (language, tone, style, complexity, syntax, effect / accuracy, morbidity rate, narrative, complexity, names of characters, number of characters, emotional states of characters, ages of characters, appearances of characters, relationships between characters, roles of characters, titles of characters, actions of characters, reactions of characters, interactions between characters, backgrounds / experiences of characters, descriptions of scenes, locations of scenes, days of scenes, dates of scenes, times of scenes, seasons of scenes, objects in scenes, events in scenes, backgrounds of scenes, etc.), text description characteristics (e.g., language, tone, style, syntax, complexity, coherence, accuracy, length, time, vocabulary, emphasis, ellipsis, feature description, etc.), and / or user characteristics (e.g., age, demographics, family history, background / experience, diagnosis, non-diagnosis, differential diagnosis, complaints, progression, presentation, medical history, frequency / type of medical appointment, etc.) from the image description task multimedia.

[0037] The generation component 210 can be trained to perform image description related classification tasks (e.g., generation of a predetermined disease state / diagnosis / non-diagnosis / differential diagnosis of a new text description, simulated text description / story / prompt / image, etc.) using the analyzed image description task multimedia related to at least one predetermined disease state through supervised and / or unsupervised learning.

[0038] The generation component 210 can generate at least one image description task story based on at least one user input prompt, a predetermined input (such as an image caption task, accuracy / validity / reliability / soundness / sensitivity, conditions and / or symptoms, keywords, summary, extracted features, etc.) and / or a network search (such as a supplementary study). The generation component 210 can generate at least one image description task image and / or a text description prompt based on the generated image description task story. In the case of multiple potential generated images, stories, and / or text description prompts, the generation component 210 can present the user with explanations and / or options (such as different narratives, complexities, advantages / disadvantages of relative diagnoses, etc.) for selection. The generation component 210 can generate simulated text descriptions for groups of healthy individuals and groups of unhealthy individuals and extract features therefrom. Based on the extracted features and / or the predetermined input, the generation component 210 can generate stories, images, text description prompts, and / or simulated text descriptions that are predicted to elicit / show distinguishable text description characteristics of the story between the group of healthy individuals and the group of individuals with a given medical condition.

[0039] For example:

[0040] Medical provider user prompt to the generation component 210:

[0041] Please assist in brainstorming descriptions of multiple scenarios to create images for cognitive evaluation. The scenario should involve multiple characters (2 - 5 people) who may be involved in interactions with the environment. For example, one character may be distracted, spill water, fall, or lead to interesting or negative outcomes such as surprising events. The scenario should include multiple objects in the environment that the characters can manipulate or interact with. The description should provide stimuli that can be interpreted and described. Locations such as a beach, a pool, or a main road also play important roles.

[0042] Generate response of component 210:

[0043] Of course. Descriptions of multiple scenarios for creating images for cognitive function evaluation with multiple characters and an interactive environment are as follows.

[0044] Location: A crowded supermarket. Description: A group of friends are shopping for groceries together. One character is distracted by a free sample display and accidentally collides with and knocks down a pyramid of canned goods. As the cans roll around the aisle, the other characters react with surprise and amusement.

[0045] Location: A crowded amusement park. Description: A family of five is riding a roller coaster together. When the coaster reaches the top, one character's hat is blown off and lands on the head of a nearby statue, startling everyone. As the character tries to retrieve the hat, the family members point and laugh.

[0046] Location: A neighborhood park with a pond. Description: A group of friends are gathered around a picnic table near the pond, engrossed in a card game. One character, who is focused on the game, accidentally knocks over a juice cup and spills it on the cards. The other characters react with surprise and annoyance and try to salvage the game.

[0047] Location: A scenic hiking trail. Description: A group of friends are exploring a beautiful trail. One character is distracted by taking a photo and accidentally steps in a muddy area, falls, and lands on their bottom. The other characters rush to help, laughing at the unexpected mishap.

[0048] Location: Beach volleyball court. Description: Two teams are engaged in an intense beach volleyball match. One of the characters is distracted by seagulls flying overhead, loses concentration, and misses an important shot. When the ball lands far from the intended target, the other characters react in surprise.

[0049] Location: A bustling square. Description: A group of friends is enjoying the atmosphere while sitting on a bench. Due to the hot weather, an ice cream cone of one of the characters starts to melt rapidly and drips on their hand. They react in surprise and quickly try to lick the melting ice cream.

[0050] Location: Bowling alley. Description: A group of friends is friendly competing in bowling. One of the characters is distracted by a flashy advertisement on the electronic score display, forgets to release the bowling ball, staggers forward, and barely avoids falling. The other characters react with surprise and laughter.

[0051] Remember that you can modify and adapt these descriptions to suit your specific needs, incorporating different elements from various scenarios to freely create unique scenarios.

[0052] Medical provider user input prompt:

[0053] Each of these situations is characterized by multiple individuals and objects, incorporates a wide range of elements such as distraction, surprise, and interaction, and all occur in specific locations.

[0054] Characteristics of Alzheimer's disease (AD) and schizophrenia:

[0055] In the case of AD, aspects such as verbosity, richness of vocabulary in terms of total word count, use of unique words, number of utterances, and even the number of characters in words that may indicate cognitive aging are evaluated. General characteristics of dementia such as repetitiveness are also considered. The Information Content Unit (ICU) is evaluated to quantify the number of objects or actions in an image. Since individuals with AD tend to use certain words more frequently, word frequency is taken into account. The characteristic difficulty in forming sentences with a large and complex syntax is also evaluated.

[0056] Regarding schizophrenia, aspects such as reduction in syntactic complexity and semantic impairments, which are evident in organizing a single proposition into a larger structure, are evaluated. Changes in parts of speech are also evaluated, such as an increase in the use of pronouns but a decrease in the use of adverbs, adjectives, and determiners (e.g., "the", "a").

[0057] In the case of schizophrenia:

[0058] In schizophrenia, a reduction in syntactic complexity is seen. Also, pay attention to semantic problems such as difficulty in organizing individual propositions into larger units of meaning. Notice the changes in parts of speech such as an increase in the use of pronouns but a decrease in adverbs, adjectives, and determiners (e.g., "the", "a").

[0059] In individuals with schizophrenia, elements such as a decrease in the complexity of sentence structure are evaluated. Also, observe any breakdown in semantics, such as problems in gathering individual statements into a broader structure. Notice the changes in parts of speech such as an increase in the use of pronouns but a decrease in adverbs, adjectives, and determiners ("the", "a", etc.).

[0060] For individuals with schizophrenia, pay attention to the simplified sentence structure. Also consider the semantic disorders seen in the difficulty of forming a broader structure from individual thoughts. Note the changes in language use, such as a preference for pronouns, but a lack of adverbs, adjectives, and determiners (e.g., "the", "a").

[0061] Healthcare provider user input prompt:

[0062] Please simulate how an AD patient would describe the first scenario.

[0063] Response of the generation component 210:

[0064] Certainly. The simulation of how an AD patient might describe the first scenario, taking into account AD-related characteristics, is as follows.

[0065] "The coffee shop, well, the coffee shop was crowded. There were people everywhere, talking, talking a lot. I was sitting at a table, well, facing the laptop. I was working, typing, typing a lot of words. And then, oh, I reached for my coffee cup but knocked it over and spilled it. It splashed all over the laptop and the table and everywhere. It was hot. It was hot. There were a lot of words. There were a lot of words but now, oh, the laptop won't work. People saw. They saw me. I was embarrassed. I was embarrassed. There were a lot of words but now it's ruined."

[0066] In this simulation, we can observe the characteristics of loquaciousness, repetition, and difficulty with richness of vocabulary and complexity of syntax. This narrative includes a focus on a crowded environment, the actions of the characters, and a negative outcome with a lack of coherence and repetition of certain phrases.

[0067] Healthcare provider user input prompt:

[0068] Please help me come up with some scene descriptions for creating images for a cognitive test. The scene should have multiple characters (2 - 5 people) who can interact with each other. Usually, one of the characters is distracted, and something interesting / bad happens (water spills, someone falls), or there is a surprising element. There should be multiple objects that the characters can interact with. The scene should provide stimuli that can explain a medical history. The location is also important (e.g., beach, pool, main road, etc.).

[0069] Response of the generation component 210:

[0070] Sure, I'd be happy to help. Some scene descriptions that might be useful for your cognitive function test are as follows.

[0071] Picnic in the park: Four friends are having a picnic in a lush green park. One of them is busy looking at the birds in the tree and doesn't notice a squirrel sneaking up on the picnic basket. The other friends are pointing at the squirrel and laughing. Meanwhile, another friend is trying to sit on a sandwich that someone dropped. There are various objects around, such as a frisbee, a picnic blanket, a coffee thermos, and a kite stuck in the tree.

[0072] Beach day: A family of five is spending the day at the beach. The father is building a sandcastle with the youngest child but is so engrossed that he doesn't notice the incoming tide that is about to wash away the castle. The mother is reading a book, and the ice cream is melting in the sun. Two older children are playing beach ball but are about to hit a sunbathing person who is sleeping. There are beach towels, a bucket and spade, seashells, and a cooler box around.

[0073] Disaster at the Ski Lodge: Three friends are at a cozy ski lodge. One is gazing out the window at the beautiful snow, distracted and pouring hot chocolate onto the table instead of into a cup. Another friend is trying to take a picture while laughing. Meanwhile, the third friend is slipping on a puddle of spilled hot chocolate. The room is strewn with ski equipment, a crackling fireplace, board games, and mugs of hot drinks.

[0074] Roadside Repair: Two friends are trying to change a flat tire on a main road. One is busy talking on the phone and giving instructions, unaware that the other is struggling with a jack that looks like it might slip at any moment. The trunk of the car is open, and suitcases, maps, a cooler box, and a spare tire are visible. A nearby sign indicates that they are only a few miles from the city.

[0075] Swimming Pool Party: Four friends are having a pool party. One is busy setting up a music player and doesn't notice that they're about to fall into the pool. The second friend is inflating a giant pool float, and the third friend is grilling hot dogs but doesn't notice that one has fallen off the grill. The fourth friend is about to jump into the pool but doesn't notice that there's an inflated pool float in the way. The scene is filled with summer decorations, pool toys, barbecue equipment, and garden furniture.

[0076] Each of these scenarios involves multiple characters and objects, elements that are distracting, surprising, and various interactions, as well as a distinct location.

[0077] Healthcare Provider User Input Prompt

[0078] The characteristics observed in patients with AD and schizophrenia are as follows. Please simulate how a patient with schizophrenia would describe the first scenario. In addition to those listed above, the list of characteristics of AD is as follows: Verbosity and lexical richness measured by total word count, unique words, number of utterances, or even the number of characters in words are predictors of cognitive aging. Repetitiveness is very common in Information content units (ICU) to quantify, for example, the number of objects or actions in an image. Word frequency: Depending on the word, people with AD tend to use some words more than others. People with AD have difficulty using large syntactically complex sentences. Regarding schizophrenia, in addition to those listed above, there is a reduction in syntactic complexity as follows. There are semantic impairments such as organizing individual propositions into larger structures. More pronouns are used as parts of speech, and fewer adverbs, adjectives, and determiners (e.g., "the", "a") are used.

[0079] Response of the generating component 210:

[0080] Sure, I can help. An example of how a patient with schizophrenia might describe the first scenario, "A picnic in the park", is as follows. This description attempts to reflect some of the language characteristics you mentioned, including reduced syntactic complexity, semantic impairments, and changes in the use of parts of speech.

[0081] "Friends. Four of them. In the park. It's green all around. One person is looking at a bird. There's a bird in the tree. There's a squirrel near it. Creeping up, yes, creeping up. The friends are laughing. Pointing at the squirrel. Another friend is trying to sit down. On top of the sandwich. The sandwich is on the ground. A picnic. There are things around. A frisbee. A blanket on the ground. Coffee in the thermos. A kite. Caught. Caught in the tree."

[0082] Note that this is a generalized description and may not accurately represent all individuals with schizophrenia. This is because symptoms and their severity can vary widely from person to person. Similarly, specific language impairments can vary based on individual differences and the course of the disease.

[0083] The generation of image description task images using an AI program 150 may include an analysis component 220. The analysis component 220 may compare simulated text descriptions (e.g., extracted features and / or distributions) from groups of healthy individuals and groups of non-healthy individuals. The analysis component 220 may determine whether the comparison of simulated text descriptions (e.g., extracted features and / or distributions) is within a predetermined threshold of difference sufficient to make a diagnosis in a group of individuals with a predetermined medical condition (e.g., within at least one predetermined threshold of accuracy, precision, effectiveness, sensitivity, reliability, etc.). The analysis component 220 may approve the generated images for use in an image description task (e.g., diagnosis / non-diagnosis of a predetermined medical condition, and / or differential diagnosis), and / or store in a repository a user input prompt, a predetermined input, an image description task / image / story / text description / extracted feature when the comparison of the extracted features of the text descriptions of the group of healthy individuals and the group of individuals with a predetermined medical condition shows a difference of at least a predetermined threshold. If the comparison of the extracted features of the text descriptions of the group of healthy individuals and the group of individuals with a predetermined medical condition does not show a difference of at least a predetermined threshold, the analysis component 220 may adjust at least one of a predetermined input, a generated story, a generated image, a generated text description prompt, and a generated simulated text description automatically or via a user request. In an embodiment, the analysis component 220 may modify the learned experimental design of an image description task predicted to improve the discrimination of the comparison of text descriptions (e.g., extracted features and / or distributions) within a predetermined threshold of difference sufficient to make a diagnosis in a group of individuals with a predetermined medical condition (e.g., within a predetermined threshold).

[0084] For example, the analysis component 220 can compare simulated text descriptions generated from a group of healthy individuals and a group of individuals with AD and / or schizophrenia. The analysis component 220 determines that the simulated text descriptions exhibit inconsistent accuracy from iteration to iteration, and thus the analysis component increases the actions and characters that occur in the story and corresponding images until reliability and accuracy thresholds are achieved.

[0085] The generation of image description task images using an AI program 150 may include a task implementation component 230. The task implementation component 230 may be displayed on an interactive interface for a user (e.g., a scientist, a medical provider, a patient, etc.). The task implementation component 230 may select (e.g., newly generate and / or obtain from a repository) an image description task (e.g., an experimental design, a story, an image, a written prompt, a simulated text description, etc.) based on a user input prompt and / or a predetermined input by the user (e.g., a scientist and / or a medical provider). The task implementation component 230 may provide the selected image description task to the user (e.g., a patient) via the interactive interface. The task implementation component 230 may provide a text description prompt and / or extract features and / or distributions from a response text description and compare them with at least one of the simulated / actual text descriptions from a group of healthy individuals and a group of individuals with a predetermined medical condition. The task implementation component 230 may provide a diagnosis / nondiagnosis / differential diagnosis and their explanations, and may provide a confidence level and / or an odds ratio, etc. When the diagnosis is unexpected (e.g., a discrepancy between the diagnosis by a medical provider and the diagnosis related to the features extracted from the patient characteristics), the task implementation component 230 may automatically and / or with the user's approval execute at least one other image description task and / or a correction. When the comparison of the extracted features of the text descriptions of the group of healthy individuals and the group of individuals with a predetermined medical condition does not show a difference of at least a predetermined threshold and / or does not meet at least one predetermined threshold of accuracy, precision, effectiveness, sensitivity, and / or reliability, the task implementation component 230 may adjust at least one of a predetermined input, a generated story (e.g., an increase / decrease in detail, complexity, length, etc.), a generated image, a generated text description prompt, and / or a generated simulated text description (e.g., weight, quality, quantity, etc.). The task implementation component 230 may provide the diagnosis and the corresponding evaluation / score to the user via the interactive interface along with an annotation / explanation.In an embodiment, the task implementation component may provide a user survey to confirm diagnosis / non-diagnosis and / or to scrutinize differential diagnosis.

[0086] For example, the task implementation component 230 uses a medical provider input prompt, a predetermined input, a generated story / image / text description prompt / a text description indicating a predetermined threshold of differences to be compared. The task implementation component 230 provides an image description task about AD and / or schizophrenia to a new patient user. The new patient user is presented with a text description prompt related to a crowded amusement park and a corresponding approved image. The new patient's text description is as follows: "A bustling square. Explanation: A group of friends is sitting on a bench. It is a release from constant loneliness. Since it usually rains, they are enjoying the atmosphere. The ice cream cone of one character starts to melt rapidly due to the hot weather and drips from the hand. It is typical. Friends react with surprise and laughter and try to lick it quickly before the melting ice cream causes even more chaos." The comparison of differences in the new patient's text description shows a difference from a predetermined threshold from the group of healthy individuals and excluded schizophrenia patients, but the tone of the new patient's text description suggests a potential differential diagnosis of depression, which is complicated by the overall similarity in the characteristic features of the text description. The task implementation component 230 then provides an image description task of a picnic in the park. The new patient's text description is as follows: "The coffee shop, yes, the coffee shop was crowded. There were people everywhere, talking. Talking too much. Very noisy. I was sitting at a table, yes, facing the laptop. Working, typing. Typing a lot of words. And then, oh, reaching for the coffee cup but tipping it over and spilling it. Typical. Spilled all over the laptop and the table and everywhere. Hot. It was hot. Ruined a nice shirt. A lot of words. There were so many words, but now, oh, the laptop won't move. People saw. They saw me. I was embarrassed. I was embarrassed and sad. There were so many words, but now it's ruined. Just like everything else." The task implementation component 230 confirms a diagnosis of depression within 90% confidence and presents the conclusion to the waiting medical provider.

[0087] Figure 3 shows a flowchart of a method for generating an image 300 for an image captioning task using AI, according to an exemplary embodiment of the concepts of the present invention.

[0088] The method for generating an image for an image captioning task using AI may include the following:

[0089] Using AI, generate a story based on at least one of a user input prompt and a predetermined input (step 302);

[0090] Using AI, generate an image based on the generated story (step 304);

[0091] Simulate a group of healthy individuals and a group of individuals having a predetermined medical condition, and use AI to generate a textual description of the generated image (step 306).

[0092] Extract diagnostic verbal features from the textual descriptions of the group of healthy individuals and the group of individuals having a predetermined medical condition (step 308);

[0093] Compare the extracted diagnostic verbal features for the textual descriptions of the group of healthy individuals and the group of individuals having a predetermined medical condition (step 310); and

[0094] When the comparison of the extracted features of the textual descriptions of the group of healthy individuals and the group of individuals having a predetermined medical condition shows a difference of at least a predetermined threshold, use the generated image in an image captioning task (step 312).

[0095] Based on the above, a computer system, method, and computer program product have been disclosed. However, numerous modifications, additions, and substitutions may be made without departing from the scope of the exemplary embodiments of the concepts of the present invention. Accordingly, the exemplary embodiments of the concepts of the present invention are disclosed by way of example and not limitation.

Claims

1. A method for generating an image description task image using AI, comprising: generating a story using AI based on at least one of a user input prompt and a predetermined input; generating an image using AI based on the generated story; simulating a group of healthy individuals and a group of individuals with a predetermined medical condition using AI to generate a textual description of the generated image; extracting diagnostic linguistic features from the textual descriptions of the group of healthy individuals and the group of individuals with the predetermined medical condition; comparing the extracted diagnostic linguistic features of the textual descriptions for the group of healthy individuals and the group of individuals with the predetermined medical condition; and using the generated image in an image description task when the comparison of the extracted features of the textual descriptions for the group of healthy individuals and the group of individuals with the predetermined medical condition shows a difference of at least a predetermined threshold A method comprising the above steps.

2. The method according to claim 1, wherein the step of using the generated image in an image description task includes approving the generated image for diagnostic use of the predetermined medical condition.

3. The method according to claim 1, wherein the step of using the generated image in an image description task includes outputting a diagnosis or non-diagnosis for the predetermined medical condition.

4. The method according to claim 1, further comprising generating a simulated textual description for at least one individual among at least one of the group of healthy individuals and the group of individuals with the predetermined medical condition.

5. The method according to any one of claims 1 to 4, further comprising generating a textual description prompt related to the generated image.

6. At least one of the generated story, the generated image, and the generated textual description prompt is designed to elicit a textual description from the group of healthy individuals and the group of individuals with the predetermined medical condition, and the comparison of the extracted diagnostic linguistic features shows a difference of at least the predetermined threshold. The method according to claim 5.

7. The method according to claim 5, further comprising adjusting at least one of the predetermined input, the generated story, the generated image, the generated text description prompt, and the generated simulated text description when the comparison of the extracted features of the text description of the group of healthy individuals and the group of individuals having the predetermined medical condition does not show a difference of at least the predetermined threshold value.

8. A computer program for generating an image for an image description task using AI, comprising program instructions for causing at least one of one or more computer processors to execute a method, the method comprising: generating a story based on at least one of a user input prompt and a predetermined input using AI; generating an image based on the generated story using AI; simulating a group of healthy individuals and a group of individuals having a predetermined medical condition using AI to generate a text description of the generated image; extracting diagnostic verbal features from the text descriptions of the group of healthy individuals and the group of individuals having the predetermined medical condition; comparing the extracted diagnostic verbal features of the text descriptions for the group of healthy individuals and the group of individuals having the predetermined medical condition; and using the generated image in an image description task when the comparison of the extracted features of the text descriptions of the group of healthy individuals and the group of individuals having the predetermined medical condition shows a difference of at least a predetermined threshold value. A computer program comprising the above steps.

9. The computer program according to claim 8, wherein the step of using the generated image in an image description task includes approving the generated image for diagnostic use of the predetermined medical condition.

10. The computer program according to claim 8, wherein the step of using the generated image in an image description task includes outputting a diagnosis or non-diagnosis for the predetermined medical condition.

11. The computer program according to claim 8, further comprising generating a simulated text description for at least one individual in at least one of the group of healthy individuals and the group of individuals having the predetermined medical condition.

12. The computer program according to any one of claims 8 to 11, further comprising the step of generating a text description prompt related to the generated image.

13. At least one of the generated story, the generated image, and the generated text description prompt is designed to elicit text descriptions from the group of healthy individuals and the group of individuals having the predetermined medical condition, and the comparison of the extracted diagnostic verbal features indicates at least a difference from the predetermined threshold value. The computer program according to claim 12.

14. When the comparison of the extracted features of the text descriptions of the group of healthy individuals and the group of individuals having the predetermined medical condition does not indicate at least a difference from the predetermined threshold value, the method further comprises the step of adjusting at least one of the predetermined input, the generated story, the generated image, the generated text description prompt, and the generated simulated text description. The computer program according to claim 12.

15. A computer system (CS) for generating an image for an image description task using AI, comprising one or more computer processors, one or more computer-readable storage media, and program instructions stored on the one or more computer-readable storage media for causing at least one of the one or more computer processors to execute a method, the method comprising: Generating a story using AI based on at least one of a user input prompt and a predetermined input; Generating an image using AI based on the generated story; Using AI to simulate a group of healthy individuals and a group of individuals having a predetermined medical condition to generate a text description of the generated image; Extracting diagnostic verbal features from the text descriptions of the group of healthy individuals and the group of individuals having the predetermined medical condition; Comparing the extracted diagnostic verbal features of the text descriptions of the group of healthy individuals and the group of individuals having the predetermined medical condition; and Using the generated image in an image description task when the comparison of the extracted features of the text descriptions of the group of healthy individuals and the group of individuals having the predetermined medical condition indicates at least a difference from a predetermined threshold value. CS comprising.

16. In the image description task, the step of using the generated image includes the step of approving the generated image for the diagnostic use of the predetermined medical condition, the CS according to claim 15.

17. In the image description task, the step of using the generated image includes the step of outputting a diagnosis or non-diagnosis for the predetermined medical condition, the CS according to claim 15.

18. The CS according to claim 15 further comprises the step of generating a simulated text description for at least one individual among at least one of the group of healthy individuals and the group of individuals having the predetermined medical condition.

19. The CS according to any one of claims 15 to 18 further comprises the step of generating a text description prompt related to the generated image.

20. At least one of the generated story, the generated image, and the generated text description prompt is designed to elicit a text description from the group of healthy individuals and the group of individuals having the predetermined medical condition, and the comparison of the extracted diagnostic verbal features shows at least a difference from the predetermined threshold value, the CS according to claim 19.