Artificial intelligence teaching system for guided learning and application thereof
By integrating a digital morphological image library with clinical case information into an intelligent collaborative platform, and combining it with an AI-guided module, the system addresses the issues of passive learning and lack of real-time guidance in existing medical education systems. This enables students to actively collaborate and develop comprehensive clinical diagnostic skills, thereby enhancing the depth and accuracy of their learning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FOSHAN UNIVERSITY
- Filing Date
- 2025-12-31
- Publication Date
- 2026-05-05
AI Technical Summary
The existing medical education system suffers from problems such as passive and singular learning methods, disconnect between content and collaborative functions, lack of real-time professional guidance and synchronization mechanisms in digital teaching, making it difficult to cultivate students' comprehensive clinical diagnostic abilities.
It provides an integrated, case-centric, intelligent collaborative learning platform that combines a high-quality morphological digital image library with complete clinical case information. Through an AI-powered intelligent guidance module, it analyzes the interactive behavior of learning groups in real time, providing context-relevant prompts and feedback to help students transform from passive observers to active problem solvers.
It enables students to actively collaborate in digital teaching, cultivates their comprehensive clinical thinking skills, enhances the depth and accuracy of learning, and provides an immersive and efficient learning experience with simulated expert guidance.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical education technology, and specifically to an artificial intelligence teaching system for guided learning and its applications. Background Technology
[0002] Morphology is the cornerstone of medical disciplines such as pathology, histology, and hematology, requiring learners to possess a high level of image recognition and diagnostic skills. Traditional teaching methods primarily rely on optical microscopes and glass slides. Learners must observe a limited number of specimens in a fixed laboratory within a limited time, which is insufficient to meet the demands of modern medical education for efficient and flexible learning. With the development of information technology, digital teaching has become a trend. Advances in digital microscopy imaging technology, by equipping traditional optical microscopes with high-resolution digital cameras, allow for the capture of high-resolution digital images of typical lesion areas and key morphological features in physical slide specimens. These carefully selected digital images have been compiled, giving rise to a vast database of digital morphology and online teaching atlases. These platforms allow students to access a large number of case images anytime, anywhere, initially addressing the problem of resource scarcity.
[0003] The most similar solutions can be divided into three categories: 1. Virtual Microscope / Digital Atlas Platforms: The core function of these platforms is to store and display massive amounts of digital slide images. Users can zoom, pan, and perform other operations on the digital images just like operating a real microscope. The platforms typically include simple text and image annotations explaining key morphological features. These platforms solve the "seeing" problem, but are essentially a one-way, individual learning model.
[0004] 2. General-purpose collaborative / gamified teaching platforms: For example, using existing commercial gamification platforms like Kahoot! to assist in histology teaching. These platforms stimulate student interest and participation through Q&A, competitions, etc., and support teamwork. However, these platforms are general-purpose tools and do not contain professional, structured medical case content. Teachers need to prepare morphological images and questions themselves and import them into these general-purpose platforms, resulting in a separation between content and collaborative functions, making it impossible to achieve in-depth collaboration based on digital slides themselves.
[0005] In summary, existing teaching systems suffer from the following drawbacks: Passive and singular learning methods: While existing digital image platforms primarily address image "accessibility," the learning method remains a passive "browsing-memorizing" model, lacking effective guidance and interaction, and failing to cultivate students' clinical diagnostic reasoning abilities. Disconnect between content and collaborative functions: Although general collaborative platforms can promote interaction, their functions do not match the needs of professional morphological observation. Students cannot perform real-time, synchronous annotation and discussion of specific areas of digital slide images on the platform, severely limiting the depth and efficiency of collaboration. Lack of complete clinical context: Most platforms only provide isolated morphological images, lacking accompanying complete clinical case information such as medical history and laboratory tests, hindering the development of students' comprehensive ability to combine morphological findings with clinical practice. Separation of intelligent guidance and collaborative learning: Existing AI tutoring systems are typically in "single-person mode," with AI interacting with a single student, unable to understand and intervene in the collaborative dynamics between multiple users. Conversely, collaborative platforms completely lack AI participation. This results in students being unable to simultaneously benefit from peer collaboration and real-time guidance from expert-level AI in one environment, leading to a fragmented learning experience. The one-way and passive limitations of synchronization mechanisms: Furthermore, existing multi-screen synchronization technologies (such as operation synchronization in remote conferencing) essentially broadcast and reproduce the host's operation instructions. The system itself lacks any ability to analyze and understand user behavior; it is merely a passive "instruction megaphone." This mechanism can be used for remote demonstrations, but it cannot intervene in or guide the learning process itself in a teaching setting, nor can it identify whether a learning group has deviated from the correct diagnostic approach. Therefore, its application in the field of education has fundamental limitations. Summary of the Invention
[0006] To address the aforementioned problems, this invention provides an AI-powered teaching system for guided learning. This system offers an integrated, case-centric, intelligent collaborative learning platform. It creates complete clinical scenarios for learners, helping them establish a direct connection between target features and real-world data, thereby cultivating comprehensive clinical thinking skills. Simultaneously, it transforms learners from passive observers into active, collaborative problem solvers. Furthermore, to address the lack of real-time professional guidance in collaborative learning, this system innovatively introduces an AI-powered intelligent guidance module. This module analyzes the interactive behavior of learning groups in real time and provides context-relevant prompts, questions, and feedback, proactively guiding students to focus on key features and correct potential misconceptions. This achieves an immersive and highly effective learning experience, simulating expert guidance.
[0007] This invention provides an AI-powered teaching system for guided learning, comprising: a database module, a case study module, a collaborative workspace module, an AI-powered intelligent guidance module, and a user interface. The database module stores, manages, and retrieves image files. The case study module associates image files with real-world data to form teaching cases. The collaborative workspace module allows multiple system users to enter the same virtual room to synchronously, share, and collaboratively analyze the same image file within the teaching cases, outputting interactive data packets to the AI-powered intelligent guidance module and receiving and displaying real-time feedback from the AI-powered intelligent guidance module. The AI-powered intelligent guidance module analyzes the interactive data packets output by the collaborative workspace module in real time, generating and providing feedback guidance information. The user interface is used to input user commands and display the guidance information.
[0008] The inventors, through their research on existing technologies, discovered that the aforementioned shortcomings stem primarily from the separation of design philosophy and technical architecture in existing solutions. Database platform designers focus on image storage and display technologies, while collaboration tool designers focus on general interactive functions. The market lacks a solution that natively integrates high-quality, structured morphological case content with a collaborative workspace deeply customized for morphological learning. Furthermore, AI tutor systems prioritize simulating one-on-one human-machine teaching, and their technical architecture fails to consider how to handle and respond to real-time, concurrent interactive data streams from multiple users, making integration into collaborative workspaces difficult. Based on this, the inventors propose the aforementioned artificial intelligence teaching system. This system provides an integrated, case-centric, intelligent collaborative learning platform that creates complete clinical scenarios for learners, helping them establish direct connections between target features and real data, thereby cultivating comprehensive clinical thinking skills. Simultaneously, it transforms learners from passive observers into proactive, collaborative problem solvers. Furthermore, to address the challenge of lacking real-time professional guidance in collaborative learning, this system innovatively introduces an AI-powered intelligent guidance module. This module analyzes the interactive behavior of learning groups in real time and provides context-relevant prompts, questions, and feedback. It proactively guides students to focus on key features and correct potential misconceptions, achieving an immersive and efficient learning experience that simulates expert guidance.
[0009] In one embodiment, the database module includes a metadata database and a file server. The relational mapping mechanism of the database module includes: storing image files in the file server, generating identifiers, storing descriptive information, identifiers, and storage paths of the image files in metadata records, and storing the metadata records in the metadata database.
[0010] In one embodiment, the identifier corresponds globally and uniquely to the image file.
[0011] In one embodiment, the image file is associated with the real data through a many-to-many relationship.
[0012] In one embodiment, the synchronization, sharing, and collaborative analysis includes: synchronized view, shared annotation, and integrated chat; the synchronized view includes: the system user's operation parameters on the image file view are broadcast in real time to the remaining system users in the same virtual room via the WebSocket protocol, triggering the client's application logic to update the view and synchronize the view; The shared annotation includes: the operation parameters of system users annotating image files are broadcast in real time to the remaining system users in the same virtual room via the WebSocket protocol, so that the annotations are displayed synchronously; The integrated chat includes: system users in the same virtual room discussing via a real-time chat box based on the WebSocket protocol.
[0013] In one embodiment, the application logic includes: JavaScript functions.
[0014] In one embodiment, the output of interactive data packets to the AI intelligent guidance module includes: when a system user views or annotates an image file, generating an interactive data packet, sending it to the application server via the WebSocket protocol, and forwarding it to the AI intelligent guidance module for real-time analysis; The process of receiving and displaying real-time feedback from the AI intelligent guidance module includes: the AI intelligent guidance module performing real-time analysis, generating guidance information, and having the application server feed it back to the collaborative workspace module via the WebSocket protocol for display.
[0015] In one embodiment, the AI intelligent guidance module includes: a multi-task deep learning visual analysis model, a parameter calculation algorithm, and a dialogue generation model; The multi-task deep learning visual analysis model includes: a convolutional neural network and multiple parallel decoding heads; the convolutional neural network is used to extract features and output them to the decoding heads, the decoding heads are used to generate a segmentation mask based on the features, and the segmentation mask extracts target parameters through a parameter calculation algorithm and passes them to the dialogue generation model; The dialogue generation model includes a unified language model interface layer, which is used to call the dialogue generation capability to generate guidance information based on the target parameters. The unified language model interface layer decouples the upper-layer AI intelligent guidance logic from the lower-layer language model.
[0016] In one embodiment, the AI intelligent guidance logic includes a group attention detection and key area deviation analysis algorithm. The execution of the group attention detection and key area deviation analysis algorithm includes the following steps: predefined key diagnosis areas, real-time aggregation of group attention, deviation analysis and triggering algorithm, and generation of guidance information; The predefined key diagnosis areas include: when the teaching case is stored in the meta-database, at least one area containing key features is marked as the key diagnosis area of the teaching case and stored in the meta-database; The real-time aggregation of group attention includes: receiving the interaction data packet, generating a dynamic two-dimensional array, simulating the formation of an attention heat map, and quantifying the duration and frequency of observation of the key diagnosis areas; The deviation analysis and triggering algorithm includes: calculating the cumulative attention obtained by each key diagnosis area, comparing it with the total attention of the entire image file in the teaching case, obtaining the proportion of attention of the key diagnosis area, and triggering an internal instruction for the key diagnosis area according to the triggering rule; The generation of guidance information includes: generating guidance information according to the internal instruction.
[0017] In one embodiment, the triggering rule includes: IF (session duration > X seconds) AND (the proportion of attention of a certain key diagnosis area < Y%) THEN trigger the guidance information for this key diagnosis area.
[0018] The above X and Y are set according to teaching needs.
[0019] The present invention also provides a morphological teaching and / or learning method, using the artificial intelligence teaching system for teaching and / or learning; the real data includes: real-world research data.
[0020] In one embodiment, the real-world research data includes case data, and the case data includes at least one of medical history, chief complaint, and laboratory test results.
[0021] In one embodiment, the parameter calculation algorithm is a quantitative morphological parameter calculation algorithm, and the quantitative morphological parameter calculation algorithm includes: calculation of nuclear-cytoplasmic ratio based on region segmentation, relative size scale with red blood cells based on region matching, and quantification of particle characteristics based on gray-level co-occurrence matrix.
[0022] In one embodiment, the descriptive information includes at least one of cell name, disease classification, specimen source, staining method, and magnification.
[0023] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by the processor, implements the method described herein.
[0024] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described thereon.
[0025] Compared with the prior art, the present invention has the following beneficial effects: This invention provides an AI-powered teaching system for guided learning and its application. This AI teaching system offers an integrated, case-centric, intelligent collaborative learning platform. The system creates complete clinical scenarios for learners, helping them establish a direct connection between target features and real-world data, thereby cultivating comprehensive clinical thinking skills. Simultaneously, it transforms learners from passive observers into active, collaborative problem solvers. Furthermore, to address the challenge of lacking real-time professional guidance in collaborative learning, this system innovatively introduces an AI-powered intelligent guidance module. This module analyzes the interactive behavior of learning groups in real time and provides context-relevant prompts, questions, and feedback, proactively guiding students to focus on key features and correct potential misconceptions, achieving an immersive and highly effective learning experience similar to expert guidance. Attached Figure Description
[0026] Figure 1 This is a flowchart of the artificial intelligence teaching system in the embodiment; Figure 2-9 This is a schematic diagram illustrating the effect of an artificial intelligence teaching system in practical application. Detailed Implementation
[0027] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Preferred embodiments of the invention are shown in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of the invention.
[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0029] Unless otherwise specified, all reagents, materials, and equipment used in this embodiment are commercially available; unless otherwise specified, all test methods are conventional test methods in this field.
[0030] Example 1 I. An artificial intelligence teaching system for guided learning.
[0031] The flowchart of the artificial intelligence teaching system is as follows: Figure 1 As shown.
[0032] User Interface (UI): The front-end uses a responsive design for the application server, adapting to various devices such as PCs and tablets. Students and teachers access the platform through a browser.
[0033] Morphological database module: Function: Store, manage, and retrieve high-resolution digital cell smear images.
[0034] Technical Implementation: The images are high-resolution digital photographs taken using a professional digital camera mounted on an optical microscope and stored on a file server in standard formats such as JPEG or TIFF. To reliably manage these images and firmly associate them with descriptive information, this system employs a mature relational database mapping mechanism. Specifically, when each image file is stored on the server, the system generates a globally unique identifier (UUID) to ensure the uniqueness of its index. Simultaneously, in the backend metadata database (such as MySQL), a record not only fully contains all descriptive information such as cell name, disease classification, specimen source, staining method, and magnification, but also stores a unique identifier pointing to the image file and its specific storage path (URI). In this way, the association between "image file" and "descriptive information" is firmly bound together through the same row of records in the database. When students search using multiple tags, the system essentially queries the database for metadata records that match the criteria. Once a match is found, the system can accurately locate and retrieve the corresponding image file from the file server using the path information stored in that record.
[0035] Case-Based Learning Module: Function: Associate images in the database with de-anonymized clinical case information to construct complete teaching cases.
[0036] Technical Implementation: This module employs a "many-to-many relationship" design pattern at the database level to establish flexible associations between images and case files. Specifically, in addition to the Cases table storing case data (such as medical history, chief complaint, and laboratory test results) and the Images table storing image metadata, the system also creates a connection table (or association table) named CaseImageAssociation. The core fields of this connection table are case_id and image_id, which serve as foreign keys pointing to the primary keys of the Cases and Images tables, respectively. When one or more images need to be associated with a single case, the system creates corresponding records in this connection table. For example, to associate two images with IDs 2025101001 and 2005101002 with case ID 5, the system only needs to insert two records into the CaseImageAssociation table: (5, 2005101001) and (5, 2005101002). Therefore, the function of "linking one or more image IDs with a case ID" is technically implemented by inserting data rows into this connection table. When a student accesses a case, the system first retrieves a list of all relevant image IDs from the connection table based on the case ID, then retrieves all corresponding image information from the Images table based on this list, and finally presents the complete case text and all relevant images together on the UI.
[0037] Collaborative Workspace Module: Function: This is the core of the invention. It allows multiple students (a group) to enter the same virtual room to perform real-time, synchronous collaborative analysis of the same digital slice, and to receive and display real-time feedback from the AI-guided module.
[0038] Technical Implementation: Synchronized View: When a user zooms or pans an image, their operation parameters (coordinates, zoom level) are broadcast in real time to all other users in the room via the WebSocket protocol. Other users' clients receive these parameters through the WebSocket connection and trigger their local JavaScript functions. These functions then call the front-end image rendering engine's API to immediately update the view to the same position and zoom level, thus achieving view synchronization. Therefore, WebSocket handles the "transmission" of parameters, while the client's "application logic" handles the "update execution" of the view.
[0039] Shared annotation: Users can annotate images using tools such as pens, arrows, and rectangles. Each annotation operation (such as drawing the coordinates and radius of a circle) is also broadcast in real time via WebSocket. The annotations and cursor positions of all group members are displayed synchronously on all users' screens.
[0040] Integrated Chat: The module embeds a WebSocket-based instant chat box for group members to conduct text discussions.
[0041] Closed-loop data interaction with the AI module: This module not only handles synchronization between users, but more importantly, it serves as the data interface for the AI guidance function. When a user performs view or annotation operations, the generated interaction data packet (containing user ID, event type, coordinate parameters, etc.) is sent to the application server via WebSocket and performs two tasks: 1) broadcasting to other users to achieve synchronization; 2) forwarding to the AI intelligent guidance module for real-time analysis. After the AI module completes its analysis, the generated guidance information (such as prompts, questions, or corrections) is pushed back to the collaborative workspace by the application server through the same WebSocket channel and displayed in real time on the user interface of all team members.
[0042] AI-Powered Guidance Module: Function: Another core feature of this invention. It analyzes user behavior data from collaborative workspaces in real time, combines it with medical knowledge relevant to the current case, and generates and pushes guiding instructions.
[0043] Technical Implementation: AI-powered intelligent guidance module: 1) Multi-Task Deep Learning Visual Analysis Model: This invention does not employ a traditional single-task classification model, but rather a multi-task deep learning visual analysis model with a specially designed structure and output mechanism. This model uses a convolutional neural network (CNN) as a shared feature extraction backbone and constructs multiple parallel task decoder heads at its top layer. This allows the system to simultaneously complete multiple analysis tasks, such as cell type identification, semantic segmentation, and anomaly detection, during a single forward propagation. Unlike conventional models that only output single-class results, this invention introduces a structured information representation mechanism at the output layer, integrating the results of each task into a highly structured JSON object in a unified data format. This object not only contains cell type and confidence information but also quantitative parameters such as cell diameter, nucleoplasmic ratio, relative size to red blood cells, anomaly feature labels, and texture feature vectors. This directly generates standardized input at the model level that can be called by the AI intelligent guidance module. Through this design, the model achieves multi-dimensional fusion output of visual recognition results, enabling the AI system to directly utilize the model-generated data for teaching guidance and intelligent feedback without additional parsing or conversion. This holistic design, from network structure to output format, significantly improves the real-time performance and scalability of the visual analysis module in teaching scenarios. The core of the model is a convolutional neural network (CNN) backbone, but multiple parallel "decoder heads" are designed at its output, enabling it to complete multiple analysis tasks simultaneously in a single forward pass and output a highly structured JSON object containing multi-dimensional information. Its main tasks include: Task A - Target Classification: Identify the specific cell type (e.g., primitive cells, promyelocytes, etc.) within an image region and provide a confidence score.
[0044] Task B - Semantic Segmentation: Perform pixel-level segmentation on the identified cells, accurately delineate the cell outline, the outline of the cell nucleus, and the cytoplasm region, and generate the corresponding segmentation mask.
[0045] Task C - Abnormal Feature Detection: Specifically trained to identify specific abnormal features of significant diagnostic importance, such as Auer rods, Döhle bodies, toxic particles, Cabot rings, etc., and output them as Boolean values or labels.
[0046] 2) Quantitative Morphological Parameter Calculation Algorithm: After the visual analysis model outputs the segmentation mask, the system then calls a series of deterministic image processing algorithms unique to this invention to perform precise quantitative calculations on the segmented regions, extracting key morphological parameters that traditional models cannot provide. The nucleus-cytoplasm ratio (N / C Ratio) is calculated based on region segmentation: by separately calculating the pixel areas of the nucleus and cytoplasm segmentation masks, the key diagnostic indicator of the nucleus-cytoplasm ratio is accurately calculated.
[0047] Region-based matching relative size scale with red blood cells: The algorithm automatically finds mature red blood cells with regular shapes near the target cell and takes their average size as a reference. By calculating the ratio of their diameters or areas, it provides a basis for diagnosis based on relative size.
[0048] Particle feature quantization based on gray-level co-occurrence matrix (GLCM): The algorithm applies texture analysis algorithms (such as gray-level co-occurrence matrix) to the cytoplasm region to quantify the texture features such as particle size, density, and color intensity, and outputs a feature vector.
[0049] After the above steps, the multi-task deep learning visual analysis model finally outputs a structured data object: These algorithms are all deterministic computation processes, meaning that under the same input image and segmentation results, they always generate completely consistent quantitative parameter outputs. After the algorithm is executed, the system integrates the calculated indicators with the classification, segmentation, and other results output by the aforementioned visual recognition model, and encapsulates them in a structured data format.
[0050] The final generated structured data object is represented in JSON format, for example: { "cell_type": "promyelocyte", "confidence": 0.92, "diameter_um": 18.5, "nc_ratio": 0.75, "relative_size_to_rbc": 2.2, "abnormal_features": ["auer_rod_present"], "granule_texture_vector": [0.8, 0.6, 0.9, ...], "segmentation_mask": "..." } This JSON object serves only as an example of a data structure, demonstrating the organization of the algorithm's output. The system uses this structured result as input to the subsequent dialogue generation model (LLM), enabling it to generate far more accurate and in-depth guided questions and feedback than conventional models. 3) Dialogue Generation Model: To achieve maximum flexibility and scalability, the system does not hardcode a connection to a specific dialogue AI. Instead, it invokes dialogue generation capabilities through a "Unified Language Model Interface." This interface layer is a software abstraction layer that decouples the upper-layer application (AI intelligent guidance logic) from the underlying specific language model implementation. This interface layer supports two configurable working modes: External API mode: This interface layer has built-in adapters for various mainstream large-scale language model (LLM) APIs, such as API adapters for Gemini, Claude, ChatGPT, and domestic models like Kimi and Deepseek. System administrators can dynamically select the API service currently in use in the backend configuration based on cost, performance, or regional availability. This adapter is responsible for automatically converting standardized requests within the system into the format required by specific API vendors and parsing the returned results. Local default model mode: To meet the needs of high data privacy or off-network deployment, the system also supports calling a locally deployed default model. This model can be an open-source language model fine-tuned with domain knowledge (such as Llama, Qwen series, etc.), deployed on the user's private server. When configured in this mode, the unified language model interface layer communicates directly with this local model service, ensuring that teaching interaction data never leaves the institution's internal network. This pluggable architecture allows the invention to flexibly respond to future technological developments, easily integrate new and better language models, and also provides a reliable private deployment solution for medical and educational institutions with special data security requirements, demonstrating significant technical advantages and commercial value.
[0051] AI-powered intelligent guidance logic: Hotspot Analysis: This function is not implemented by simply calling external APIs, but is based on a unique "group attention monitoring and key area deviation analysis" algorithm developed in this invention. This algorithm is executed within a dedicated "hotspot analysis" function. The specific implementation is as follows: Predefined Key Diagnostic Area (KDA): When each teaching case is stored in the database, experts or teachers will pre-mark one or more rectangular areas in the image that contain key positive / negative features and store them as KDAs in the metadata of that case.
[0052] Real-time aggregation of group attention: During collaborative learning, "heatmap analysis" receives and aggregates view data packets (including view center coordinates and zoom level) from all group members in real time. The system dynamically maintains a two-dimensional array on the server side for this collaborative session, simulating an "attention heatmap" to quantify the duration and frequency of observation of each region on the image.
[0053] Deviation Analysis and Triggering Algorithm: The module periodically (e.g., every 10 seconds) calculates the cumulative attention gained by each predefined KDA and compares it with the total attention of the entire image to arrive at a "KDA attention percentage". The triggering algorithm includes a preset triggering rule, such as: IF (session duration > 90 seconds) AND (attention percentage of a certain KDA < 5%) THEN trigger a guiding instruction for that KDA. This rule ensures that the system only intervenes "intelligently" when the group ignores key information for an extended period, rather than engaging in ineffective harassment.
[0054] Generation of the prompt instruction: The algorithm described above triggers an internal instruction containing context, such as {"trigger_event": "ignored_kda", "kda_id": "kda_01"}. Based on this instruction, the system can then selectively: (a) directly retrieve a prompt from a pre-defined "prompt corpus" associated with the KDA (e.g., "It is recommended to observe the cell morphology in the lower right corner of the image."); or (b) submit the structured instruction ("Students ignored the key area in the lower right corner") as a prompt to the Large Language Model (LLM) API, which will then generate a more dynamic and diverse prompt.
[0055] II. Summary.
[0056] The technical problems this embodiment aims to solve are: 1) Traditional teaching models are limited by physical resources (such as microscopes and slides) and time and space, resulting in insufficient opportunities for student practice and low learning efficiency; 2) Existing digital teaching resources are mostly static image libraries, lacking effective interactive and collaborative mechanisms, which easily leads to passive learning by students and makes it difficult to cultivate comprehensive clinical diagnostic thinking; 3) Even in collaborative learning environments, due to the lack of real-time guidance from experts, student groups may deviate from the focus of discussions, miss key morphological features, or form incorrect consensus on observed phenomena, thereby limiting the depth and accuracy of learning. Existing technical solutions cannot provide scalable, intelligent, real-time guidance to solve this problem.
[0057] Meanwhile, this embodiment aims to overcome the shortcomings of existing technologies, such as the disconnect between teaching content and collaborative tools, passive learning methods, and a lack of effective guidance, by providing an integrated, case-centered, intelligent collaborative learning platform. To achieve this, this embodiment first deeply integrates a high-quality morphological digital image library with complete clinical case information to construct a structured teaching case library. This creates a complete clinical context for learners, aiming to help them establish a direct link between morphological features and disease diagnosis, thereby cultivating their comprehensive clinical thinking ability. Building on this, to transform learners from passive observers to active, collaborative problem solvers, this embodiment further develops a collaborative workspace tailored for morphological diagnosis and supporting real-time synchronous operation by multiple users. Finally, and crucially, to address the lack of real-time professional guidance in collaborative learning, the platform innovatively introduces an AI intelligent guidance module. This module can analyze the interactive behavior of learning groups in real time (such as jointly observed areas and labeled content) and provide highly context-relevant prompts, questions, and feedback, thereby proactively guiding students to focus on key features and correct potential misconceptions, achieving an immersive and efficient learning effect that simulates expert guidance.
[0058] Specifically, the artificial intelligence teaching system in this embodiment achieves the following advantages and technical effects: 1. Deep native integration of teaching content and collaborative tools.
[0059] In this embodiment, structured clinical cases (images + medical history + laboratory data) are seamlessly integrated with morphology-customized collaborative tools (synchronous views, shared annotations) into a single platform and interface. In contrast, existing technologies separate databases from collaborative tools (such as Kahoot! or general chat software). These two are separate, requiring users to switch between multiple applications and preventing direct, rich-media collaboration on images.
[0060] The difference is that this embodiment realizes "collaboration within a case", while the existing technology can only achieve "collaboration around a case", and the depth and efficiency of collaboration are completely different.
[0061] 2. Supports a morphological collaboration mechanism that enables real-time synchronous annotation and view control by multiple users.
[0062] In this embodiment, WebSocket technology enables real-time synchronization of panning and zooming operations performed by one person on the same high-resolution digital slice, while all team members are observing the same slice. Furthermore, all annotations and pen marks are shared in real time. In existing technologies, virtual microscope platforms are operated by a single person. General collaboration tools (such as shared whiteboards) can share annotations, but they do not support smooth, synchronized scaling and panning of WSI digital slices of terabyte (TB) size.
[0063] Difference: This embodiment solves the technical challenge of enabling real-time synchronous interaction among multiple users on high-resolution medical images, creating an online experience that is close to "multiple people sharing a single microscope".
[0064] 3. A closed-loop processing mechanism that deeply integrates group collaborative data flow with AI analysis guidance.
[0065] In this embodiment, an innovative data processing closed loop is constructed. The system not only broadcasts user collaborative interaction data (such as view location and annotations) via WebSocket to achieve synchronization among users, but more importantly, it provides this real-time data stream, representing collective wisdom and challenges, as input to the backend AI analysis engine. The AI engine analyzes this data in real time, generating guidance information (such as prompts, corrections, and questions) highly relevant to the current collaborative context. This "intelligent" information is then pushed back to the collaborative group via WebSocket, dynamically and in real-time influencing and improving the learning process. This constitutes a complete intelligent teaching cycle of "perception-analysis-guidance-feedback."
[0066] Existing multi-user synchronization systems only transmit data to enable one client to view / perform actions from another; the data itself is not analyzed or understood by the system, representing a one-way broadcast of instructions. Furthermore, existing AI tutoring systems typically only process the behavior of a single user and lack the ability to process and analyze real-time interactive data from multiple groups.
[0067] The key difference lies in the fact that this embodiment endows collaborative data streams with entirely new technical meaning and applications—namely, as real-time input for AI mentors. It deeply couples the previously separate technical fields of "multi-user synchronization" and "AI analysis" through an innovative closed-loop data architecture, achieving a technological leap from passive "information display synchronization" to proactive "intelligent process guidance."
[0068] 4. Significantly enhances learning initiative and participation.
[0069] This advantage is achieved through the collaborative workspace module of this embodiment and its built-in real-time communication server based on the WebSocket protocol. This technology broadcasts and synchronizes annotations and view control commands from multiple users in real time, ensuring that the actions of all group members are visible to each other, thus creating an interactive environment where collaborative analysis is essential. This fundamentally differs from the passive, single-person browsing or non-integrated general-purpose tool model of existing technologies, technically transforming students from information receivers into active explorers.
[0070] 5. Effectively cultivate comprehensive clinical thinking.
[0071] This advantage is achieved at the database level through the case learning module in this embodiment. This module establishes structured data associations, linking clinical case entries containing medical history and test results with one or more morphological image entries, and presenting them integrated on a single user interface. This differs from existing technologies that only provide isolated images; it technically forces the combination of text and image information, thereby guiding students to perform comprehensive analysis and effectively training their clinical reasoning abilities.
[0072] 6. It has achieved standardized and scalable expert-level real-time guidance.
[0073] This advantage is achieved through the AI-powered intelligent guidance module and its closed-loop data interaction mechanism with the collaborative workspace in this embodiment. This technique utilizes a CNN model trained with expert knowledge to replace some real-time guidance tasks that previously required human experts. Compared to existing technologies that rely on limited teacher resources for itinerant guidance, the AI-powered intelligent guidance module in this embodiment can provide high-quality, standardized guidance to countless learning groups simultaneously 24 / 7, solving the fundamental problems of scarce and unevenly distributed high-quality teaching resources and exhibiting extremely high scalability.
[0074] 7. Significantly improved the depth and accuracy of learning.
[0075] This advantage is achieved through the real-time analysis and feedback technology based on multi-person collaborative behavior in this embodiment. The AI module can instantly detect and correct collective misconceptions that may occur in group discussions, or prompt participants to pay attention to overlooked key information. This is fundamentally different from the learning model in existing technologies where students may "mislead each other" or "jointly miss" information. This embodiment technically creates a collaborative environment with an "expert safety net," ensuring that collaborative learning always develops in the right and in-depth direction.
[0076] Experimental Example 1 Label verification: When a student labels a cell and attempts to name it, the AI module receives the coordinates of the labeled area, calls the CNN model for recognition, and verifies the student's answer. If the student labels correctly, it is affirmed and supplemented with relevant knowledge; if incorrect, it provides corrective prompts: "The cell you labeled is more consistent with the characteristics of 'promyelocyte'. Please note the characteristics of its nucleus and cytoplasm." Question generation: The AI can automatically generate heuristic questions based on the cells currently being observed, such as: "Please note the nucleocytoplasmic ratio of this cell. What does this usually imply?" and push the questions to the group's chat box.
[0077] Data Flow: The overall data flow in this experimental example demonstrates the clear division of labor and efficient collaboration between modules. First, user interactions on the client-side UI are sent to the application server via a persistent WebSocket connection, transmitting standardized "multi-user collaborative interaction data packets." This application server, built on a Node.js environment using the Express.js framework and the Socket.IO library, acts as the "central hub" or "business logic layer" of the entire system. It manages all collaborative sessions, handles user permissions, and broadcasts received interaction data packets in real time to other users within the same session for synchronization.
[0078] Meanwhile, for interactive data requiring AI analysis, the application server calls the RESTful API of a separate AI inference server via HTTP. This AI inference server is a dedicated, high-performance computing service node built using Python and based on the TorchServe or TensorFlow Serving framework. Its core function is to load and run pre-trained CNN visual recognition models. The "algorithm" executed internally is the **forward propagation** process of the CNN model, receiving image data as input and outputting structured JSON results such as cell classification and coordinates.
[0079] After the AI inference server completes its calculations, it returns the analysis results to the application server via an API. Finally, based on the returned structured results, the application server calls the LLM API to generate introductory text and pushes the complete introductory information back to all users' client UIs via WebSocket.
[0080] Therefore, a complete intelligent guidance data flow loop is as follows: Client UI -> WebSocket -> Application Server (Node.js) -> API Call -> AI Inference Server (Python / TorchServe) -> API Response -> Application Server (Node.js) -> WebSocket -> Client UI. The application server, as the system's business logic hub, is responsible for collaborative session management, data forwarding, and interaction control with the AI module; the AI inference server, as a specialized computing node, undertakes high-performance computing tasks such as CNN model inference and quantitative image analysis. This layered architecture, composed of the application server and the AI inference server working together, ensures the real-time nature and context-relevance of guidance information while achieving high cohesion, low coupling, and good scalability in the system structure.
[0081] The schematic diagram of the effect of the artificial intelligence teaching system of the present invention in practical application is shown below. Figure 2-9 As shown, by Figure 2-9 The demonstration shows that the system can accommodate several users in the same virtual room. When one user performs view operations (such as zooming) or annotation operations, other users can also view and annotate synchronously. Users can also communicate and ask questions to the "AI teaching assistant" presented by the AI intelligent guidance module. The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0082] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. An artificial intelligence teaching system for guided learning, characterized in that, include: The system comprises a database module, a case study module, a collaborative workspace module, an AI-powered intelligent guidance module, and a user interface. The database module stores, manages, and retrieves image files. The case study module associates image files with real-world data to create teaching cases. The collaborative workspace module allows multiple system users to enter the same virtual room to synchronously, share, and collaboratively analyze the same image file within the teaching cases, outputting interactive data packets to the AI-powered intelligent guidance module and receiving and displaying its real-time feedback. The AI-powered intelligent guidance module analyzes the interactive data packets output by the collaborative workspace module in real time, generating and providing guidance information. The user interface is used to input user commands and display the guidance information.
2. The artificial intelligence teaching system according to claim 1, characterized in that, The database module includes a metadata database and a file server. The relational mapping mechanism of the database module includes: storing image files in the file server, generating identifiers, and storing the descriptive information, identifiers, and storage paths of the image files in metadata records, which are then stored in the metadata database.
3. The artificial intelligence teaching system according to claim 1, characterized in that, The image file and the real data are associated through a many-to-many relationship.
4. The artificial intelligence teaching system according to claim 1, wherein the synchronization, sharing, and collaborative analysis includes: Synchronized views, shared annotations, and integrated chat; The synchronized view includes: the system user's operation parameters for viewing the image file are broadcast in real time to the remaining system users in the same virtual room via the WebSocket protocol, triggering the client's application logic to update the view and synchronize the view; The shared annotation includes: the operation parameters of system users annotating image files are broadcast in real time to the remaining system users in the same virtual room via the WebSocket protocol, so that the annotations are displayed synchronously; The integrated chat includes: system users in the same virtual room discussing via a real-time chat box based on the WebSocket protocol.
5. The artificial intelligence teaching system according to claim 1, characterized in that, The process of outputting interactive data packets to the AI intelligent guidance module includes: when a system user views or annotates an image file, generating an interactive data packet, sending it to the application server via the WebSocket protocol, and forwarding it to the AI intelligent guidance module for real-time analysis; The process of receiving and displaying real-time feedback from the AI intelligent guidance module includes: the AI intelligent guidance module performing real-time analysis, generating guidance information, and having the application server feed it back to the collaborative workspace module via the WebSocket protocol for display.
6. The artificial intelligence teaching system according to claim 1, characterized in that, The AI intelligent guidance module includes: a multi-task deep learning visual analysis model, a parameter calculation algorithm, and a dialogue generation model; The multi-task deep learning visual analysis model includes: a convolutional neural network and multiple parallel decoding heads; the convolutional neural network is used to extract features and output them to the decoding heads, the decoding heads are used to generate a segmentation mask based on the features, and the segmentation mask extracts target parameters through a parameter calculation algorithm and passes them to the dialogue generation model; The dialogue generation model includes a unified language model interface layer, which is used to call the dialogue generation capability to generate guidance information based on the target parameters. The unified language model interface layer decouples the upper-layer AI intelligent guidance logic from the lower-layer language model.
7. The artificial intelligence teaching system according to claim 6, characterized in that, The AI intelligent guidance logic includes a group attention detection and key region deviation analysis algorithm. The execution of the group attention detection and key region deviation analysis algorithm includes the following steps: Pre-definition of key diagnostic regions, real-time aggregation of group attention, deviation analysis and triggering algorithm, generation of guidance information; The predefinition of the key diagnostic region includes: when the teaching case is stored in the metadata database, at least one region containing key features is identified as the key diagnostic region of the teaching case and stored in the metadata database; The real-time aggregation of group attention includes: receiving the interactive data packet, generating a dynamic two-dimensional array, simulating the formation of an attention heatmap, and quantifying the duration and frequency of observation of key diagnostic areas; The deviation analysis and triggering algorithm includes: calculating the cumulative attention obtained by each key diagnostic region, comparing it with the total attention of the entire image file in the teaching case to obtain the attention ratio of the key diagnostic region, and triggering internal instructions for the key diagnostic region according to the triggering rules; The generation of boot information includes: generating boot information according to the internal instructions.
8. A method for teaching and / or learning morphology, characterized in that, Teaching and / or learning are conducted using the artificial intelligence teaching system described in any one of claims 1-7; the real data includes: real-world research data.
9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by the processor, implements the method as described in claim 8.
10. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method of claim 8.