An interactive image report bidirectional generation terminal
Patent Information
- Application Number
- CN202610881856.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-25
AI Technical Summary
若系统在诊断段中自动补入未经确认的毫米级测量值、精细解剖定位、时间描述或复查建议,则容易形成临床上不可接受的幻觉式补全
[0037]相对于现有技术,本发明具有如下优点,本发明首先将影像报告生成从“整段替代”转化为“逐轮联动修订”,使医生始终能够围绕当前编辑内容控制另一段文本的更新范围,从而提高对报告内容的掌控能力;其次,本发明通过后台差分和前端修订渲染机制保留了原始文本与候选文本之间的清晰边界,使人机协作过程更符合临床审阅习惯;再次,本发明通过标准化对象、多模型约束和高风险片段屏蔽机制抑制文本幻觉扩散,减少未经确认的测量值、定位描述和复查建议被自动传播到另一段中的概率;进一步地,本发明将医生的真实交互行为沉淀为强化学习更新信号,使模型具有面向真实业务流程的持续优化能力;最后,本发明在 RK3588 终端上实现网页后端、本地状态服务和量化模型推理的统一部署,使系统能够在院内局域网、床旁便携设备和弱网环境中稳定运行,具有明确的实用价值、部署价值和推广意义。
Smart Images

Figure CN122822202A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a terminal, specifically an interactive image report bidirectional generation terminal, belonging to the technical fields of medical imaging, artificial intelligence, natural language processing, medical information systems, and edge computing terminals. Background Technology
[0002] The process of creating medical imaging reports is highly interactive and involves step-by-step verification. In clinical practice, doctors do not always write an entire report from scratch in one go. Instead, they often write a relatively clear part first, and then supplement, revise, and verify other parts. Existing automated text generation solutions typically abstract report writing into a single-round task of generating entire paragraphs, directly outputting the entire paragraph or even the whole report after obtaining some context. While such solutions can provide readable text, for doctors, there is often a lack of clear interactive boundaries regarding what content the system has modified, why it has modified it, and what content is merely a suggestion rather than a final conclusion. Therefore, it is difficult to seamlessly integrate with the actual clinical writing process.
[0003] The descriptive and diagnostic sections in imaging reports have a continuous bidirectional semantic constraint relationship. When doctors modify the descriptive section, they often want the system to provide synchronous updates only for the diagnostic section; when doctors first determine the diagnostic conclusion, they often need the system to supplement the matching descriptive text. Existing systems have low support for this bidirectional linkage, partial updates, and round-by-round confirmation, which easily leads to problems such as excessive rewriting of entire sections, repeated fluctuations of irrelevant content, and direct overwriting of doctors' existing text, resulting in insufficient system controllability.
[0004] In medical settings, the security of text generation is also a critical issue. If the system automatically inserts unverified millimeter-level measurements, detailed anatomical locations, time descriptions, or review suggestions into the diagnostic section, it can easily create clinically unacceptable hallucination-like completion. Meanwhile, application environments such as hospital intranets, bedside terminals, and portable workstations have a clear need for local deployment capabilities. Traditional inference methods relying on high-power servers or cloud graphics processors have limitations in terms of network stability, privacy protection, deployment costs, and device portability. Therefore, there is an urgent need for a complete product solution that is edge-deployable, text-interaction-centric, and balances generation quality, auditability, and local operability. Summary of the Invention
[0005] This invention addresses the technical problems existing in the prior art by providing an interactive bidirectional image report generation terminal. The method takes the image report text currently being edited by the doctor as the core processing object, and uses the descriptive and diagnostic sections in the image report as mutually constrained dual text state units. Through technologies such as session state modeling, bidirectional condition generation, differential revision rendering, hallucination suppression, interactive feedback learning, and edge deployment scheduling, an image report writing assistance product suitable for hospital environments is realized.
[0006] To achieve the above objectives, the technical solution of the present invention is as follows: an interactive image report bidirectional generation terminal, the terminal comprising a front-end interactive interface, a session state management module, a standardized object generation module, a bidirectional generation module, a differential calculation and revision rendering module, an illusion suppression module, a behavior recording and reinforcement learning update module, and an RK3588 end-side deployment module. The front-end interactive interface is used to support real-time editing of descriptive and diagnostic segments by doctors; the session state management module is used to maintain the confirmation text, candidate text, and historical trajectory of the current case in each round of interaction; the standardized object generation module is used to stably express the current examination context; the bidirectional generation module is used to perform linked generation between descriptive and diagnostic segments; the differential calculation and revision rendering module is used to convert the complete target segment text generated by the model into revision traces that can be reviewed by the front end; the illusion suppression module is used to block the propagation of high-risk segments; the behavior recording and reinforcement learning update module is used to convert the doctor's acceptance and rejection behaviors into subsequent version update signals; the RK3588 end-side deployment module is used to complete the joint operation of web page backend, local state service, and large model inference in a single portable device.
[0007] In the session state management module, the writing session is represented as a continuously evolving dual-text state, denoted as the first... Wheel interaction state is
[0008]
[0009] in, Indicates the first Wheel of interaction status, and These represent the descriptive text and diagnostic text that have been confirmed by the doctors in the current round, respectively. and These represent the candidate descriptive text and candidate diagnostic text generated by the model, respectively. Indicates the direction of generation in this round. Indicates conversation history and interaction patterns; superscript and These represent the descriptive and diagnostic sections, respectively. When a doctor edits the descriptive text, the system instructs... When a doctor edits a diagnostic text, the system instructs... Therefore, the system operates according to the logic of "modifying one side and linking the other side" in any round, so that the description segment and the diagnostic segment are always in a dynamic state of mutual constraint.
[0010] To improve the consistency of generation conditions, this invention calls multiple dedicated language models to construct standardized objects before each round of generation. The standardized object generation module includes BERT for inspection location standardization, BERT for inspection modality standardization, BERT for enhancement scheme standardization, and BERT for template indexing. Each sub-model extracts constraint information directly related to report generation from the current inspection context text and generates a unified object through a fusion function. ,Right now
[0011]
[0012] in, Indicates the first Standardized objects, Indicates the current inspection context text. , , and These represent the BERT encoders for the inspection location, inspection modality, enhancement scheme, and template index normalization subtasks, respectively. Indicates the fusion function; subscript , , , These correspond to parts, modalities, enhancements, and templates, respectively. Through this mechanism, the system maps freely written inspection contexts into stable generative condition representations, thereby providing a unified semantic anchor for subsequent bidirectional generation.
[0013] In the generation phase, a decoder-based large language model is used as the backbone network for text generation. The backbone network has a causal decoding structure, and its initial hidden state satisfies...
[0014]
[0015] in, This represents the initial hidden state of the decoder's input layer. Represents word embedding mapping, Represents the positional encoding mapping. Current round generation context. It consists of standardized objects, the current round source segment text, the previous round target segment text, session history, and security constraint information.
[0016]
[0017] in, This indicates the context assembly function. This indicates the source text that the doctors actually modified in this round. This refers to the target text from the previous round. Indicates a security constraint mask; superscript and These represent the source segment and the target segment in this round, respectively, with their specific values determined by the generation direction. Decision. Regarding the first... Layer One point of attention, there is
[0018]
[0019] in, and They represent the first Layer and first Hidden state, , , They represent the first Layer The query, key, and value representation of each attention head. , , These represent the corresponding linear projection matrices, This indicates the output of the attention head. This represents the attention mask formed by the causal attention mask and the security constraints. Represents the dimension of the key vector. Indicates the total number of attention heads. Indicates the first Multi-head attention output projection matrix, This indicates a splicing operation. Indicates a feedforward network. Let represent the normalized exponential function. The generation probability of the target segment candidate text satisfies ...
[0020]
[0021] System retrieve
[0022]
[0023] This is the result returned in this round. The parameter is The conditional probabilities given by the generative model. Represents the candidate target segment word sequence. This indicates the length of the sequence. Indicates the first Each word element, Indicates the first The prefix sequence preceding each lexical unit This means selecting the sequence with the highest conditional probability from all candidate sequences. Indicates the first The model generates candidate text for the target segment. It's important to emphasize that the model returns the complete target segment text generated based on the current session context, not directly outputting revision markers; revision marks are calculated separately by the backend after the model output. If the doctor confirms the target segment text in the current round as... The model's candidate text is Then the system executes
[0024]
[0025] in, Indicates the first The difference between the target segment confirmation text and the model candidate text is calculated. This is a text difference operator. The backend is based on... Insertion and deletion marks are generated and sent to the front-end editor for rendering. What doctors see on the interface are the revision marks calculated by the system, so they can still retain final control over the original text.
[0026] To further improve response efficiency and text stability in edge scenarios, this invention introduces a local rewrite priority strategy in the bidirectional generation module process. The system calculates the editing ratio based on the scope of modifications made by the doctor in the current round.
[0027]
[0028] in, Indicates the first Rotate editing ratio, This represents the set of valid edited segments in this round. Indicates the total length of the effective edited segment. Indicates the length of the source text segment. This is used to avoid the denominator being zero when the source segment is empty. When the semantic anchor point is located, the system constructs a local rewrite context for the target segment, generating only the segments directly related to the current change; when Then, the system switches to full-segment generation. This strategy reduces irrelevant fluctuations in small-scale editing scenarios and lowers the overhead of repetitive computations on edge devices.
[0029] The hallucination suppression module includes a hallucination suppression and non-propagating fragment masking mechanism to suppress the spread of errors between the descriptive and diagnostic segments. The system maintains a set of high-risk fragments based on standardized objects and contextual constraints. The system's content includes unconfirmed measurements, detailed anatomical locations, time-related information, and follow-up recommendations. It uses contextual consistency rules to constrain and screen the generated results. If a high-risk segment appears in the model's candidate text that is inconsistent with the current round of evidence, the system marks it as non-propagable content and presents it as high-risk during the revision rendering stage, while simultaneously blocking its direct writing. This mechanism helps the system avoid several typical errors, including automatically adding unconfirmed millimeter-level values to the diagnosis section, automatically rewriting left and right side information in another section, and automatically adding follow-up periods and recommendations when doctor confirmation is lacking. The system can also set a confirmation content locking mechanism. For measurements, left and right side locations, key diagnostic conclusions, and follow-up recommendations already confirmed by the doctor, the system records the confirmation identifier and source text segment in the session state. During subsequent generation, the confirmed content participates in context assembly as a strong constraint, and the model cannot arbitrarily rewrite it without receiving further editing instructions from the doctor. For newly generated content from the model that has not yet been confirmed by doctors, the system marks it as a candidate, allowing it to be displayed only as a suggestion on the front end. It is not permitted to use it as confirmed evidence to drive the generation of another text segment. This prevents candidate content from being mistakenly disseminated as factual evidence in multiple rounds of interaction.
[0030] In the behavior recording and reinforcement learning update module, the system records behavioral signals such as acceptance, rejection, deletion, dwell time, and high-risk withdrawal during the doctor's review and revision process, and converts them into reward signals required for subsequent model version updates. Let the first... Round of interaction rewards are
[0031]
[0032] in, Indicates the first Round of interactive rewards, Indicates the strength of the revision adoption. Indicates the retention strength of candidate texts in the final report. This indicates the editing workload incurred by the doctor in correcting the output of that round. Indicates the trigger strength of high-risk segments. , , and These represent the non-negative weight coefficients of the four types of behavioral signals mentioned above. Based on the reward signals, the system constructs offline reinforcement learning samples, enabling subsequent deployment versions to gradually shift towards generating outputs that are more easily accepted by doctors, have a lower editing burden, and are more secure, thereby forming a truly closed-loop optimization capability oriented towards clinical use scenarios.
[0033] In the RK3588 edge deployment module, the terminal uses the RK3588 chip as the main controller, and the computing unit consists of an NPU and a CPU. Web page backend services, session state management, standardized object generation, text difference calculation, and large model inference are all deployed within the same terminal, without relying on a dedicated GPU. The terminal loads a lightweight image reporting model that has been distilled, pruned, and quantized. The model parameter size is controlled to the order of 0.5B, using grouped symmetric 4-bit weight quantization and 8-bit activation representation. The backbone decoding layer is mapped to the NPU for execution, while the CPU is responsible for word segmentation, state machine scheduling, BERT standardized model invocation, difference post-processing, and web page service response. For the autoregressive inference process, the terminal reuses the key-value cache corresponding to the unchanged context of the previous round, and its update relationship satisfies...
[0034]
[0035] in, and They represent the first Reusable key and value caches and These square brackets represent the key cache blocks and value cache blocks that have been added or changed in this round, respectively. This indicates that the sequence is spliced along the sequence dimension, thereby reducing redundant calculations. The terminal also employs a template-first, then-complete short decoding strategy. First, the structural skeleton and fixed sentence structure of the target segment are determined by a standardized object. Then, short sequence completion is performed only on missing segments to shorten the average generation length. To adapt to long-term operating conditions, the system monitors the chip temperature in real time. ,when At that time, the maximum generation length per round was reduced from 192 words to 96 words, while maintaining the local rewrite priority strategy to control power consumption and heat generation.
[0036] The above solution features a complete productized technical solution built around the image report writing process, a dual-text state modeling approach with description and diagnosis segments as core objects, and an interactive bidirectional generation mechanism that automatically switches the generation direction based on the current editing segment. This solution employs a standardized object generation mechanism composed of multiple BERT sub-models, which maps the check context to stable generation constraints and template anchors. The solution utilizes a collaborative link where "the model returns the complete target segment text, the system calculates the difference in the background, and the front-end renders the revision traces," rather than directly using the model as a difference output device. The solution adopts a local rewrite priority strategy and a corresponding high-risk segment masking mechanism, enabling the system to suppress hallucination propagation during text-linked generation. It also features a closed-loop optimization method that transforms doctor acceptance, rejection, deletion, and risk withdrawal behaviors into reinforcement learning update signals. The solution is implemented in a hardware and software integrated manner on the RK3588 terminal, including a 0.5B lightweight model, 4-bit weight quantization, and KV cache. The interactive image reporting system, terminal equipment, and corresponding software implementation process comprised of the above modules: reuse, template-first-fill-later short decoding strategy, co-deployment of web page backend and model inference, and generation length control mechanism based on temperature and power status.
[0037] Compared to existing technologies, this invention has the following advantages: First, it transforms image report generation from "whole-segment replacement" to "round-by-round linked revision," enabling doctors to always control the update scope of another segment of text around the currently edited content, thereby improving their control over the report content. Second, through background differential and front-end revision rendering mechanisms, this invention preserves clear boundaries between the original text and candidate text, making the human-computer collaboration process more in line with clinical review habits. Third, through standardized objects, multi-model constraints, and high-risk segment masking mechanisms, this invention suppresses the spread of text illusions, reducing the probability of unconfirmed measurements, location descriptions, and review suggestions being automatically propagated to another segment. Furthermore, this invention precipitates doctors' real interactive behaviors as reinforcement learning update signals, enabling the model to have continuous optimization capabilities oriented towards real business processes. Finally, this invention achieves unified deployment of web backend, local state service, and quantitative model inference on the RK3588 terminal, enabling the system to operate stably in hospital LANs, bedside portable devices, and weak network environments, demonstrating clear practical value, deployment value, and promotional significance. Attached Figure Description
[0038] Figure 1 This is a schematic diagram of the overall system structure.
[0039] Figure 2 The flowchart is for the core data objects and state machine.
[0040] Figure 3 Here is the flowchart for the bidirectional generation method.
[0041] Figure 4 Diagram of the deployment structure for the RK3588 terminal. Detailed Implementation
[0042] To enhance understanding of the present invention, the embodiments will be described in detail below with reference to the accompanying drawings.
[0043] Example 1: See Figures 1-4 An interactive image report bidirectional generation terminal is disclosed. The terminal comprises a front-end interactive interface, a session state management module, a standardized object generation module, a bidirectional generation module, a differential calculation and revision rendering module, an illusion suppression module, a behavior recording and reinforcement learning update module, and an RK3588 edge deployment module. The front-end interactive interface is used for doctors to edit descriptive and diagnostic segments in real time. The session state management module maintains the confirmation text, candidate text, and historical trajectory of the current case in each round of interaction. The standardized object generation module provides a stable representation of the current examination context. The bidirectional generation module performs linked generation between descriptive and diagnostic segments. The differential calculation and revision rendering module converts the complete target segment text generated by the model into revision traces that can be reviewed by the front end. The illusion suppression module blocks the propagation of high-risk segments. The behavior recording and reinforcement learning update module converts the doctor's acceptance and rejection behaviors into subsequent version update signals. The RK3588 edge deployment module enables the joint operation of a web backend, local state service, and large model inference within a single portable device.
[0044] In the session state management module, the writing session is represented as a continuously evolving dual-text state, denoted as the first... Wheel interaction state is
[0045]
[0046] in, and These represent the descriptive text and diagnostic text that have been confirmed by the doctors in the current round, respectively. and These represent the candidate descriptive text and candidate diagnostic text generated by the model, respectively. Indicates the direction of generation in this round. The system displays the session history and interaction trajectory, and when a doctor edits descriptive text, the system instructs... When a doctor edits a diagnostic text, the system instructs... Therefore, the system operates according to the logic of "modifying one side and linking the other side" in any round, so that the description segment and the diagnostic segment are always in a dynamic state of mutual constraint.
[0047] To improve the consistency of generation conditions, this invention calls multiple dedicated language models to construct standardized objects before each round of generation. The standardized object generation module includes BERT for inspection location standardization, BERT for inspection modality standardization, BERT for enhancement scheme standardization, and BERT for template indexing. Each sub-model extracts constraint information directly related to report generation from the current inspection context text and generates a unified object through a fusion function. ,Right now
[0048]
[0049] in, Indicates the current inspection context text. , , and BERT encoders corresponding to different standardized subtasks, The fusion function represents the mechanism described above, which maps the freely written inspection context into a stable representation of the generation conditions, thereby providing a unified semantic anchor for subsequent bidirectional generation.
[0050] In the generation phase, a decoder-based large language model is used as the backbone network for text generation. The backbone network has a causal decoding structure, and its initial hidden state satisfies...
[0051]
[0052] in, Represents word embedding mapping, Indicates positional encoding, current round generation context It consists of standardized objects, the current round source segment text, the previous round target segment text, session history, and security constraint information.
[0053]
[0054] in, This indicates the source text that the doctors actually modified in this round. This refers to the target text from the previous round. Represents the security constraint mask, for the first Layer One point of attention, there is
[0055]
[0056] The generation probability of the target segment candidate text satisfies
[0057]
[0058] System retrieve
[0059]
[0060] As for the results returned in this round, it's important to emphasize that the model returns the complete target segment text generated based on the current session context, not directly outputting revision marks; revision marks are calculated separately by the backend after the model outputs, if the target segment text confirmed by the doctor in this round is... The model's candidate text is Then the system executes
[0061]
[0062] in For text difference operators, the backend is based on Insertion and deletion marks are generated and sent to the front-end editor for rendering. What doctors see on the interface are the revision marks calculated by the system, so they can still retain final control over the original text.
[0063] To further improve response efficiency and text stability in edge scenarios, this invention introduces a local rewrite priority strategy in the bidirectional generation module process. The system calculates the editing ratio based on the scope of modifications made by the doctor in the current round.
[0064]
[0065] in, This represents the set of valid edited segments in this round. Indicates the length of the source text segment, when When the semantic anchor point is located, the system constructs a local rewrite context for the target segment, generating only the segments directly related to the current change; when Then, the system switches to full-segment generation. This strategy reduces irrelevant fluctuations in small-scale editing scenarios and lowers the overhead of repetitive computations on edge devices.
[0066] The hallucination suppression module incorporates a hallucination suppression and non-propagable fragment masking mechanism to suppress the propagation of errors between descriptive and diagnostic segments. Based on standardized objects and contextual constraints, the system maintains a set of high-risk fragments, including unconfirmed measurements, detailed anatomical locations, temporal expressions, and follow-up recommendation text units. The generated results are constrained and screened according to contextual consistency rules. If a high-risk fragment appears in the model's candidate text that is inconsistent with the current round of evidence conditions, the system marks that fragment as non-propagable content and presents it as high-risk during the revision rendering stage, while simultaneously blocking its direct linkage. This mechanism helps the system avoid several typical errors, including automatically adding unconfirmed millimeter-level values to the diagnostic segment, automatically rewriting left and right side information in another segment, and automatically adding follow-up periods and recommendations when physician confirmation is lacking. The system can also set up a confirmed content locking mechanism. For measurements, left and right side locations, key diagnostic conclusions, and follow-up recommendations already confirmed by the doctor, the system records the confirmation identifier and source text segment in the session state. During subsequent generation, the confirmed content serves as a strong constraint in the context assembly, and the model cannot arbitrarily rewrite it without further editing instructions from the doctor. For newly generated content that has not yet been confirmed by the doctor, the system marks it as a candidate state, allowing it only to be displayed as a suggestion on the front end, and prohibiting it from being used as confirmed evidence to drive the generation of another text segment. This prevents candidate content from being mistakenly disseminated as factual evidence repeatedly in multiple rounds of interaction.
[0067] In the behavior recording and reinforcement learning update module, the system records behavioral signals such as acceptance, rejection, deletion, dwell time, and high-risk withdrawal during the doctor's review and revision process, and converts them into reward signals required for subsequent model version updates. Let the first... Round of interaction rewards are
[0068]
[0069] in, Indicates the strength of the revision adoption. Indicates the retention strength of candidate texts in the final report. This indicates the editing workload incurred by the doctor in correcting the output of that round. This indicates the trigger strength of high-risk fragments. Based on the above reward signals, the system constructs offline reinforcement learning samples, enabling subsequent deployment versions to gradually shift towards generating outputs that are more easily accepted by doctors, have a lower editing burden, and are more secure, thereby forming a closed-loop optimization capability truly geared towards clinical use scenarios.
[0070] In the RK3588 edge deployment module, the terminal uses the RK3588 chip as the main controller, and the computing unit consists of an NPU and a CPU. Web page backend services, session state management, standardized object generation, text difference calculation, and large model inference are all deployed within the same terminal, without relying on a dedicated GPU. The terminal loads a lightweight image reporting model that has been distilled, pruned, and quantized. The model parameter size is controlled to the order of 0.5B, using grouped symmetric 4-bit weight quantization and 8-bit activation representation. The backbone decoding layer is mapped to the NPU for execution, while the CPU is responsible for word segmentation, state machine scheduling, BERT standardized model invocation, difference post-processing, and web page service response. For the autoregressive inference process, the terminal reuses the key-value cache corresponding to the unchanged context of the previous round, and its update relationship satisfies...
[0071]
[0072] This reduces redundant calculations. The terminal also employs a template-first, then-complete short decoding strategy. First, the structural skeleton and fixed sentence structure of the target segment are determined by a standardized object. Then, short-sequence completion is performed only on missing segments to shorten the average generation length. To adapt to the long-term operating conditions of portable devices, the system monitors the chip temperature in real time. and terminal power ,when or At that time, the maximum generation length per round was reduced from 192 words to 96 words, while maintaining the local rewrite priority strategy to control power consumption and heat generation.
[0073] It should be noted that the above embodiments are not intended to limit the scope of protection of the present invention. Equivalent transformations or substitutions made based on the above technical solutions all fall within the scope of protection of the claims of the present invention.
Claims
1. An interactive image report bidirectional generation terminal, characterized in that, The terminal consists of a front-end interactive interface, a session state management module, a standardized object generation module, a bidirectional generation module, a differential calculation and revision rendering module, an illusion suppression module, a behavior recording and reinforcement learning update module, and an RK3588 edge deployment module. The front-end interactive interface is used for doctors to edit descriptive and diagnostic segments in real time; the session state management module is used to maintain the confirmation text, candidate text, and historical trajectory of the current case in each round of interaction; the standardized object generation module is used to stably express the current examination context; the bidirectional generation module is used to perform linked generation between descriptive and diagnostic segments; the differential calculation and revision rendering module is used to convert the complete target segment text generated by the model into revision traces that can be reviewed by the front end; the illusion suppression module is used to block the propagation of high-risk segments; the behavior recording and reinforcement learning update module is used to convert doctors' acceptance and rejection behaviors into subsequent version update signals; and the RK3588 edge deployment module is used to complete the joint operation of web backend, local state service, and large model inference in a single portable device.
2. The interactive image report bidirectional generation terminal according to claim 1, characterized in that, In the session state management module, the writing session is represented as a continuously evolving dual-text state, let the first state be... Wheel interaction state is in, Indicates the first Wheel of interaction status, and These represent the descriptive text and diagnostic text that have been confirmed by the doctors in the current round, respectively. and These represent the candidate descriptive text and candidate diagnostic text generated by the model, respectively. Indicates the direction of generation in this round. Indicates conversation history and interaction patterns; superscript and These represent the descriptive and diagnostic sections, respectively. When a doctor edits the descriptive text, the system instructs... When a doctor edits a diagnostic text, the system instructs... .
3. The interactive image report bidirectional generation terminal according to claim 2, characterized in that, The standardized object generation module includes BERT for inspection location standardization, BERT for inspection modality standardization, BERT for enhancement scheme standardization, and BERT for template indexing. Each sub-model extracts constraint information directly related to report generation from the current inspection context text and generates a unified object through a fusion function. ,Right now in, Indicates the first Standardized objects, Indicates the current inspection context text. , , and These represent the BERT encoders for the inspection location, inspection modality, enhancement scheme, and template index normalization subtasks, respectively. Indicates the fusion function; subscript , , , These correspond to parts, modalities, enhancements, and templates, respectively. Through this mechanism, the system maps freely written inspection contexts into stable generation condition representations, thereby providing a unified semantic anchor for subsequent bidirectional generation. In the generation phase, a decoder-based large language model is used as the backbone network for text generation. The backbone network has a causal decoding structure, and its initial hidden state satisfies... in, This represents the initial hidden state of the decoder's input layer. Represents word embedding mapping, Represents the positional encoding mapping, the current round's generation context. It consists of standardized objects, the current round source segment text, the previous round target segment text, session history, and security constraint information. in, This indicates the context assembly function. This indicates the source text that the doctors actually modified in this round. This refers to the target text from the previous round. Indicates a security constraint mask; superscript and These represent the source segment and the target segment in this round, respectively, with their specific values determined by the generation direction. Decision. Regarding the first... Layer One point of attention, there is in, and They represent the first Layer and first Hidden state, , , They represent the first Layer The query, key, and value representation of each attention head. , , These represent the corresponding linear projection matrices, This indicates the output of the attention head. This represents the attention mask formed by the combination of causal attention mask and security constraints. Represents the dimension of the key vector. Indicates the total number of attention heads. Indicates the first Multi-head attention output projection matrix, This indicates a splicing operation. Indicates a feedforward network. Let represent the normalized exponential function. The generation probability of the target segment candidate text satisfies ... System retrieve As part of the results returned in this round, The parameter is The conditional probabilities given by the generative model. Represents the candidate target segment word sequence. This indicates the length of the sequence. Indicates the first Each word element, Indicates the first The prefix sequence preceding each lexical unit This indicates that the sequence with the highest conditional probability is selected from all candidate sequences. Indicates the first The target segment candidate text generated by the round model is returned as a complete target segment text based on the current session context, rather than directly outputting revision marks. Revision traces are calculated separately by the backend after the model output; if the doctor confirms the target segment text in the current round is... The candidate text for the model is Then the system executes in, Indicates the first The difference between the target segment confirmation text and the model candidate text is calculated. For text difference operators, the backend is based on Insertion and deletion marks are generated and sent to the front-end editor for rendering. What doctors see on the interface are the revision marks calculated by the system, but they still retain final control over the original text.
4. The interactive image report bidirectional generation terminal according to claim 3, characterized in that, A local rewrite priority strategy is introduced in the bidirectional generation module process. The system calculates the editing ratio based on the scope of modifications made by the doctor in the current round. in, Indicates the first Rotate editing ratio, This represents the set of valid edited segments in this round. Indicates the total length of the effective edited segment. Indicates the length of the source text segment. This is used to avoid the denominator being zero when the source segment is empty. When the semantic anchor point is located, the system constructs a local rewrite context for the target segment, generating only the segments directly related to the current change; when Then, the system switches to full-segment generation. This strategy reduces irrelevant fluctuations in small-scale editing scenarios and lowers the overhead of repetitive computations on edge devices.
5. The interactive image report bidirectional generation terminal according to claim 2, characterized in that, The hallucination suppression module includes a hallucination suppression and non-propagating fragment masking mechanism to suppress the spread of errors between descriptive and diagnostic segments. The system maintains a set of high-risk fragments based on standardized objects and contextual constraints. The content includes unconfirmed measurement values, detailed anatomical locations, time expressions, and text units for follow-up recommendations. The generated results are constrained and screened according to contextual consistency rules. If a high-risk segment appears in the model's candidate text that is inconsistent with the current round of evidence conditions, the system marks that segment as non-automatically propagable content and presents it as high-risk during the revision rendering stage, while simultaneously blocking its direct linkage and writing. This mechanism helps the system avoid several typical errors, including automatically adding unconfirmed millimeter-level values to the diagnosis segment, automatically rewriting left and right side information in another segment, and automatically adding follow-up periods and recommendations when doctor confirmation is lacking. The system can also set a confirmation content locking mechanism. For measurement values, left and right side locations, key diagnostic conclusions, and follow-up recommendations already confirmed by the doctor, the system records their confirmation identifier and source text segment in the session state. During subsequent generation, confirmed content participates in context assembly as a strong constraint condition, and the model cannot arbitrarily rewrite it without receiving further editing instructions from the doctor. For newly generated content that has not yet been confirmed by the doctor, the system marks it as a candidate state, only allowing it to be displayed as a suggestion on the front end, and not allowing it to continue driving the generation of another text segment as confirmed evidence. This prevents candidate content from being mistakenly presented as factual evidence and repeatedly disseminated during multiple rounds of interaction.
6. The interactive image report bidirectional generation terminal according to claim 2, characterized in that, In the behavior recording and reinforcement learning update module, the system records behavioral signals such as acceptance, rejection, deletion, dwell time, and high-risk withdrawal during the doctor's review and revision process, and converts them into reward signals required for subsequent model version updates. Let the first... Round of interaction rewards are in, Indicates the first Round of interactive rewards, Indicates the strength of the revision adoption. Indicates the retention strength of candidate texts in the final report. This indicates the editing workload incurred by the doctor in correcting the output of that round. Indicates the trigger strength of high-risk segments. , , and These represent the non-negative weight coefficients of the four types of behavioral signals mentioned above. Based on the reward signals, the system constructs offline reinforcement learning samples, enabling subsequent deployment versions to gradually shift towards generating outputs that are more easily accepted by doctors, have a lower editing burden, and are more secure, thereby forming a truly closed-loop optimization capability oriented towards clinical use scenarios.
7. The interactive image report bidirectional generation terminal according to claim 2, characterized in that, In the RK3588 edge deployment module, the terminal uses the RK3588 chip as the main controller, and the computing unit consists of an NPU and a CPU. Web page backend services, session state management, standardized object generation, text difference calculation, and large model inference are all deployed within the same terminal, without relying on a dedicated GPU. The terminal loads a lightweight image reporting model that has been distilled, pruned, and quantized, with model parameters controlled at the 0.5B level. It uses grouped symmetric 4-bit weight quantization and 8-bit activation representation, mapping the backbone decoding layer to the NPU for execution. The CPU is responsible for word segmentation, state machine scheduling, BERT standardized model invocation, difference post-processing, and web page service response. For the autoregressive inference process, the terminal reuses the key-value cache corresponding to the unchanged context of the previous round, and its update relationship satisfies... in, and They represent the first Reusable key and value caches and These square brackets represent the key cache blocks and value cache blocks that have been added or changed in this round, respectively. This indicates that the sequence is spliced along the sequence dimension, thereby reducing redundant calculations. The terminal also adopts a template-first, then-padding short decoding strategy. First, the structural skeleton and fixed sentence pattern of the target segment are determined by the standardized object. Then, short sequence completion is performed only on the missing segments to shorten the average generation length. To adapt to the long-term operating conditions of portable devices, the system monitors the chip temperature in real time. ,when At that time, the maximum generation length per round was reduced from 192 words to 96 words, while maintaining the local rewrite priority strategy to control power consumption and heat generation.