Image Recognition-Based Popular Science Display and Interactive System for Radiotherapy Equipment
By improving the loss function, search strategy, and feature extraction algorithm, the accuracy and stability issues in radiotherapy equipment identification and popular science interaction were resolved, achieving high-precision image recognition and interactive display, thus improving the popular science effect and user experience of radiotherapy equipment.
Patent Information
- Application Number
- CN202511096694.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Existing technologies struggle to achieve high-precision, stable, and interactive image recognition in scenarios involving radiotherapy equipment identification and science popularization. In particular, when equipment structures are complex, obstructed, angles change, or boundaries are blurred, traditional detection algorithms struggle to accurately locate and identify targets, and are insensitive to small targets, affecting model convergence.
By employing improved loss function design, search strategy optimization, feature extraction structure, and target selection algorithm, including PWIoU loss function, dynamic anchor box generation, dense connection structure, residual skip connection mechanism, and improved clustering-nonmaximum suppression algorithm, combined with multimodal display and user profile adaptation, a high-precision image recognition and intelligent interaction system is constructed.
It significantly improves the image recognition accuracy and robustness of radiotherapy equipment, making it suitable for scenarios such as medical science popularization exhibition halls and virtual exhibition halls. It provides an immersive communication experience and assists in doctor-patient communication, enhancing the public's awareness and understanding of radiotherapy equipment.
Smart Images

Figure CN120599313B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical science popularization and display technology, and in particular to an interactive system for popularizing and displaying radiotherapy equipment based on image recognition. Background Technology
[0002] With the development of modern medical technology, radiotherapy has become one of the core methods in cancer treatment, with representative equipment including high-precision devices such as proton therapy machines, linear accelerators, and CyberKnife. Although these devices are widely used in the medical field, there is still a significant gap in public understanding of their structural principles, working methods, and technological advantages, which affects the efficiency of science popularization and doctor-patient communication.
[0003] In recent years, image recognition and deep learning technologies have achieved remarkable results in medical image analysis, particularly in target detection and image semantic understanding. However, directly applying these technologies to the identification of radiotherapy equipment and in educational interactive scenarios still faces the following technical challenges:
[0004] The complex and diverse structure of the equipment, along with the occlusion, angle changes, and blurred boundaries often present in the images, make it difficult for traditional detection algorithms to accurately locate and identify targets. The proportion of targets in radiotherapy equipment is often small, and traditional IoU loss is not sensitive to small targets during training, resulting in gradient instability and affecting model convergence. The recognition task not only needs to classify accurately but also needs to drive the subsequent display of multimedia interactive content, which places higher demands on both detection accuracy and efficiency.
[0005] Therefore, there is an urgent need to build an integrated solution for image recognition and popular science display that combines accuracy, stability, and interactive capabilities for radiotherapy equipment identification scenarios. Summary of the Invention
[0006] This invention aims to address the needs of image recognition and popular science display for radiotherapy equipment by proposing a complete and innovative algorithm framework. Through comprehensive improvements in loss function design, search strategy optimization, feature extraction structure innovation, and target selection algorithm, this framework successfully constructs a high-precision image recognition and intelligent interactive display system for radiotherapy equipment. It possesses good versatility, robustness, and practical value, and is applicable to various scenarios such as medical popular science exhibition halls, virtual exhibition halls, and intelligent guide systems.
[0007] To achieve the above objectives, the following technical solution is adopted:
[0008] An image recognition-based interactive system for popularizing and demonstrating radiotherapy equipment includes:
[0009] The data acquisition module is used to collect physical images, 3D models and technical documents of radiotherapy equipment, and to build an image database containing equipment type tags and a knowledge base containing multi-dimensional popular science content.
[0010] The image recognition module is used to input the acquired images into the deep learning model. Through neural networks combined with improved feature extraction and target localization algorithms, it extracts the key semantic features of radiotherapy equipment in the image, and retains the optimal bounding box and removes redundancy based on the improved clustering-nonmaximum suppression algorithm, so as to achieve accurate identification and category judgment of radiotherapy equipment.
[0011] The science popularization and interaction module matches knowledge base content based on the recognition results, displays device information in a multimodal format, and adapts the content depth based on user profiles;
[0012] The deployment integration module supports deployment on online platforms and offline terminals, and triggers recognition and interaction functions through image scanning or uploading.
[0013] Furthermore, the image recognition module includes:
[0014] The target localization unit adopts a dynamic anchor frame generation mechanism and optimizes localization accuracy by fusing F1 scores and dynamic difficulty-aware loss functions.
[0015] The feature extraction unit employs a multi-scale feature extraction network with a dense connection structure and a residual skip connection mechanism, combined with sub-pixel convolutional upsampling to preserve image details;
[0016] The result filtering unit, based on the improved clustering-nonmaximum suppression algorithm, removes redundant prediction boxes through clustering and confidence penalty mechanisms, and outputs the final recognition result.
[0017] Furthermore, the target positioning unit includes:
[0018] The classification target calculation sub-unit uses the F1 score as the classification optimization target to balance precision and recall and solve the problem of positive and negative sample imbalance.
[0019] The regression target calculation subunit uses a dynamic difficulty-aware loss function as the regression optimization target. It dynamically adjusts the gradient weights of the bounding box regression based on the sample anomaly degree and monotonic focusing coefficient to enhance the learning ability for small targets and blurred edge samples.
[0020] The dual-objective fusion mechanism weights and superimposes the F1 score and the dynamic difficulty-aware loss function to form a multi-objective joint loss function; gradient backpropagation is performed based on the joint loss function to simultaneously update the parameters of the classification network and the regression network, thereby achieving collaborative optimization of classification results and localization results.
[0021] The anchor frame optimization strategy uses a nonlinear convergence factor to dynamically adjust the search range, performing global exploration in the early stage of iteration and local fine-tuning convergence in the later stage.
[0022] The anchor box generation mechanism predicts the offset by constraining the Sigmoid function and dynamically outputs the device bounding box coordinates.
[0023] Furthermore, the dynamic difficulty-aware loss function dynamically adjusts the gradient weights of the bounding box regression based on the sample anomaly degree and the monotonic focusing coefficient, including:
[0024] Calculate the anomaly β of the current sample, which is the ratio of the IoU loss value of the current sample to the average IoU loss value;
[0025] When β>1, it is identified as a difficult sample and the gradient gain is amplified to enhance the learning of small targets and blurry edge samples.
[0026] When β < 1, it is considered a simple sample and the gradient gain is suppressed to prevent overfitting;
[0027] The gradient gain of difficult samples is exponentially amplified by a monotonic focusing coefficient γ.
[0028] Furthermore, the feature extraction unit includes:
[0029] The data augmentation unit performs rotation, scaling, and copy-paste processing on the input image;
[0030] The encoding-decoding structure extracts multi-scale semantic features in the downsampling stage and restores spatial resolution through sub-pixel convolution in the upsampling stage.
[0031] The residual connection channel overlays shallow features with the reconstructed feature map to prevent gradient vanishing.
[0032] Furthermore, the result filtering unit includes:
[0033] The overlap relationship between predicted bounding boxes is calculated based on the intersection-over-union matrix;
[0034] An iterative suppression mechanism eliminates redundant boxes that overlap with high-confidence boxes round by round using a mask matrix;
[0035] The confidence penalty mechanism uses an exponential decay function to reduce the false suppression rate of occluded targets.
[0036] Furthermore, the science popularization interactive module includes:
[0037] The multi-dimensional knowledge display unit provides equipment structure and principles, treatment process animations, technology comparison charts, and AR scene fusion demonstrations.
[0038] The rich media output unit supports video playback, simultaneous output of text, images and voice, and natural language question-and-answer interaction.
[0039] The user profile adaptation unit differentiates between teenagers, patients, and medical staff, and dynamically adjusts the content presentation format and depth.
[0040] Furthermore, the AR scene fusion demonstration specifically includes:
[0041] Image anchor point alignment technology based on feature point matching or plane detection is used to stably overlay 3D models or dynamic path animations onto the identified device images or real-world scenes.
[0042] For proton therapy devices, the trajectory of the beam from the accelerator to the tumor target area is visualized.
[0043] Furthermore, the deployment integration module includes:
[0044] Offline deployment involves configuring touch terminals and AR devices in physical exhibition halls to support QR code scanning and recognition of physical exhibits;
[0045] The online integration method provides image upload and recognition services through WeChat mini programs and virtual exhibition halls.
[0046] Furthermore, the radiotherapy equipment includes a proton therapy device, a linear accelerator, a CyberKnife, and a Gamma Knife.
[0047] Compared with the prior art, the present invention achieves the following beneficial effects:
[0048] 1. This invention proposes the PWIoU loss function: Based on the traditional IoU, a "dynamic difficulty-aware mechanism" is introduced. Through the combined effect of "anomaly degree" and "monotonically focused coefficient", the gradient contribution is dynamically adjusted, which significantly enhances the model's learning ability for difficult samples such as small targets and blurred edges, alleviates the gradient instability problem during training, and effectively suppresses overfitting to simple samples, thereby improving the overall detection robustness and accuracy.
[0049] 2. This invention improves the gray wolf optimization algorithm: To avoid the anchor box search process getting stuck in local optima, an improved gray wolf optimization strategy is proposed, and a "dynamic nonlinear convergence factor" based on an exponential function is designed to realize an adaptive adjustment mechanism from global coarse exploration to local fine convergence, which greatly improves the anchor box training effect and is particularly suitable for image localization tasks of medical devices with complex structures.
[0050] 3. This invention optimizes the image feature extraction structure: it introduces a "dense connection structure" and a "residual skip connection mechanism" in the network design, combined with "multi-level sub-pixel convolution upsampling", to achieve multi-scale deep semantic modeling while preserving image edge and texture details, effectively improving the model's adaptability to multi-angle and multi-resolution input images, and providing high-quality feature representation for subsequent recognition and matching.
[0051] 4. This invention constructs an improved clustering nonmaximum suppression algorithm: for scenes with occlusion and dense targets, an integrated target screening scheme combining clustering grouping, iterative suppression, and a score penalty mechanism is designed. This algorithm not only effectively removes redundant candidate boxes but also enhances the ability to retain occluded true targets through an exponential decay mechanism. Simultaneously, it supports parallel matrix computation, ensuring that the model maintains high accuracy while possessing high inference efficiency.
[0052] In summary, through comprehensive improvements in loss function design, search strategy optimization, feature extraction structure innovation, and target selection algorithm, this invention successfully constructs a high-precision image recognition and intelligent interactive display system for radiotherapy equipment. It possesses good versatility, robustness, and practical value, and is applicable to various scenarios such as medical science popularization exhibition halls, virtual exhibition halls, and intelligent guide systems.
[0053] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of the present invention, nor is it intended to restrict the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0054] The above and other features, advantages, and aspects of the various embodiments of the present invention will become more apparent from the accompanying drawings and the following detailed description. The drawings are provided for a better understanding of the invention and are not intended to limit the invention. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0055] Figure 1 This is a schematic diagram of a module of a radiotherapy equipment popular science display and interactive system based on image recognition according to an embodiment of the present invention;
[0056] Figure 2 This is a schematic diagram of the architecture of an image recognition-based radiotherapy equipment popular science display and interactive system according to an embodiment of the present invention. Detailed Implementation
[0057] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0058] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0059] Figure 1 This is a schematic diagram of a module of a radiotherapy equipment popular science display and interactive system based on image recognition according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the architecture of an image recognition-based radiotherapy equipment popular science display and interactive system according to an embodiment of the present invention. Figure 1 and Figure 2 As shown, a radiotherapy equipment popular science demonstration and interactive system 100 based on image recognition includes:
[0060] The data acquisition and construction module 110 is used to acquire physical images, 3D models and technical documents of radiotherapy equipment, and to build an image database containing equipment type tags and a knowledge base containing multi-dimensional popular science content; among which, radiotherapy equipment includes proton therapy devices, linear accelerators, CyberKnife and Gamma Knife.
[0061] S1. Requirements Analysis and Data Acquisition
[0062] S1.1: Target User and Application Scenario Positioning
[0063] To build a radiotherapy equipment image recognition and popular science display system with practical application value and social dissemination power, we must first clarify the system's core service targets and typical application scenarios.
[0064] The system's target users mainly include the following three categories:
[0065] The general public, especially teenagers, patients and their families, have significant knowledge gaps regarding radiotherapy and are an important target group for science popularization.
[0066] Medical professionals, including radiation oncologists, nurses, and equipment engineers, can use the system to assist in patient communication and equipment demonstrations.
[0067] Science education institutions, such as health education centers, hospital publicity departments, museums and science and technology museums, are used for daily or special displays of knowledge related to radiotherapy equipment.
[0068] Based on user profiles and survey results, the system needs to meet the following core requirements:
[0069] Equipment identification and interactive presentation: Based on image recognition technology, it automatically identifies radiotherapy equipment (such as proton therapy devices, linear accelerators, CyberKnife, etc.) and links multimedia resources to achieve interactive knowledge display.
[0070] Multi-dimensional information presentation: Supports the display of relevant knowledge content from multiple dimensions such as structural principles, treatment process, indication analysis, and equipment comparison, enhancing user understanding.
[0071] Immersive communication experience: By linking image recognition with AR / VR technology, we provide a more immersive and visual communication method, improving user engagement and recall rate.
[0072] Doctor-patient communication aid: Used for communication between doctors and patients, to increase patients' trust and acceptance of treatment equipment, and to alleviate information asymmetry between doctors and patients.
[0073] The system covers two main application scenarios: offline and online.
[0074] Offline scenarios: such as hospital radiotherapy centers, health science exhibition halls, etc., deploy interactive display terminals (touch screens, AR glasses, etc.) and guide users to scan and recognize by combining physical equipment or models.
[0075] Online scenarios include the hospital's official website science popularization section, embedded pages in WeChat official accounts, WeChat mini programs, and virtual exhibition halls (VR version), meeting the needs for remote access and convenient science popularization. Users can upload or scan QR codes to identify images and obtain matching content.
[0076] S1.2: Medical Equipment Data Acquisition
[0077] Building the image recognition model and knowledge-driven engine required for the system relies on abundant image and text data. Data acquisition is carried out in three levels:
[0078] (1) Equipment material collection
[0079] Image data: Organize on-site shooting to collect high-definition image materials of various radiotherapy equipment from different angles and under different operating conditions, ensuring that the images contain key structural features, brand logos, etc.
[0080] 3D modeling: Based on real equipment images and structural drawings, create high-fidelity 3D models (such as the Elekta Unity accelerator and the IBA proton therapy system) for AR / VR scene display.
[0081] Technical data collection: The system collects equipment manuals, clinical instructions, application reports, etc., extracts key technical parameters (such as acceleration method, radiation type, dose range, etc.), and constructs a structured attribute library.
[0082] Text compilation: Collect and organize various radiotherapy science popularization materials released by authoritative institutions, such as health lecture manuscripts and public service announcement scripts, to facilitate the formation of multimodal knowledge content that matches images.
[0083] (2) User behavior survey
[0084] Using methods such as questionnaires, in-depth interviews, and online log analysis, this study aimed to understand the knowledge levels, information needs, and content preferences of different user types regarding radiotherapy equipment. For example, teenagers preferred animated explanations and anthropomorphic metaphors; medical staff focused on technical parameters and performance indicators; and ordinary patients were more concerned about the treatment process and safety. The survey results will guide subsequent content tag design, interaction method selection, and user profile construction.
[0085] S1.3: Image and Knowledge Base Construction
[0086] To support subsequent recognition model training, image matching, and knowledge display, a structured image database and a knowledge content base need to be established:
[0087] (1) Image database construction
[0088] Image resource integration: Centralized management of acquired radiotherapy equipment images (including physical images, operation interfaces, component close-ups, etc.) and generated 3D modeling resources;
[0089] Image annotation and classification: Fine-grained labeling of images, including equipment type, manufacturer name, key parts (such as beam pipes, radiators, etc.);
[0090] Data update mechanism: Establish an update mechanism to regularly introduce new equipment images (such as robotic radiotherapy systems, high-energy gamma knives, etc.) to keep the data fresh and forward-looking.
[0091] (2) Construction of popular science knowledge base
[0092] Knowledge content structuring: The collected popular science materials and technical documents are processed in a structured manner to create a multi-dimensional knowledge graph (equipment introduction, principle explanation, indication matching, frequently asked questions and answers, etc.);
[0093] Integrating credible content: Introducing authoritative information sources (such as radiotherapy guidelines, medical journals, and hospital clinical cases) to ensure the scientific validity and accuracy of the knowledge content;
[0094] Content-adaptive tag design: Based on the user type tag content depth (such as patient perspective, medical staff perspective, and adolescent perspective), differentiated push of knowledge content can be achieved.
[0095] Upon completion of this module, it will provide a solid data and knowledge foundation for subsequent image recognition training and interactive display, ensuring that the system has comprehensive support capabilities in terms of recognition accuracy, content authority, and user interaction experience.
[0096] The image recognition module 120 is used to input the acquired image into the deep learning model. Through the neural network combined with the improved feature extraction and target localization algorithm, the key semantic features of the radiotherapy equipment in the image are extracted. Based on the improved clustering-nonmaximum suppression algorithm, the optimal box is retained and redundancy is removed, so as to achieve accurate identification and category judgment of the radiotherapy equipment.
[0097] S2, Image Recognition
[0098] The image recognition module 120 is the core technology of the entire system. It aims to accurately identify and classify radiotherapy equipment in input images using a deep learning model, providing crucial support for subsequent multimedia presentations and knowledge sharing. Considering the challenges of diverse image types, complex shooting conditions, and intricate target structures in real-world applications, the system supports multiple image sources and interaction modes, comprehensively covering real-world usage needs.
[0099] S2.1: Constructing an intelligent image recognition task
[0100] Step S2 involves intelligent recognition of three types of input images: real-world equipment images, planar carrier images, and user-uploaded images. The system uses a neural network combined with improved feature extraction and target localization algorithms to extract key semantic features of the radiotherapy equipment in the images and achieve robust recognition from multiple angles and resolutions.
[0101] Equipment entity image recognition: These images are mainly taken from real scenes in hospital showrooms, radiotherapy centers, or equipment installation areas. The recognition function identifies the complete radiotherapy equipment entity in the image, such as proton therapy equipment, linear accelerators, CyberKnife, Gamma Knife, etc.
[0102] Planar carrier image recognition: These images come from promotional materials posted in hospitals or science popularization venues, such as display boards, brochures, and guide maps. The task is to identify images of radiotherapy equipment in these promotional materials, such as identifying the CyberKnife device illustrated in "10 Popular Science Posters".
[0103] User-uploaded image recognition: During a hospital visit, users can take pictures of radiotherapy equipment in their department using their mobile phones (such as equipment displays outside the waiting area or equipment in the treatment room); users can also upload historical photos (such as photos taken during hospital visits or when accompanied by family members).
[0104] S2.2: Setting the core algorithm
[0105] S2.2.1: Image Target Localization
[0106] Image target localization is a fundamental step in the image recognition module of this system. Its main task is to accurately identify the main area of the radiotherapy equipment in diverse background environments, while eliminating interference factors such as exhibition hall settings and people. Considering the complex shape of radiotherapy equipment and the fact that the target area is often obscured or small in scale, the system adopts a dynamic anchor frame generation mechanism and designs a multi-objective optimization strategy that integrates F1 score and PWIoU loss, which significantly improves localization accuracy and model robustness.
[0107] Furthermore, the image recognition module 120 includes:
[0108] The target localization unit 121 adopts a dynamic anchor frame generation mechanism and optimizes the localization accuracy by fusing the F1 score and the dynamic difficulty-aware loss function.
[0109] To address the issue of insufficient adaptability of traditional anchor box presets in object detection, this step dynamically generates anchor boxes that better fit the dataset through a multi-objective optimization algorithm, improving localization accuracy and regression efficiency. This step integrates F1 score and PW10U loss as optimization objectives to balance detection accuracy and bounding box regression quality.
[0110] Furthermore, the target positioning unit 121 includes:
[0111] The classification target calculation subunit 1211 uses the F1 score as the classification optimization target to balance precision and recall and solve the problem of positive and negative sample imbalance.
[0112] The F1 score takes into account both precision and recall. In cases of imbalance between positive and negative samples, it can more realistically reflect the overall performance of the recognition model, balancing classification accuracy and coverage.
[0113] The regression target calculation subunit 1212 adopts a dynamic difficulty-aware loss function as the regression optimization target. It dynamically adjusts the gradient weights of the bounding box regression based on the sample anomaly degree and monotonic focusing coefficient to enhance the learning ability for small targets and blurred edge samples.
[0114] PWIoU loss introduces a dynamic difficulty-aware mechanism on top of traditional IoU, effectively mitigating gradient instability in target recognition. This method further utilizes "anomaly degree" to... The "monotonically focused coefficient γ" enhances the model's learning ability for difficult samples (such as blurry images and small-scale targets) while suppressing overfitting to simple samples.
[0115] Furthermore, the dynamic difficulty-aware loss function dynamically adjusts the gradient weights of bounding box regression based on the sample anomaly degree and the monotonic focusing coefficient. This includes: calculating the anomaly degree β of the current sample, which is the ratio of the IoU loss value of the current sample to the average IoU loss value; when β>1, it is identified as a difficult sample and the gradient gain is amplified to enhance the learning of small targets and blurred edge samples; when β<1, it is identified as a simple sample and the gradient gain is suppressed to prevent overfitting; and the gradient gain of difficult samples is exponentially amplified by the monotonic focusing coefficient γ.
[0116] PWIoU loss The definition is as follows:
[0117]
[0118]
[0119]
[0120]
[0121] 0 of which, Basic loss term; Distance factor: Based on the center distance and size of the predicted box and the ground truth box, the loss weight is dynamically adjusted; This suppresses gradient dominance in high-quality anchor frames and balances positive and negative samples. : Coordinates of the center point of the prediction box; : Coordinates of the center point of the true bounding box; : Width and height of the ground truth bounding box, used to normalize the distance error; γ: Monotonic focusing coefficient, amplifying the gradient contribution of difficult samples; : The IoU loss value of the current sample (reflecting the overlap error between the predicted box and the ground truth box); : Dynamic normalization factor, usually the average IoU loss of the current batch or during training, used to standardize the loss value of the current sample and alleviate gradient vanishing. α: Exponential adjustment factor, controls the decay rate of gradient gain (hyperparameter), controlling the rate of amplification / suppression (steepness). β: Measures the "abnormality" of the current sample, non-monotonically adjusting the gradient gain to adapt to dynamic sample distribution. β>1 indicates that the sample loss is higher than average, belonging to difficult samples (such as small or blurry targets), requiring increased gradient gain. β<1 indicates that the sample loss is lower than average, belonging to simple samples (such as large or clear targets), requiring reduced gradient contribution. δ: Balance point parameter, used to set the threshold for gradient adjustment (usually set to 1, i.e., defaulting to the average loss as the benchmark).
[0122] when When r=1, the gradient gain remains unchanged, and the corresponding sample is the average difficulty.
[0123] when For difficult samples, r>1, and the larger β or α, the larger γ, the stronger the gradient amplification effect (exponential), and the more the model pays attention to such samples.
[0124] when (Simple samples), r<1, gradient gain is suppressed to avoid the model overfitting to simple samples.
[0125] The dual-objective fusion mechanism 1213 weightedly superimposes the F1 score and the dynamic difficulty-aware loss function to form a multi-objective joint loss function. :
[0126]
[0127] Where λ is the balancing weight (0 < λ < 1), used to adjust the optimization emphasis between classification accuracy and positioning accuracy; Represents the classification loss based on the F1 score; This represents the regression loss based on improved IoU.
[0128] Gradient backpropagation is performed based on the joint loss function to simultaneously update the parameters of the classification network and the regression network, thereby achieving collaborative optimization of classification and localization results.
[0129] The anchor frame optimization strategy 1214 adopts a nonlinear convergence factor to dynamically adjust the search range, performs global exploration in the early stage of iteration, and performs local fine convergence in the later stage.
[0130] Improved Convergence Factor and Dynamic Adjustment Strategy: Traditional anchor box optimization methods are prone to getting stuck in local optima when faced with complex image backgrounds or multi-scale targets due to fixed search strategies and weak local exploration capabilities. This leads to problems such as anchor boxes failing to cover the real target or repeated localization. To address this, this invention introduces an improved Grey Wolf optimization algorithm into the anchor box search process and designs a dynamic nonlinear convergence factor. To enhance its adaptive capability, the convergence factor dynamically adjusts the search range based on a nonlinear function, specifically as follows:
[0131] Traditional gray wolf optimization algorithms are prone to getting trapped in local optima. This can be addressed by using a nonlinear adjustment factor. Enhance global search capabilities:
[0132]
[0133] Where ρ: convergence control factor for the current iteration; t: current iteration number; T: maximum iteration number; μ: adjustment strength coefficient (empirical value is 2); e: base of the natural logarithm, approximately equal to 2.71828, used for the exponential function to ensure that the adjustment factor changes in a non-linear manner.
[0134] The anchor box optimization strategy dynamically adjusts the convergence factor during training, and its value transitions nonlinearly from global search (ρ≈2) to local fine-tuning (ρ≈0) as the iteration progress t / T progresses.
[0135] This function is used in the early iteration phase (i.e.) When the convergence control factor ρ is relatively small, it approaches 2. At this point, the optimization strategy emphasizes global exploration, capturing multiple possible locations through a large-scale search, which helps avoid missing targets or getting trapped in local minima. In the later iteration stage (i.e., t→T), the convergence control factor ρ gradually approaches 0, and the optimization process shifts from coarse search to refined local optimization, focusing on refining the anchor frame boundaries and positions to improve the final positioning accuracy.
[0136] This mechanism effectively balances the search breadth and convergence speed in anchor box training, enabling the system to quickly locate the approximate target area in the early stages of training and to finely adjust the anchor box parameters in the later stages of training. It is suitable for radiotherapy equipment scenarios where there is occlusion, complex structure, or blurred contours in the image.
[0137] Anchor frame generation mechanism 1215 predicts the offset by constraining the Sigmoid function and dynamically outputs the device bounding box coordinates.
[0138] Dynamic Anchor Box Generation: The quality of anchor box generation directly affects the accuracy and recall of the image recognition module. Traditional anchor box methods rely on fixed size presets, often resulting in poor matching or candidate box redundancy when dealing with multi-scale and multi-angle targets. To improve adaptability and efficiency, the system designs a dynamic anchor box generation mechanism based on network prediction and parameter offset, which can flexibly adjust the anchor box size and position under different input images and target features. The core formula of this mechanism is as follows:
[0139]
[0140] in, : The coordinates of the center point of the target bounding box (final coordinates after dynamic adjustment). : The width and height of the target bounding box (final size after dynamic scaling); The predicted offset value of the center point coordinates output by the neural network. The neural network outputs predicted scaling values for width and height, with four parameters representing relative adjustment values. : The coordinates of the center point of the initial anchor frame (preset position); : The width and height of the initial anchor frame (preset dimensions); The Sigmoid function constrains the output to... The interval is used to ensure the rationality of the anchor frame coordinates and dimensions.
[0141] By linking the anchor frame optimization strategy 1214 with the anchor frame generation mechanism 1215, the anchor frames are continuously fine-tuned and optimized using loss feedback during the training process, achieving adaptive optimization throughout the entire process from coarse-grained initialization to fine-grained localization.
[0142] Feature extraction unit 122 employs a multi-scale feature extraction network with a dense connection structure and residual skip connection mechanism, combined with sub-pixel convolution upsampling to preserve image details;
[0143] S2.2.2: Image Feature Extraction
[0144] Furthermore, the feature extraction unit 122 includes: a data augmentation unit that performs rotation, scaling, and copy-paste processing on the input image; an encoder-decoder structure that extracts multi-scale semantic features in the downsampling stage and restores spatial resolution through sub-pixel convolution in the upsampling stage; and a residual connection channel that superimposes shallow features with the reconstructed feature map to prevent gradient vanishing.
[0145] After image target localization in step S2.2.1, the feature extraction unit 122 further performs feature extraction on the identified target. Image feature extraction aims to extract key visual semantic information from the radiotherapy equipment image to support subsequent recognition and matching tasks. To improve the model's adaptability to diverse image inputs, the system first expands the sample size and simulates changes in the equipment under different shooting angles and environmental conditions through classic data augmentation techniques such as rotation, scaling, cropping, and copy-pasting.
[0146] In the feature extraction network structure, the system employs a dense connection and residual skip connection mechanism, combined with a multi-level sub-pixel convolution upsampling strategy, to effectively improve the model's ability to restore image details and express deep features. The activation function is PReLU to enhance the model's ability to model nonlinear structures. The convolution kernel size is uniformly 3×3, and the number of channels is configured according to the network layer, using options such as (64,16) and (16,16) to balance extraction depth and computational efficiency. The specific process is described below:
[0147]
[0148]
[0149] in, : Feature map of the (n+1)th layer; : No. Subpixel convolutional upsampling operations in the layer are used to restore resolution; : Feature maps (residual information) passed in parallel with the current layer; The reconstructed feature map (or output feature map) of the nth layer is the image feature map after sub-pixel convolution upsampling during the image feature extraction process. Shallow feature maps connected to residuals The result obtained after addition; C: the original input feature map; : Feature map after downsampling at layer n; Downsampling operations at the nth layer (such as convolution, max pooling, etc.).
[0150] The above structure implements an "encode-decode" processing flow: the encoding stage extracts high-level semantic features at multiple scales through downsampling; the decoding stage uses sub-pixel convolution to restore spatial details and avoid information loss; residual connections effectively prevent gradient vanishing in deep networks while enhancing feature fidelity.
[0151] Ultimately, the feature extraction unit 122 can accurately extract key visual features of the radiotherapy equipment, such as the equipment's outline, functional components, and brand logos, and perform high-dimensional feature matching with pre-labeled standard images in the database. The model supports multi-view recognition (such as front view, side view, and images in operation), and combines a 3D model library for viewpoint compensation to ensure comprehensiveness and accuracy of recognition.
[0152] The result filtering unit 123, based on the improved clustering nonmaximum suppression algorithm, removes redundant prediction boxes through clustering grouping and confidence penalty mechanism, and outputs the final recognition result.
[0153] S2.2.3: Output the final recognition result
[0154] After completing the target localization (step S2.2.1) and feature extraction (step S2.2.2), the model generates multiple predicted bounding boxes for each image. Each box contains the predicted category (e.g., proton therapy device, linear accelerator, CyberKnife, etc.), confidence score, and coordinate information. Due to the existence of redundant prediction results (e.g., multiple overlapping boxes pointing to the same target), step S2.2.3 introduces an improved clustering nonmaximum suppression algorithm. Through clustering, iterative suppression, and score penalty mechanisms, the optimal box is retained and redundancy is removed, ensuring that the final target detection result is accurate and efficient.
[0155] Furthermore, the result filtering unit 123 includes: calculating the overlap relationship between predicted boxes based on the intersection-union matrix; an iterative suppression mechanism that removes redundant boxes overlapping with high-confidence boxes round by round using a mask matrix; and a confidence penalty mechanism that uses an exponential decay function to reduce the false suppression rate of occluded targets. Specifically:
[0156] First, construct the IoU (Intersection over Union) matrix between the predicted bounding boxes, and then perform grouping and clustering on the candidate boxes with high overlap:
[0157] Where X: the output upper triangular IoU matrix; Extract the upper triangular matrix, ignoring the diagonal. To avoid double counting; Represents an N×N dimensional IoU matrix, with matrix elements Store the pairwise overlap relationships between all bounding boxes; N: the total number of predicted bounding boxes in a single image; : No. and the One prediction box; Intersection over Union (IoU) is the ratio of the intersection area to the union area of two candidate boxes. It measures the degree of overlap. IoU = intersection area / union area (range [0,1], the larger the value, the higher the overlap).
[0158] The iterative suppression process is based on the IoU matrix and retains candidate boxes with high confidence and low redundancy through a masking mechanism, performing the following steps:
[0159]
[0160] Where m: the number of iterations, initially m=1, and the maximum number of iterations is... ; : Retained tags from the previous iteration ( Indicates reservation. (Indicates removal); : Diagonal mask matrix, by Generate a function to retain the currently active frame; : Diagonal matrix generation operation; : Suppression matrix, representing the set of candidates that significantly overlap with the reserved box; The maximum intersection-union ratio (MUNR) for each column is used to determine whether there is significant overlap. : Maximum value operation; Threshold (set to 0.5), if If so, then suppress the box; It is the candidate box retention label vector of the m-th iteration.
[0161] The iterative suppression mechanism gradually removes redundant bounding boxes through multiple rounds of filtering. In the m-th round, a mask matrix is generated based on the retained results from the (m-1)-th round, and the mask matrix is then adjusted according to the overlap threshold. Update the current round's retained flags.
[0162] To avoid retaining low-overlapping but high-confidence interference boxes, the system introduces an exponential decay mechanism to dynamically penalize the confidence of overlapping boxes:
[0163]
[0164] in, : The j-th candidate box The original confidence level; : Corrected confidence level after penalty; Confidence penalty factor (set to 0.2, empirical value), specifically represents the temperature coefficient that controls the severity of the penalty and the rate of decay. The smaller the value, the more severe the penalty. Candidate boxes in round m and The IoU value between them; This represents the set of indices of all other candidate boxes i whose IoU with candidate box j is greater than or equal to the threshold ε and is not j itself; Exponential decay function, box j highly overlapping A penalty is imposed on box i, the higher the overlap ( The larger the value of the exponential decay function, the smaller the penalty factor (the smaller the value of the exponential decay function), leading to... The more it is reduced, the more true targets are retained that are falsely detected due to occlusion; Indicates the candidate box All overlapping boxes A cumulative penalty is applied, and the penalty is stacked when multiple boxes overlap (e.g., a gamma knife overlapped by two redundant boxes). .
[0165] The specific process of the improved clustering nonmaximum suppression algorithm is as follows:
[0166] 1. Input: Set of prediction boxes The corresponding confidence score IoU threshold ε (typically 0.5), temperature coefficient τ (typically 0.2).
[0167] 2. Initialize and retain the tag vector b = [1, 1, ..., 1] (length N).
[0168] 3. Construct the IoU matrix: Calculate the IoU between all pairs of predicted boxes to obtain an N×N matrix M, where... Extract its strict upper triangular part (excluding the diagonal) to obtain matrix X.
[0169] 4. Iteration suppression:
[0170] Loop from m = 1 to N:
[0171] If all elements in b are 0, then exit the loop;
[0172] Find the confidence level within the set of boxes currently marked as reserved (b = 1). Highest candidate box index ;
[0173] For all currently marked as reserved, i.e., j, satisfy b j The box where j == 1 and j != k :calculate (or obtained from pre-calculated M) If iou >= ε, then set b. j = 0 (suppression box) And apply confidence penalty: update box confidence level .
[0174] 5. Output: All b i Candidate boxes with ==1 and its (possibly penalized) confidence level As the final test result.
[0175] For example: When a linear accelerator is detected multiple times: Initial state: Generate 3 overlapping boxes ( ), Round 1: (Highest score) Retained ,inhibition Termination: Output a unique box. This solves the problem of repeated detection of radiotherapy equipment. It also preserves obscured equipment; for example, when a CyberKnife is partially obscured: traditional NMS directly deletes the low-resolution bounding box, resulting in target loss. This invention uses an algorithm that corrects this through confidence penalty. Still satisfied .
[0176] This mechanism is particularly suitable for situations where there are occluded or structurally similar regions in medical device images, and can significantly improve the model's ability to preserve the real target.
[0177] The entire process in step S2.2.3 supports parallel matrix computation. The IoU matrix processing adopts an upper triangular approach to avoid redundant calculations and significantly reduce computational complexity. The measured average inference time can be controlled at 7.1ms / image, meeting the performance requirements of real-time interactive scenarios. Finally, after cluster suppression and confidence adjustment, the system will output the optimal radiotherapy device identification results, including: device category, precise location (boundary coordinates), and confidence score. These identification results will directly drive subsequent science popularization content display and multimedia interaction (see step S3 for details).
[0178] The science popularization interaction module 130 matches knowledge base content based on recognition results, displays device information in a multimodal format, and adapts the content depth based on user profiles;
[0179] S3, Interactive Science Popularization Content
[0180] After completing the recognition and classification of radiotherapy equipment images, the system will automatically link relevant knowledge content and achieve a "what you see is what you get" intelligent science popularization interactive experience through multimodal display and personalized push mechanism. This module focuses on multi-dimensional knowledge coverage, diversified communication media, and multi-user profile adaptation to comprehensively improve the public's awareness of radiotherapy equipment.
[0181] Furthermore, the science popularization interactive module 130 includes:
[0182] The multidimensional knowledge display unit 131 provides equipment structure principles, treatment process animations, technology comparison charts, and AR scene fusion demonstrations. The AR scene fusion demonstration specifically involves using image anchor point alignment technology based on feature point matching (such as ORB, SIFT) or plane detection to stably overlay 3D models or dynamic path animations onto the identified equipment images or real-world scenes. For proton therapy equipment, it visualizes the trajectory of the beam from the accelerator to the tumor target area.
[0183] S3.1: Multi-dimensional knowledge display
[0184] To meet the diverse needs of users at different levels for understanding medical knowledge, the system presents radiotherapy equipment-related content in a comprehensive manner, from basic information to clinical applications, and then to technological comparisons and future trends. Specifically, this includes:
[0185] (1) Basic information display
[0186] After identifying the device, the system will display its name, function, and structural composition. It will provide easy-to-understand explanations of the technical principles, such as a visual demonstration of how a proton beam is precisely focused on the tumor site, helping the public understand its characteristics that distinguish it from traditional X-rays. The system will also showcase the research and development background and industry information, such as the application of the first batch of proton radiotherapy systems in China, enhancing users' confidence in the technological strength of domestically produced equipment.
[0187] (2) Interpretation of clinical application
[0188] Based on the identified device type, the system will automatically match its applicable disease (e.g., linear accelerators are suitable for lung cancer and nasopharyngeal carcinoma, while CyberKnife is suitable for small intracranial tumors). It displays a complete treatment process animation, showing the entire process from patient location, treatment plan development, treatment to follow-up, enhancing patient informed consent and participation. It also includes real clinical cases or doctor-patient interview clips to enhance the credibility and relatability of the content.
[0189] (3) Analysis of technological advantages
[0190] The system uses charts and comparative animations to highlight the performance comparisons of different devices in terms of treatment precision, side effects, treatment duration, and the breadth of indications. For example, it showcases the dose control advantages of the IBA Proteus PLUS proton therapy system in pediatric cancer treatment, emphasizing its "controllable, precise, and safe" characteristics. It also supports interactive switching between "traditional radiotherapy vs. novel radiotherapy" scenarios, enhancing user engagement.
[0191] (4) AR scene fusion display
[0192] Using augmented reality (AR) technology, when a user scans an image with a mobile device, the system automatically overlays a 3D model or dynamic animation onto the original image. For example, after recognizing an image of a proton therapy device, an animated path demonstrating "the proton beam being released from the accelerator, passing through the beam pipe, and focusing on the tumor" is overlaid. The system precisely aligns virtual content with real-world images based on image anchor points, achieving a seamless integration of physical display and virtual content, enhancing the immersive experience.
[0193] The rich media output unit 132 supports video playback, simultaneous output of text, images and voice, and natural language question-and-answer interaction.
[0194] S3.2: Rich Media Content Formats
[0195] To adapt to different terminals and usage scenarios, the system uses a combination of various media carriers to display content, thereby improving the dissemination effect and user acceptance:
[0196] (1) Video / Animation Playback
[0197] The system includes built-in short video resources such as "Equipment Operating Principles," "Treatment Simulation," and "Expert Explanations," which users can click to play. For example, a video explaining how a linear accelerator generates and controls high-energy rays provides an intuitive demonstration of complex physical processes.
[0198] (2) Image and voice fusion output
[0199] The system supports automatic synchronization of text and image summaries with expert-recorded audio explanations, helping users with different reading habits access content. Users can choose between "patient's perspective interpretation" or "doctor's explanation mode" to achieve multi-verbal expression of the same knowledge point.
[0200] (3) Intelligent interactive question and answer
[0201] Based on an NLP semantic recognition model, the system allows users to ask questions via voice or text (such as "Is hospitalization required during CyberKnife treatment?"), and returns matching text / image or video responses. The backend links to authoritative content libraries (such as remote treatment guidelines, patient health education materials, and interpretations of health poverty alleviation policies) to enable intelligent knowledge search and precise recommendations.
[0202] User profile adaptation unit 133 distinguishes between teenagers, patients, and medical staff, and dynamically adjusts the content presentation format and depth.
[0203] S3.3 Personalized Content Adaptation
[0204] To enhance user engagement and content relevance, the system intelligently adjusts the depth and format of content presentation based on user profiles, achieving the goal of "tailored and precise science popularization":
[0205] (1) Teen users
[0206] It offers engaging and gamified content, such as comparing radiotherapy equipment to "tumor snipers" or "minimally invasive radiation weapons," and showcasing proton therapy equipment as a "proton rocket launch site." Interactive simulation animations help users understand the process of proton beams aiming at and striking tumor targets. Cartoonish 3D models and interactive Q&A challenges further stimulate learning interest and participation.
[0207] (2) Patient users
[0208] Focusing on practical issues of concern to patients, such as treatment side effects, procedural safety, and pre- and post-treatment precautions, we provide real patient cases, post-operative recovery guidelines, and answers to frequently asked questions to enhance trust between doctors and patients.
[0209] (3) Medical staff
[0210] The equipment's technical parameters, user interface, performance indicators, and upgrade path are showcased. The latest clinical research findings, guideline interpretations, and research progress are provided as reference materials for physicians' science popularization and teaching.
[0211] The deployment integration module 140 supports deployment on online platforms and offline terminals, and triggers recognition and interaction functions through image scanning or uploading.
[0212] S4 Deployment and Integration
[0213] To achieve widespread application and multi-scenario adaptation of the system, two implementation methods are designed: offline deployment and online platform integration, to achieve application effects that cover all scenarios and reach all people.
[0214] Furthermore, the deployment integration module 140 includes: offline deployment, configuring touch terminals and AR devices in physical exhibition halls to support QR code recognition of physical exhibits; and online integration, providing image upload and recognition services through WeChat mini programs and virtual exhibition halls.
[0215] S4.1: Offline Deployment
[0216] Interactive devices such as touchscreen kiosks, AR recognition terminals, or electronic guide screens are deployed in hospital radiotherapy centers and health science exhibition halls. The system integrates with physical display boards, equipment prototypes, or graphic information panels, supporting methods such as QR code scanning, photo recognition, and voice interaction to display recognition results and science information on-site. This three-pronged approach—image recognition, AR content, and physical devices—enhances the realism and interactivity of science education.
[0217] S4.2: Online Platform Integration
[0218] The system is embedded into platforms such as the hospital's official website's "Science Popularization Column," WeChat official account menu, and patient service mini-program, providing the public with access to health science popularization anytime, anywhere. Users can upload images via their mobile phones (such as photos of equipment taken during medical visits), and the system automatically recognizes and returns matching content pages, improving ease of use. A VR virtual exhibition hall version has been developed to adapt to online activities such as "National Science and Technology Week" and "Cancer Prevention and Control Publicity Week," enabling immersive display of equipment knowledge and remote science popularization education, effectively expanding the scope and application scenarios of science popularization.
[0219] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0220] It should also be noted that, in the embodiments of this application, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0221] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined in the embodiments of this application may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown in this application, but is to be accorded the widest scope consistent with the principles and novel features disclosed in the embodiments of this application.
Claims
1. A radiotherapy equipment popular science demonstration and interactive system based on image recognition, characterized in that, include: The data acquisition module is used to collect physical images, 3D models and technical documents of radiotherapy equipment, and to build an image database containing equipment type tags and a knowledge base containing multi-dimensional popular science content. The image recognition module is used to input the acquired images into the deep learning model. Through neural networks combined with improved feature extraction and target localization algorithms, it extracts the key semantic features of radiotherapy equipment in the image, and retains the optimal bounding box and removes redundancy based on the improved clustering nonmaximum suppression algorithm, so as to achieve accurate identification and category judgment of radiotherapy equipment. The image recognition module includes: The target localization unit adopts a dynamic anchor frame generation mechanism and optimizes localization accuracy by fusing F1 scores and dynamic difficulty-aware loss functions. The feature extraction unit employs a multi-scale feature extraction network with a dense connection structure and a residual skip connection mechanism, combined with sub-pixel convolutional upsampling to preserve image details; The result filtering unit, based on the improved clustering nonmaximum suppression algorithm, removes redundant prediction boxes through clustering grouping and confidence penalty mechanism, and outputs the final recognition result. The target positioning unit includes: The classification target calculation sub-unit uses the F1 score as the classification optimization target to balance precision and recall and solve the problem of positive and negative sample imbalance. The regression target calculation subunit uses a dynamic difficulty-aware loss function as the regression optimization target. It dynamically adjusts the gradient weights of the bounding box regression based on the sample anomaly degree and monotonic focusing coefficient to enhance the learning ability for small targets and blurred edge samples. The dual-objective fusion mechanism weights and superimposes the F1 score and the dynamic difficulty-aware loss function to form a multi-objective joint loss function; gradient backpropagation is performed based on the joint loss function to simultaneously update the parameters of the classification network and the regression network, thereby achieving collaborative optimization of classification results and localization results. The anchor frame optimization strategy uses a nonlinear convergence factor to dynamically adjust the search range, performing global exploration in the early stage of iteration and local fine-tuning convergence in the later stage. Among them, the search range is dynamically adjusted using a nonlinear convergence factor, including: Where ρ: convergence control factor for the current iteration; t: current iteration number; T: maximum iteration number; μ: adjustment strength coefficient; e: base of the natural logarithm, used for exponential functions to ensure that the adjustment factor changes in a non-linear manner; The anchor box generation mechanism predicts the offset by constraining the Sigmoid function and dynamically outputs the device bounding box coordinates. The science popularization and interaction module matches knowledge base content based on the recognition results, displays device information in a multimodal format, and adapts the content depth based on user profiles; The deployment integration module supports deployment on online platforms and offline terminals, and triggers recognition and interaction functions through image scanning or uploading.
2. The system according to claim 1, characterized in that, The dynamic difficulty-aware loss function dynamically adjusts the gradient weights of bounding box regression based on the sample anomaly degree and the monotonic focusing coefficient, including: Calculate the anomaly β of the current sample, which is the ratio of the IoU loss value of the current sample to the average IoU loss value; When β>1, it is identified as a difficult sample and the gradient gain is amplified to enhance the learning of small targets and blurry edge samples. When β < 1, it is considered a simple sample and the gradient gain is suppressed to prevent overfitting; The gradient gain of difficult samples is exponentially amplified by a monotonic focusing coefficient γ.
3. The system according to claim 1, characterized in that, The feature extraction unit includes: The data augmentation unit performs rotation, scaling, and copy-paste processing on the input image; The encoding-decoding structure extracts multi-scale semantic features in the downsampling stage and restores spatial resolution through sub-pixel convolution in the upsampling stage. The residual connection channel overlays shallow features with the reconstructed feature map to prevent gradient vanishing.
4. The system according to claim 1, characterized in that, The result filtering unit includes: The overlap relationship between predicted bounding boxes is calculated based on the intersection-over-union matrix; An iterative suppression mechanism eliminates redundant boxes that overlap with high-confidence boxes round by round using a mask matrix; The confidence penalty mechanism uses an exponential decay function to reduce the false suppression rate of occluded targets.
5. The system according to claim 1, characterized in that, The science popularization interactive module includes: The multi-dimensional knowledge display unit provides equipment structure and principles, treatment process animations, technology comparison charts, and AR scene fusion demonstrations. The rich media output unit supports video playback, simultaneous output of text, images and voice, and natural language question-and-answer interaction. The user profile adaptation unit differentiates between teenagers, patients, and medical staff, and dynamically adjusts the content presentation format and depth.
6. The system according to claim 5, characterized in that, The AR scene fusion demonstration is as follows: Image anchor point alignment technology based on feature point matching or plane detection is used to stably overlay 3D models or dynamic path animations onto the identified device images or real-world scenes. For proton therapy equipment, the trajectory of the beam from the accelerator to the tumor target area is visualized.
7. The system according to claim 1, characterized in that, The deployment integration module includes: Offline deployment involves configuring touch terminals and AR devices in physical exhibition halls to support QR code scanning and recognition of physical exhibits; The online integration method provides image upload and recognition services through WeChat mini programs and virtual exhibition halls.
8. The system according to any one of claims 1-7, characterized in that, The radiotherapy equipment includes a proton therapy device, a linear accelerator, a CyberKnife, and a Gamma Knife.
Citation Information
Patent Citations
Popular science knowledge interactive displaying method and device
CN106815885A
Faster RCNN-based solar cell panel fault identification method
CN117036659A