Gynecological disease early screening and diagnosis method based on big data analysis

By employing a spatiotemporal fusion frame extraction algorithm, a human-machine collaborative annotation platform, a YOLOv8 model enhanced with CBAM, and a three-level knowledge quality control system, the problem of insufficient multimodal data integration in existing AI-assisted diagnostic systems for gynecological disease screening and diagnosis has been solved, achieving high-precision data support and improved medical efficiency.

CN121034592APending Publication Date: 2025-11-28HEFEI DVL ELECTRON CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511127934.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing AI-assisted diagnostic systems lack the integration of multimodal data in gynecological disease screening and diagnosis, leading to biased risk assessment and higher surgical risks. Furthermore, the lack of effective quality control mechanisms results in a higher risk of misdiagnosis.

Method used

By employing a deep-perception-based dual-stream neural network architecture to achieve deep perception of surgical scenes, a spatiotemporal fusion algorithm for temporal features, and combining a human-machine collaborative annotation platform, a CBAM-enhanced YOLOv8 model, and a three-level knowledge quality control system, the entire chain of surgical video processing, from intelligent processing to annotation quality control, is innovated, providing high-precision data support for AI surgical navigation.

Benefits of technology

It significantly reduces the risk of misdiagnosis, improves medical efficiency, enhances labeling quality and detection accuracy, ensures medical compliance, and achieves system sustainability and efficient data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121034592A_ABST
    Figure CN121034592A_ABST
Patent Text Reader

Abstract

The invention discloses a gynecological disease early screening and diagnosis method based on big data analysis, and relates to the technical field of medical information technology and intelligent diagnosis. Deep perception of an operation scene is realized through a double-flow neural network architecture; the spatial feature flow adopts a lightweight convolutional network to extract an instrument form and an organ anatomical structure, and the time sequence action flow analyzes a continuous frame optical flow field based on 3D convolution to capture operation continuity; a physiological parameter linkage mechanism is introduced, when the monitor detects that the bleeding amount suddenly increases or the blood oxygen saturation degree suddenly drops, the frame extraction frequency is automatically increased to 30 fps from conventional 5 fps, and a Gaussian filter is used in the preprocessing stage. According to the method, full-chain innovation of operation videos from intelligent processing to labeling quality control is realized, high-precision data support is provided for AI operation navigation, misdiagnosis risks are remarkably reduced, and medical efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of medical information technology and intelligent diagnosis technology, and particularly relates to a gynecological disease early screening and diagnosis method based on big data analysis. BACKGROUND

[0002] The gynecological disease early screening and diagnosis method mainly focuses on the cross application of medical big data and artificial intelligence technology, needs to be carried out around the technical bottleneck of traditional screening means, the limitation of existing intelligent diagnosis and the challenge of multi-modal data fusion, and the intelligent diagnosis often relies on AI assisted diagnosis system.

[0003] The existing AI assisted diagnosis system is developed based on a single data source, such as electronic colposcope image analysis in cervical cancer screening or ultrasonic image recognition model of ovarian cancer. Although such system can improve the efficiency of a specific link, it lacks the integration of multi-modal data, leading to one-sided risk assessment and high surgical risk, and therefore has room for improvement. SUMMARY

[0004] The purpose of the present application is to solve the problems in the prior art and provide a gynecological disease early screening and diagnosis method based on big data analysis. The method has the advantages of realizing the whole chain innovation from intelligent processing to labeling quality control of surgical video, providing high-precision data support for AI surgical navigation, significantly reducing the risk of misdiagnosis and improving medical efficiency.

[0005] In order to achieve the above purpose, the present application adopts the following technical scheme:

[0006] A gynecological disease early screening and diagnosis method based on big data analysis comprises the following steps:

[0007] Step 1: Dynamic key frame extraction

[0008] The depth perception of the surgical scene is realized through a double-flow neural network architecture. The spatial feature flow adopts a lightweight convolutional network to extract the shape of the instrument and the anatomical structure of the organ, and the time series action flow is based on 3D convolution to analyze the continuous frame optical flow field to capture the operation continuity. The physiological parameter linkage mechanism is introduced, when the monitor detects a sudden increase in bleeding volume or a sharp drop in blood oxygen saturation, the frame extraction frequency is automatically increased from the normal 5fps to 30fps. In the preprocessing stage, a Gaussian filter is used.

[0009] Step 2: Human-computer collaborative high-precision labeling platform

[0010] To address the issues of annotation efficiency and quality, a two-stage pre-annotation-correction workflow was developed. In the pre-annotation stage, an improved MaskR-CNN model was employed, integrating deformable convolutions to handle blood occlusion scenarios. Adaptive ROI sampling enhanced the ability to identify small target lesions, achieving a device detection accuracy of mAP@0.94. In the correction stage, an intelligent assistance system was designed. When a physician modifies a certain type of annotation, it automatically pushes historical similar case reference schemes and initiates a three-party expert arbitration process for disputed annotations. This mechanism improves annotation efficiency to 18 frames / minute, reduces manual intervention by 85%, and increases label consistency to 98%.

[0011] Step 3: YOLOv8 detection model optimized for surgical scenarios

[0012] ECA and CBAM attention mechanism modules are embedded in the standard YOLOv8 architecture. These two modules are located at the end of the backbone network and in the feature fusion layer. They enhance key feature channels through channel attention and focus on dangerous operation areas using spatial attention. For special surgical interferences, reflection suppression and dynamic smoke simulation enhancement strategies are developed: the former adds a mirror-like highlight effect to the training data of metal instruments, and the latter generates progressive smoke occlusion based on a fluid dynamics model. The improved model achieves an AP value of 0.67 in the detection of small instruments such as suture needles and achieves an intraoperative bleeding point recognition rate of over 91%.

[0013] Step Four: Knowledge-Driven Three-Tier Quality Control System

[0014] To ensure medical compliance, a hierarchical verification framework is constructed. The basic layer uses a rule engine to enforce checks on the integrity of mandatory organs and instruments. The intermediate layer uses geometric constraint algorithms to verify spatial rationality, such as detecting whether instruments abnormally penetrate organ entities. The highest layer relies on a surgical knowledge graph for logical reasoning. When the system detects "unlabeled hemostatic clips but vascular transection operations," an alarm is automatically triggered. The knowledge graph contains a relationship network of 28 types of organs, 12 types of instruments, and 50 operations, such as the risk link that "electric hook operations may damage the hepatic artery," enabling the quality control model to intercept 98.5% of medical logic errors.

[0015] Step 5: Closed-loop system evolution mechanism

[0016] The system employs a real-time feedback-driven workflow. Labeling errors detected by the quality control module are automatically extracted from relevant video clips and packaged as training samples. When the accumulated number of errors reaches 50 or the model performance decreases by 2%, the incremental learning engine is triggered to update the detection model. The update process strictly adheres to medical device specifications—the new model must process 100 surgical videos in parallel with the old version, and deployment is only permitted after statistical verification of performance improvement. The system adopts a cloud-edge-device collaborative architecture: edge devices perform real-time frame extraction and pre-labeling; the cloud GPU cluster processes model training; and the operating room terminal provides visualized quality control reports.

[0017] The present invention is further configured such that the Gaussian filtering formula is:

[0018] Here, G(x,y) is the value of the Gaussian filter at point (x,y), and is the standard deviation. The Gaussian function is mainly used to construct the filter weights to reduce image noise and details, so as to facilitate subsequent edge detection and feature extraction operations.

[0019] The present invention is further configured such that the forward propagation formula of the convolutional layer in step three is:

[0020] Output=σ(Weight*Input+Bias);

[0021] Where σ represents the ReLU activation function and * represents the convolution operation, this formula is used many times in image recognition and feature extraction for the calculation of convolution.

[0022] The present invention is further configured such that, in step four, a support vector machine is used for surgical path planning, and the optimization formula of the support vector machine is:

[0023]

[0024] subjecttoy (i) (w T x (i) +b)≥1-ξ i ,ξ i ≥0;

[0025] Here, w and b are the parameters of the hyperplane, C is the regularization parameter, and C is the slack variable.

[0026] The present invention is further configured such that the dynamic keyframe extraction adopts a spatiotemporal fusion frame extraction algorithm, which adopts a dual-stream parallel processing framework; the spatial branch is based on a lightweight MobileNetV3 network, inputs a single frame of RGB image, and extracts the morphological and texture features of the instrument and organ through depthwise separable convolution; the temporal branch inputs 10 consecutive frames of optical flow field, and analyzes the motion trajectory of the instrument and the tissue deformation pattern by a 3D-ResNet18 network.

[0027] The present invention is further configured such that the workflow of the MaskR-CNN model includes the following steps:

[0028] Step 1: Feature Extraction: The SwinTransformer backbone network generates multi-scale feature maps;

[0029] Step 2: Occlusion handling: Add deformable convolution before the ROIAlign layer to improve the instrument recognition rate in blood occlusion scenarios, mAP@0.5 from 0.78 to 0.91;

[0030] Step 3: Boundary optimization: Add an EdgeRefineNet subnetwork to refine the instrument edges through three levels of dilated convolutions.

[0031] The present invention is further configured such that the YOLOv8 detection model needs to be trained, and the gradient calculation formula of the backpropagation algorithm is as follows when training the deep learning model:

[0032]

[0033] Where E is the loss function, is the weight between neurons i and j in layer l, is the weighted input of neuron i in layer l, is the error term of neuron i in layer l, and is the activation output of neuron j in layer (l-1).

[0034] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method described above.

[0035] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.

[0036] The beneficial effects of this invention are as follows: This invention achieves full-chain innovation in surgical video from intelligent processing to annotation quality control through spatiotemporal fusion frame extraction algorithm, human-machine collaborative annotation platform, CBAM-enhanced YOLOv8 model, three-level knowledge quality control system and closed-loop optimization system, providing high-precision data support for AI surgical navigation, significantly reducing the risk of misdiagnosis and improving medical efficiency. Attached Figure Description

[0037] Figure 1 This is a schematic diagram of the spatiotemporal fusion frame extraction algorithm for an early screening and diagnosis method for gynecological diseases based on big data analysis proposed in this invention.

[0038] Figure 2 This is a schematic diagram of the human-computer collaborative annotation process for an early screening and diagnosis method for gynecological diseases based on big data analysis proposed in this invention.

[0039] Figure 3 This is a schematic diagram of the YOLOv8+ECA+CBAM detection model for an early screening and diagnosis method for gynecological diseases based on big data analysis proposed in this invention.

[0040] Figure 4 This is a schematic diagram of the three-level quality control model for an early screening and diagnosis method for gynecological diseases based on big data analysis proposed in this invention.

[0041] Figure 5This is a schematic diagram of the closed-loop optimization system for an early screening and diagnosis method for gynecological diseases based on big data analysis proposed in this invention.

[0042] Figure 6 This is a schematic diagram of the improved Yolov8 network structure for an early screening and diagnosis method for gynecological diseases based on big data analysis proposed in this invention. Detailed Implementation

[0043] The technical solution of this patent will be further described in detail below with reference to specific embodiments.

[0044] The embodiments of this patent are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this patent, and should not be construed as limiting this patent.

[0045] Reference Figures 1-6 A method for early screening and diagnosis of gynecological diseases based on big data analysis includes the following steps:

[0046] Step 1: Dynamic Keyframe Extraction

[0047] Deep perception of the surgical scene is achieved through a dual-stream neural network architecture; the spatial feature flow uses a lightweight convolutional network to extract the instrument morphology and organ anatomical structure, and the temporal action flow is based on 3D convolutional analysis of continuous frame optical flow field to capture the continuity of operation; a physiological parameter linkage mechanism is introduced, and when the monitor detects a sudden increase in bleeding or a sudden drop in blood oxygen saturation, the frame sampling frequency is automatically increased from the usual 5fps to 30fps. In the preprocessing stage, a Gaussian filter is used.

[0048] Step Two: Human-Machine Collaborative High-Precision Annotation Platform

[0049] To address the issues of annotation efficiency and quality, a two-stage pre-annotation-correction workflow was developed. In the pre-annotation stage, an improved MaskR-CNN model was employed, integrating deformable convolutions to handle blood occlusion scenarios. Adaptive ROI sampling enhanced the ability to identify small target lesions, achieving a device detection accuracy of mAP@0.94. In the correction stage, an intelligent assistance system was designed. When a physician modifies a certain type of annotation, it automatically pushes historical similar case reference schemes and initiates a three-party expert arbitration process for disputed annotations. This mechanism improves annotation efficiency to 18 frames / minute, reduces manual intervention by 85%, and increases label consistency to 98%.

[0050] Step 3: YOLOv8 detection model optimized for surgical scenarios

[0051] ECA and CBAM attention mechanism modules are embedded in the standard YOLOv8 architecture. These two modules are located at the end of the backbone network and in the feature fusion layer. They enhance key feature channels through channel attention and focus on dangerous operation areas using spatial attention. For special surgical interferences, reflection suppression and dynamic smoke simulation enhancement strategies are developed: the former adds a mirror-like highlight effect to the training data of metal instruments, and the latter generates progressive smoke occlusion based on a fluid dynamics model. The improved model achieves an AP value of 0.67 in the detection of small instruments such as suture needles and achieves an intraoperative bleeding point recognition rate of over 91%.

[0052] Step Four: Knowledge-Driven Three-Tier Quality Control System

[0053] To ensure medical compliance, a hierarchical verification framework is constructed. The basic layer uses a rule engine to enforce checks on the integrity of mandatory organs and instruments. The intermediate layer uses geometric constraint algorithms to verify spatial rationality, such as detecting whether instruments abnormally penetrate organ entities. The highest layer relies on a surgical knowledge graph for logical reasoning. When the system detects "unlabeled hemostatic clips but vascular transection operations," an alarm is automatically triggered. The knowledge graph contains a relationship network of 28 types of organs, 12 types of instruments, and 50 operations, such as the risk link that "electric hook operations may damage the hepatic artery," enabling the quality control model to intercept 98.5% of medical logic errors.

[0054] Step 5: Closed-loop system evolution mechanism

[0055] The system employs a real-time feedback-driven workflow. Labeling errors detected by the quality control module are automatically extracted from relevant video clips and packaged as training samples. When the accumulated number of errors reaches 50 or the model performance decreases by 2%, the incremental learning engine is triggered to update the detection model. The update process strictly adheres to medical device specifications—the new model must process 100 surgical videos in parallel with the old version, and deployment is only permitted after statistical verification of performance improvement. The system adopts a cloud-edge-device collaborative architecture: edge devices perform real-time frame extraction and pre-labeling; the cloud GPU cluster processes model training; and the operating room terminal provides visualized quality control reports.

[0056] The Gaussian filtering formula is:

[0057] Here, G(x,y) is the value of the Gaussian filter at point (x,y), and is the standard deviation. The Gaussian function is mainly used to construct the filter weights to reduce image noise and details, so as to facilitate subsequent edge detection and feature extraction operations.

[0058] The formula for forward propagation of a convolutional layer is:

[0059] Output=σ(Weight*Input+Bias);

[0060] Where σ represents the ReLU activation function and * represents the convolution operation, this formula is used many times in image recognition and feature extraction for the calculation of convolution.

[0061] Support Vector Machines (SVMs) are used for surgical path planning. The optimization formula for SVMs is as follows:

[0062]

[0063] subjecttoy (i) (w T x (i) +b)≥1-ξ i ,ξ i ≥0;

[0064] Here, w and b are the parameters of the hyperplane, C is the regularization parameter, and C is the slack variable.

[0065] Spatiotemporal fusion frame extraction algorithm

[0066] A dual-stream parallel processing framework is adopted. The spatial branch is based on a lightweight MobileNetV3 network. It takes a single frame of RGB image as input and extracts the morphology of the instrument (such as the bending angle of the electric hook) and the texture features of the organ (such as the distribution of blood vessels on the surface of the liver) through depthwise separable convolution. The temporal branch takes 10 consecutive frames of optical flow field (generated by TV-L1 algorithm) as input and analyzes the motion trajectory of the instrument (such as the displacement velocity of the needle holder) and the tissue deformation pattern (such as the frequency of intestinal peristalsis) by a 3D-ResNet18 network.

[0067] Core Innovations:

[0068] ① Dynamic gating fusion: A sigmoid weighted module is designed. When the movement speed of the device is greater than 15 pixels / frame (corresponding to the actual movement speed greater than 3cm / s), the temporal branch weight is automatically increased to 0.85 to ensure the integrity of fast motion capture.

[0069] ② Physiological parameter linkage: Real-time analysis of monitor data stream; when bleeding volume suddenly increases by more than 30 ml / min or blood oxygen saturation drops by more than 5% within 5 minutes, the frame rate is increased from 5 fps to 30 fps.

[0070] ③ Buffer preloading: Using an LSTM surgical stage classifier (with a recognition accuracy of 94%), 20 frames of raw data are pre-buried at critical stages (such as exposure of the gallbladder triangle) to avoid loss of the start frame of the operation.

[0071] Processing flow:

[0072] ① Spatial feature extraction and temporal motion analysis are performed simultaneously after video input;

[0073] ② The fusion module calculates the spatiotemporal confidence score (range 0-1);

[0074] ③ When the score is >0.9 and remains above 5 frames for a period of time, high-speed frame skipping is initiated;

[0075] ④ The output frame includes a timestamp, operation type label, and trigger physiological parameter value.

[0076] Technical specifications: In 200 laparoscopic surgeries, the recall rate of critical operation frames was 99.2%, the redundant frame removal rate was 82.7%, and the processing latency was ≤50ms / frame (4K resolution).

[0077] Human-machine collaborative annotation platform

[0078] Pre-annotation engine: Based on an improved Mask R-CNN architecture:

[0079] ① Feature extraction: The SwinTransformer backbone network generates multi-scale feature maps;

[0080] ②Occlusion handling: Add deformable convolution before the ROIAlign layer to improve the instrument recognition rate in blood occlusion scenarios (mAP@0.5 from 0.78→0.91);

[0081] ③ Boundary optimization: Add the EdgeRefineNet subnetwork to refine the instrument edges through three levels of dilated convolution.

[0082] Doctor Interaction System:

[0083] ① Intelligent correction recommendation: Build a knowledge base for annotation correction. When a doctor modifies the annotation of "cystic artery", the system will automatically push correction schemes for similar anatomical structures.

[0084] ② Tripartite arbitration mechanism: The disputed labeling automatically initiates a joint review request, which is independently evaluated by the chief surgeon and two experts with associate senior titles or above, and the final label is determined by majority voting.

[0085] Workflow:

[0086] 1. Input keyframes into the pre-annotated model to generate initial labels;

[0087] II. Doctor's Interface Loading Results:

[0088] ①Accept labels: Directly insert into the database;

[0089] ② Correct annotation: Trigger intelligent recommendation;

[0090] ③ Added annotation: Initiate arbitration process.

[0091] 3. Corrected data is sent back to the training set in real time, triggering incremental updates of the model at night.

[0092] Technical benefits: Annotation efficiency reaches 18 frames / minute, manual intervention is reduced by 85%, and annotation consistency is improved to 98.3%.

[0093] ECA+CBAM Enhanced YOLOv8 Surgical Detection Model

[0094] ECA embedding location: located between the end of the CSPDarknet53 backbone and the PANet feature fusion layer, CBAM is embedded in the Head.

[0095] Surgical enhancement strategies:

[0096] ①Reflection suppression: Add dynamic specular synthesis to the training data: randomly generate overexposed areas on the instrument surface, covering 15%-30% of the pixels;

[0097] ② Smoke simulation: Physically simulated smoke is generated based on the Navier-Stokes equation and graded according to visibility (light: transparency > 0.7, heavy: < 0.3).

[0098] Training optimization:

[0099] ① Loss function: FocalLoss is used to address instrument size imbalance (the needle pixel ratio is only 0.01%).

[0100] ② Multi-scale training: Input resolution progressively scaled from 640×640 to 1280×1280

[0101] Performance validation: On the cholecystectomy test set, the pin detection AP@0.5 reached 0.67, the intraoperative bleeding point recognition rate was 91.2%, and the inference speed was 142 FPS (Tesla T4 GPU).

[0102] When training a deep learning model, the gradient calculation formula for the backpropagation algorithm is:

[0103]

[0104] Where E is the loss function, is the weight between neurons i and j in layer l, is the weighted input of neuron i in layer l, is the error term of neuron i in layer l, and is the activation output of neuron j in layer (l-1).

[0105] Three-level knowledge-driven quality control model

[0106] L1 Basic Validation: The rule engine validates 28 mandatory elements, such as the "common hepatic duct" and "cystic artery" labels that must be included in cholecystectomy, and an alarm is triggered immediately if they are missing.

[0107] L2 space logic:

[0108] ① Penetration detection: An alarm is triggered when the IoU between the instrument and organ mask is greater than 0.05 and lasts for more than 3 frames, such as when scissors abnormally enter the common bile duct;

[0109] ② Distance warning: Real-time monitoring of the Euclidean distance between dangerous instruments and critical organs; a warning is generated when the distance is less than 10 pixels.

[0110] L3 healthcare compliance:

[0111] ① Knowledge graph construction: It includes 402 relationship chains covering 12 types of instruments, 28 organs, and 50 types of operations;

[0112] ②Logic reasoning engine: Based on the Drools rule base, it performs causal chain verification and issues an alarm when it finds a broken blood vessel end that is not labeled with a hemostatic clip.

[0113] Error handling mechanism:

[0114] ① Generate a quality control report: label the error type (L1 / L2 / L3 level) + associated video time period (accurate to the second);

[0115] ② Output correction suggestions: such as "The common hepatic duct needs to be added at 32 minutes and 15 seconds";

[0116] ③ Automatically package error samples: including error frames, error tags, and gold standard references.

[0117] Closed-loop optimization system

[0118] Error-driven evolution:

[0119] I. Sample Capture: After the quality control model identifies an error, it automatically extracts 30 seconds of video before and after the error point (including labeled data and monitor records).

[0120] II. Incremental learning trigger: Updates are initiated when any one of the following conditions is met:

[0121] ① Accumulate 50 new erroneous samples;

[0122] ②The error rate increased by 2% in 100 consecutive surgeries;

[0123] ③ New surgical types (such as initial prostatectomy).

[0124] Model update process:

[0125] I. Fine-tuning training: Based on the original model, new samples are trained using a cosine annealing learning rate (initial value 1e-4);

[0126] II. A / B Testing: Comparison of key performance indicators (KPIs) of 100 surgical videos processed in parallel by the old and new models:

[0127] ①Missing rate of key instruments;

[0128] ②Number of false positive alarms;

[0129] ③ Level 3 logic error rate.

[0130] Statistical validation: Performance improvement was confirmed by McNemar test (p < 0.05). The system was then released in tiers.

[0131] ① Secondary Update: Push directly to edge devices

[0132] ②Major Update: Requires review by the Clinical Ethics Committee

[0133] Cloud-edge-device collaboration:

[0134] ① Edge layer: Jetson AGXOrin performs real-time frame extraction and pre-annotation;

[0135] ② Terminal: The operating room interactive screen displays the quality control heat map.

[0136] This invention achieves full-chain innovation in surgical video processing, from intelligent processing to annotation quality control, through a spatiotemporal fusion frame extraction algorithm, a human-machine collaborative annotation platform, a CBAM-enhanced YOLOv8 model, a three-level knowledge quality control system, and a closed-loop optimization system. It provides high-precision data support for AI surgical navigation, significantly reduces the risk of misdiagnosis, and improves medical efficiency.

[0137] This invention achieves breakthrough improvements in five dimensions: data processing efficiency, annotation quality, detection accuracy, medical compliance assurance, and system sustainability. In the data processing stage, traditional fixed-interval frame extraction schemes result in over 90% redundant frames. In contrast, the spatiotemporal fusion frame extraction algorithm of this invention uses a dual-stream neural network to dynamically perceive key surgical nodes, increasing the effective frame capture rate to 99.2%, compressing data volume by 82%, and reducing the processing time for a single surgical video from an average of 8 hours to 40 minutes. Regarding annotation quality, existing fully manual annotation methods have an efficiency of only 2 frames per minute, and label consistency is only 65%-70% due to differences in physician experience. The human-machine collaborative platform of this invention, combining a pre-annotation engine and intelligent correction recommendation, increases the annotation speed to 18 frames per minute, reduces manual intervention by 85%, and achieves label consistency of 98.3%, providing a high-purity data foundation for model training.

[0138] I. In terms of detection model performance, mainstream general-purpose target detectors such as YOLOv5 perform poorly in complex surgical scenarios: the AP@0.5 for small instruments such as sutures is only 0.45, and the missed detection rate of intraoperative bleeding points exceeds 24%. In contrast, the ECA+CBAM enhanced YOLOv8 of this invention, through a channel-spatial dual attention mechanism combined with surgery-specific enhancement strategies (reflection suppression and smoke simulation), improves the AP@0.5 for suture detection to 0.67, achieves a bleeding point recognition rate of 91.2%, and maintains a real-time performance of 142 FPS in 4K video streams. Medical compliance quality control is a crucial aspect completely lacking in traditional technologies—existing methods rely solely on visual inspection by physicians, failing to detect deeper errors such as instruments penetrating organs or contradictions in operational procedures.

[0139] II. Regarding the sustainable evolution of the system, traditional models become fixed once deployed and cannot adapt to new surgical procedures or equipment iterations. The closed-loop optimization system of this invention automatically triggers incremental learning through quality control errors, and pushes updates after confirmation by A / B testing and McNemar's test, achieving an average of 2 model iterations per month. Cross-operative generalization tests show that when expanding from cholecystectomy to prostatectomy, the fluctuation of key instrument detection is <3%, which is significantly better than the >15% performance degradation of traditional methods.

[0140] III. Medical Compliance Quality Control: From Nothing to Something. Traditional methods lack automated quality control, and deep-seated medical logic errors (such as cutting blood vessels without using hemostatic clips) rely entirely on visual inspection by physicians, resulting in a missed detection rate exceeding 30%. This invention constructs a three-level knowledge-driven quality control system based on a knowledge graph containing 402 medical logic relationships. It achieves three-layer verification: L1 basic integrity, L2 spatial rationality, and L3 medical logic, intercepting 98.5% of potential annotation errors and avoiding the risk of misjudgment caused by data defects in AI surgical navigation systems.

[0141] Hierarchy Detect content Core technology L1 Mandatory item integrity Rule engine (28 organs / instruments) L2 Spatial rationality Geometric constraints (IoU > 0.05 alarm) L3 Medical logic compliance Knowledge graph reasoning (402 relationship chains)

[0142] This invention represents a comprehensive breakthrough compared to existing technologies: the keyframe capture rate jumps from 68% to 99.2%, and the annotation efficiency is improved by 9 times to 18 frames / minute; the ECA+CBAM-enhanced YOLOv8 model improves the needle detection accuracy (AP@0.5) to 0.67; the three-level knowledge quality control system intercepts 98.5% of medical logic errors; the closed-loop system supports an average of 2 iterations per month, with performance fluctuations across surgical procedures of <3%. Clinical evidence shows that it shortens the time of a single operation by 8.5 minutes, laying a high-precision data foundation for AI surgical navigation and promoting the development of intelligent surgery.

[0143] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method described above.

[0144] A computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of the above-described method.

[0145] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for early screening and diagnosis of gynecological diseases based on big data analysis, characterized in that, Includes the following steps: Step 1: Dynamic Keyframe Extraction Achieving deep perception of surgical scenarios through a dual-stream neural network architecture; Spatial feature flow uses lightweight convolutional networks to extract instrument morphology and organ anatomical structure, while temporal action flow is based on 3D convolutional analysis of continuous frame optical flow fields to capture operational continuity. A physiological parameter linkage mechanism is introduced. When the monitor detects a sudden increase in bleeding or a sudden drop in blood oxygen saturation, the frame rate is automatically increased from the normal 5fps to 30fps. A Gaussian filter is used in the preprocessing stage. Step Two: Human-Machine Collaborative High-Precision Annotation Platform To address the issues of annotation efficiency and quality, a two-stage workflow of pre-annotation and correction was developed. In the pre-annotation stage, an improved MaskR-CNN model is used, which integrates deformable convolution to deal with blood occlusion scenarios. Adaptive ROI sampling enhances the ability to identify small target lesions, enabling the instrument detection accuracy to reach mAP@0.

94. In the correction stage, an intelligent auxiliary system is designed to automatically push historical similar case reference schemes when physicians modify a certain type of annotation, and initiate a three-party expert arbitration process for disputed annotations. This mechanism improves annotation efficiency to 18 frames per minute, reduces manual intervention by 85%, and increases label consistency to 98%. Step 3: YOLOv8 detection model optimized for surgical scenarios ECA and CBAM attention mechanism modules are embedded in the standard YOLOv8 architecture. The two modules are located at the end of the backbone network and in the feature fusion layer. They enhance key feature channels through channel attention and focus on dangerous operation areas using spatial attention. For special surgical interference, reflective suppression and dynamic smoke simulation enhancement strategies are developed: the former adds the mirror highlight effect of metal instruments to the training data, and the latter generates progressive smoke occlusion based on the fluid dynamics model. The improved model achieved an AP value of 0.67 in the detection of small suture instruments and a bleeding point recognition rate of over 91% during surgery. Step Four: Knowledge-Driven Three-Tier Quality Control System To ensure medical compliance, a hierarchical verification framework is constructed. The basic layer uses a rule engine to enforce checks on the integrity of mandatory organs and instruments. The intermediate layer uses geometric constraint algorithms to verify spatial rationality, such as detecting whether instruments abnormally penetrate organ entities. The highest layer relies on a surgical knowledge graph for logical reasoning. When the system detects "unlabeled hemostatic clips but vascular transection operations," an alarm is automatically triggered. The knowledge graph contains a relationship network of 28 types of organs, 12 types of instruments, and 50 operations, such as the risk link of "electric hook operation may damage the hepatic artery," enabling the quality control model to intercept 98.5% of medical logic errors. Step 5: Closed-loop system evolution mechanism Design a real-time feedback-driven workflow; The quality control module will automatically extract relevant video clips and package them as training samples when it detects annotation errors. When the cumulative number of errors reaches 50 or the model performance drops by 2%, the incremental learning engine will be triggered to update the detection model. The update process strictly follows medical device specifications—the new model must process 100 surgical videos in parallel with the old version, and the performance improvement must be statistically verified before deployment. The system adopts a cloud-edge-device collaborative architecture: edge devices perform real-time frame extraction and pre-annotation. The cloud-based GPU cluster processes model training; the operating room terminal provides visualized quality control reports.

2. The method for early screening and diagnosis of gynecological diseases based on big data analysis according to claim 1, characterized in that, The Gaussian filtering formula is as follows: Here, G(x,y) is the value of the Gaussian filter at point (x,y), and is the standard deviation. The Gaussian function is mainly used to construct the filter weights to reduce image noise and details, so as to facilitate subsequent edge detection and feature extraction operations.

3. The method for early screening and diagnosis of gynecological diseases based on big data analysis according to claim 1, characterized in that, The formula for forward propagation of the convolutional layer in step three is: Output=σ(Weight*Input+Bias); Where σ represents the ReLU activation function and * represents the convolution operation, this formula is used many times in image recognition and feature extraction for the calculation of convolution.

4. The method for early screening and diagnosis of gynecological diseases based on big data analysis according to claim 1, characterized in that, Step four uses a support vector machine for surgical path planning. The optimization formula for the support vector machine is as follows: subjectto y (i) (w T x (i) +b)≥1-ξ i ,ξ i ≥0; Here, w and b are the parameters of the hyperplane, C is the regularization parameter, and C is the slack variable.

5. The method for early screening and diagnosis of gynecological diseases based on big data analysis according to claim 1, characterized in that, The dynamic keyframe extraction employs a spatiotemporal fusion frame extraction algorithm, which uses a dual-stream parallel processing framework. The spatial branch is based on a lightweight MobileNetV3 network, which takes a single frame of RGB image as input and extracts the morphological and texture features of the instrument and organ through depthwise separable convolution. The temporal branch takes 10 consecutive frames of optical flow field as input and analyzes the instrument's motion trajectory and tissue deformation pattern using a 3D-ResNet18 network.

6. The method for early screening and diagnosis of gynecological diseases based on big data analysis according to claim 5, characterized in that, The workflow of the Mask R-CNN model includes the following steps: Step 1: Feature Extraction: The SwinTransformer backbone network generates multi-scale feature maps; Step 2: Occlusion handling: Add deformable convolution before the ROIAlign layer to improve the instrument recognition rate in blood occlusion scenarios, mAP@0.5 from 0.78 to 0.91; Step 3: Boundary optimization: Add an EdgeRefineNet subnetwork to refine the instrument edges through three levels of dilated convolutions.

7. The method for early screening and diagnosis of gynecological diseases based on big data analysis according to claim 6, characterized in that, The YOLOv8 detection model needs to be trained. When training the deep learning model, the gradient calculation formula for the backpropagation algorithm is: Where E is the loss function, is the weight between neurons i and j in layer l, is the weighted input of neuron i in layer l, is the error term of neuron i in layer l, and is the activation output of neuron j in layer (l-1).

8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.