Artificial intelligence-based operating room active data collection platform
Patent Information
- Application Number
- CN202610837338.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-10
- Publication Date
- 2026-08-18
AI Technical Summary
[0003]现有技术存在两方面显著缺点:一方面,缺乏能够兼容多数据来源格式协议的统一平台,不同设备产生的术野视频、内镜影像、超声图像、电子病历等数据因格式标准不统一形成信息孤岛,导致多源数据无法有效整合与协同处理,难以实现手术过程数据的全面采集与综合分析;另一方面,病灶区域识别、手术关键信息提取等核心环节过度依赖专家经验,缺乏基于人工智能与海量数据的主动识别与分析能力,不仅增加了专家工作负担,还可能因人为判断差异影响识别准确性与一致性,无法充分发挥数据价值以支撑手术效率提升与医疗质量优化
[0015] Beneficial Effects: This invention proposes an AI-based active data collection platform for the operating room. It comprehensively collects surgical-related visual data through a multi-channel visual information capture module. A deep learning real-time recognition and tracking module enables active identification of instruments, anatomical structures, and lesion areas, tracking surgical steps and personnel information and evaluating collaboration efficiency. A multi-source data fusion processing module constructs a spatiotemporal graph, a surgical event pattern recognition module forms a structured timeline, and a surgical knowledge graph construction module achieves deep data association and semantic understanding. The core cross-device data compatibility and adaptation module directly breaks down information silos caused by different countries and devices, ensuring compatibility with the format protocols of multi-source data such as surgical field videos, endoscopic images, ultrasound images, and electronic medical records. This completely solves the problems of incompatibility and ineffective integration of multiple data source formats in existing technologies. Meanwhile, the platform uses artificial intelligence technology as its core and deep learning models to analyze and process massive surgical data, enabling the proactive identification and extraction of key information such as lesion areas. This replaces the traditional model that relies too heavily on expert experience, reducing the workload of experts and improving the accuracy and consistency of identification. It fully explores the value of surgical data, providing strong support for improving surgical efficiency and optimizing medical quality, and comprehensively makes up for the shortcomings of existing technologies in multi-source data integration and intelligent identification of core information.
Smart Images

Figure CN122598897A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of operating room data collection technology, and more particularly to an artificial intelligence-based active operating room data collection platform. Background Technology
[0002] As a core scenario for medical treatment, the operating room generates process data that includes crucial information such as surgical procedures, instrument usage, personnel collaboration, and lesion characteristics. This data is of significant value for improving medical quality, optimizing technology, and conducting medical research. Currently, during surgery, multiple high-definition cameras and surgical field cameras capture visual information such as video, instrument operation, and personnel interaction, while simultaneously generating multi-source data streams including audio and equipment operation data. Furthermore, surgical data involves various types of information, including endoscopic images, ultrasound images, and electronic medical records. With the development of artificial intelligence technology, using deep learning models to automatically identify surgical instruments, anatomical structures, and lesion areas, track key operational steps, and construct surgical spatiotemporal maps and knowledge graphs to achieve deep data association and semantic understanding has become an important direction for improving surgical efficiency and reducing medical risks. However, differences in data formats from different countries and brands of equipment, as well as the need for multi-source data integration, place higher demands on the platform's compatibility and data processing capabilities.
[0003] Existing technologies have two significant drawbacks: First, they lack a unified platform compatible with multiple data source format protocols. Data such as surgical field videos, endoscopic images, ultrasound images, and electronic medical records generated by different devices form information silos due to inconsistent format standards, making it difficult to effectively integrate and collaboratively process multi-source data and achieve comprehensive collection and analysis of surgical process data. Second, core aspects such as lesion area identification and extraction of key surgical information rely excessively on expert experience and lack the ability to proactively identify and analyze based on artificial intelligence and massive amounts of data. This not only increases the workload of experts but may also affect the accuracy and consistency of identification due to differences in human judgment, failing to fully leverage the value of data to support improved surgical efficiency and optimized medical quality. Summary of the Invention
[0004] In order to overcome the shortcomings and deficiencies of existing technologies, this invention provides an artificial intelligence-based active data collection platform for operating rooms.
[0005] The technical solution adopted in this invention is an artificial intelligence-based operating room active data collection platform, characterized by comprising: a multi-channel visual information capture module, a deep learning real-time recognition and tracking module, a multi-source data fusion and processing module, a surgical event pattern recognition module, a surgical knowledge graph construction module, and a cross-device data compatibility and adaptation module. The multi-channel visual information capture module captures surgical videos, instrument operations, and personnel interactions through multiple high-definition cameras in the operating room and a surgical field camera. The deep learning real-time recognition and tracking module receives the output information from the multi-channel visual information capture module, identifies surgical instruments, anatomical structures, and lesion areas through a deep learning model, tracks and calibrates operational steps, identifies the identities, locations, and calibration actions of surgical team members, and evaluates team collaboration efficiency. The multi-source data fusion and processing module receives the output data from the deep learning real-time recognition and tracking module, fuses video, audio, and equipment data streams to construct a spatiotemporal map of the surgical process, and the surgical event pattern recognition module... The module receives the output graph from the multi-source data fusion processing module, defines and labels surgical events through pattern recognition algorithms, and forms a structured surgical timeline. The surgical knowledge graph construction module receives the output timeline from the surgical event pattern recognition module, and constructs a knowledge graph containing surgical procedures, steps, instruments, anatomy, complications, personnel entities and relationships based on massive surgical data, performing deep data association and semantic understanding. The cross-device data compatibility and adaptation module connects with each module, integrates information silos and data incompatibility issues of different countries and different devices, and is compatible with the format protocols of multiple data sources such as surgical field videos, endoscopic images, ultrasound images, and electronic medical records.
[0006] Furthermore, the deep learning real-time recognition and tracking module includes an instrument anatomical structure recognition unit, a surgical step tracking unit, a personnel information recognition unit, and a collaboration efficiency evaluation unit. The instrument anatomical structure recognition unit receives visual information output from the multi-channel visual information capture module and uses a deep learning model to extract and match features of the surgical instruments' morphological characteristics, spatial position, and anatomical markers and tissue boundaries to achieve accurate recognition of surgical instruments and anatomical structures. Simultaneously, it performs multi-dimensional feature analysis on the grayscale distribution, texture features, and morphological contours of the lesion area to actively identify the lesion area. The surgical step tracking unit receives the recognition results output from the instrument anatomical structure recognition unit and, combined with the temporal characteristics and logical connections of the surgical operation, identifies the starting point, execution point, and other parameters of the surgical operation. The system continuously monitors and records the surgical procedure's progress and completion status, forming a dynamic tracking link. The personnel information identification unit receives visual information about personnel interactions from the multi-channel visual information capture module, extracts facial features, limb movement features, and clothing identification features of surgical team members, and confirms member identities through feature comparison. Simultaneously, it obtains the three-dimensional coordinate information of members in the operating room based on a visual positioning algorithm, and records the amplitude, execution frequency, and movement correlation of calibrated actions. The collaboration efficiency evaluation unit receives the step tracking data output by the surgical procedure tracking unit and the personnel movement data output by the personnel information identification unit, constructs a team collaboration behavior matrix, and completes a quantitative evaluation of team collaboration efficiency by analyzing action response time, operational coordination, and information transmission path.
[0007] Furthermore, the multi-source data fusion processing module includes a video data stream processing unit, an audio data stream processing unit, an equipment data stream processing unit, and a spatiotemporal graph construction unit. The video data stream processing unit receives surgical video-related data output by the deep learning real-time recognition and tracking module, performs feature extraction and inter-frame correlation analysis on video frames, extracts spatial and temporal features of the surgical scene, and forms structured video data. The audio data stream processing unit collects voice commands, equipment operation sounds, and environmental sound data in the operating room, extracts frequency features, amplitude features, and temporal features of the sound through audio feature extraction algorithms, and completes the structured processing of the audio data. The equipment data stream processing unit collects the operating parameters and sensor data of the surgical equipment, performs format parsing and feature filtering on the data, extracts calibration equipment data related to the surgical process, and forms standardized equipment data. The spatiotemporal graph construction unit receives the structured data output by the video data stream processing unit, the audio data stream processing unit, and the equipment data stream processing unit, establishes a data association model with time as the horizontal axis and spatial location as the vertical axis, maps the data of each dimension to the corresponding spatiotemporal nodes, and constructs a spatiotemporal graph including multi-dimensional information of the entire surgical process.
[0008] Furthermore, the surgical event pattern recognition module includes an event feature extraction unit, an event definition unit, an event marking unit, and a timeline generation unit. The event feature extraction unit receives the spatiotemporal map output by the multi-source data fusion processing module, extracts the mutation features, correlation features, and temporal features of the data in the map, and filters out feature vectors related to the calibrated surgical events. The event definition unit establishes an event feature library based on the feature vectors output by the feature extraction unit through a pattern recognition algorithm, defines the feature thresholds and correlation rules for different types of calibrated surgical events, and forms a standardized event recognition standard. The event marking unit marks the nodes in the spatiotemporal map that meet the characteristics of the calibrated surgical events according to the recognition standard output by the event definition unit, and records the start time, duration, and associated data of the events. The timeline generation unit receives the marking results output by the event marking unit, sorts the marked events in chronological order, and integrates the surgical steps, instrument usage, and personnel action information associated with the events to form a structured surgical timeline.
[0009] Furthermore, in the surgical knowledge graph construction module, the entity relationship weight calculation model is as follows: ,in, Representing entities With entity Relationship weights Representing entities In the Feature values of each feature dimension Representing entities In the Feature values of each feature dimension Indicates the total number of feature dimensions. Indicates the time decay coefficient. Representing entities The corresponding surgical time point, Representing entities The corresponding surgical time point, This represents the correlation strength adjustment coefficient. Representing entities With entity The frequency of co-occurrence.
[0010] Furthermore, in the cross-device data compatibility adaptation module, the data format conversion model is as follows: ,in, This represents the target format data after conversion. Indicates the number of data source types. Indicates the first Format conversion matrix of data source class Indicates the first The raw data of the data source. Indicates the first Format offset of the data source class Indicates the global format calibration coefficient. This indicates the format compatibility compensation amount.
[0011] Furthermore, in the deep learning real-time recognition and tracking module, the probability model for lesion region recognition is as follows: ,in Indicates input data The probability of identifying the corresponding lesion area. Indicates the number of feature extraction layers. Indicates the first Weight coefficients of layer features Indicates input data In the The feature mapping function of the layer, This indicates the identification bias term.
[0012] Furthermore, in the surgical event pattern recognition module, the event matching degree calculation model is as follows: ,in, Indicates the event to be identified With standard events The degree of matching, Indicates the number of event feature dimensions. Indicates the event to be identified In the dimensional eigenvalues Represents standard events In the dimensional eigenvalues Indicates the weight of time similarity. Indicates the event to be identified With standard events Time similarity, Indicates the weight of the difference in event length. Indicates the event to be identified The duration, Represents standard events The duration.
[0013] Furthermore, in the multi-source data fusion processing module, the spatiotemporal node data fusion model is as follows: Show time Fusion of spatial location node data, Indicates the weights for video data fusion. express time Video data of spatial location, Indicates the weights for audio data fusion. express time Audio data of spatial location, Indicates the weight of device data fusion. express time Spatial location device data, express time Spatial location fusion error adjustment factor, Indicates the number of dimensions in the device data. express Time of the first Dimensional device data.
[0014] An AI-based active data collection platform for the operating room operates through the following steps: First, multiple high-definition cameras and surgical field cameras in the operating room capture surgical videos, instrument operations, and personnel interactions at preset frame rates and resolutions, simultaneously acquiring audio signals and surgical equipment data streams while maintaining data timestamp synchronization. Second, visual information is input into a deep learning model. Convolutional neural networks and recurrent neural networks extract instrument and anatomical structure features, an attention mechanism focuses on the lesion area, and visual positioning technology obtains personnel spatial coordinates to construct a collaborative behavior evaluation index system. Third, multi-source data formats are parsed, and calibrated feature parameters are extracted. The process involves six steps: First, establishing association mapping rules, aligning data according to time axis and spatial location, and constructing a three-dimensional spatiotemporal information map. Second, mining the spatiotemporal map features, defining and calibrating surgical event identification standards through cluster analysis and rule matching, sorting and labeling events by time, and integrating data to form a structured time axis. Third, collecting massive amounts of surgical data and structured time axis data, extracting entity information, mining associations through entity relationship extraction algorithms, establishing attribute and rule bases, and constructing a multi-dimensional knowledge graph. Fourth, designing multi-protocol compatible interfaces, converting and adapting data from different sources through format parsing algorithms, establishing a compatible standard library, and ensuring smooth data transmission and unified processing.
[0015] Beneficial Effects: This invention proposes an AI-based active data collection platform for the operating room. It comprehensively collects surgical-related visual data through a multi-channel visual information capture module. A deep learning real-time recognition and tracking module enables active identification of instruments, anatomical structures, and lesion areas, tracking surgical steps and personnel information and evaluating collaboration efficiency. A multi-source data fusion processing module constructs a spatiotemporal graph, a surgical event pattern recognition module forms a structured timeline, and a surgical knowledge graph construction module achieves deep data association and semantic understanding. The core cross-device data compatibility and adaptation module directly breaks down information silos caused by different countries and devices, ensuring compatibility with the format protocols of multi-source data such as surgical field videos, endoscopic images, ultrasound images, and electronic medical records. This completely solves the problems of incompatibility and ineffective integration of multiple data source formats in existing technologies. Meanwhile, the platform uses artificial intelligence technology as its core and deep learning models to analyze and process massive surgical data, enabling the proactive identification and extraction of key information such as lesion areas. This replaces the traditional model that relies too heavily on expert experience, reducing the workload of experts and improving the accuracy and consistency of identification. It fully explores the value of surgical data, providing strong support for improving surgical efficiency and optimizing medical quality, and comprehensively makes up for the shortcomings of existing technologies in multi-source data integration and intelligent identification of core information. Attached Figure Description
[0016] Figure 1 This is a diagram showing the system module composition of the present invention; Figure 2 This is a flowchart of the system operation steps of the present invention. Detailed Implementation
[0017] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0018] like Figure 1 As shown, the AI-based operating room proactive data collection platform includes: a multi-channel visual information capture module, a deep learning real-time recognition and tracking module, a multi-source data fusion and processing module, a surgical event pattern recognition module, a surgical knowledge graph construction module, and a cross-device data compatibility and adaptation module. The multi-channel visual information capture module captures surgical videos, instrument operations, and personnel interactions through multiple high-definition cameras in the operating room and a surgical field camera. The deep learning real-time recognition and tracking module receives the output information from the multi-channel visual information capture module, identifies surgical instruments, anatomical structures, and lesion areas through a deep learning model, tracks and calibrates operational steps, identifies the identities, locations, and calibration actions of surgical team members, and evaluates team collaboration efficiency. The multi-source data fusion and processing module receives the output data from the deep learning real-time recognition and tracking module, fuses video, audio, and equipment data streams to construct a spatiotemporal map of the surgical process, and the surgical event pattern recognition module... The module receives the output graph from the multi-source data fusion processing module, defines and labels surgical events through pattern recognition algorithms, and forms a structured surgical timeline. The surgical knowledge graph construction module receives the output timeline from the surgical event pattern recognition module, and constructs a knowledge graph containing surgical procedures, steps, instruments, anatomy, complications, personnel entities and relationships based on massive surgical data, performing deep data association and semantic understanding. The cross-device data compatibility and adaptation module connects with each module, integrates information silos and data incompatibility issues of different countries and different devices, and is compatible with the format protocols of multiple data sources such as surgical field videos, endoscopic images, ultrasound images, and electronic medical records.
[0019] The multi-channel visual information capture module serves as the core foundation for the platform's data acquisition. It constructs a comprehensive visual acquisition network by strategically deploying 8-12 high-definition cameras and 2-4 surgical field cameras within the operating room. The high-definition cameras utilize 4K resolution (3840×2160 pixels) and support a frame rate of 60 frames per second. The surgical field cameras are equipped with 10x optical zoom lenses and macro shooting capabilities, covering a focal length range of 5-50 mm, enabling precise capture of the surgical incision interior and instrument manipulation details. During implementation, cameras are deployed in key locations such as the four corners of the operating room ceiling, both sides of the operating table, and above the instrument table. The surgical field cameras are fixed to the end of the surgical microscope or laparoscopic equipment using sterile brackets, simultaneously capturing the entire surgical video stream, instrument opening and closing movements, grip postures, and interactive footage such as the positioning, movement, and gestures of surgical team members. All acquisition devices support real-time data transmission and adopt a dual-link transmission mode of HDMI 2.1 interface and Gigabit Ethernet to ensure that the visual information transmission delay is controlled within 50 milliseconds. This provides high-fidelity and high-timeliness raw data support for subsequent modules. Its significance lies in breaking through the limitations of traditional single-view acquisition, realizing the capture of visual information in the surgical scene without blind spots and in all dimensions, and providing a sufficient data foundation for subsequent intelligent recognition and analysis.
[0020] The deep learning real-time recognition and tracking module receives high-definition video streams and operation screen data transmitted from multiple visual information capture modules. It processes the data based on a pre-trained deep learning network architecture, which includes 12 convolutional layers, 4 recurrent neural network layers, and 3 fully connected layers. The input image is cropped and resized to 256×256 pixels before feature extraction. In practice, the convolutional layers first extract 128-dimensional feature vectors of surgical instruments, including edge contours and texture details. These are then combined with 86-dimensional feature parameters such as skeletal markers and tissue color of anatomical structures to accurately identify over 30 commonly used surgical instruments, such as forceps and hemostats, and over 20 core anatomical structures, such as the liver and heart. The recognition response time is less than 100 milliseconds. Simultaneously, for lesion areas, multi-scale feature fusion technology is used to analyze the grayscale distribution range, texture density, and curvature changes of the morphological contour, enabling proactive identification of early, small lesions. In terms of surgical step tracking, 15-20 key operation steps, such as incision and suturing, are continuously monitored based on temporal feature matching technology. The execution time of each step is recorded (accuracy to 0.1 seconds) and the operation trajectory coordinates (error ≤ 1 mm). By extracting 106 facial feature points, 24 limb joint motion parameters, and the unique identification code of the surgical gown, the identities of 3-8 surgical team members are confirmed. The coordinate information of the members in the three-dimensional space of the operating room is obtained with the help of binocular visual positioning technology (X, Y, Z axis accuracy ≤ 5 cm). The amplitude range (0-180 degrees), execution frequency (0.5-5 times / minute), and time interval between actions of key actions are recorded. Finally, by constructing a multi-dimensional collaboration evaluation matrix, the team's action response speed, operation coordination gap and other indicators are quantitatively calculated to form a collaboration efficiency evaluation result.
[0021] The multi-source data fusion processing module receives the recognition results, tracking data, and evaluation information output by the deep learning real-time recognition and tracking module. Simultaneously, it connects to audio acquisition devices and data streams from various surgical equipment within the operating room. The audio acquisition utilizes an omnidirectional microphone array (4-6 microphone units), supporting a 48kHz sampling rate and 24-bit depth audio capture, covering a frequency range of 20Hz-20kHz, accurately capturing audio information such as doctor's voice commands and equipment alarm sounds. The equipment data stream includes parameters such as the magnification of the surgical microscope, the air pressure value of the laparoscope, and the power output of the electrosurgical unit, with a data update frequency of 10 times / second, transmitted using the DICOM 3.0 medical data standard format. During implementation, the video data undergoes inter-frame differential processing and keyframe extraction, extracting one keyframe every three frames to form structured video data. Simultaneously, the audio data undergoes spectral analysis and feature extraction to separate the speech signal from environmental noise. The equipment data stream is subjected to outlier detection and effective data filtering, retaining operational parameters directly related to the surgical procedure. Subsequently, a spatiotemporal alignment algorithm was employed, using UTC timestamps as a benchmark (with millisecond-level accuracy), to synchronously associate video keyframes, audio feature segments, and device parameter data along the timeline. A three-dimensional spatial coordinate system was established with the center of the operating table as the origin, mapping data of each dimension to corresponding spatiotemporal nodes. This constructed a three-dimensional spatiotemporal map of the surgical process, including time, space, and data attributes. The map data storage adopted a distributed database, supporting more than 1,000 data entries written and queried per second. Its core implementation lies in breaking the independent storage mode of different types of data, realizing the organic integration and correlation mapping of multi-source data, and providing a unified data carrier for subsequent surgical event identification.
[0022] The surgical event pattern recognition module uses the spatiotemporal map output by the multi-source data fusion processing module as its processing object. It detects and labels key surgical events based on an improved pattern recognition algorithm, which includes three core steps: feature extraction, rule matching, and event determination. During implementation, feature mining is first performed on the data nodes in the spatiotemporal map to extract abrupt change features (such as sudden changes in equipment power or a sharp increase in movement amplitude), correlation features (such as the positional association between instrument operation and anatomical structures), and temporal features (such as the time sequence regularity of continuous movements), constructing a 64-dimensional event feature vector. Subsequently, using the event feature library established during the training phase, feature thresholds for different types of key surgical events are defined. For example, the feature threshold for surgical incision events is set as instrument-skin contact time ≥ 3 seconds and movement amplitude ≥ 45 degrees; the feature threshold for hemostasis events is set as electrosurgical power ≥ 30W and duration ≥ 5 seconds. Simultaneously, event association rules are set to clarify the sequential logic and dependencies between events. Based on the above standards, each data node in the spatiotemporal graph is matched and judged one by one. Nodes that meet the characteristics of key surgical events are marked, and data such as the start time (accurate to 0.1 seconds), duration, associated instrument number, personnel identity, and anatomical location of the event are recorded. Finally, the timeline generation unit sorts all marked events in chronological order, integrates the surgical step description, instrument usage records, personnel action data, and other information corresponding to each event, and forms a structured surgical timeline. The timeline is stored in XML format, and each event node includes more than 20 attribute fields, supporting multi-dimensional retrieval by time interval, event type, etc. Its value lies in transforming disordered surgical process data into an ordered and traceable event sequence, providing structured data support for the construction of surgical knowledge graphs.
[0023] The surgical knowledge graph construction module receives structured surgical timeline data output by the surgical event pattern recognition module. Combined with over 100,000 historical surgical cases accumulated on the platform, it employs a step-by-step implementation process of entity extraction, relationship mining, and graph construction. During implementation, natural language processing and structured data parsing technologies are first used to extract six major categories of entities from the surgical timeline and historical data: surgical procedures, surgical steps, surgical instruments, anatomical structures, complications, and surgical team members. The surgical procedure entity includes over 500 common surgical types from 12 departments such as general surgery and orthopedics. The surgical step entity includes over 200 standard steps throughout the entire process, including preoperative preparation, intraoperative procedures, and postoperative care. Each entity includes 15-20 descriptive items such as name, code, and attribute parameters. Subsequently, entity relationship extraction technology is used to mine the relationships between different entities, such as the inclusion relationship between surgical procedures and corresponding surgical steps, the adaptation relationship between surgical steps and required instruments, the positional relationship between anatomical structures and lesion areas, and the execution relationship between personnel and operational steps. Relationship strength is characterized using a quantification method, calculated based on indicators such as co-occurrence frequency and logical relevance. Based on the extracted entities and relationships, a multi-dimensional surgical knowledge graph is constructed using a graph database. The graph includes three core components: entity nodes, relationship edges, and attribute information. It supports path querying, association analysis, and semantic reasoning between entities. The graph is updated daily, incorporating newly added surgical data. A distributed storage architecture is used to ensure data security and access efficiency. The core of its implementation lies in achieving deep association and semantic understanding of surgical data, breaking down data silos, and providing knowledge support for clinical decision-making, surgical teaching, and medical research.
[0024] The cross-device data compatibility adaptation module serves as the platform's underlying compatibility support, establishing a bidirectional data interaction channel with the aforementioned modules. It addresses data compatibility issues across different countries and brands of equipment through three main technical approaches: hardware interface adaptation, software protocol conversion, and data format standardization. During implementation, the hardware is equipped with a 16-channel multi-protocol interface conversion unit, supporting 8 video interfaces including HDMI, DVI, and S-Video; 6 data interfaces including RJ45 and USB 3.0; and 5 medical-specific interfaces including HL7 and DICOM. This allows direct connection to surgical equipment, imaging equipment, and electronic medical record systems from countries such as the United States, Germany, and Japan. On the software side, a multi-protocol parsing engine is developed, incorporating parsing rules for over 20 internationally recognized data protocols. This engine can parse and convert over 10 formats of surgical field videos (MP4, AVI, etc.), 8 formats of endoscopic and ultrasound images (DICOM, BMP, etc.), and 6 formats of electronic medical records (XML, JSON, etc.). During data processing, the format of data from different sources is first identified. Based on features such as data header identifiers and field structures, the data type and original format are determined. Then, data is converted according to the platform's unified data standards. Video data is uniformly converted to H.265 encoding format, with a resolution standardized to 1920×1080 pixels and a frame rate of 30 frames per second. Image data is uniformly converted to 16-bit grayscale image format, with resolutions adapted to two standards: 512×512 pixels and 1024×1024 pixels. Text-based medical record data is uniformly converted to UTF-8 encoded structured data tables. Simultaneously, a data compatibility standard library is established, and the format protocol information of newly added devices is updated in real time. A data verification mechanism ensures the integrity and accuracy of the converted data, with a verification pass rate of ≥99.8%. The significance of this implementation lies in breaking down information barriers between different devices and systems, achieving seamless access and unified processing of multi-source data such as surgical field videos, endoscopic images, ultrasound images, and electronic medical records, providing underlying compatibility assurance for the platform's entire data processing workflow.
[0025] Preferably, the deep learning real-time recognition and tracking module includes an instrument anatomical structure recognition unit, a surgical step tracking unit, a personnel information recognition unit, and a collaboration efficiency evaluation unit. The instrument anatomical structure recognition unit receives visual information output from a multi-channel visual information capture module and uses a deep learning model to extract and match features of the surgical instruments' morphological characteristics, spatial position, and anatomical markers and tissue boundaries to achieve accurate recognition of surgical instruments and anatomical structures. Simultaneously, it performs multi-dimensional feature analysis on the grayscale distribution, texture features, and morphological contours of the lesion area to actively identify the lesion area. The surgical step tracking unit receives the recognition results output from the instrument anatomical structure recognition unit and, combined with the temporal characteristics and logical connections of the surgical operation, identifies the starting point, execution point, and other parameters of the surgical operation. The system continuously monitors and records the surgical procedure's progress and completion status, forming a dynamic tracking link. The personnel information identification unit receives visual information about personnel interactions from the multi-channel visual information capture module, extracts facial features, limb movement features, and clothing identification features of surgical team members, and confirms member identities through feature comparison. Simultaneously, it obtains the three-dimensional coordinate information of members in the operating room based on a visual positioning algorithm, and records the amplitude, execution frequency, and movement correlation of calibrated actions. The collaboration efficiency evaluation unit receives the step tracking data output by the surgical procedure tracking unit and the personnel movement data output by the personnel information identification unit, constructs a team collaboration behavior matrix, and completes a quantitative evaluation of team collaboration efficiency by analyzing action response time, operational coordination, and information transmission path.
[0026] Specifically, the deep learning real-time recognition and tracking module comprises four functional units, which work together to accurately identify and track surgical-related information. The instrument anatomy recognition unit receives 4K resolution, 60 frames per second visual data from the multi-channel visual information capture module. It extracts 128-dimensional feature vectors for surgical instruments and 86-dimensional feature parameters for anatomical structures using a deep learning model composed of 12 convolutional layers. Feature matching is performed on over 30 commonly used surgical instruments and over 20 types of core anatomical structures. Simultaneously, multi-scale feature fusion technology is used to analyze the grayscale distribution, texture density, and morphological contour curvature changes in the lesion area. The entire recognition process has a response time controlled within 100 milliseconds, ensuring timeliness and accuracy. The surgical step tracking unit, based on the temporal feature matching capability of a 4-layer recurrent neural network, connects with the recognition results from the instrument anatomy recognition unit. It continuously monitors the start point, execution process, and end state of 15-20 key surgical operation steps, recording the execution time (accuracy up to 0.1 seconds) and operation trajectory coordinates (error ≤ 1 mm) for each step, forming a complete step tracking chain. The personnel information identification unit identifies 3-8 surgical team members by extracting 106 facial feature points, 24 motion parameters of limb joints, and the unique identification code of their surgical gowns. It uses binocular vision positioning technology to obtain the members' three-dimensional spatial coordinates (X, Y, and Z axis accuracy ≤ 5 cm), simultaneously recording the amplitude range (0-180 degrees), execution frequency (0.5-5 times / minute), and time intervals between key movements. The collaboration efficiency evaluation unit receives step data from the surgical procedure tracking unit and motion data from the personnel information identification unit, constructing a multi-dimensional collaborative behavior matrix. Through quantitative analysis of indicators such as action response time, operational coordination, and information transmission path, it generates an objective evaluation result of team collaboration efficiency, providing data support for optimizing the surgical process.
[0027] Preferably, the multi-source data fusion processing module includes a video data stream processing unit, an audio data stream processing unit, an equipment data stream processing unit, and a spatiotemporal graph construction unit. The video data stream processing unit receives surgical video-related data output by the deep learning real-time recognition and tracking module, performs feature extraction and inter-frame correlation analysis on video frames, extracts spatial and temporal features of the surgical scene, and forms structured video data. The audio data stream processing unit collects voice commands, equipment operation sounds, and environmental sound data in the operating room, extracts frequency features, amplitude features, and temporal features of the sound through audio feature extraction algorithms, and completes the structured processing of the audio data. The equipment data stream processing unit collects the operating parameters and sensor data of the surgical equipment, performs format parsing and feature filtering on the data, extracts calibration equipment data related to the surgical process, and forms standardized equipment data. The spatiotemporal graph construction unit receives the structured data output by the video data stream processing unit, the audio data stream processing unit, and the equipment data stream processing unit, establishes a data association model with time as the horizontal axis and spatial location as the vertical axis, maps the data of each dimension to the corresponding spatiotemporal nodes, and constructs a spatiotemporal graph including multi-dimensional information of the entire surgical process.
[0028] Specifically, the four units of the multi-source data fusion processing module operate in an orderly manner according to the data processing flow, ensuring the effective integration of multi-source data. The video data stream processing unit receives structured visual data output from the deep learning real-time recognition and tracking module. Based on the inter-frame difference algorithm, it processes the 4K resolution, 60 frames / second video stream, filtering effective data at a frequency of one key frame every three frames. Simultaneously, it extracts the spatial and temporal features of the video frames, transforming the unstructured video data into structured data including 256×256 pixel feature maps. The audio data stream processing unit uses an omnidirectional microphone array composed of 4-6 microphone units to collect audio data with a 48kHz sampling rate and 24-bit depth, covering the 20Hz-20kHz frequency range. It employs spectrum analysis technology to extract the frequency, amplitude, and temporal features of the sound, separating effective audio such as doctor's voice commands and equipment operation sounds from environmental noise, completing the structured conversion of the audio data. The equipment data stream processing unit collects equipment parameters such as the magnification of the surgical microscope, the air pressure value of the laparoscopy, and the power output of the electrosurgical unit. It receives raw data in DICOM 3.0 format at an update frequency of 10 times per second, removes invalid data through an outlier detection algorithm, and filters out key parameters directly related to the surgical process to form standardized equipment data. The spatiotemporal map construction unit, as the core integration unit, receives structured video, audio, and equipment data output from the above three units. It achieves data time alignment based on UTC timestamps (accurate to the millisecond level), establishes a three-dimensional spatial coordinate system with the center of the operating table as the origin, and maps the data of each dimension to the corresponding spatiotemporal nodes to construct a three-dimensional spatiotemporal map including time, space, and data attributes. The map is stored in a distributed database, supporting more than 1,000 data entries written and queried per second, providing a unified data carrier for subsequent event recognition.
[0029] Preferably, the surgical event pattern recognition module includes an event feature extraction unit, an event definition unit, an event marking unit, and a timeline generation unit. The event feature extraction unit receives the spatiotemporal map output by the multi-source data fusion processing module, extracts the mutation features, correlation features, and temporal features of the data in the map, and filters out feature vectors related to the calibrated surgical events. The event definition unit establishes an event feature library based on the feature vectors output by the feature extraction unit through a pattern recognition algorithm, defines the feature thresholds and correlation rules for different types of calibrated surgical events, and forms a standardized event recognition standard. The event marking unit marks the nodes in the spatiotemporal map that meet the characteristics of the calibrated surgical events according to the recognition standard output by the event definition unit, and records the start time, duration, and associated data of the events. The timeline generation unit receives the marking results output by the event marking unit, sorts the marked events in chronological order, and integrates the surgical steps, instrument usage, and personnel action information associated with the events to form a structured surgical timeline.
[0030] Specifically, the surgical event pattern recognition module comprises four units that, through progressive processing, achieve precise definition and timeline construction of surgical events. The event feature extraction unit takes the spatiotemporal map output by the multi-source data fusion processing module as input and employs feature mining algorithms to extract abrupt change features (such as sudden changes in equipment power or a sharp increase in movement amplitude), correlation features (such as the correlation between instruments and anatomical structures), and temporal features (such as the regularity of continuous movement time sequences) from the map, constructing a 64-dimensional event feature vector to provide core feature support for event recognition. The event definition unit, based on the massive surgical data accumulated during the training phase, establishes a feature library including various key surgical events. Through pattern recognition algorithms, it sets feature thresholds and association rules for different events, clarifying the feature judgment criteria for events and ensuring the standardization and uniformity of event definitions. The event labeling unit, according to the recognition criteria set by the event definition unit, matches and judges each data node in the spatiotemporal map one by one, accurately labeling nodes that meet the characteristics of key surgical events, and recording more than 20 attribute information items such as the event's start time (accurate to 0.1 seconds), duration, associated instrument number, personnel identity, and anatomical location, ensuring the completeness of the labeling information. The timeline generation unit sorts all marked key surgical events in chronological order, integrates the associated information such as surgical step descriptions, instrument usage records, and personnel action data corresponding to each event, and constructs a structured surgical timeline using XML format. It supports multi-dimensional retrieval by time interval, event type, and other dimensions, transforming disordered surgical process data into an ordered and traceable sequence of events, providing a structured data foundation for the construction of surgical knowledge graphs, while ensuring the convenience of data query and use.
[0031] Preferably, in the surgical knowledge graph construction module, the entity relationship weight calculation model is as follows: ,in, Representing entities With entity Relationship weights Representing entities In the Feature values of each feature dimension Representing entities In the Feature values of each feature dimension Indicates the total number of feature dimensions. Indicates the time decay coefficient. Representing entities The corresponding surgical time point, Representing entities The corresponding surgical time point, This represents the correlation strength adjustment coefficient. Representing entities With entity The frequency of co-occurrence.
[0032] Specifically, the entity relationship weight calculation step in the surgical knowledge graph construction module is based on massive surgical data and includes key steps such as entity feature extraction, time factor calibration, and association strength adjustment. First, feature dimension information for each of the two types of related entities is extracted from the surgical timeline and historical data. The total number of feature dimensions is set to 120-150, including entity attributes, functional associations, usage scenarios, and other aspects. These features are then quantified into calculable numerical forms. Next, the corresponding products of the two types of entities on each feature dimension are calculated, and the sum is obtained to obtain the basic association value. Simultaneously, a time decay coefficient is introduced, with a value ranging from 0.01 to 0.05, to weaken the association strength between entities with long time intervals, with the time node accurate to the minute level of the surgical operation. In the denominator, the sum of squares of the feature dimension values for each type of entity is calculated and the square root is taken. This sum is then combined with the product of the entity co-occurrence frequency and the association strength adjustment coefficient (value ranging from 0.1 to 0.3) to form a normalization factor. The relationship weights between entities are obtained by calculating the ratio of the basic association value to the normalized processing factor. The weight values are controlled between 0 and 1, with higher values indicating stronger associations. This calculation method achieves accurate quantification of entity relationships by integrating feature matching degree, temporal correlation, and co-occurrence frequency. This ensures that the entity associations in the constructed knowledge graph conform to the actual surgical logic, providing reliable support for semantic understanding and deep association of data.
[0033] Preferably, in the cross-device data compatibility adaptation module, the data format conversion model is as follows: ,in, This represents the target format data after conversion. Indicates the number of data source types. Indicates the first Format conversion matrix of data source class Indicates the first The raw data of the data source. Indicates the first Format offset of the data source class Indicates the global format calibration coefficient. This indicates the format compatibility compensation amount.
[0034] Specifically, the data format conversion function of the cross-device data compatibility and adaptation module needs to cover the processes of format parsing, conversion calculation, and calibration optimization for multiple types of data sources. First, the number of data source types is clearly defined, including 8-10 common data types such as surgical field videos, endoscopic images, ultrasound images, and electronic medical records. A dedicated format conversion matrix and format offset are preset for each data source type. The conversion matrix dimension is set to 16×16 to 32×32 based on data complexity, and the offset value ranges from 0 to 50, used to correct systematic biases between different formats. After parsing the formats of various types of raw data, linear operations are performed according to the corresponding conversion matrix and offset to obtain preliminary converted data. Then, a global format calibration coefficient is introduced, with a value ranging from 0.95 to 1.05, to unify the standard deviation of the formats of various data types. Finally, a format compatibility compensation amount is added. The compensation amount is dynamically adjusted according to the degree of data format difference, ranging from 0.1 to 2.0, to ensure that the converted data is fully compatible with the platform's preset formats. The entire conversion process is executed at a processing speed of 100-200 data entries per second, with a format conversion accuracy rate of over 99.8%. Through multi-step calculations and dynamic calibration, the format of data sources from different countries and devices is unified, providing underlying technical support for the integration and smooth transmission of multi-source data.
[0035] Preferably, in the deep learning real-time recognition and tracking module, the lesion region recognition probability model is as follows: ,in Indicates input data The probability of identifying the corresponding lesion area. Indicates the number of feature extraction layers. Indicates the first Weight coefficients of layer features Indicates input data In the The feature mapping function of the layer, This indicates the identification bias term.
[0036] Specifically, the lesion region identification probability calculation of the deep learning real-time recognition and tracking module relies on a multi-layer feature extraction and probability transformation model to ensure the accuracy of the recognition results. The feature extraction layers are set to 8-12 layers, with each layer mining features from different dimensions of the input data. The first layer focuses on basic grayscale features, the middle layers progressively extract complex features such as texture, shape, and edges, and the last two layers perform multi-feature fusion and optimization. Each layer's features are assigned a unique weight coefficient, obtained through training on massive amounts of labeled lesion data, with values ranging from 0.05 to 0.2. The weight coefficient for core feature layers is higher than that for ordinary feature layers. The input data is processed through the feature mapping function of each layer, converting the extracted features into quantized values. The product of the quantized feature values of all layers and their corresponding weight coefficients is accumulated, and then the recognition bias term (ranging from -1.0 to 1.0) is subtracted to obtain the comprehensive feature score. This score is then substituted into the probability transformation function, and through non-linear operations, the score is mapped to a recognition probability between 0 and 1. The closer the probability value is to 1, the higher the probability that the region corresponding to the input data is a lesion. The entire calculation process has a response time of 50-80 milliseconds, and the recognition probability threshold is set at 0.75. If the value is higher than this threshold, it is determined to be a lesion area. Through multi-layer feature fusion and accurate probability calculation, the active recognition of lesion areas is achieved, replacing the traditional manual judgment mode.
[0037] Preferably, in the surgical event pattern recognition module, the event matching degree calculation model is as follows: ,in, Indicates the event to be identified With standard events The degree of matching, Indicates the number of event feature dimensions. Indicates the event to be identified In the dimensional eigenvalues Represents standard events In the dimensional eigenvalues Indicates the weight of time similarity. Indicates the event to be identified With standard events Time similarity, Indicates the weight of the difference in event length. Indicates the event to be identified The duration, Represents standard events The duration.
[0038] Specifically, the event matching degree calculation of the surgical event pattern recognition module achieves accurate event matching through multi-dimensional feature comparison, time similarity calibration, and length difference correction. The number of event feature dimensions is set to 40-60, including key dimensions such as event triggering conditions, duration features, associated data, and executing entity. Values for the event to be identified and the standard event are extracted for each feature dimension, and the products of the corresponding dimension values are calculated and summed to obtain the basic feature matching value. A time similarity weight (value 0.3-0.5) is introduced. By calculating the percentage of overlap between the event to be identified and the standard event on the time axis, time similarity is obtained. The product of time similarity and the weight is used as a time calibration term and added to the basic feature matching value. The denominator is the maximum sum of the squares of the feature dimension values of the event to be identified and the standard event. Simultaneously, the product of the event length difference weight (value 0.2-0.4) and the absolute difference in duration between the two types of events is added to form a normalization factor. The event matching degree is obtained by calculating the ratio of the numerator to the denominator. The matching degree ranges from 0 to 1, and a matching degree threshold of 0.8 is set. If the value is higher than this threshold, the event to be identified is considered to be consistent with the standard event. This calculation method comprehensively considers feature matching degree, temporal correlation, and length consistency to ensure the accuracy of key surgical event identification and provide a reliable basis for the construction of structured surgical timelines.
[0039] Preferably, in the multi-source data fusion processing module, the spatiotemporal node data fusion model is as follows: Show time Fusion of spatial location node data, Indicates the weights for video data fusion. express time Video data of spatial location, Indicates the weights for audio data fusion. express time Audio data of spatial location, Indicates the weight of device data fusion. express time Spatial location device data, express time Spatial location fusion error adjustment factor, Indicates the number of dimensions in the device data. express Time of the first Dimensional device data.
[0040] Specifically, the spatiotemporal node data fusion calculation of the multi-source data fusion processing module achieves organic data integration through multi-type data weighting, error adjustment, and dynamic calibration. Fusion weights are assigned to video data, audio data, and device data, with weight values ranging from 0.3 to 0.5, dynamically adjusted according to data importance; video data typically has a higher weight than audio and device data. The quantized values of video, audio, and device data at the same spatiotemporal node are extracted, multiplied by their respective fusion weights, and summed to obtain the basic fused data. The fusion error adjustment factor ranges from 0.01 to 0.03, dynamically adjusted based on the accuracy of the data acquisition equipment and the degree of environmental interference. The number of dimensions for device data is set to 15-25, including key dimensions such as equipment operating parameters and sensor data. The absolute difference between the quantized values of video data and the quantized values of device data in each dimension is calculated, added by 1, and the logarithm is taken. All logarithmic results are accumulated, and this accumulated result is multiplied by the fusion error adjustment factor to obtain the error adjustment term. The basic fusion data is added to the error adjustment term to obtain the final spatiotemporal node fusion data. The accuracy of the fusion data is controlled within 0.01. Through multi-data weighting and dynamic error adjustment, it is ensured that the fused data can comprehensively and accurately reflect the surgical scene information of the corresponding spatiotemporal node, providing high-quality data support for the construction of surgical spatiotemporal maps.
[0041] like Figure 2 As shown, an AI-based active data collection platform for the operating room operates through the following steps: First, multiple high-definition cameras and surgical field cameras in the operating room capture surgical videos, instrument operations, and personnel interactions at preset frame rates and resolutions, simultaneously acquiring audio signals and surgical equipment data streams while maintaining data timestamp synchronization. Second, visual information is input into a deep learning model, which extracts instrument and anatomical structure features through convolutional neural networks and recurrent neural networks, focuses attention on the lesion area using an attention mechanism, and obtains personnel spatial coordinates using visual positioning technology to construct a collaborative behavior evaluation index system. Third, the platform parses multi-source data formats and extracts calibrated feature parameters. The process involves six steps: First, establishing association mapping rules, aligning data according to time axis and spatial location, and constructing a three-dimensional spatiotemporal information map. Second, mining the spatiotemporal map features, defining and calibrating surgical event identification standards through cluster analysis and rule matching, sorting and labeling events by time, and integrating data to form a structured time axis. Third, collecting massive amounts of surgical data and structured time axis data, extracting entity information, mining associations through entity relationship extraction algorithms, establishing attribute and rule bases, and constructing a multi-dimensional knowledge graph. Fourth, designing multi-protocol compatible interfaces, converting and adapting data from different sources through format parsing algorithms, establishing a compatible standard library, and ensuring smooth data transmission and unified processing.
[0042] This AI-based active data collection platform for the operating room comprehensively acquires visual data of the surgical scene through a multi-channel visual information capture module. A deep learning-based real-time recognition and tracking module accurately identifies instruments, anatomical structures, and lesion areas, tracks surgical steps, and assesses team collaboration efficiency. A spatiotemporal graph constructed by a multi-source data fusion processing module and a structured timeline formed by a surgical event pattern recognition module enable visualization and orderly presentation of the surgical process. A surgical knowledge graph construction module deeply mines entity relationships, endowing data with semantic understanding capabilities, while a cross-device data compatibility and adaptation module provides underlying support for end-to-end data processing. These interconnected modules ensure comprehensive data acquisition, accurate identification, efficient data integration, and in-depth application.
[0043] Addressing the issues of incompatible data formats and information silos from multiple sources, the platform utilizes a cross-device data compatibility and adaptation module with multi-protocol compatible interfaces. This allows for format conversion and adaptation of data from different sources and countries, including surgical field videos, endoscopic images, ultrasound images, and electronic medical records. A unified compatibility standard library is established to enable smooth transmission and integrated processing of various data types, completely breaking down data barriers. To address the drawbacks of relying on expert experience in core processes such as lesion identification, the platform uses a deep learning model as its core. Through analysis and training on massive amounts of surgical data, it proactively identifies and extracts information such as lesion areas, surgical instruments, and key procedures, replacing traditional manual judgment. This reduces the workload of experts, improves the accuracy and consistency of identification results, fully releases the value of surgical data, and provides strong support for improving medical quality.
[0044] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," "link," and "fix" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0045] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An artificial intelligence based operating room proactive data collection platform, characterized in that, include: Multi-channel visual information capture module, deep learning real-time recognition and tracking module, multi-source data fusion and processing module, surgical event pattern recognition module, surgical knowledge graph construction module, and cross-device data compatibility and adaptation module; The multi-channel visual information capture module captures surgical videos, instrument operations, and visual information of personnel interaction through multiple high-definition cameras in the operating room and a surgical field camera. The deep learning real-time recognition and tracking module receives output information from the multi-channel visual information capture module, identifies surgical instruments, anatomical structures, and lesion areas through a deep learning model, tracks and calibrates operational steps, identifies the identities, locations, and calibration actions of surgical team members, and evaluates team collaboration efficiency. The multi-source data fusion processing module receives output data from the deep learning real-time recognition and tracking module, and fuses video, audio, and device data streams to construct a spatiotemporal map of the surgical process. The surgical event pattern recognition module receives the map output from the multi-source data fusion processing module, defines and labels surgical events through pattern recognition algorithms, and forms a structured surgical timeline. The surgical knowledge graph construction module receives the timeline output from the surgical event pattern recognition module, and constructs a knowledge graph containing surgical procedures, steps, instruments, anatomy, complications, personnel entities, and relationships based on massive surgical data, enabling deep data association and semantic understanding. The cross-device data compatibility and adaptation module connects with each module, integrating information silos and data incompatibility issues between different countries and devices, and is compatible with the format protocols of multiple data sources, including surgical field videos, endoscopic images, ultrasound images, and electronic medical records.
2. The artificial intelligence-based active data collection platform for the operating room according to claim 1, characterized in that, The deep learning real-time recognition and tracking module includes an instrument anatomical structure recognition unit, a surgical step tracking unit, a personnel information recognition unit, and a collaboration efficiency evaluation unit. The instrument anatomical structure recognition unit receives visual information output from a multi-channel visual information capture module and uses a deep learning model to extract and match features of the surgical instruments' morphological characteristics, spatial location, and anatomical markers and tissue boundaries, achieving accurate recognition of surgical instruments and anatomical structures. Simultaneously, it performs multi-dimensional feature analysis on the grayscale distribution, texture features, and morphological contours of the lesion area for active lesion area recognition. The surgical step tracking unit receives the recognition results output from the instrument anatomical structure recognition unit and, combined with the temporal characteristics and logical connections of the surgical operation, identifies the starting point of the surgical operation and the execution steps. The system continuously monitors and records the progress and completion status of the surgical procedure, forming a dynamic tracking link. The personnel information identification unit receives the visual information of personnel interaction output by the multi-channel visual information capture module, extracts the facial features, limb movement features, and clothing identification features of the surgical team members, and confirms the members' identities through feature comparison. At the same time, it obtains the three-dimensional coordinate information of the members in the operating room based on the visual positioning algorithm, and records the movement amplitude, execution frequency, and movement correlation of the calibrated movements. The collaboration efficiency evaluation unit receives the step tracking data output by the surgical step tracking unit and the personnel movement data output by the personnel information identification unit, constructs a team collaboration behavior matrix, and completes the quantitative evaluation of team collaboration efficiency by analyzing the action response time, the degree of operational coordination, and the information transmission path.
3. The artificial intelligence-based active data collection platform for the operating room according to claim 1, characterized in that, The multi-source data fusion processing module includes a video data stream processing unit, an audio data stream processing unit, an equipment data stream processing unit, and a spatiotemporal graph construction unit. The video data stream processing unit receives surgical video-related data output by the deep learning real-time recognition and tracking module, performs feature extraction and inter-frame correlation analysis on video frames, extracts spatial and temporal features of the surgical scene, and forms structured video data. The audio data stream processing unit collects voice commands, equipment operation sounds, and environmental sound data from the operating room, extracts frequency features, amplitude features, and temporal features of the sounds through audio feature extraction algorithms, and completes the structured processing of the audio data. The equipment data stream processing unit collects the operating parameters and sensor data of the surgical equipment, performs format parsing and feature filtering on the data, extracts calibration equipment data related to the surgical process, and forms standardized equipment data. The spatiotemporal graph construction unit receives the structured data output by the video data stream processing unit, the audio data stream processing unit, and the equipment data stream processing unit, establishes a data association model with time as the horizontal axis and spatial location as the vertical axis, maps the data of each dimension to the corresponding spatiotemporal nodes, and constructs a spatiotemporal graph including multi-dimensional information of the entire surgical process.
4. The artificial intelligence-based active data collection platform for the operating room according to claim 1, characterized in that, The surgical event pattern recognition module includes an event feature extraction unit, an event definition unit, an event marking unit, and a timeline generation unit. The event feature extraction unit receives the spatiotemporal map output by the multi-source data fusion processing module, extracts the mutation features, correlation features, and temporal features of the data in the map, and filters out feature vectors related to the calibrated surgical events. The event definition unit establishes an event feature library based on the feature vectors output by the feature extraction unit, and defines the feature thresholds and correlation rules for different types of calibrated surgical events to form a standardized event recognition standard. The event marking unit marks the nodes in the spatiotemporal map that meet the characteristics of the calibrated surgical events according to the recognition standard output by the event definition unit, and records the start time, duration, and associated data of the events. The timeline generation unit receives the marking results output by the event marking unit, sorts the marked events in chronological order, and integrates the surgical steps, instrument usage, and personnel action information associated with the events to form a structured surgical timeline.
5. The artificial intelligence-based active data collection platform for the operating room according to claim 1, characterized in that, In the surgical knowledge graph construction module, the entity relationship weight calculation model is as follows: ,in, Representing entities With entity Relationship weights Representing entities In the Feature values of each feature dimension Representing entities In the Feature values of each feature dimension Indicates the total number of feature dimensions. Indicates the time decay coefficient. Representing entities The corresponding surgical time point, Representing entities The corresponding surgical time point, This represents the correlation strength adjustment coefficient. Representing entities With entity The frequency of co-occurrence.
6. The artificial intelligence-based active data collection platform for the operating room according to claim 1, characterized in that, In the cross-device data compatibility adaptation module, the data format conversion model is as follows: ,in, This represents the target format data after conversion. Indicates the number of data source types. Indicates the first Format conversion matrix of data source class Indicates the first The raw data of the data source. Indicates the first Format offset of the data source class Indicates the global format calibration coefficient. This indicates the format compatibility compensation amount.
7. The artificial intelligence-based active data collection platform for the operating room according to claim 1, characterized in that, In the deep learning real-time recognition and tracking module, the probability model for lesion area recognition is as follows: ,in Indicates input data The probability of identifying the corresponding lesion area. Indicates the number of feature extraction layers. Indicates the first Weight coefficients of layer features Indicates input data In the The feature mapping function of the layer, This indicates the identification bias term.
8. The artificial intelligence-based active data collection platform for the operating room according to claim 1, characterized in that, In the surgical event pattern recognition module, the event matching degree calculation model is as follows: ,in, Indicates the event to be identified With standard events The degree of matching, Indicates the number of event feature dimensions. Indicates the event to be identified In the dimensional eigenvalues Represents standard events In the dimensional eigenvalues Indicates the weight of time similarity. Indicates the event to be identified With standard events Time similarity, Indicates the weight of the difference in event length. Indicates the event to be identified The duration, Represents standard events The duration.
9. The artificial intelligence-based active data collection platform for the operating room according to claim 1, characterized in that, In the multi-source data fusion processing module, the spatiotemporal node data fusion model is as follows: Show time Fusion of spatial location node data, Indicates the weights for video data fusion. express time Video data of spatial location, Indicates the weights for audio data fusion. express time Audio data of spatial location, Indicates the weight of device data fusion. express time Spatial location device data, express time Spatial location fusion error adjustment factor, Indicates the number of dimensions in the device data. express Time of the first Dimensional device data.
10. The artificial intelligence-based active data collection platform for the operating room according to any one of claims 1-9, characterized in that, The platform operates through the following steps: First, it captures surgical videos, instrument operations, and personnel interactions using multiple high-definition cameras in the operating room and surgical field cameras at preset frame rates and resolutions, simultaneously acquiring audio signals and surgical equipment data streams while maintaining data timestamp synchronization. Second, it inputs visual information into a deep learning model, extracting instrument and anatomical structure features through convolutional neural networks and recurrent neural networks, focusing attention on the lesion area using an attention mechanism, and obtaining personnel spatial coordinates using visual positioning technology to construct a collaborative behavior evaluation index system. Third, it parses multi-source data formats, extracts and calibrates feature parameters, establishes association mapping rules, and then... The first step involves aligning the timeline with spatial location data to construct a three-dimensional spatiotemporal map. The second step involves mining the spatiotemporal map features, defining and calibrating surgical event identification standards through cluster analysis and rule matching, sorting and labeling events by time, and integrating the data to form a structured timeline. The third step involves collecting massive amounts of surgical data and structured timeline data, extracting entity information, mining relationships through entity relationship extraction algorithms, establishing attribute and rule bases, and constructing a multi-dimensional knowledge graph. The fourth step involves designing multi-protocol compatible interfaces, using format parsing algorithms to convert and adapt data from different sources, establishing a compatible standard library, and ensuring smooth data transmission and unified processing.