Distributed storage and high-concurrency retrieval system of endoscopic video data
Patent Information
- Application Number
- CN202610870015.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2046-06-16
AI Technical Summary
[0003]现有的医学影像数据管理系统在应对内镜视频这一特殊数据形态时,其底层架构设计与视频数据固有特性之间存在根本性的适配缺陷;具体而言,内镜视频数据兼具体量庞大、生成速率持续且波动显著、时序关联结构复杂以及多模态信息深度交织等特性,而传统集中式存储体系所采用的静态容量规划和中心化调度模式难以动态适应这种高增速、强关联的数据形态,随着视频数据的持续累积,存储系统容易出现节点负载分布失衡、读写热点过度集中以及横向扩展能力受限等问题,跨节点间的数据分片策略与一致性保障机制难以兼顾系统效率与数据可靠性;更为关键的是,术中实时辅助决策、紧急远程会诊和快速病例回溯等临床场景需要多用户同时对特定视频片段进行精确定位与快速调取,现有检索体系主要依赖预设元数据进行静态索引匹配,缺乏对视频时序结构及内容语义的深层组织能力,在面对高并发访问请求时检索路径冗长、资源竞争加剧、响应延迟大幅上升,无法满足临床紧急场景对时效性的严格要求;这一架构适配性缺陷最终严重制约了海量内镜视频数据资源在智能化手术辅助中的深度应用与价值释放
本发明通过基于自适应分片策略在场景切换特征点位置进行语义边界切分,使分片边界与视频内容的自然语义边界对齐,避免将时序关联紧密的视频片段强制拆分导致的跨节点检索开销增大,同时通过采样各存储节点的容量占用率、读写队列深度与网络带宽利用率生成节点状态评估向量并计算负载均衡权重因子,按照动态负载均衡规则将视频分片数据映射至目标存储节点,实现视频分片在存储集群中的均衡分布,有效缓解节点负载分布失衡与读写热点过度集中的问题,提升存储系统的整体吞吐能力与横向扩展能力;根据视频结构描述信息与分片位置索引记录构建包含时序定位索引层与内容特征索引层的多层级复合索引结构,时序定位索引层采用时间区间树结构组织以支持高效的时间范围查询,内容特征索引层采用多维空间划分结构组织以支持内容语义相似性检索,通过跨层关联指针实现两个索引层之间的双向链接,使系统能够根据检索条件类型灵活选择单层访问模式或双层联合访问模式,充分利用时序结构与内容语义的深层组织能力提升检索定位精度;对并发检索请求进行请求特征解析与路由路径规划,按照节点网络拓扑距离对目标存储节点进行分组并规划独立的请求分发路径,将检索任务分发至对应存储节点执行局部检索运算,通过分布式并行处理缩短检索路径并降低资源竞争,有效提升高并发场景下的检索响应效率;对局部检索结果集按照时序关联顺序进行跨节点聚合运算,通过时序对齐处理与重复分片去重保证聚合结果的时序连续性与数据完整性,并基于请求优先级序列对聚合后的检索结果进行排序输出,确保紧急临床场景下的检索请求能够优先获得响应,从而满足术中实时辅助决策与紧急远程会诊等临床场景对时效性的严格要求,充分释放海量内镜视频数据资源在智能化手术辅助中的应用价值。
Smart Images

Figure CN122412646B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data retrieval technology, and more specifically, to a distributed storage and high-concurrency retrieval system for endoscopic video data. Background Technology
[0002] With the widespread application of minimally invasive surgical techniques and the in-depth advancement of digital medical systems, endoscopic surgery, represented by laparoscopy and digestive endoscopy, has become a core component of modern clinical diagnosis and treatment. The popularization of ultra-high-definition imaging systems and multimodal fluorescence acquisition technology has resulted in a large amount of high-resolution video data generated by each surgery. This video data fully records the entire surgical process and has important application value in scenarios such as postoperative quality assessment, clinical teaching and training, multidisciplinary joint consultation, and the development of intelligent auxiliary diagnostic systems. In order to fully realize the potential of massive endoscopic video resources, medical institutions need to build a data management platform that combines large-scale storage capacity and efficient retrieval performance.
[0003] Existing medical image data management systems suffer from fundamental incompatibility between their underlying architecture and the inherent characteristics of endoscopic video data, a unique data format. Specifically, endoscopic video data is characterized by its massive volume, continuous and fluctuating generation rate, complex temporal correlations, and deep interweaving of multimodal information. Traditional centralized storage systems, with their static capacity planning and centralized scheduling, struggle to dynamically adapt to this high-growth, highly correlated data format. As video data accumulates, storage systems are prone to issues such as unbalanced node load distribution, excessive concentration of read / write hotspots, and limited horizontal scalability. Furthermore, cross-node data sharding strategies and consistency... The existing security mechanisms struggle to balance system efficiency and data reliability. More critically, clinical scenarios such as real-time intraoperative decision support, emergency remote consultations, and rapid case retrieval require multiple users to simultaneously locate and quickly retrieve specific video segments. Existing retrieval systems primarily rely on static index matching using pre-defined metadata, lacking the ability to deeply organize the temporal structure and semantic content of videos. When faced with high-concurrency access requests, the retrieval path becomes lengthy, resource competition intensifies, and response latency increases significantly, failing to meet the stringent timeliness requirements of emergency clinical scenarios. This architectural adaptability deficiency ultimately severely restricts the in-depth application and value release of massive endoscopic video data resources in intelligent surgical assistance.
[0004] In view of this, the present invention proposes a distributed storage and high-concurrency retrieval system for endoscopic video data to solve the above problems. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of the prior art and to achieve the above objectives, the present invention provides the following technical solution: Data acquisition module: acquires the endoscopic video data stream to be stored, performs frame-level temporal parsing on the endoscopic video data stream, extracts temporal correlation markers and scene switching feature points in the video frame sequence, and generates video structure description information based on the temporal correlation markers and scene switching feature points; Data sharding module: Based on a preset adaptive sharding strategy, the endoscopic video data stream is semantically segmented to obtain video shard data. The video shard data is then mapped to the target storage nodes in the distributed storage cluster according to dynamic load balancing rules, and a sharding location index record is established on each target storage node. Structure building module: Constructs a multi-level composite index structure based on video structure description information and segment location index records. A bidirectional link between the temporal positioning index layer and the content feature index layer is realized through cross-layer association pointers. The multi-level composite index structure includes a temporal positioning index layer and a content feature index layer. The retrieval module receives a queue of concurrent retrieval requests, performs request feature parsing and routing path planning on each retrieval request in the queue, and distributes each retrieval request to the corresponding storage node based on a multi-level composite index structure to perform local retrieval operations and obtain a local retrieval result set. The results output module performs cross-node aggregation operations on the local search result set according to the temporal association order, sorts the aggregated search results based on the request priority sequence, and outputs the final search response data.
[0006] The technical effects and advantages of the distributed storage and high-concurrency retrieval system for endoscopic video data of this invention are as follows: This invention employs an adaptive sharding strategy to perform semantic boundary segmentation at scene switching feature points, aligning shard boundaries with the natural semantic boundaries of the video content. This avoids the increased cross-node retrieval overhead caused by forcibly splitting temporally related video segments. Simultaneously, it generates node status evaluation vectors and calculates load balancing weight factors by sampling the capacity utilization, read / write queue depth, and network bandwidth utilization of each storage node. Following dynamic load balancing rules, it maps video shard data to target storage nodes, achieving a balanced distribution of video shards across the storage cluster. This effectively alleviates the problems of unbalanced node load distribution and excessive concentration of read / write hotspots, improving the overall throughput and horizontal scalability of the storage system. A multi-level composite index structure, comprising a temporal positioning index layer and a content feature index layer, is constructed based on video structure description information and shard location index records. The temporal positioning index layer uses a time interval tree structure to support efficient time range queries, while the content feature index layer uses a multi-dimensional spatial partitioning structure to support content semantic similarity retrieval. Cross-layer association pointers facilitate communication between the two index layers. Bidirectional links enable the system to flexibly select between single-layer and dual-layer joint access modes based on the type of search criteria, fully leveraging the deep organizational capabilities of temporal structure and content semantics to improve search and location accuracy. For concurrent search requests, request feature parsing and routing path planning are performed. Target storage nodes are grouped according to the node network topology distance, and independent request distribution paths are planned. Search tasks are distributed to corresponding storage nodes for local search operations. Distributed parallel processing shortens the search path and reduces resource contention, effectively improving search response efficiency in high-concurrency scenarios. Local search result sets are aggregated across nodes according to temporal association order. Temporal alignment and duplicate sharding deduplication ensure the temporal continuity and data integrity of the aggregated results. The aggregated search results are sorted and output based on request priority sequence, ensuring that search requests in emergency clinical scenarios receive priority responses. This meets the stringent timeliness requirements of clinical scenarios such as intraoperative real-time decision support and emergency remote consultations, fully releasing the application value of massive endoscopic video data resources in intelligent surgical assistance. Attached Figure Description
[0007] Figure 1 This is a schematic diagram of the distributed storage and high-concurrency retrieval system for endoscopic video data according to the present invention. Detailed Implementation
[0008] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0009] Example 1 Please see Figure 1 As shown, the distributed storage and high-concurrency retrieval system for endoscopic video data in this embodiment includes: Data acquisition module: Acquires the endoscopic video data stream to be stored, performs frame-level temporal parsing on the endoscopic video data stream, extracts temporal correlation markers and scene switching feature points in the video frame sequence, and generates video structure description information based on the temporal correlation markers and scene switching feature points.
[0010] Endoscopic surgery video data originates from medical endoscopic equipment such as laparoscopes, gastroscopes, and colonoscopes. During the procedure, it is acquired in real-time by a high-definition imaging system and transmitted to a data management platform. In this embodiment, the endoscopic video data stream uses H.264 or H.265 encoding format, supports resolutions from 1920×1080 to 4096×2160, and has a frame rate range of 25 frames / second to 60 frames / second. The endoscopic video data stream enters the data receiving module of the storage system through a dedicated transmission channel on the hospital's intranet.
[0011] Because endoscopic surgery videos are characterized by strong temporal continuity and frequent scene changes, different surgical stages (such as lesion localization, resection, and hemostasis) exhibit significant differences in visual content. Directly storing and retrieving the entire video stream would lead to chaotic data organization and low retrieval efficiency. Therefore, it is necessary to perform frame-level temporal analysis on the endoscopic video data stream, identify scene transition boundaries, and extract structured descriptive information from the video to provide a foundation for subsequent semantic segmentation and index construction.
[0012] Preferably, in some possible implementations of the embodiments of the present invention, the frame-level temporal parsing and video structure description information generation method includes: establishing a frame sequence receiving buffer for the endoscope video data stream; sequentially writing continuously received video frame data into the frame sequence receiving buffer according to the video sampling clock; and assigning a frame sequence number and a timestamp identifier to each video frame; extracting adjacent frame groups in the frame sequence receiving buffer using a sliding window method; performing inter-frame difference measurement calculation on the video frames within the adjacent frame groups; marking the current frame position as a scene switching feature point when the inter-frame difference measurement value exceeds a preset scene switching threshold; dividing the frame sequence into temporal continuous segments according to the scene switching feature points; extracting visual content description vectors for the video frames within each temporal continuous segment; encapsulating the start and end timestamps, frame sequence number range, and visual content description vectors of the temporal continuous segment into temporal association markers; and generating video structure description information based on the temporal association markers of all temporal continuous segments.
[0013] In this embodiment, a frame sequence receiving buffer is established within the data receiving module. The buffer adopts a circular queue structure, and its capacity is set to 3 to 5 times the video frame rate. For a video stream with a frame rate of 30 frames per second, the buffer capacity is set to 120 frames, corresponding to a 4-second video data buffering capacity. This capacity setting is based on the following principles: firstly, it is necessary to ensure sufficient buffer depth to cope with network transmission jitter; secondly, it is necessary to avoid excessive memory consumption due to an excessively large buffer.
[0014] When video frame data enters the buffer, the system assigns a unique frame sequence number and timestamp identifier to each frame according to the video sampling clock. The frame sequence number uses a 64-bit unsigned integer, starting from the first frame when the data stream begins to be received, and incrementing sequentially; the timestamp identifier uses the Unix timestamp format, accurate to the millisecond level.
[0015] A sliding window is set in the frame sequence receiving buffer, with a window size ranging from 5 to 10 frames. In this embodiment, a window size of 7 frames is selected. The sliding window slides in the buffer with a step size of 1 frame. After each slide, an inter-frame difference measurement calculation is performed on adjacent frame groups within the window.
[0016] The specific method for calculating the inter-frame difference metric is as follows: Two adjacent frames within the window are converted to grayscale images. The similarity between the grayscale vectors constructed from the grayscale values of the two grayscale images is calculated as a structural similarity index. Similarity calculation methods include conventional methods such as cosine similarity calculation. Simultaneously, color histograms are extracted from the two frames, and the Bach distance between the vectors constructed from the histogram data is calculated. The inter-frame difference metric value is expressed by the formula: In the formula, For the first Frame and the Inter-frame difference measurement between frames; For the first Frame and the The structural similarity index of the frame image, with a value ranging from 0 to 1; For the first Frame and the The Bach distance of the frame color histogram, with a value ranging from 0 to 1; This is a weighting coefficient, with a value ranging from 0.5 to 1. In this embodiment, it is set... The weighting factor is 0.6. This weighting factor is set based on the fact that structural changes are usually more significant than color changes when switching scenes in endoscopic videos.
[0017] Preset scene switching threshold The threshold was set to 0.35, with a range of [0.25, 0.50]. Statistical analysis of endoscopic surgery videos revealed that the mean difference metric between adjacent frames within the same surgical scene was 0.08, with a standard deviation of 0.05; the mean difference metric between adjacent frames in different surgical scenes was 0.52, with a standard deviation of 0.12. Setting the threshold to 0.35 effectively distinguishes between frame changes within a scene and scene transitions, keeping the false detection rate below 3%. At that time, the first The frame is marked as a scene transition feature point, and the frame number and timestamp of the feature point are recorded.
[0018] The frame sequence is divided into multiple temporally continuous segments based on the detected scene transition feature points. If detected... If there are scene switching feature points, the frame sequence is divided into: Each temporally continuous segment consists of a sequence of video frames that are consecutive over a period of time and have consistent scene content.
[0019] For each continuous temporal segment, a visual content description vector is extracted, and samples are uniformly taken from the continuous temporal segments. One keyframe The value range is 3 to 10, and this embodiment sets... The value is 5; a 2048-dimensional feature vector is extracted for each keyframe using a pre-trained deep convolutional neural network (such as ResNet-50); The feature vectors of each keyframe are averaged using pooling to obtain the visual content description vector for that temporal continuous segment. The start and end timestamps, frame number range, and visual content description vector of the temporal continuous segment are encapsulated into temporal association tags. These temporal association tags for all temporal continuous segments are then organized chronologically to generate video structure description information. This video structure description information fully records the temporal structure and content features of the video, providing a data foundation for subsequent semantic segmentation and index construction.
[0020] Data sharding module: Based on a preset adaptive sharding strategy, the endoscopic video data stream is semantically segmented to obtain video shard data. The video shard data is then mapped to the target storage nodes in the distributed storage cluster according to dynamic load balancing rules, and a shard location index record is established on each target storage node.
[0021] Traditional video segmentation strategies typically use fixed durations or sizes for segmentation. This approach ignores the semantic boundaries of video content, easily splitting a complete surgical scene into multiple segments. This necessitates cross-segment concatenation during subsequent retrieval, increasing retrieval complexity and response latency. Therefore, this invention employs an adaptive segmentation strategy, using scene transition feature points as semantic boundaries for segmentation, ensuring that each segment contains semantically complete video content.
[0022] Meanwhile, the load status of each node in a distributed storage cluster changes dynamically over time. If a static sharding mapping rule is used, some nodes may become overloaded while others remain idle, affecting the overall system performance. Therefore, it is necessary to dynamically adjust the sharding mapping strategy based on the real-time status of each node to achieve load balancing.
[0023] The position sequence of all scene transition feature points is read from the video structure description information and used as a semantic boundary candidate set. Assuming detection... For each scene switching feature point, the semantic boundary candidate set is... ,in Indicates the first The frame sequence number of the scene switching feature point.
[0024] Because the scene switching frequency in endoscopic surgery videos is uneven, and some scenes are too short (e.g., rapid instrument changes) or too long (e.g., prolonged observation and waiting), directly using all scene switching points as segmentation boundaries would result in excessively large differences in segment size. Therefore, it is necessary to set segmentation granularity constraints for filtering.
[0025] The preset fragmentation granularity constraints include: Minimum fragment duration Set to 10 seconds, with a value range of [5 seconds, 30 seconds], to avoid generating too many fragmented small segments; Maximum fragment duration Set to 120 seconds, with a value range of [60 seconds, 300 seconds], to avoid individual fragments being too large and affecting transmission and retrieval efficiency; Minimum number of fragment frames Set to 300 frames per second (corresponding to a 10-second video length of 30 frames per second); Maximum number of fragment frames Set to 3600 frames per second (corresponding to a 120-second video length at 30 frames per second).
[0026] The specific method for screening based on fragment granularity constraints is as follows: 1. Initialize the set of valid fragment boundaries (Including the start position of the video); 2. Set the starting position of the current fragment. ; 3. Traverse each boundary point in the semantic boundary candidate set. : Calculate the time interval between the current boundary point and the start position of the partition. ;like Skip that boundary point and continue checking the next one; like ,Will join in ,renew ; like ,exist Force insertion at the fragment boundary and update and re-examine ; 4. Add the end of the video. .
[0027] Through the above screening, a set of effective partitioning boundaries is obtained, wherein the duration between adjacent boundaries satisfies the partitioning granularity constraint condition.
[0028] Perform boundary alignment segmentation on the endoscopic video data stream according to the effective segmentation boundary set. Let the effective segmentation boundary set be... The video is then divided into... The first fragment, the... Each fragment contains the frame sequence number. arrive All video frames.
[0029] Each segmented data is encapsulated as an independent video fragment data unit, and the fragment unique identifier is generated in UUID format to ensure global uniqueness in a distributed environment.
[0030] A distributed storage cluster consists of multiple storage nodes. This embodiment assumes that the cluster includes... One storage node, The value range is 8 to 64. Real-time status sampling is performed on each storage node, with a sampling period set to 500 milliseconds to 2000 milliseconds. In this embodiment, a sampling period of 1000 milliseconds is selected.
[0031] The node status information obtained through sampling includes: storage capacity utilization: the ratio of currently used storage capacity to total storage capacity, with a value range of [0,1]; read / write queue depth: the number of read / write requests currently waiting to be processed, with a value range of non-negative integers; and network bandwidth utilization: the ratio of current network transmission bandwidth to maximum available bandwidth, with a value range of [0,1]. These three indicators are combined to form a node status evaluation vector.
[0032] Preferably, in some possible implementations of the embodiments of the present invention, the method for calculating the load balancing weight factor includes: extracting each component from the node state evaluation vector, normalizing each component according to the dimension conversion rule to obtain standardized state components; assigning differentiated weight coefficients to the standardized state components, linearly combining the weighted standardized state components to obtain the node comprehensive load index, and performing a reverse mapping operation on the node comprehensive load index to generate the load balancing weight factor.
[0033] The rules for dimension conversion are as follows: Storage capacity utilization: already a dimensionless ratio, no conversion required, directly used as a standardized component; Read / write queue depth: normalized using the maximum value; Network bandwidth utilization: already a dimensionless ratio, directly used as a standardized component.
[0034] The weighting coefficients are allocated as follows: storage capacity utilization rate weight 0.3, read / write queue depth weight 0.4, and network bandwidth utilization rate weight 0.3. This weighting is based on the following: read / write queue depth directly reflects the current processing pressure of a node and has the greatest impact on the response time of sharded storage, therefore it has the highest weight; storage capacity and network bandwidth both have a significant impact on long-term storage planning, and their weights are roughly equal.
[0035] The overall node load index is obtained by weighting and summing storage capacity utilization, read / write queue depth, and network bandwidth utilization; the load balancing weight factor is obtained by performing a reverse mapping operation on the overall node load index. In the formula, For the first The load balancing weight factor for each storage node ranges from (0,1), and the sum of the weight factors for all nodes is 1. For the first The overall load index of each storage node. For the first The overall load index of each storage node; For the temperature parameter, to control the concentration of weight allocation, this embodiment sets... The value is 3.0, and the range is [1.0, 5.0]. When When the load is larger, nodes with lower loads receive higher weights, and sharding mapping tends to concentrate on low-load nodes; when When the weights are smaller, the weight distribution is more even.
[0036] Target storage nodes are selected for each video segment data unit based on load balancing weighting factors. The selection method employs a weighted random sampling strategy: for the first... Each video segment is divided into several parts, and a target storage node is randomly selected based on the load balancing weight factor of each node. Video segment data units are transmitted to the target storage node via inter-node communication channels and written to a queue. A reliable transmission protocol (such as TCP or gRPC) is used during transmission, and data integrity is verified after transmission. The mapping relationship between the unique segment identifier and the target storage node address is recorded in the segment registry. The segment registry is implemented using a distributed key-value store (such as a Redis cluster or etcd) to support high-concurrency read and write access.
[0037] It should be noted that this embodiment uses a single-replica storage strategy for illustration; in actual production environments, a multi-replica strategy (such as 3 replicas) can be used to store shards on multiple nodes simultaneously to improve data reliability. The selection of replica nodes also follows the load balancing principle.
[0038] Structure building module: Constructs a multi-level composite index structure based on video structure description information and segment location index records, and realizes bidirectional link between the temporal positioning index layer and the content feature index layer through cross-layer association pointers.
[0039] Existing video retrieval systems typically rely on single-dimensional indexes (such as time indexes or keyword indexes), making it difficult to efficiently support both time-range retrieval and content similarity retrieval simultaneously. The retrieval needs for endoscopic surgery videos are diverse: clinicians may need to quickly locate surgical records within a specific time period (time-series retrieval), may need to find historical cases similar to specific surgical scenarios (content retrieval), or may need to find video clips with specific characteristics within a specific time range (compound retrieval).
[0040] Therefore, this invention constructs a multi-level composite index structure that includes a temporal positioning index layer and a content feature index layer, and realizes bidirectional links between the two index layers through cross-layer association pointers, supporting flexible multi-condition retrieval.
[0041] The temporal range information of each continuous temporal segment is extracted from the video structure description information, including the start and end timestamps. The temporal localization index layer is organized using a time interval tree structure, with the time range as the index key and the segment's unique identifier as the index pointer value.
[0042] Preferably, in some possible implementations of the embodiments of the present invention, the method for constructing a time interval tree structure includes: using the time range of each continuous segment in the time domain as interval nodes, constructing a balanced interval tree data structure, storing the start time value, end time value, interval span value, and branch pointers pointing to child nodes in each interval node; performing interval overlap optimization processing on the interval tree data structure, associating and marking interval nodes with time overlap relationships, and adding an overlapping association linked list in the interval nodes to support fast traversal of overlapping intervals.
[0043] The data structure definition of interval tree nodes is: { interval_start: The start time of the time interval (millisecond timestamp). interval_end: The end time of the time interval (millisecond timestamp). interval_span: Interval span (milliseconds). max_end: The maximum end time of all intervals in the subtree rooted at this node. shard_id: The unique identifier for the corresponding shard. left_child: Pointer to the left child node. right_child: Pointer to the right child node parent: pointer to the parent node. overlap_list: An overlapping linked list (stores references to other intervals that overlap with the current interval in time). content_ptr: A pointer to the content feature association (pointing to the corresponding node in the content feature index layer); The interval tree is constructed using incremental insertion: for each new time interval, starting from the root node, if the start time of the new interval is less than the start time of the current node's interval, recursively insert the interval into the left subtree; otherwise, recursively insert it into the right subtree. After insertion, a balancing operation (such as red-black tree rotation) is performed to ensure the tree's height is [value missing]. ,in This represents the total number of nodes in the interval. Interval overlap optimization: When inserting a new interval, check the overlap relationship between the new interval and existing intervals, and add the overlapping intervals to each other's overlapping association list.
[0044] Visual content description vectors for each temporal continuous segment are extracted from the video structure description information. The content feature index layer is organized using a multi-dimensional spatial partitioning structure. In this embodiment, a KD-tree is selected as the index structure, with the visual content description vector as the index key and the segment unique identifier as the index pointer value.
[0045] The construction method of KD tree is as follows: adopt the median splitting strategy, recursively select the dimension with the largest variance as the splitting dimension at each level, use the median of all data points in the dimension as the splitting threshold, divide the data points into left and right parts, and recursively construct subtrees.
[0046] A content feature association pointer is embedded in each interval node of the temporal positioning index layer. This pointer points to a KD-tree node in the content feature index layer that has the same unique fragment identifier. Simultaneously, a temporal positioning association pointer is embedded in each KD-tree node of the content feature index layer, pointing to the corresponding interval node in the temporal positioning index layer.
[0047] The process of establishing cross-layer association pointers: After completing the construction of the two index layers, all fragmented data units are traversed. For each fragment, the corresponding interval node is found in the time-series positioning index layer, and the corresponding KD-tree node is found in the content feature index layer. The addresses of the two nodes are set as each other's association pointers.
[0048] Generate index version identifiers for both the time-series positioning index layer and the content feature index layer. The index version identifiers adopt semantic version number format (such as "v1.2.3") or timestamp format to track the index update history and ensure index consistency in a distributed environment.
[0049] The data structure for index management information is defined as follows: { time_index_version: Time-series positioning index layer version identifier. content_index_version: Version identifier for the content feature index layer. total_shards: The total number of shards contained in the index. time_index_root: The address of the root node of the index layer in time-series localization. content_index_root: The address of the root node of the content feature index layer. create_time: Index creation time. last_update_time: Last update time index_checksum: Index integrity checksum; Index management information is stored in the index control center. The index control center uses a distributed storage system with master-slave replication (such as ZooKeeper or Consul) to ensure high availability of index metadata.
[0050] Step S4: Receive the concurrent retrieval request queue, perform request feature parsing and routing path planning for each retrieval request in the concurrent retrieval request queue, and distribute each retrieval request to the corresponding storage node to perform local retrieval operations based on the multi-level composite index structure to obtain the local retrieval result set.
[0051] In clinical applications, multiple doctors may simultaneously access endoscopic video data. For example, postoperative quality assessment requires reviewing surgical recordings, multidisciplinary consultations require retrieving relevant case studies, and residency training requires reviewing teaching cases. These concurrent retrieval requests place stringent demands on the system's responsiveness. Traditional serial processing methods cannot meet the timeliness requirements of high-concurrency scenarios; therefore, an efficient request scheduling and distributed retrieval mechanism needs to be designed.
[0052] Establish a concurrent retrieval request receiving channel, using a message queue middleware (such as Kafka or RabbitMQ) to receive retrieval requests from various clients. The receiving channel supports multi-protocol access (HTTP / REST, gRPC, WebSocket) to adapt to different types of client applications.
[0053] Perform format validation on each retrieval request entering the receiving channel. The validation includes: 1. Request header integrity: Check if necessary authentication information, request ID, timestamp, and other fields exist; 2. Request body structure: Verify that the JSON / XML format is correct and that required fields are complete; 3. Parameter validity: Check whether the time range parameter is a valid timestamp, and whether the feature vector dimension is correct, etc.
[0054] Search requests that pass format validation are written to the search request waiting queue in the order of arrival. The waiting queue is implemented using a priority queue, which supports scheduling based on request priority.
[0055] Search requests are extracted in batches from the search request waiting queue, with the batch size set between 16 and 64; in this embodiment, a batch size of 32 is selected. The search criteria fields carried by each search request are parsed, including time range search criteria and content feature search criteria.
[0056] Preferably, in some possible implementations of the embodiments of the present invention, the rules for determining the index access strategy are as follows: Scenario 1: When the search criteria field only contains time range search conditions, the index access strategy is determined to be a time-series positioning index layer single-level access mode. In this mode, it is only necessary to perform a range query operation in the time range tree to find all range nodes that overlap with the target time range.
[0057] Scenario 2: When the search criteria field only contains content feature search conditions, the index access strategy is determined to be a single-layer access mode for the content feature index layer. In this mode, only a nearest neighbor query operation needs to be performed in the KD-tree to find the vector with the highest similarity to the query vector. 1 node Greater than zero.
[0058] Scenario 3: When the search criteria field includes both time range search conditions and content feature search conditions, the index access strategy is determined to be a two-layer joint access mode. In this mode, the candidate fragment set within the time range is first determined through the time-series positioning index layer, and then the candidate fragment set is filtered for content features in the content feature index layer through cross-layer association pointers. The two-layer joint access mode effectively reduces the search space for content retrieval and improves retrieval efficiency.
[0059] Index traversal operations are performed in a multi-level composite index structure based on a defined index access strategy.
[0060] For a single-level access mode of the time-series positioning index layer, the range query is as follows: Starting from the root node of the interval tree, check if the time interval of the current node overlaps with the query time range; if they overlap, add the segmentation identifier of the current node to the result set; if the start time of the query time range is less than or equal to the maximum end time of the left subtree, recursively query the left subtree; if the end time of the query time range is greater than or equal to the start time of the current node's interval, recursively query the right subtree; return the result set.
[0061] For the single-layer access mode of the content feature index layer, the nearest neighbor query algorithm adopts the BBF algorithm implemented by the priority queue, which limits the maximum number of accessed nodes to 200 to 500 in order to balance query accuracy and response speed.
[0062] For the two-layer joint access mode, firstly, a range query is performed to obtain a candidate shard set. Then, the content feature index nodes corresponding to each shard in the candidate shard set are traversed (quickly located through cross-layer association pointers). The cosine similarity between the query vector and the feature vectors of each node is calculated, and shards with similarity exceeding a threshold are filtered out. After obtaining the set of unique identifiers of shards that meet the search conditions, they are converted into a set of target storage node addresses according to the shard registry.
[0063] The target storage node address set is optimized for routing. The target storage nodes are grouped according to their network topology distance, calculated based on network RTT (Real-Time Tolerance) probe results. Nodes with an RTT difference of less than 5 milliseconds are grouped together, and an independent request distribution path is planned for each node group.
[0064] The retrieval task description is distributed to the corresponding target storage nodes along the planned request distribution path via the inter-node communication channel (using the gRPC protocol). After receiving the retrieval task, each target storage node performs fragment data retrieval operations locally, reads the target fragment data from its local storage, extracts matching video segment information according to the retrieval conditions, and returns the partial retrieval results.
[0065] Step S5: Perform cross-node aggregation operations on the local search result set according to the temporal association order, and sort and output the aggregated search results based on the request priority sequence to generate the final search response data.
[0066] Because distributed retrieval distributes query tasks to multiple storage nodes for parallel execution, the local retrieval results returned by each node are scattered and may be out of order. Therefore, they need to be aggregated and organized at the central node to generate a complete and ordered retrieval response. Simultaneously, in high-concurrency scenarios, different retrieval requests have different levels of urgency (e.g., urgent consultation requests should take priority over regular query requests), requiring response sorting based on request priority to ensure timely responses to critical business needs. An aggregation result collection buffer is established, employing a double-buffering mechanism to improve collection efficiency. The buffer capacity is set to 2 to 3 times the expected number of concurrent requests; in this embodiment, the buffer capacity is set to 1000. The local retrieval result sets returned by each target storage node are received. The received local retrieval result sets are grouped and stored according to the source node identifier for subsequent deduplication and aggregation processing.
[0067] Extract the video fragment data and its corresponding timestamp information from each local search result set, and perform time-series alignment processing on the video fragment data from different storage nodes according to the timestamp information.
[0068] Preferably, in some possible implementations of the embodiments of the present invention, the time alignment processing method includes: extracting the start timestamp and end timestamp from each video segment data, constructing a segment time interval mapping table with the start timestamp and end timestamp; arranging each video segment data in ascending order according to the start timestamp based on the segment time interval mapping table, verifying the time continuity of adjacent segments after arrangement, inserting gap markers when there are gaps in the time intervals of adjacent segments, and performing overlapping interval alignment correction when the time intervals of adjacent segments overlap.
[0069] Temporal continuity verification: For adjacent fragments sorted in ascending order of their start timestamps and ,test i's end timestamp and start timestamp Relationship: If Then there exists a time gap, the length of which is Insert gap markers into the resulting sequence; if Then there is a time overlap, and the overlap length is... Perform overlapping interval alignment correction: The effective start time is adjusted to This avoids repeated playback of video clips and generates a temporally ordered aggregated segment sequence.
[0070] Perform duplicate shard detection on the time-ordered aggregated shard sequence. Since the same shard may be stored on multiple nodes due to the replication mechanism, different nodes may return the same shard result. Detection method: Traverse the aggregated shard sequence, construct a hash table using the shard's unique identifier as the key, and if an entry with the same shard's unique identifier is found, it is determined to be a duplicate shard.
[0071] Perform merging processing on duplicate shards: retain the shard record with the highest relevance score (if it is a content retrieval), or retain the shard record from the node with the lowest response latency (if it is a time-series retrieval), and generate a deduplicated aggregated retrieval result set.
[0072] The aggregated search results are prioritized based on the request priority identifier of each search request. Request priorities are divided into 10 levels (level 1 to 10, with level 10 being the highest), and different priorities correspond to different clinical scenarios. Priority 9 to 10: Emergency scenarios such as emergency consultations and real-time intraoperative decision support; Priority levels 5 to 8: routine scenarios such as postoperative quality assessment and routine case review; Priority 1 to 4: Non-emergency scenarios such as teaching and training, scientific research and statistics.
[0073] Sorting rules: First, requests are sorted in descending order of priority. If priorities are the same, they are sorted in ascending order of arrival time (first come, first served). The aggregated search results are then encapsulated into a search response message according to priority. The final search response data is output to the request initiator through the search response channel. The response channel supports both synchronous and asynchronous modes: for requests with high real-time requirements, synchronous mode is used to return results directly; for requests with large amounts of result data, asynchronous mode is used to return the task ID first, and the client obtains the complete results through polling or WebSocket push.
[0074] This embodiment establishes a structured description of the video through frame-level temporal parsing and scene transition detection. An adaptive fragmentation strategy based on semantic boundaries ensures the integrity of fragmented content, and a dynamic load balancing mechanism achieves efficient utilization of storage resources. A multi-level composite index structure and cross-layer associated pointer design support flexible multi-condition retrieval. Distributed parallel retrieval and cross-node aggregation mechanisms ensure response efficiency in high-concurrency scenarios, and a priority scheduling strategy ensures timely response to critical business requests. This invention effectively solves the architectural adaptability defects of traditional medical image management systems when dealing with endoscopic video data, providing an efficient and reliable data management infrastructure for clinical diagnosis and treatment, teaching and training, and scientific research applications.
[0075] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A distributed storage and high-concurrency retrieval system for endoscopic video data, characterized in that, include: Data acquisition module: acquires the endoscopic video data stream to be stored, performs frame-level temporal parsing on the endoscopic video data stream, extracts temporal correlation markers and scene switching feature points in the video frame sequence, and generates video structure description information based on the temporal correlation markers and scene switching feature points; The data sharding module performs semantic boundary segmentation on the endoscopic video data stream based on a preset adaptive sharding strategy, obtaining video shard data. This video shard data is then mapped to target storage nodes in a distributed storage cluster according to dynamic load balancing rules, and a shard location index record is established on each target storage node. The module includes: reading the location sequence of scene transition feature points from the video structure description information; using this sequence as a semantic boundary candidate set; filtering the semantic boundary candidate set based on preset sharding granularity constraints to obtain a set of effective shard boundaries; and performing boundary alignment sharding on the endoscopic video data stream according to the effective shard boundary set. Each subsequent data segment is encapsulated into an independent video fragment data unit, and a unique fragment identifier and fragment metadata description are generated for each video fragment data unit. The current storage capacity utilization, read / write queue depth, and network bandwidth utilization of each storage node in the distributed storage cluster are sampled to generate a node status evaluation vector. The load balancing weight factor of each storage node is calculated based on the node status evaluation vector. According to the load balancing weight factor, a target storage node is selected for each video fragment data unit. The data of the video fragment data unit transmitted to the corresponding target storage node is written into the queue, and the mapping relationship between the fragment unique identifier and the target storage node address is recorded in the fragment registry. Structure building module: Constructs a multi-level composite index structure based on video structure description information and segment location index records. A bidirectional link between the temporal positioning index layer and the content feature index layer is realized through cross-layer association pointers. The multi-level composite index structure includes a temporal positioning index layer and a content feature index layer. The retrieval module receives a queue of concurrent retrieval requests, performs request feature parsing and routing path planning on each retrieval request in the queue, and distributes each retrieval request to the corresponding storage node based on a multi-level composite index structure to perform local retrieval operations and obtain a local retrieval result set. The results output module performs cross-node aggregation operations on the local search result set according to the temporal association order, sorts the aggregated search results based on the request priority sequence, and outputs the final search response data.
2. The distributed storage and high-concurrency retrieval system of endoscopic video data according to claim 1, characterized in that, Frame-level temporal analysis is performed on the endoscopic video data stream to extract temporal correlation markers and scene transition feature points from the video frame sequence. Based on these markers, video structural description information is generated, including: A frame sequence receiving buffer is established for the endoscopic video data stream. The continuously received video frame data is sequentially written into the frame sequence receiving buffer according to the video sampling clock, and a frame sequence number and timestamp identifier are assigned to each video frame. In the frame sequence receiving buffer, adjacent frame groups are extracted in a sliding window manner. Inter-frame difference measurement is performed on the video frames in the adjacent frame groups. When the inter-frame difference measurement value exceeds the preset scene switching threshold, the current frame position is marked as a scene switching feature point, and the timestamp and frame number corresponding to the scene switching feature point are recorded. The frame sequence is divided into temporal continuous segments based on scene switching feature points. Visual content description vectors are extracted from the video frames in each temporal continuous segment. The start and end timestamps, frame number ranges, and visual content description vectors of the temporal continuous segments are encapsulated into temporal association tags. Video structure description information is generated based on the temporal association tags of all temporal continuous segments.
3. The distributed storage and high-concurrency retrieval system for endoscopic video data according to claim 2, characterized in that, The load balancing weight factor for each storage node is calculated based on the node state evaluation vector, including: The storage capacity utilization rate component, read / write queue depth component, and network bandwidth utilization rate component are extracted from the node state evaluation vector. Each component is normalized according to the preset unit conversion rules to obtain the standardized state components. Differentiated weight coefficients are assigned to the standardized state components. The weighted standardized state components are then linearly combined to obtain the node comprehensive load index. A reverse mapping operation is then performed on the node comprehensive load index to generate load balancing weight factors.
4. The distributed storage and high-concurrency retrieval system for endoscopic video data according to claim 3, characterized in that, A multi-level composite index structure is constructed based on video structure description information and segment location index records. A bidirectional link between the temporal location index layer and the content feature index layer is achieved through cross-layer association pointers, including: Extract the time range information of each continuous segment in the time domain from the video structure description information, and construct a temporal positioning index layer according to the time order. The temporal positioning index layer is organized using a time interval tree structure, with the time range as the index key value and the segment unique identifier as the index pointing value. Visual content description vectors of each continuous temporal segment are extracted from the video structure description information. A content feature index layer is constructed according to the vector space distance relationship. The content feature index layer is organized using a multi-dimensional space partitioning structure, with the visual content description vector as the index key and the segment unique identifier as the index pointing value. Content feature association pointers are embedded in each index node of the temporal positioning index layer. The content feature association pointers point to the index nodes in the content feature index layer that have the same fragment unique identifier. At the same time, temporal positioning association pointers are embedded in each index node of the content feature index layer. The temporal positioning association pointers point to the corresponding index nodes in the temporal positioning index layer. Index version identifiers are generated for the time-series positioning index layer and the content feature index layer, respectively. The index version identifiers and index structure metadata are encapsulated into index management information, and the index management information is stored in the index control center.
5. The distributed storage and high-concurrency retrieval system for endoscopic video data according to claim 4, characterized in that, The time-series location index layer is organized using a time interval tree structure, including: Using the time range of each continuous segment in the time domain as interval nodes, a balanced interval tree data structure is constructed. In each interval node, the start time value, end time value, interval span value, and branch pointer to the child node are stored. The interval tree data structure is optimized by performing interval overlap processing, which involves associating interval nodes with time overlap relationships and adding an overlapping association list to the interval nodes.
6. The distributed storage and high-concurrency retrieval system for endoscopic video data according to claim 5, characterized in that, The system receives a queue of concurrent retrieval requests, performs request feature parsing and routing path planning for each request, and distributes each retrieval request to the corresponding storage node for local retrieval operations based on a multi-level composite index structure, including: Establish a concurrent retrieval request receiving channel, verify the request format of retrieval requests entering through the concurrent retrieval request receiving channel, and write the verified retrieval requests into the retrieval request waiting queue according to their arrival time. Retrieve retrieval requests in batches from the retrieval request waiting queue, parse the retrieval condition fields carried by each retrieval request, including time range retrieval conditions and content feature retrieval conditions, and determine the corresponding index access strategy based on the type combination of the retrieval condition fields; Based on the index access strategy, index traversal operations are performed in a multi-level composite index structure to obtain a set of shard unique identifiers that meet the search conditions. The set of shard unique identifiers is then converted into a set of target storage node addresses according to the shard registry. The target storage node address set is optimized by routing path optimization. The target storage nodes are grouped according to the network topology distance of the nodes, and an independent request distribution path is planned for each node group. For each retrieval request, a retrieval task description body is generated. The retrieval task description body is then distributed to the corresponding target storage node along the planned request distribution path through the inter-node communication channel. The target storage node then performs the fragmented data retrieval operation locally and returns the local retrieval results.
7. The distributed storage and high-concurrency retrieval system for endoscopic video data according to claim 6, characterized in that, The corresponding index access strategy is determined based on the combination of data types of the search criteria fields, including: When the search criteria field contains only time range search criteria, the index access strategy is determined to be a time-series positioning index layer single-level access mode. When the search criteria field contains only content feature search criteria, the index access strategy is determined to be the content feature index layer single-level access mode. When the search criteria field contains both time range search criteria and content feature search criteria, the index access strategy is determined to be a two-layer joint access mode. In the two-layer joint access mode, the candidate fragment set within the time range is first determined by the time-series positioning index layer, and then the content feature index layer filters the candidate fragment set by cross-layer association pointers.
8. The distributed storage and high-concurrency retrieval system for endoscopic video data according to claim 7, characterized in that, The local search result set is aggregated across nodes according to temporal association order, and the aggregated search results are sorted and output based on the request priority sequence, including: Establish an aggregation result collection buffer to receive local search result sets returned by each target storage node, and group and store the received local search result sets according to the source node identifier; Extract the video fragment data that was found in the search results from each local search set and its corresponding timestamp information. Perform time-series alignment processing on the video fragment data from different storage nodes according to the timestamp information to generate a time-ordered aggregated fragment sequence. Perform duplicate fragment detection and deduplication on the temporally ordered aggregated fragment sequence, merge duplicate fragment data with the same unique fragment identifier, and generate an aggregated retrieval result set; The aggregated search results are sorted by priority based on the request priority identifier of each search request. The aggregated search results are then encapsulated into a search response message according to the priority order and the final search response data is output to the request initiator through the search response channel.
9. The distributed storage and high-concurrency retrieval system for endoscopic video data according to claim 8, characterized in that, Time-series alignment of video fragment data from different storage nodes is performed based on timestamp information, including: Extract the start and end timestamps from each video segment data, and construct a segment time interval mapping table using the start and end timestamps; Based on the segmented time interval mapping table, the video segment data is sorted in ascending order according to the start timestamp. The temporal continuity of adjacent segments is verified. When there is a gap between the time intervals of adjacent segments, a gap marker is inserted. When the time intervals of adjacent segments overlap, the overlapping interval alignment correction is performed.
Citation Information
Patent Citations
Video acceleration processing method under distributed computing framework
CN119854517A
Intelligent database video retrieval method based on big data technology
CN120216722A