Scene reconstruction method and device, electronic device, and storage medium

By jointly parsing multimodal surgical image data and semantic guidance command text, and combining visual and semantic optimization techniques, a high-precision 3D scene representation is generated, which solves the problem of reconstruction accuracy under complex dynamic interference in surgical scenes and achieves accurate localization and semantic understanding of key entities.

CN122391466APending Publication Date: 2026-07-14HONG KONG POLYTECHNIC UNIVERSITY (SHENZHEN) FRONTIER TECHNOLOGY INNOVATION CENTER CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HONG KONG POLYTECHNIC UNIVERSITY (SHENZHEN) FRONTIER TECHNOLOGY INNOVATION CENTER CO LTD
Filing Date
2026-03-18
Publication Date
2026-07-14

Smart Images

  • Figure CN122391466A_ABST
    Figure CN122391466A_ABST
Patent Text Reader

Abstract

The application provides a scene reconstruction method and device, electronic equipment and storage medium, and belongs to the technical field of medical image processing. The method comprises the following steps: performing semantic visual analysis according to a semantic guidance instruction text and a plurality of multi-modal surgery image data to obtain visual attribute description text data of a target entity; performing scene reconstruction according to the plurality of multi-modal surgery image data to obtain initial surgery scene three-dimensional representation data; performing boundary feature extraction according to the plurality of multi-modal surgery image data to obtain boundary feature data corresponding to each multi-modal surgery image data; performing boundary visual optimization on the initial surgery scene three-dimensional representation data according to the boundary feature data corresponding to the plurality of multi-modal surgery image data; and performing semantic optimization on the visually optimized surgery scene three-dimensional representation data according to the visual attribute description text data to obtain target surgery scene three-dimensional representation data. The application can improve the accuracy of scene reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical image processing technology, and in particular to a scene reconstruction method and apparatus, electronic device and storage medium. Background Technology

[0002] In the field of medical image processing technology, scenes can be reconstructed from surgical video data to form a three-dimensional scene representation describing the surgical process, thus providing visualization support for surgical navigation and postoperative review. However, the three-dimensional scene representation obtained through this scene reconstruction method is mainly used to reflect the overall geometry and appearance of the scene. For key entities in the scene, such as surgical instruments, specific anatomical structures, or surgical materials, there are still significant limitations in identifying and understanding their type and location.

[0003] Currently, some methods in the industry have improved upon these scene reconstruction approaches by introducing language descriptions during the reconstruction process, enabling the generated 3D scene representation to reflect certain semantic information of entities. However, when dealing with dynamic disturbances in surgical scenes, such as tissue deformation, instrument movement, and visual occlusion, these methods struggle to maintain an accurate and stable correlation between semantic information and the dynamic geometric structure of entities. This leads to feature drift, reducing the reliability of semantic understanding of entities and thus affecting the accuracy of scene reconstruction. Therefore, improving the accuracy of scene reconstruction in surgical scenes with complex dynamic disturbances is a pressing technical problem that needs to be solved in the industry. Summary of the Invention

[0004] The main objective of this application is to propose a scene reconstruction method, apparatus, electronic device, and storage medium, which aims to improve the accuracy of scene reconstruction in surgical scenarios with complex dynamic interference.

[0005] To achieve the above objectives, a first aspect of this application proposes a scenario re-implementation method, the method comprising: Acquire multiple multimodal surgical image data corresponding to the target surgical scene arranged in chronological order, as well as semantic guidance instruction text corresponding to the target entity in the target surgical scene; Semantic visual analysis is performed based on the semantic guidance instruction text and the multiple multimodal surgical image data to obtain the visual attribute description text data of the target entity; Scene reconstruction is performed based on the multiple multimodal surgical image data to obtain initial three-dimensional representation data of the surgical scene; Boundary features are extracted from the multiple multimodal surgical image data to obtain boundary feature data corresponding to each multimodal surgical image data. Based on the boundary feature data corresponding to the multiple multimodal surgical image data, the initial surgical scene 3D representation data is visually optimized to obtain visually optimized surgical scene 3D representation data. Based on the visual attribute description text data, semantic optimization is performed on the visually optimized 3D representation data of the surgical scene to obtain the target 3D representation data of the surgical scene.

[0006] In some embodiments, the step of semantically optimizing the visually optimized 3D representation data of the surgical scene based on the visual attribute description text data to obtain the target 3D representation data of the surgical scene includes: Semantic feature rendering is performed on the visually optimized 3D representation data of the surgical scene to obtain semantic rendering feature data. Semantic similarity data is obtained by performing semantic similarity analysis on the semantic rendering feature data and the visual attribute description text data. Based on the semantic similarity data, the visually optimized 3D representation data of the surgical scene is semantically optimized to obtain the 3D representation data of the target surgical scene.

[0007] In some embodiments, the semantic rendering feature data includes multiple pixel feature data, and the semantic similarity analysis based on the semantic rendering feature data and the visual attribute description text data to obtain semantic similarity data includes: Obtain multiple Gaussian point feature data from the visually optimized 3D representation data of the surgical scene; For each pixel feature data, a contribution weight analysis is performed on the multiple Gaussian point feature data based on the pixel feature data to obtain the rendering contribution weight coefficient of each Gaussian point feature data. For each pixel feature data, feature aggregation is performed based on the multiple Gaussian point feature data and the rendering contribution weight coefficient of each Gaussian point feature data to obtain the aggregated semantic features of the pixel feature data. Semantic similarity data is obtained by performing semantic similarity analysis on the aggregated semantic features of the multiple pixel feature data and the visual attribute description text data.

[0008] In some embodiments, the step of performing boundary visual optimization on the initial 3D representation data of the surgical scene based on the boundary feature data corresponding to the plurality of multimodal surgical image data to obtain visually optimized 3D representation data of the surgical scene includes: Visual residual data is obtained by performing visual residual analysis based on the boundary feature data and the initial three-dimensional representation data of the surgical scene; The initial 3D representation data of the surgical scene is visually optimized based on the visual residual data to obtain visually optimized 3D representation data of the surgical scene.

[0009] In some embodiments, the step of performing visual residual analysis based on the boundary feature data and the initial three-dimensional representation data of the surgical scene to obtain visual residual data includes: The feature dimensions are stitched together based on the boundary feature data and the initial three-dimensional representation data of the surgical scene to obtain three-dimensional boundary stitching feature data; The three-dimensional boundary stitching feature data is subjected to nonlinear feature transformation to obtain nonlinear mapping feature data; The visual residual data is obtained by performing residual mapping on the nonlinear mapping feature data.

[0010] In some embodiments, the step of performing semantic visual parsing based on the semantic guidance instruction text and the plurality of multimodal surgical image data to obtain visual attribute description text data of the target entity includes: For each multimodal surgical image data, mask segmentation is performed based on the multimodal surgical image data to obtain a mask segmentation image of the target entity; For each multimodal surgical image data, semantic visual attribute detection is performed based on the semantic guidance instruction text, the multimodal surgical image data, and the mask segmentation image to obtain the visual attribute description text data of the target entity.

[0011] In some embodiments, after semantically optimizing the visually optimized 3D representation data of the surgical scene based on the visual attribute description text data to obtain the target 3D representation data of the surgical scene, the method further includes: Obtain a target entity query instruction for the target surgical scenario; Semantic feature rendering is performed based on the three-dimensional representation data of the target surgical scene to obtain semantic feature data of the target scene; Semantic matching is performed based on the target entity query command and the target scene semantic feature data to obtain the location information of the target entity; Based on the three-dimensional representation data of the target surgical scene and the positioning information of the target entity, a three-dimensional visualization rendering is performed to obtain the three-dimensional visualization result of the target entity.

[0012] To achieve the above objectives, a second aspect of this application provides a scene reconstruction apparatus, the apparatus comprising: The data acquisition unit is used to acquire multiple multimodal surgical image data arranged in chronological order corresponding to the target surgical scene, as well as semantic guidance instruction text corresponding to the target entity in the target surgical scene; The visual analysis unit is used to perform semantic visual analysis based on the semantic guidance instruction text and the multiple multimodal surgical image data to obtain the visual attribute description text data of the target entity; The scene reconstruction unit is used to reconstruct the scene based on the multiple multimodal surgical image data to obtain initial three-dimensional representation data of the surgical scene; The feature extraction unit is used to extract boundary features based on the multiple multimodal surgical image data to obtain boundary feature data corresponding to each of the multimodal surgical image data. The visual optimization unit is used to perform boundary visual optimization on the initial three-dimensional representation data of the surgical scene based on the boundary feature data corresponding to the multiple multimodal surgical image data, so as to obtain the visually optimized three-dimensional representation data of the surgical scene. The semantic optimization unit is used to perform semantic optimization on the visually optimized 3D representation data of the surgical scene based on the visual attribute description text data, so as to obtain the target 3D representation data of the surgical scene.

[0013] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.

[0014] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.

[0015] The scene reconstruction method, apparatus, electronic device, and storage medium proposed in this application acquire multiple multimodal surgical image data arranged in chronological order corresponding to a target surgical scene, as well as semantic guidance instruction text corresponding to target entities in the target surgical scene. Then, semantic visual parsing is performed based on the semantic guidance instruction text and the multiple multimodal surgical image data to obtain visual attribute description text data of the target entities. Next, scene reconstruction is performed based on the multiple multimodal surgical image data to obtain initial 3D representation data of the surgical scene. Further, boundary features are extracted based on the multiple multimodal surgical image data to obtain boundary feature data corresponding to each multimodal surgical image data. Then, boundary visual optimization is performed on the initial 3D representation data of the surgical scene based on the boundary feature data corresponding to the multiple multimodal surgical image data to obtain visually optimized 3D representation data of the surgical scene. Finally, semantic optimization is performed on the visually optimized 3D representation data of the surgical scene based on the visual attribute description text data to obtain the 3D representation data of the target surgical scene. Thus, this embodiment of the application achieves joint parsing of semantic guidance instruction text and image data to generate attribute description text for guiding semantic optimization, enabling the reconstruction results to accurately locate and identify specific entities; by performing geometric optimization on the initial 3D representation through boundary features, the geometric fidelity of key structures such as instrument edges and tissue boundaries is enhanced, and semantic drift in dynamic surgical scenes is suppressed; then the attribute description text is injected into the optimized 3D representation, so that the reconstruction results have semantic understanding capabilities while carrying geometric information, realizing the improvement from appearance reconstruction to semantic reconstruction. That is, this embodiment of the application can improve the accuracy of scene reconstruction in surgical scenes with complex dynamic interference. Attached Figure Description

[0016] Figure 1 This is a flowchart of the scenario re-method provided in the embodiments of this application; Figure 2 yes Figure 1 The flowchart of step S106 in the process; Figure 3 yes Figure 2 The flowchart of step S202 in the text; Figure 4 yes Figure 1 The flowchart of step S105 in the process; Figure 5 yes Figure 4 The flowchart of step S401 in the process; Figure 6 yes Figure 1 The flowchart of step S102 in the document; Figure 7 yes Figure 1 Flowchart of the steps following step S106; Figure 8This is a schematic diagram of the scene reconstruction device provided in the embodiments of this application; Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0018] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0020] Scene reconstruction methods play a crucial role in applications such as surgical navigation, postoperative review, and medical training. By reconstructing 3D surgical video data, intuitive visual support can be provided for clinical decision-making. However, existing surgical scene reconstruction methods, after generating a 3D representation, typically only reflect the overall geometric structure and appearance of the scene, making it difficult to effectively identify and understand key entities such as the type, location, and state of surgical instruments and anatomical structures. For example, existing technologies use dynamic view synthesis techniques to reconstruct surgical scenes in 4D, such as Deform3DGS and EndoRDGS based on Gaussian splashing. These methods focus on improving the visual fidelity of geometry and appearance, but the generated 3D representation lacks semantic understanding of entities within the scene (such as instruments and anatomical structures), failing to meet the needs of intraoperative perception for entity recognition and localization. Alternatively, related technologies also enhance semantic information by introducing language descriptions during reconstruction, such as 4DLangSplat. However, when faced with complex dynamic factors in surgical scenes, such as tissue deformation, instrument movement, visual occlusion, and smoke interference, semantic features are prone to drift and struggle to maintain a stable association with the geometric structure of entities. This reduces the reliability of semantic understanding of entities and affects the overall accuracy of scene reconstruction. Therefore, embodiments of this application provide a scene reconstruction method and apparatus, electronic device, and storage medium, aiming to improve the accuracy of scene reconstruction in surgical scenes with complex dynamic interference.

[0021] The scene re-enhancing method provided in this application relates to the field of medical image processing technology. The scene re-enhancing method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the scene re-enhancing method, but is not limited to the above forms.

[0022] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0023] It should be noted that in all specific embodiments of this application, when processing is required based on object information or historical data related to the object, the object's permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of an object, separate permission or consent from the object is obtained through pop-ups or redirection to a confirmation page. Only after obtaining the object's separate permission or consent is the necessary object-related data required for the proper functioning of these embodiments acquired.

[0024] Figure 1 This is an optional flowchart of the scenario re-method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S106: Step S101: Obtain multiple multimodal surgical image data arranged in chronological order corresponding to the target surgical scene, and semantic guidance instruction text corresponding to the target entity in the target surgical scene; Step S102: Perform semantic visual analysis based on semantic guidance instruction text and multiple multimodal surgical image data to obtain visual attribute description text data of the target entity; Step S103: Reconstruct the scene based on multiple multimodal surgical image data to obtain initial three-dimensional representation data of the surgical scene; Step S104: Extract boundary features based on multiple multimodal surgical image data to obtain boundary feature data corresponding to each multimodal surgical image data. Step S105: Perform boundary visual optimization on the initial three-dimensional representation data of the surgical scene based on the boundary feature data corresponding to multiple multimodal surgical image data to obtain visually optimized three-dimensional representation data of the surgical scene. Step S106: Based on the visual attribute description text data, perform semantic optimization on the visually optimized 3D representation data of the surgical scene to obtain the target 3D representation data of the surgical scene.

[0025] Steps S101 to S106 as illustrated in this embodiment involve acquiring multiple multimodal surgical image data arranged chronologically corresponding to the target surgical scene, and semantic guidance instruction text corresponding to the target entity in the target surgical scene; then performing semantic visual parsing based on the semantic guidance instruction text and the multiple multimodal surgical image data to obtain visual attribute description text data of the target entity; next, reconstructing the scene based on the multiple multimodal surgical image data to obtain initial 3D representation data of the surgical scene; further, extracting boundary features based on the multiple multimodal surgical image data to obtain boundary feature data corresponding to each multimodal surgical image data; then performing boundary visual optimization on the initial 3D representation data of the surgical scene based on the boundary feature data corresponding to the multiple multimodal surgical image data to obtain visually optimized 3D representation data of the surgical scene; finally, performing semantic optimization on the visually optimized 3D representation data of the surgical scene based on the visual attribute description text data to obtain the 3D representation data of the target surgical scene. Thus, this embodiment of the application achieves joint parsing of semantic guidance instruction text and image data to generate attribute description text for guiding semantic optimization, enabling the reconstruction results to accurately locate and identify specific entities; by performing geometric optimization on the initial 3D representation through boundary features, the geometric fidelity of key structures such as instrument edges and tissue boundaries is enhanced, and semantic drift in dynamic surgical scenes is suppressed; then the attribute description text is injected into the optimized 3D representation, so that the reconstruction results have semantic understanding capabilities while carrying geometric information, realizing the improvement from appearance reconstruction to semantic reconstruction. That is, this embodiment of the application can improve the accuracy of scene reconstruction in surgical scenes with complex dynamic interference.

[0026] In step S101 of some embodiments, the target surgical scene can refer to a specific surgical procedure for which scene reconstruction is to be performed. For example, the target surgical scene can be a thoracoscopic surgical procedure or a laparoscopic surgical procedure. It is understood that the type of target surgical scene can be adjusted according to actual needs. Multimodal surgical image data can refer to a set of image data collected during a specific surgical procedure that reflects multidimensional information of the surgical scene. For example, multimodal surgical image data can be a combination of RGB image data and depth image data corresponding to the collected target surgical scene; or, multimodal surgical image data can also be a combination of RGB image data and infrared image data corresponding to the collected target surgical scene. It is understood that the specific composition and quantity of multimodal surgical image data can be adjusted according to actual needs, but it must be ensured that the multiple multimodal surgical image data are arranged in chronological order. Target entity can refer to a key object in the surgical scene that needs to be semantically understood. For example, the target entity can be a surgical instrument, a specific anatomical structure, or surgical materials. It is understood that the specific type of target entity can be adjusted according to actual needs. Semantic guidance instruction text can refer to natural language prompts used for specific-dimensional semantic analysis of the target entity. For example, the semantic guidance instruction text could be "intraoperative environment sensing," used to instruct the analysis of the relative positional relationships of the target entity in conjunction with the surrounding tissue environment; or, the semantic guidance instruction text could be "action priority," used to instruct the priority analysis of the target entity's motion state. It is understood that the specific content of the semantic guidance instruction text can be adjusted according to actual needs.

[0027] In step S102 of some embodiments, semantic visual analysis can refer to the process of semantically identifying and extracting the visual representation of a target entity based on the analysis dimensions indicated by the semantic guidance instruction text and the visual information presented by the multimodal surgical image data, in order to generate structured descriptive text. Visual attribute description text data can refer to structured text data obtained after semantic visual analysis of the target entity, used to describe the visual attributes of the target entity. For example, if the target entity is an ultrasonic scalpel and the semantic guidance instruction text is "action priority," then the visual attribute description text data could be "ultrasonic scalpel cutting"; or, if the target entity is a gripper and the semantic guidance instruction text is "intraoperative environment sensing," then the visual attribute description text data could be "gripper close to tissue."

[0028] In step S103 of some embodiments, scene reconstruction can refer to the process of generating a description of the three-dimensional structure of the surgical scene based on the visual information contained in multiple multimodal surgical image data using three-dimensional reconstruction technology. The initial three-dimensional representation data of the surgical scene can refer to the three-dimensional data representation obtained after scene reconstruction of multiple multimodal surgical image data, used to describe the preliminary geometric structure of the surgical scene. For example, the initial three-dimensional representation data of the surgical scene can be generated by three-dimensional reconstruction of multiple multimodal surgical image data, producing a dynamic three-dimensional representation composed of multiple Gaussian primitives, where the trajectory of each Gaussian primitive is parametrically modeled through a deformation network based on Gaussian radial basis functions, enabling it to undergo smooth deformation over time to simulate non-rigid motion in the surgical scene; or, the initial three-dimensional representation data of the surgical scene can also be generated by three-dimensional reconstruction of multiple multimodal surgical image data, producing a dynamic three-dimensional representation composed of multiple Gaussian primitives, where the position and orientation of each Gaussian primitive at different times are dynamically updated through interpolation to adapt to motion changes in the surgical scene.

[0029] It should be noted that each Gaussian primitive in the initial 3D representation data of the surgical scene can include geometric attributes such as position, rotation, scaling, and opacity, used to characterize the spatial distribution and visual features of a local region in the target surgical scene at a specific time. Each Gaussian primitive can be assigned a unified latent vector data, which includes visual subspace data for carrying visual structural information and linguistic subspace data for carrying linguistic semantic information.

[0030] In step S104 of some embodiments, boundary feature extraction can refer to the process of identifying and extracting structural boundary information such as the edges and contours of target entities in the multimodal surgical image data. Boundary feature data can refer to feature data obtained after boundary feature extraction of the multimodal surgical image data, used to characterize the structural boundary information such as the edges and contours of target entities in the image. For example, if the multimodal surgical image data contains an ultrasonic scalpel, the boundary feature data can be feature information used to characterize the edge contour of the ultrasonic scalpel; or, if the multimodal surgical image data contains anatomical tissue, the boundary feature data can be feature information used to characterize the boundary line of the anatomical tissue.

[0031] In step S105 of some embodiments, boundary visual optimization can refer to the process of adjusting the initial 3D representation data of the surgical scene based on boundary feature data to enhance the clarity of the target entity boundaries. The visually optimized 3D representation data of the surgical scene can refer to the 3D data representation obtained after boundary visual optimization, which has better clarity of the target entity boundaries. For example, if the initial 3D representation data of the surgical scene consists of multiple Gaussian primitives, each Gaussian primitive is assigned a unified latent vector data, and this unified latent vector data includes visual subspace data for carrying visual structural information, then the visually optimized 3D representation data of the surgical scene can be obtained by optimizing and adjusting the visual subspace data based on the boundary feature data.

[0032] In step S106 of some embodiments, semantic optimization can refer to adjusting the visually optimized 3D representation data of the surgical scene based on the visual attribute description text data, so that the target entity in the 3D representation is associated with the corresponding semantic attributes. The target surgical scene 3D representation data can refer to the 3D data representation obtained after semantic optimization, which simultaneously carries the geometric information and semantic attributes of the target entity. For example, if the visually optimized 3D representation data of the surgical scene consists of multiple Gaussian primitives, each Gaussian primitive is assigned a unified latent vector data, and the unified latent vector data includes linguistic subspace data for carrying linguistic semantic information, then the target surgical scene 3D representation data can be obtained by optimizing and adjusting the linguistic subspace data based on the visual attribute description text data.

[0033] Please see Figure 2 In some embodiments, step S106 may include, but is not limited to, steps S201 to S203: Step S201: Perform semantic feature rendering on the visually optimized 3D representation data of the surgical scene to obtain semantic rendering feature data; Step S202: Perform semantic similarity analysis based on semantic rendering feature data and visual attribute description text data to obtain semantic similarity data; Step S203: Based on the semantic similarity data, perform semantic optimization on the visually optimized 3D representation data of the surgical scene to obtain the 3D representation data of the target surgical scene.

[0034] In step S201 of some embodiments, semantic feature rendering can refer to the process of mapping semantically related feature information in the visually optimized 3D representation data of the surgical scene to a 2D image plane. Semantic rendering feature data can refer to data obtained after semantic feature rendering of the visually optimized 3D representation data of the surgical scene, used to characterize the semantic features corresponding to each position on the 2D image plane. For example, semantic rendering feature data can be generated by weighted projection of the semantic features carried by each spatial point in the 3D representation data to produce semantic feature vector data corresponding to each pixel on the 2D image plane; or, semantic rendering feature data can also be obtained by differentiating the semantic features carried by each spatial point in the 3D representation data according to their projection position and visibility to obtain semantic feature data corresponding to each position on the 2D image plane. It is understood that the specific generation method of semantic rendering feature data can be adjusted according to actual needs.

[0035] In step S202 of some embodiments, semantic similarity analysis can refer to the process of calculating the matching degree between semantic rendering feature data and visual attribute description text data. Semantic similarity data can refer to the data obtained after semantic similarity analysis, used to characterize the degree of matching between semantic rendering feature data and visual attribute description text data.

[0036] In step S203 of some embodiments, semantic optimization can refer to adjusting the visually optimized 3D representation data of the surgical scene according to the matching degree indicated by the semantic similarity data, so that the semantic features corresponding to the target entity tend to be consistent with the semantic attributes described by the visual attribute description text data. The target surgical scene 3D representation data can refer to the 3D data representation obtained after semantic optimization, in which the semantic features corresponding to the target entity tend to be consistent with the semantic attributes described by the visual attribute description text data. For example, if the semantic similarity data indicates that the blade area of ​​the target entity in the visually optimized 3D representation data of the surgical scene has a high degree of matching with the action attribute "cutting", then the semantic feature distribution of this area remains unchanged during the optimization process, thereby maintaining its accurate association with the visual attribute description text data; or, if the semantic similarity data indicates that the handle area of ​​the target entity is incorrectly associated with the action attribute "cutting", resulting in a low matching degree, then the semantic feature distribution of this area is adjusted during the optimization process, so that it is associated with the type attribute "ultrasonic scalpel" rather than with the action attribute, thereby correcting the incorrect binding of semantic features.

[0037] It is understood that, in this embodiment, semantic feature rendering is first performed on the visually optimized 3D representation data of the surgical scene to obtain semantic rendering feature data that characterizes the semantic features corresponding to each position on the 2D image plane. Then, semantic similarity analysis is performed on the semantic rendering feature data and the visual attribute description text data to obtain semantic similarity data that characterizes the degree of matching between the two. Subsequently, the visually optimized 3D representation data of the surgical scene is adjusted according to the degree of matching indicated by the semantic similarity data, so that the semantic features corresponding to the target entity are consistent with the semantic attributes described by the visual attribute description text data, thus obtaining the target surgical scene 3D representation data. In this way, the matching status of each region of the target entity in the 3D representation with the desired semantic attributes can be accurately identified through semantic similarity data. Regions with a high degree of matching maintain their semantic feature distribution, while regions with a low degree of matching are adjusted accordingly. This reduces the occurrence of incorrect binding or drift of semantic features in different parts of the entity, enabling the semantic attributes such as the type, action, and position of the target entity to be accurately attached to the corresponding regions in 3D space, thereby improving the accuracy of semantic reconstruction.

[0038] Please see Figure 3 In some embodiments, the semantic rendering feature data includes multiple pixel feature data, and step S202 may include, but is not limited to, steps S301 to S304: Step S301: Obtain multiple Gaussian point feature data from the visually optimized 3D representation data of the surgical scene; Step S302: For each pixel feature data, perform contribution weight analysis on multiple Gaussian point feature data based on the pixel feature data to obtain the rendering contribution weight coefficient of each Gaussian point feature data. Step S303: For each pixel feature data, feature aggregation is performed based on multiple Gaussian point feature data and the rendering contribution weight coefficient of each Gaussian point feature data to obtain the aggregated semantic features of the pixel feature data. Step S304: Perform semantic similarity analysis on the aggregated semantic features of multiple pixel feature data and the visual attribute description text data to obtain semantic similarity data.

[0039] In step S301 of some embodiments, the semantic rendering feature data includes multiple pixel feature data. Each pixel feature data is used to indicate the aggregated semantic features at the corresponding pixel location on the two-dimensional image plane represented by the semantic rendering feature data. Gaussian point feature data can refer to the feature data carried by each Gaussian primitive in the visually optimized 3D representation data of the surgical scene, used to characterize the semantic attributes at that location.

[0040] In step S302 of some embodiments, contribution weight analysis can refer to the process of determining the rendering contribution of each Gaussian point feature data in the visually optimized 3D representation data of the surgical scene to the pixel position corresponding to that pixel feature data, for each pixel feature data. The rendering contribution weight coefficient can refer to a numerical value obtained after performing contribution weight analysis on the Gaussian point feature data, used to characterize the rendering contribution of that Gaussian point feature data to the specific pixel feature data. For example, the rendering contribution weight coefficient can be calculated by weighting the opacity of the Gaussian primitive corresponding to the Gaussian point feature data during projection and its distance from the pixel center; or, the rendering contribution weight coefficient can also be determined by comprehensively considering the visibility, depth order, and projection weight of the Gaussian primitive corresponding to the Gaussian point feature data during rendering. It is understood that the specific calculation method of the rendering contribution weight coefficient can be adjusted according to actual needs.

[0041] In step S303 of some embodiments, the aggregated semantic feature can refer to the feature data used to characterize the final semantic information of the pixel position, obtained by weighting and aggregating the rendering contribution weight coefficients of multiple Gaussian point feature data corresponding to each pixel feature data. For example, the aggregated semantic feature can be obtained by multiplying each Gaussian point feature data by its corresponding rendering contribution weight coefficient and then summing the results; or, the aggregated semantic feature can also be obtained by selecting the top K (e.g., K=20) Gaussian point feature data after sorting the contribution weight coefficients and then performing weighted aggregation. It is understood that the specific aggregation method of the aggregated semantic feature can be adjusted according to actual needs.

[0042] In step S304 of some embodiments, semantic similarity analysis can refer to the process of calculating the matching degree between aggregated semantic features of multiple pixel feature data and visual attribute descriptive text data. Semantic similarity data can refer to data obtained after performing semantic similarity analysis on the aggregated semantic features of multiple pixel feature data and visual attribute descriptive text data, used to characterize the degree of matching between the two. For example, semantic similarity data can be obtained by calculating the cosine similarity between the aggregated semantic features and the text features of the visual attribute descriptive text data; or, semantic similarity data can also be obtained by calculating the Euclidean distance between the aggregated semantic features and the text features of the visual attribute descriptive text data. It is understood that the specific calculation method of semantic similarity can be adjusted according to actual needs.

[0043] Understandably, this embodiment first performs semantic feature rendering on the visually optimized 3D representation data of the surgical scene to obtain semantic rendering feature data representing the semantic information corresponding to each pixel position on the 2D image plane. Then, it performs contribution weight analysis on multiple pixel feature data within the semantic rendering feature data to determine the rendering contribution weight coefficient of each Gaussian point feature data. Subsequently, it performs weighted aggregation on the Gaussian point feature data based on the rendering contribution weight coefficient to obtain aggregated semantic features of each pixel feature data. Finally, it performs semantic optimization on the visually optimized 3D representation data of the surgical scene based on the semantic similarity analysis results between the aggregated semantic features and the visual attribute description text data to obtain the target 3D representation data of the surgical scene. In this way, contribution weight analysis can accurately identify Gaussian point feature data that contributes significantly to the semantic rendering of each pixel position, and weighted aggregation can yield more representative aggregated semantic features. Semantic similarity analysis then guides the semantic optimization process, thereby reducing information loss and erroneous associations of semantic features during rendering and optimization, and improving the matching accuracy between semantic features and target entities.

[0044] Please see Figure 4 In some embodiments, step S105 may include, but is not limited to, steps S401 to S402: Step S401: Perform visual residual analysis based on boundary feature data and initial 3D representation data of the surgical scene to obtain visual residual data; Step S402: Visually optimize the initial 3D representation data of the surgical scene based on the visual residual data to obtain visually optimized 3D representation data of the surgical scene.

[0045] In step S401 of some embodiments, visual residual analysis can refer to the process of determining the difference between the visual representation of the initial three-dimensional representation of the surgical scene in the target entity boundary region and the expected boundary represented by the boundary feature data, based on boundary feature data and initial three-dimensional representation data of the surgical scene. Visual residual data can refer to data obtained after performing visual residual analysis on the initial three-dimensional representation data of the surgical scene, used to characterize the magnitude and direction of adjustment required for the initial three-dimensional representation data of the surgical scene in the target entity boundary region.

[0046] In step S402 of some embodiments, visual optimization can refer to the process of correcting the initial 3D representation data of the surgical scene according to the adjustment magnitude and direction indicated by the visual residual data, so that the visual representation of the target entity boundary region approaches the desired boundary represented by the boundary feature data. The visually optimized 3D representation data of the surgical scene can refer to the 3D data representation obtained after visually optimizing the initial 3D representation data of the surgical scene, where the clarity of the target entity boundary is improved. For example, the visually optimized 3D representation data of the surgical scene can be obtained by superimposing the visual residual data onto the visual subspace data of the corresponding Gaussian primitives in the initial 3D representation data of the surgical scene; or, the visually optimized 3D representation data of the surgical scene can also be obtained by adjusting the geometric properties or visual features of each Gaussian primitive in the initial 3D representation data of the surgical scene according to the visual residual data. It is understood that the specific implementation of visual optimization can be adjusted according to actual needs.

[0047] Understandably, in this embodiment, visual residual analysis is first performed based on boundary feature data and initial 3D representation data of the surgical scene to obtain visual residual data that characterizes the adjustment magnitude and direction of the target entity boundary region. Then, visual optimization is performed on the initial 3D representation data of the surgical scene based on the visual residual data to obtain visually optimized 3D representation data of the surgical scene. In this way, the visual residual data can accurately guide the correction of the initial 3D representation, making the visual representation of the target entity boundary region approach the desired boundary characterized by the boundary feature data. This effectively improves the clarity and geometric fidelity of the target entity boundary, providing a more accurate geometric basis for subsequent semantic optimization.

[0048] Please see Figure 5 In some embodiments, step S401 may also include, but is not limited to, steps S501 to S503: Step S501: Based on the boundary feature data and the initial three-dimensional representation data of the surgical scene, feature dimensions are stitched together to obtain three-dimensional boundary stitching feature data; Step S502: Perform nonlinear feature transformation on the three-dimensional boundary stitching feature data to obtain nonlinear mapping feature data; Step S503: Perform residual mapping on the nonlinear mapping feature data to obtain visual residual data.

[0049] In step S501 of some embodiments, the 3D boundary stitching feature data can refer to the joint feature data obtained by fusing and stitching the boundary feature data with the feature information of the target entity boundary region in the initial 3D representation data of the surgical scene along the feature dimension. For example, the 3D boundary stitching feature data can be obtained by stitching the boundary feature data with the visual subspace data, color feature data, and temporal position encoding information of the corresponding spatial points in the initial 3D representation data of the surgical scene along the feature channel; or, the 3D boundary stitching feature data can also be obtained by stitching the boundary feature data with the geometric attribute features of Gaussian primitives and the temporal encoding information of the current moment in the initial 3D representation data of the surgical scene. It is understood that the specific stitching method of the 3D boundary stitching feature data can be adjusted according to actual needs.

[0050] In step S502 of some embodiments, nonlinear feature transformation can refer to the process of performing nonlinear mapping on the 3D boundary stitching feature data to extract higher-order, more expressive feature information. Nonlinear mapping feature data can refer to intermediate feature data used to characterize higher-order feature information, obtained after performing nonlinear feature transformation on the 3D boundary stitching feature data. For example, nonlinear mapping feature data can be obtained by inputting the 3D boundary stitching feature data into a lightweight MLP module for nonlinear transformation; or, nonlinear mapping feature data can also be obtained by fusing the 3D boundary stitching feature data with visual subspace features and then inputting it into a multilayer perceptron network for transformation. It is understood that the specific implementation of the nonlinear feature transformation can be adjusted according to actual needs.

[0051] In step S503 of some embodiments, residual mapping can refer to the process of further transforming nonlinear mapping feature data to generate residual information for adjusting the initial 3D representation data of the surgical scene. Visual residual data can refer to data obtained after residual mapping of the nonlinear mapping feature data, used to characterize the magnitude and direction of adjustment required in the target entity boundary region of the initial 3D representation data of the surgical scene. For example, visual residual data can be obtained by performing a linear transformation on the nonlinear mapping feature data to generate residual values ​​for superimposing onto the original color features; or, visual residual data can also be obtained by performing convolution processing on the nonlinear mapping feature data to generate residual features for adjusting the visual subspace data. It is understood that the specific implementation of residual mapping can be adjusted according to actual needs.

[0052] Understandably, in this embodiment, the boundary feature data is first concatenated with the feature information of the target entity boundary region in the initial 3D representation data of the surgical scene, resulting in 3D boundary concatenated feature data that integrates boundary information and 3D representation features. Then, a nonlinear feature transformation is performed on the 3D boundary concatenated feature data to extract higher-order nonlinear mapping feature data. Finally, residual mapping is performed on the nonlinear mapping feature data to generate visual residual data. In this way, multi-source information can be fused through feature concatenation, higher-order features can be extracted using nonlinear transformation, and more accurate adjustment amounts can be generated through residual mapping. This allows the visual residual data to more accurately reflect the required correction magnitude and direction of the target entity boundary region, providing a more accurate adjustment basis for subsequent visual optimization.

[0053] Please see Figure 6 In some embodiments, step S102 includes, but is not limited to, steps S601 to S602: Step S601: For each multimodal surgical image data, perform mask segmentation based on the multimodal surgical image data to obtain the mask segmentation image of the target entity; Step S602: For each multimodal surgical image data, semantic visual attribute detection is performed based on the semantic guidance instruction text, the multimodal surgical image data, and the mask segmentation image to obtain the visual attribute description text data of the target entity.

[0054] In step S601 of some embodiments, mask segmentation can refer to the process of identifying and separating the pixel region corresponding to the target entity in the image based on the visual information contained in the multimodal surgical image data. The mask-segmented image can refer to a binary image or probability map obtained by masking the multimodal surgical image data, used to mark the specific pixel position of the target entity in the image. For example, the mask-segmented image can be obtained by inputting the multimodal surgical image data into a SAM segmentation model; or, the mask-segmented image can also be obtained by performing feature extraction and pixel classification on the multimodal surgical image data. It is understood that the specific method of generating the mask-segmented image can be adjusted according to actual needs.

[0055] In step S602 of some embodiments, semantic visual attribute detection can refer to the process of semantically identifying and extracting the visual representation of a target entity based on the analysis dimension indicated by the semantic guidance instruction text, combined with the visual information presented by multimodal surgical image data and masked segmentation images, to generate structured descriptive text. Visual attribute descriptive text data can refer to structured text data obtained after semantic visual attribute detection of the target entity, used to describe the specific attributes of the target entity in terms of type, action, position, or other visual dimensions. For example, if the target entity is an ultrasonic scalpel, the semantic guidance instruction text is "action priority," and the masked segmentation image marks the area where the ultrasonic scalpel is located, then the visual attribute descriptive text data can be "ultrasonic scalpel cutting" obtained by guiding the Qwen2.5-VL model to focus on the analysis of this area based on the masked image; or, if the target entity is a clamp, the semantic guidance instruction text is "intraoperative environment sensing," and the masked segmentation image marks the clamp and its surrounding tissue area, then the visual attribute descriptive text data can be "clamp close to tissue" obtained by guiding the GPT-4V model to focus on the contact relationship between the clamp and surrounding tissue based on the masked image. It is understandable that the specific content of the visual attribute description text data can be determined jointly based on the analysis dimensions of the semantic guidance instruction text and the visual information presented by the multimodal surgical image data.

[0056] Understandably, this embodiment first performs mask segmentation on each multimodal surgical image data to obtain a mask segmentation image used to mark the specific pixel positions of the target entity in the image; then, based on the analysis dimensions indicated by the semantic guidance instruction text, semantic visual attribute detection is performed by combining the visual information presented by the multimodal surgical image data and the mask segmentation image to obtain visual attribute description text data describing the specific attributes of the target entity in dimensions such as type, action, and position. In this way, the spatial range of the target entity can be limited by the mask segmentation image, reducing the interference of background noise on semantic analysis, enabling semantic visual attribute detection to focus on the target entity region, thereby extracting visual attributes that match the semantic guidance instruction, generating structured description text data, and providing more reliable data support for subsequent semantic optimization.

[0057] Please see Figure 7 In some embodiments, after step S106, steps S701 to S704 may also be included, but are not limited to: Step S701: Obtain the target entity query instruction for the target surgical scenario; Step S702: Perform semantic feature rendering based on the 3D representation data of the target surgical scene to obtain semantic feature data of the target scene; Step S703: Perform semantic matching based on the target entity query command and the target scene semantic feature data to obtain the location information of the target entity; Step S704: Perform 3D visualization rendering based on the 3D representation data of the target surgical scene and the positioning information of the target entity to obtain the 3D visualization result of the target entity.

[0058] In step S701 of some embodiments, the target entity query instruction may refer to natural language text input by the user that describes the target entity to be retrieved or located in the three-dimensional representation data of the target surgical scene. For example, the target entity query instruction may be "ultrasonic scalpel cutting on the right side" or "grasping forceps close to the tissue". It is understood that the specific content of the target entity query instruction can be adjusted according to actual needs.

[0059] In step S702 of some embodiments, semantic feature rendering can refer to the process of mapping semantically related feature information in the three-dimensional representation data of the target surgical scene to a two-dimensional image plane. Target scene semantic feature data can refer to data obtained after semantic feature rendering of the three-dimensional representation data of the target surgical scene, used to characterize the semantic features corresponding to each position on the two-dimensional image plane.

[0060] In step S703 of some embodiments, semantic matching can refer to the process of calculating the similarity between the text features of the target entity query instruction and the semantic feature data of the target scene to determine the specific location of the target entity in the scene. The location information of the target entity can refer to data obtained after semantic matching of the semantic feature data of the target scene, used to indicate the specific location or region of the target entity on the two-dimensional image plane. For example, the location information of the target entity can be a binary location mask obtained by thresholding a similarity activation map generated based on semantic matching; or, the location information of the target entity can also be the bounding box coordinates containing the target entity determined by semantic matching. It is understood that the specific form of the location information of the target entity can be adjusted according to actual needs.

[0061] In step S704 of some embodiments, 3D visualization rendering can refer to the process of generating an image or model that displays the 3D morphology of the target entity based on the 3D geometric information and the positioning information of the target entity in the 3D representation data of the target surgical scene. The 3D visualization result of the target entity can refer to the visualization data obtained after 3D visualization rendering of the 3D representation data of the target surgical scene, used to display the morphology and position of the target entity in 3D space. For example, the 3D visualization result of the target entity can be a 3D image highlighting the target entity area after rendering the 3D representation data of the target surgical scene; or, the 3D visualization result of the target entity can also be a 3D point cloud or mesh model reconstructed after extracting the Gaussian primitives corresponding to the target entity from the 3D representation data of the target surgical scene. It is understood that the specific form of the 3D visualization result of the target entity can be adjusted according to actual needs.

[0062] It is understood that the embodiments of this application first obtain a target entity query instruction for the target surgical scene; then, semantic feature rendering is performed based on the three-dimensional representation data of the target surgical scene to obtain target scene semantic feature data used to characterize the semantic features of each position on the two-dimensional image plane; furthermore, semantic matching is performed based on the target entity query instruction and the target scene semantic feature data to obtain target entity positioning information used to indicate the specific location of the target entity in the scene; finally, three-dimensional visualization rendering is performed based on the three-dimensional representation data of the target surgical scene and the target entity positioning information to obtain the three-dimensional visualization result of the target entity. In this way, semantic information in the three-dimensional representation can be mapped to the two-dimensional plane through semantic feature rendering, and then the natural language query and scene content can be accurately aligned based on semantic matching. This supports users to flexibly retrieve target entities in an open-vocabulary manner and present the search results intuitively through three-dimensional visualization, effectively improving the interactivity and understandability of the surgical scene reconstruction results, and providing more intuitive semantic support for applications such as intraoperative perception and postoperative review.

[0063] It should be noted that the scene reconstruction method proposed in this application transforms semantic guidance instructions into fine-grained visual attribute description text through semantic visual parsing, providing accurate supervision signals for semantic optimization. This enables the scene reconstruction results to support open-vocabulary natural language retrieval containing attributes such as type, action, and location. Furthermore, it enhances the geometric fidelity of instrument edges and tissue boundaries through boundary feature extraction and visual optimization. Combined with spatial structure priors and temporal modulation, it improves resistance to complex dynamic interferences such as surgical smoke, occlusion, and tissue deformation, reducing dynamic artifacts. Finally, it injects the visual attribute description text into a three-dimensional representation through semantic optimization. Combined with semantic graph rendering and weighted aggregation, it ensures that semantic features maintain a stable association with moving target entities, reducing semantic feature drift in dynamic scenes. Thus, it achieves a unity of high-fidelity reconstruction and accurate semantic retrieval in surgical scenes with complex dynamic interference, meeting the needs of intraoperative perception for entity recognition, localization, and interaction, and improving the accuracy of scene reconstruction.

[0064] Please see Figure 8 This application also provides a scene reconstruction apparatus that can implement the above-mentioned scene reconstruction method. The apparatus includes: The data acquisition unit 801 is used to acquire multiple multimodal surgical image data arranged in chronological order corresponding to the target surgical scene, as well as semantic guidance instruction text corresponding to the target entity in the target surgical scene. The visual parsing unit 802 is used to perform semantic visual parsing based on semantic guidance instruction text and multiple multimodal surgical image data to obtain visual attribute description text data of the target entity; The scene reconstruction unit 803 is used to reconstruct the scene based on multiple multimodal surgical image data to obtain initial three-dimensional representation data of the surgical scene; The feature extraction unit 804 is used to extract boundary features based on multiple multimodal surgical image data to obtain boundary feature data corresponding to each multimodal surgical image data. The visual optimization unit 805 is used to perform boundary visual optimization on the initial three-dimensional representation data of the surgical scene based on the boundary feature data corresponding to multiple multimodal surgical image data, so as to obtain the visually optimized three-dimensional representation data of the surgical scene. The semantic optimization unit 806 is used to perform semantic optimization on the visually optimized 3D representation data of the surgical scene based on the visual attribute description text data, so as to obtain the target 3D representation data of the surgical scene.

[0065] The specific implementation of this scene reconstruction device is basically the same as the specific implementation of the scene reconstruction method described above, and will not be repeated here.

[0066] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described scenario re-implementation method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0067] Please see Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 901 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 902 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 902 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and called by the processor 901 to execute the scenario re-execution method of the embodiments of this application. The input / output interface 903 is used to implement information input and output; The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904); The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.

[0068] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described scenario re-method.

[0069] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0070] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0071] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0072] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0073] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0074] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0075] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0076] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0077] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0078] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0079] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0080] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A scene reconstruction method, characterized in that, The method includes: Acquire multiple multimodal surgical image data corresponding to the target surgical scene arranged in chronological order, as well as semantic guidance instruction text corresponding to the target entity in the target surgical scene; Semantic visual analysis is performed based on the semantic guidance instruction text and the multiple multimodal surgical image data to obtain the visual attribute description text data of the target entity; Scene reconstruction is performed based on the multiple multimodal surgical image data to obtain initial three-dimensional representation data of the surgical scene; Boundary features are extracted from the multiple multimodal surgical image data to obtain boundary feature data corresponding to each multimodal surgical image data. Based on the boundary feature data corresponding to the multiple multimodal surgical image data, the initial surgical scene 3D representation data is visually optimized to obtain visually optimized surgical scene 3D representation data. Based on the visual attribute description text data, semantic optimization is performed on the visually optimized 3D representation data of the surgical scene to obtain the target 3D representation data of the surgical scene.

2. The method according to claim 1, characterized in that, The step of semantically optimizing the visually optimized 3D representation data of the surgical scene based on the visual attribute description text data to obtain the target 3D representation data of the surgical scene includes: Semantic feature rendering is performed on the visually optimized 3D representation data of the surgical scene to obtain semantic rendering feature data. Semantic similarity data is obtained by performing semantic similarity analysis on the semantic rendering feature data and the visual attribute description text data. Based on the semantic similarity data, the visually optimized 3D representation data of the surgical scene is semantically optimized to obtain the 3D representation data of the target surgical scene.

3. The method according to claim 2, characterized in that, The semantic rendering feature data includes multiple pixel feature data. The semantic similarity analysis performed based on the semantic rendering feature data and the visual attribute description text data to obtain semantic similarity data includes: Obtain multiple Gaussian point feature data from the visually optimized 3D representation data of the surgical scene; For each pixel feature data, a contribution weight analysis is performed on the multiple Gaussian point feature data based on the pixel feature data to obtain the rendering contribution weight coefficient of each Gaussian point feature data. For each pixel feature data, feature aggregation is performed based on the multiple Gaussian point feature data and the rendering contribution weight coefficient of each Gaussian point feature data to obtain the aggregated semantic features of the pixel feature data. Semantic similarity data is obtained by performing semantic similarity analysis on the aggregated semantic features of the multiple pixel feature data and the visual attribute description text data.

4. The method according to claim 1, characterized in that, The step of performing boundary visual optimization on the initial 3D representation data of the surgical scene based on the boundary feature data corresponding to the multiple multimodal surgical image data to obtain visually optimized 3D representation data of the surgical scene includes: Visual residual data is obtained by performing visual residual analysis based on the boundary feature data and the initial three-dimensional representation data of the surgical scene; The initial 3D representation data of the surgical scene is visually optimized based on the visual residual data to obtain visually optimized 3D representation data of the surgical scene.

5. The method according to claim 4, characterized in that, The step of performing visual residual analysis based on the boundary feature data and the initial 3D representation data of the surgical scene to obtain visual residual data includes: The feature dimensions are stitched together based on the boundary feature data and the initial three-dimensional representation data of the surgical scene to obtain three-dimensional boundary stitching feature data; The three-dimensional boundary stitching feature data is subjected to nonlinear feature transformation to obtain nonlinear mapping feature data; The visual residual data is obtained by performing residual mapping on the nonlinear mapping feature data.

6. The method according to claim 1, characterized in that, The step of performing semantic visual analysis based on the semantic guidance instruction text and the multiple multimodal surgical image data to obtain visual attribute description text data of the target entity includes: For each multimodal surgical image data, mask segmentation is performed based on the multimodal surgical image data to obtain a mask segmentation image of the target entity; For each multimodal surgical image data, semantic visual attribute detection is performed based on the semantic guidance instruction text, the multimodal surgical image data, and the mask segmentation image to obtain the visual attribute description text data of the target entity.

7. The method according to claim 1, characterized in that, After semantically optimizing the visually optimized 3D representation data of the surgical scene based on the visual attribute description text data to obtain the target 3D representation data of the surgical scene, the method further includes: Obtain a target entity query instruction for the target surgical scenario; Semantic feature rendering is performed based on the three-dimensional representation data of the target surgical scene to obtain semantic feature data of the target scene; Semantic matching is performed based on the target entity query command and the target scene semantic feature data to obtain the location information of the target entity; Based on the three-dimensional representation data of the target surgical scene and the positioning information of the target entity, a three-dimensional visualization rendering is performed to obtain the three-dimensional visualization result of the target entity.

8. A scene reconstruction device, characterized in that, The device includes: The data acquisition unit is used to acquire multiple multimodal surgical image data arranged in chronological order corresponding to the target surgical scene, as well as semantic guidance instruction text corresponding to the target entity in the target surgical scene; The visual analysis unit is used to perform semantic visual analysis based on the semantic guidance instruction text and the multiple multimodal surgical image data to obtain the visual attribute description text data of the target entity; The scene reconstruction unit is used to reconstruct the scene based on the multiple multimodal surgical image data to obtain initial three-dimensional representation data of the surgical scene; The feature extraction unit is used to extract boundary features based on the multiple multimodal surgical image data to obtain boundary feature data corresponding to each of the multimodal surgical image data. The visual optimization unit is used to perform boundary visual optimization on the initial three-dimensional representation data of the surgical scene based on the boundary feature data corresponding to the multiple multimodal surgical image data, so as to obtain the visually optimized three-dimensional representation data of the surgical scene. The semantic optimization unit is used to perform semantic optimization on the visually optimized 3D representation data of the surgical scene based on the visual attribute description text data, so as to obtain the target 3D representation data of the surgical scene.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the scene reconstruction method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the scene reconstruction method according to any one of claims 1 to 7.