Training method of operation visual field evaluation model, operation visual field evaluation method and equipment

By training the surgical field evaluation model and using image cropping technology, the problem of difficulty in evaluating doctors' critical safety field operation awareness is solved, and the accurate classification and evaluation of surgical videos is achieved, and surgical safety and teaching efficiency are improved.

CN120107848APending Publication Date: 2025-06-06SUN YAT SEN MEMORIAL HOSPITAL SUN YAT SEN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510142352.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Hospitals lack effective tools to regulate doctors' intraoperative operations, making it difficult to evaluate whether doctors have a critical safety field of operation awareness. Especially for young doctors and clinical medical students, if they lack clinical experience in surgery, it is difficult to acquire critical safety field of operation awareness.

Method used

A training method for surgical field evaluation model is provided. By obtaining sample video frames and target labeling information, image cropping is performed, multiple sub-pictures are obtained, and model training is performed based on these samples to obtain a surgical field evaluation model for determining the surgical field evaluation results of input video frames.

Benefits of technology

It improves the accuracy of surgical field evaluation of video frames, enhances the classification accuracy of surgical videos, and can assist in assessing whether doctors have a critical safety field operation awareness, and improves surgical safety and teaching convenience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107848A_ABST
    Figure CN120107848A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of medicine, and provides a training method of an operation visual field evaluation model, and an operation visual field evaluation method and equipment, which can perform model training based on a sample video frame, target annotation information of the sample video frame and a plurality of sub-images corresponding to the sample video frame to obtain the operation visual field evaluation model. And determining an operation visual field evaluation result of the video frame by using the operation visual field evaluation model, determining a video clip of which the operation visual field accords with the operation condition in the operation video based on the operation visual field evaluation result of the video frame, and determining a first operation visual field evaluation result based on the time distribution condition of the video clip in the operation video. According to the embodiment of the invention, the operation visual field evaluation model can extract image details, the accuracy of operation visual field evaluation of the video frame can be improved, then the classification accuracy of the operation video can be improved, and when the operation visual field evaluation model is applied to the gall bladder resection, whether a doctor has key safe visual field operation consciousness can be evaluated; and the operation safety and the teaching convenience are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of medical technology, and in particular relates to a training method for a surgical visual field assessment model, a surgical visual field assessment method and a device. Background Art

[0002] The surgical field of view refers to the scope and content that the doctor can directly see during the operation. Whether the shape of the surgical object can be clearly presented in the surgical field of view determines the reliability of the operation and is one of the key factors in reducing postoperative complications. At present, hospitals lack effective tools to standardize doctors' intraoperative operations and it is difficult to evaluate whether doctors have the awareness of key safety field operations. For young doctors and clinical medical students, due to the lack of clinical surgical experience, it is also difficult to acquire the awareness of key safety field operations. Summary of the invention

[0003] The embodiments of the present application provide a training method, a surgical field of view assessment method and a device for a surgical field of view assessment model, which can improve the accuracy of surgical field of view assessment of video frames, and then improve the classification accuracy of surgical videos. When applied to cholecystectomy, it can assist in assessing whether the doctor has key safety field of view operation awareness, thereby improving surgical safety and teaching convenience.

[0004] A first aspect of an embodiment of the present application provides a training method for a surgical field of view assessment model, comprising: obtaining a sample video frame and target annotation information of the sample video frame; performing image cropping processing on the sample video frame to obtain multiple sub-images corresponding to the sample video frame, wherein when the sample video frame belongs to a positive sample, the surgical object is completely presented in at least part of the corresponding multiple sub-images; performing model training based on the sample video frame, the multiple sub-images corresponding to the sample video frame, and the target annotation information of the sample video frame to obtain a surgical field of view assessment model, wherein the surgical field of view assessment model is used to determine a surgical field of view assessment result of an input video frame.

[0005] In an embodiment of the present application, a model is trained based on a sample video frame, a plurality of sub-images obtained by cropping the sample video frame, and target annotation information of the sample video frame to obtain a surgical field of view assessment model for determining a surgical field of view assessment result of an input video frame. The surgical field of view assessment model can be enabled to extract image details, thereby improving the accuracy of surgical field of view assessment of video frames.

[0006] In some embodiments of the first aspect, the image cropping processing is performed on the sample video frame to obtain multiple sub-images corresponding to the sample video frame, including any of the following: taking the short side of the sample video frame as the side length of the sub-image, taking the upper left corner pixel of the sample video frame as the upper left corner pixel of the sub-image, and cropping a square area in the sample video frame to obtain a sub-image; taking the short side of the sample video frame as the side length of the sub-image, taking the upper right corner pixel of the sample video frame as the upper right corner pixel of the sub-image, and cropping a square area in the sample video frame to obtain a sub-image; taking the short side of the sample video frame as the side length of the sub-image, taking the midline of the sample video frame as the midline of the sub-image, and cropping a square area in the sample video frame to obtain a sub-image.

[0007] In some implementations of the first aspect, before performing image cropping processing on the sample video frame to obtain multiple sub-images corresponding to the sample video frame, it also includes: determining a black border area of ​​the sample video frame; and cropping the black border area in the sample video frame.

[0008] In some embodiments of the first aspect, determining the black-border area of ​​the sample video frame includes: determining the pixel mean corresponding to each column of pixels in the sample video frame, the pixel mean corresponding to each column of pixels being the mean of the pixel values ​​of a preset number of pixels located in the middle of the column of pixels; taking the pixel column whose pixel mean is greater than the preset value as the boundary of the black-border area, and determining the black-border area based on the boundary.

[0009] In some embodiments of the first aspect, determining the black border area of ​​the sample video frame includes: obtaining boundaries of the black border areas respectively corresponding to multiple sampled video frames in a sample video segment to which the sample video frame belongs; and determining the black border area based on the average of the boundaries of the black border areas respectively corresponding to the multiple sampled video frames.

[0010] A second aspect of an embodiment of the present application provides a method for evaluating a surgical field of view, including: obtaining a surgical video to be evaluated; determining a video segment in the surgical video in which the surgical field of view meets surgical conditions; and determining a first evaluation result of the surgical field of view of the surgical video based on the time distribution of the video segment in the surgical video.

[0011] In an embodiment of the present application, by determining the video segments in the surgical video to be evaluated whose surgical field of view meets the surgical conditions, and determining a first evaluation result of the surgical field of view of the surgical video based on the time distribution of the video segments in the surgical video, the surgical video can be classified in combination with the time domain information of the video segments in which the surgical field of view meets the surgical conditions, which helps to improve the classification accuracy of the surgical video. When applied to cholecystectomy, it can evaluate whether the doctor has the awareness of key safety field operation, thereby improving the safety of the operation and the convenience of teaching.

[0012] In some embodiments of the second aspect, determining the video segment in the surgical video whose surgical field meets the surgical conditions includes: inputting each video frame to be evaluated in the surgical video into a surgical field evaluation model, the surgical field evaluation model is obtained by model training based on sample video frames, multiple sub-images corresponding to the sample video frames, and target annotation information of the sample video frames, the multiple sub-images corresponding to the sample video frames are obtained by image cropping the sample video frames, and when the sample video frames belong to positive samples, the surgical object is completely presented in at least part of the corresponding multiple sub-images; obtaining the surgical field evaluation results of each video frame to be evaluated output by the surgical field evaluation model; and determining the video segment according to the surgical field evaluation results of each video frame to be evaluated.

[0013] In some embodiments of the second aspect, the multiple sub-images corresponding to the sample video frame include any one of the following: a sub-image obtained by cutting out a square area in the sample video frame with the short side of the sample video frame as the sub-image side length and the upper left corner pixel of the sample video frame as the upper left corner pixel of the sub-image; a sub-image obtained by cutting out a square area in the sample video frame with the short side of the sample video frame as the sub-image side length and the upper right corner pixel of the sample video frame as the upper right corner pixel of the sub-image; a sub-image obtained by cutting out a square area in the sample video frame with the short side of the sample video frame as the sub-image side length and the midline of the sample video frame as the midline of the sub-image.

[0014] In some implementations of the second aspect, before performing image cropping processing on the sample video frame, the sample video frame undergoes black border cropping processing, and the black border cropping processing is used to crop the black border area in the sample video frame.

[0015] In some embodiments of the second aspect, the boundary of the black border area is a pixel column whose pixel mean is greater than a preset value, wherein the pixel mean corresponding to each column of pixels is the average of the pixel values ​​of a preset number of pixel points located in the middle of the column of pixels.

[0016] In some implementations of the second aspect, the boundary of the black border area is an average of boundaries of the black border areas respectively corresponding to a plurality of sampled video frames, and the plurality of sampled video frames all originate from a sample video segment to which the sample video frame belongs.

[0017] In some embodiments of the second aspect, determining a first evaluation result of the surgical field of view of the surgical video based on the time distribution of the video clips in the surgical video includes: comparing the video clips in the surgical video in pairs to determine video clip pairs whose total number of frames is greater than a frame number threshold and whose frame number ratio is greater than a ratio threshold, the total number of frames being the number of frames contained in the two video clips and between the two video clips, and the frame number ratio being the ratio of the total number of video frames that meet the surgical conditions between the two video clips to the total number of frames; if the number of video clip pairs is greater than the number threshold, determining that the first evaluation result of the surgical field of view is that the surgical video meets the surgical conditions.

[0018] In some embodiments of the second aspect, after determining a first evaluation result of the surgical field of the surgical video based on the time distribution of the video clip in the surgical video, the method further includes: in response to a user's input operation on a second evaluation result of the surgical field of the surgical video, respectively displaying the first evaluation result of the surgical field and the second evaluation result of the surgical field on the same interface, and / or, based on the first evaluation result of the surgical field and the second evaluation result of the surgical field, displaying the evaluation accuracy of the surgical user.

[0019] A third aspect of an embodiment of the present application provides a training device for a surgical field of view assessment model, comprising: a sample acquisition unit, used to acquire sample video frames and target annotation information of the sample video frames; an image cropping unit, used to perform image cropping processing on the sample video frames to obtain multiple sub-images corresponding to the sample video frames, wherein when the sample video frames belong to positive samples, the surgical objects are completely presented in at least part of the corresponding multiple sub-images; a model training unit, used to perform model training based on the sample video frames, the multiple sub-images corresponding to the sample video frames, and the target annotation information of the sample video frames to obtain a surgical field of view assessment model, wherein the surgical field of view assessment model is used to determine the surgical field of view assessment result of the input video frame.

[0020] A surgical field of view assessment device provided in the fourth aspect of an embodiment of the present application includes: a video acquisition unit, used to acquire a surgical video to be evaluated; a segment analysis unit, used to determine a video segment in the surgical video in which the surgical field of view meets the surgical conditions; and a surgical field of view assessment unit, used to determine a first assessment result of the surgical field of view of the surgical video based on the time distribution of the video segment in the surgical video.

[0021] A fifth aspect of an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the training method for the surgical field assessment model as described in any one of the first aspects are implemented, or, when the processor executes the computer program, the steps of the surgical field assessment method as described in any one of the first aspects are implemented.

[0022] A sixth aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the computer program implements the steps of the training method of the above-mentioned surgical field of view assessment model, or when the computer program is executed by a processor, the computer program implements the steps of the above-mentioned surgical field of view assessment method.

[0023] A seventh aspect of an embodiment of the present application provides a computer program product, which, when the computer program product is run on an electronic device, enables the electronic device to execute the steps of the training method for the above-mentioned surgical field of view assessment model, or, when the computer program product is run on an electronic device, enables the electronic device to execute the steps of the above-mentioned surgical field of view assessment method.

[0024] The beneficial effects of the above-mentioned third to seventh aspects of the embodiments can be referred to the description of the first and second aspects, and this application will not elaborate on them. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0026] Figure 1 It is a schematic diagram of the implementation flow of the training method of the surgical visual field assessment model provided in the embodiment of the present application;

[0027] Figure 2 is a first schematic diagram of a sample video frame provided in an embodiment of the present application;

[0028] Figure 3 is a second schematic diagram of a sample video frame provided in an embodiment of the present application;

[0029] Figure 4 is a schematic diagram of cutting out the black border area provided in an embodiment of the present application;

[0030] Figure 5 It is a schematic diagram of the implementation flow of the surgical visual field assessment method provided in the embodiment of the present application;

[0031] Figure 6 It is a schematic diagram of a specific implementation process of determining a video clip in a surgical video where the surgical field meets the surgical conditions provided in an embodiment of the present application;

[0032] Figure 7 Schematic diagram of the structure of the surgical field evaluation model provided in the embodiment of the present application;

[0033] Figure 8 This is a schematic diagram of a specific implementation process for determining a first evaluation result of a surgical field provided in an embodiment of the present application;

[0034] Fig. 9 is a schematic diagram of a surgical video provided in an embodiment of the present application;

[0035] Fig.10 It is a structural schematic diagram of a training device for a surgical visual field assessment model provided in an embodiment of the present application;

[0036] Fig.11 is a schematic diagram of the structure of a surgical visual field assessment device provided in an embodiment of the present application;

[0037] Fig.12 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0038] In order to make the purpose, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work are protected by the present application.

[0039] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.

[0040] In the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0041] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0042] The surgical field of view refers to the scope and content that the doctor can directly see during the operation. Whether the surgical object can be clearly presented in the surgical field of view determines the reliability of the operation and is one of the key factors in reducing postoperative complications. Taking laparoscopic cholecystectomy (LC) as an example, LC surgery is a standard surgical method for treating benign gallbladder lesions, and bile duct injury (BDI) is a serious complication after LC surgery, posing a huge threat to the patient's health. Many guidelines recommend preventing BDI by achieving a critical view of safety (CVS) during surgery. However, at present, hospitals lack effective tools to standardize doctors' intraoperative operations, and it is difficult to assess whether doctors have a critical view of safety operation awareness. For young doctors and clinical medical students, it is also difficult to acquire a critical view of safety operation awareness due to the lack of clinical surgical experience.

[0043] In view of this, the present application proposes a training method for a surgical field of view assessment model and a surgical field of view assessment method, which can improve the accuracy of surgical field of view assessment of video frames and accurately classify surgical videos. When applied to cholecystectomy, it can assist in assessing whether the doctor has key safety field of view operation awareness, thereby improving surgical safety and teaching convenience.

[0044] In order to illustrate the technical solution of the present application, a specific embodiment is provided below for illustration.

[0045] Please refer to Figure 1 , Figure 1 The present invention provides a schematic diagram of the implementation process of a training method for a surgical visual field assessment model, which can be applied to electronic devices. The electronic devices can refer to medical devices such as endoscopes, or smart devices such as computer devices and tablet computers, which are not limited in the present invention.

[0046] Specifically, the training method of the surgical field evaluation model may include the following steps S101 to S103.

[0047] Step S101, obtaining sample video frames and target annotation information of the sample video frames.

[0048] In the implementation of the present application, the sample video frame is a video frame used as a sample to train the surgical field evaluation model. The target annotation information of the sample video frame is a classification annotation of whether the surgical field operation presented by the sample video frame meets the surgical requirements. For example, when the sample video frame is a video frame of an LC surgical video, the target annotation information of the sample video frame may indicate whether the sample video frame has CVS.

[0049] This application does not restrict the method of obtaining sample video frames and their target annotation information. Sample video frames can be extracted from surgical videos actually shot by medical equipment (such as endoscope equipment), or generated by artificial intelligence (AI). Target annotation information can be obtained by manual annotation or automatic annotation.

[0050] It should be noted that in order to improve the reliability of the surgical field evaluation model, the number of sample video frames can be multiple, and include positive samples and negative samples. Positive samples indicate that the surgical field operation presented by the sample video frame meets the surgical requirements, and negative samples indicate that the surgical field operation presented by the sample video frame does not meet the surgical requirements. Subsequently, the surgical field evaluation model trained using positive and negative samples has better generalization ability.

[0051] Step S102: performing image cropping processing on the sample video frame to obtain a plurality of sub-images corresponding to the sample video frame.

[0052] In the implementation manner of the present application, the image cropping process can be implemented by equal segmentation, cropping according to preset pixel positions, etc., and the present application does not impose any restrictions on this. When the sample video frame belongs to a positive sample, the surgical object is completely presented in at least part of the corresponding multiple sub-images. In this way, the image details presented in at least part of the sub-images in the sample video frame of the positive sample can include the morphology of the surgical object, thereby helping the surgical field evaluation model learn the key features to determine whether the surgical field operation presented in the sample video frame meets the surgical requirements.

[0053] Step S103 , performing model training based on the sample video frame, multiple sub-images corresponding to the sample video frame, and target annotation information of the sample video frame to obtain a surgical field of view assessment model.

[0054] Specifically, a sample video frame and multiple sub-images corresponding to the sample video frame are used as model inputs, and the model is trained with the goal of minimizing the error between the model output and the target annotation information, so that a surgical field evaluation model can be obtained. The obtained surgical field evaluation model can be used to determine the surgical field evaluation result of the input video frame, that is, the surgical field operation presented by the video frame input to the surgical field evaluation model meets the surgical requirements.

[0055] Taking the LC surgery scenario as an example, model training is performed based on sample video frames of LC surgery, multiple sub-images obtained by cropping the sample video frames of LC surgery, and target annotation information of whether the sample video frames have CVS. Thus, a surgical field evaluation model for evaluating whether a video frame has CVS can be obtained.

[0056] It should be noted that the above-mentioned surgical field evaluation model can be a convolutional neural network (CNN) model, a recurrent neural network (RNN) model, a generative adversarial network (GAN) model or other existing deep learning network models, and this application does not impose any restrictions on this.

[0057] In an embodiment of the present application, a model is trained based on a sample video frame, a plurality of sub-images obtained by cropping the sample video frame, and target annotation information of the sample video frame to obtain a surgical field of view assessment model for determining a surgical field of view assessment result of an input video frame. The surgical field of view assessment model can be enabled to extract image details, thereby improving the accuracy of surgical field of view assessment of video frames.

[0058] In some embodiments of the present application, acquiring the sample video frame may include: acquiring a sample video, extracting a sample video segment from the sample video, and extracting a sample video frame from the sample video segment.

[0059] In the starting video frame of the sample video segment, the surgical object first appears and is separated from the surrounding tissues. In the ending video frame of the sample video segment, the surgical object is clamped.

[0060] For a detailed description of LC surgery, please refer to Figure 2 , Figure 2 Figure a in the figure is the starting video frame, in which the cystic duct and cystic artery are initially exposed and separated. Figure 2 Figure b is the end video frame, in which any one of the cystic duct and the cystic artery is clamped.

[0061] In some embodiments of the present application, the sample video clip can be extracted based on the manually marked timestamp. Specifically, the clinician can evaluate the sample video, mark the starting video frame and the ending video frame of the sample video clip, and the marked information is recorded in the file in the form of a timestamp. Correspondingly, the electronic device can extract the sample video clip from the sample video according to the timestamp recorded in the file. It should be noted that the present application does not limit the format of the above-mentioned file, for example, it can be in txt format.

[0062] In other embodiments of the present application, the starting video frame and the ending video frame can be determined by image recognition, and then the sample video clip can be extracted from the sample video. Specifically, by identifying the characteristic points of the surgical object and the characteristic points of the surrounding tissue, it can be determined whether the surgical object appears and is separated from the surrounding tissue, and then the starting video frame can be determined; by identifying the characteristic points of the surgical object and the characteristic points of the surgical instrument, it can be determined whether the surgical object is clamped by the surgical instrument, and then the starting video frame can be determined.

[0063] In this way, the negative impact of interfering video frames such as the endoscope entry and endoscope retraction stages on the model training process can be reduced, thereby improving the reliability of the surgical field of view assessment model.

[0064] In some embodiments of the present application, extracting sample video frames from the sample video clip may include: extracting sample video frames from the sample video clip based on a preset frame rate.

[0065] The preset frame rate may be set based on actual needs, for example, to 60 frames per second.

[0066] In some embodiments of the present application, the picture similarity between adjacent sample video frames extracted from the sample video segment is lower than a similarity threshold.

[0067] Specifically, the image similarity between adjacent sample video frames can be ensured to be below the similarity threshold by manual screening or by calculating the image similarity between the video frames of the sample video clips. In this way, the problem of a large number of similar training samples being repeated due to the similarity of adjacent frames can be avoided, which helps to improve the generalization ability of the surgical field assessment model.

[0068] In some embodiments of the present application, obtaining target annotation information of a sample video frame may include: using annotation information of a sample video segment to which the sample video frame belongs as initial annotation information of the sample video frame, and correcting the initial annotation information of the sample video frame belonging to the positive sample to obtain target annotation information of the sample video frame.

[0069] Specifically, the annotation information of the sample video clip can be obtained by manual marking or by image recognition, which is not limited in this application. If the surgical object is clearly visible in the sample video clip, for example, the free cystic duct and cystic artery can be identified, the surgical field of the sample video clip is considered to meet the surgical requirements.

[0070] Based on the annotation information of the sample video clip, the sample video frames extracted from the sample video clip can be divided into a positive sample group or a negative sample group. Since there may be some sample video frames in the positive sample group that fail to present the surgical object, the data can be cleaned and the initial annotation information of these interfering frames can be corrected to be divided into the negative sample group.

[0071] In some embodiments of the present application, a binary classification result of whether the surgical field of the sample video frame meets the surgical requirements can be obtained to correct the initial annotation information.

[0072] Specifically, the binary classification result can be determined based on whether the surgical object is recognized in the sample video frame, whether the surgical object is blocked, the image quality of the sample video frame, and whether the surgical object is separated from surrounding tissues in the sample video frame.

[0073] For example, if the sample video frame of the LC surgery satisfies at least one of the following conditions, it can be determined that the sample video frame does not have CVS: (1) the cystic duct and the cystic artery are blocked; (2) the sample video frame is blurred; (3) the color difference between the cystic duct and the cystic artery and the surrounding tissue is less than a color difference threshold; (4) the cystic duct and the cystic artery are not fully freed; (5) the cystic duct and the cystic artery are outside the field of view of the sample video frame.

[0074] In some implementations of the present application, a binary classification result provided by an industry expert may be obtained. If the binary classification result provided by the industry expert is different from the initial annotation information, the initial annotation information is corrected.

[0075] In other embodiments of the present application, the binary classification results provided by multiple clinicians can be obtained. If the binary classification results provided by multiple clinicians are consistent and different from the initial annotation information, the initial annotation information is corrected. If the binary classification results provided by multiple clinicians are inconsistent, the final binary classification result can be determined by voting, weighting, etc., and if the final binary classification result is different from the initial annotation information, the initial annotation information is corrected.

[0076] In this way, on the one hand, the sample video frames are annotated based on the annotation information of the sample video clips, and the initial annotation information of the sample video frames belonging to the positive samples is corrected. There is no need to annotate video frame by frame, which helps to improve the annotation efficiency. On the other hand, the accuracy of data annotation can be improved through correction, thereby improving the reliability of the surgical field of view assessment model.

[0077] In some embodiments of the present application, before performing image cropping processing on the sample video frame to obtain multiple sub-images corresponding to the sample video frame, it may also include: determining the black border area of ​​the sample video frame and cropping the black border area in the sample video frame.

[0078] In this way, the black border area can be prevented from interfering with the learning of the surgical field of view assessment model, thereby improving the reliability of the surgical field of view assessment model.

[0079] In some implementations of the present application, determining the black border region of the sample video frame may include: performing contour recognition on the sample video frame to obtain contour information of the black border region, and determining the black border region based on the contour information.

[0080] The contour information may refer to contour lines or contour points.

[0081] Specifically, by performing contour recognition on the sample video frame to obtain a contour line, the image area between the contour line and the left and right boundary lines of the sample video frame can be used as the black edge area. Alternatively, by performing contour recognition on the sample video frame to obtain contour points of the contour line and the upper and lower boundary lines of the sample video frame, the image area between the line connecting the contour points and the left and right boundary lines of the sample video frame can be used as the black edge area.

[0082] In some other embodiments of the present application, determining the black border region of the sample video frame may include: determining the pixel mean corresponding to each column of pixels in the sample video frame, the pixel mean corresponding to each column of pixels being the mean of the pixel values ​​of a preset number of pixels located in the middle of the column of pixels. The pixel column whose pixel mean is greater than the preset value is used as the boundary of the black border region, and the black border region is determined based on the boundary.

[0083] Specifically, for each sample video frame, each column of pixels can be traversed from the left and right sides to the middle, and the average of the pixel values ​​of a preset number of pixels in the middle of each column of pixels can be calculated. The pixel value can include three channels of RGB. If the pixel average is greater than the preset value, it is considered that the boundary between the black edge area and the key feature area is reached, otherwise, continue to traverse to the middle, and finally obtain the two boundaries of the black edge area. At this time, the image area between the two boundaries and the left and right boundary lines of the sample video frame can be used as the black edge area.

[0084] The preset number and the preset value may be set according to actual conditions, for example, to 30 and 20 respectively.

[0085] Please refer to Figure 3 In Figure a, some sample video frames have white watermarks in the upper left or lower right black border area. Selecting a preset number of pixels in the middle for average calculation can avoid the interference of white watermarks. At the same time, please refer to Figure 3In Figure b, there may be a large number of dark areas on the left and right sides of the key feature area of ​​some sample video frames. The above method can prevent these dark areas from being included in the black edge area. The black edge area in the sample video frame is cropped. The cropped image can effectively exclude the black edge area and retain the key feature area, which is conducive to enabling the surgical field assessment model to learn useful features during model training.

[0086] In some embodiments of the present application, determining the black border area of ​​the sample video frame may include: obtaining boundaries of black border areas corresponding to multiple sampled video frames in a sample video segment to which the sample video frame belongs, and determining the black border area based on the average of the boundaries of the black border areas corresponding to the multiple sampled video frames.

[0087] Specifically, for each sample video segment, multiple sample video frames can be selected, and the boundary of the black border area of ​​each sample video frame can be determined using the method provided in the previous text. The average of the boundaries of the black border areas corresponding to the multiple sample video frames is used as the boundary of the black border area of ​​each sample video frame extracted from the edge of the entire black border area. Then, the image area between the two boundaries and the left and right boundary lines of the sample video frame can be used as the black border area.

[0088] In some embodiments of the present application, the above-mentioned multiple sampled video frames may be the first N (N>1) frames of the sample video clip, or any N frames of the sample video clip. Specifically, the frame spacing between the multiple sampled video frames and the middle frame of the sample video clip may be less than the frame number threshold, that is, the video frames close to the middle frame of the sample video clip are selected as the sampled video frames. Since the video frames close to the middle frame are mostly video frames during the surgical operation, the boundary between the black border area and the key feature area is more obvious, which helps to accurately extract the boundary of the black border area.

[0089] The black border area is determined based on the average of the boundaries of the black border areas corresponding to multiple sampled video frames, which can ensure the accuracy of the black border area and further enhance the robustness of the surgical field evaluation model.

[0090] In some implementations of the present application, the sample video frame is subjected to image cropping processing to obtain multiple sub-images corresponding to the sample video frame, including any of the following:

[0091] 1. Using the short side of the sample video frame as the side length of the sub-image and the upper left corner pixel of the sample video frame as the upper left corner pixel of the sub-image, a square area is cropped out from the sample video frame to obtain the sub-image;

[0092] 2. Using the short side of the sample video frame as the side length of the sub-image and the upper right corner pixel of the sample video frame as the upper right corner pixel of the sub-image, a square area is cropped out from the sample video frame to obtain the sub-image;

[0093] 3. Using the short side of the sample video frame as the side length of the sub-image and the midline of the sample video frame as the midline of the sub-image, a square area is cropped out from the sample video frame to obtain a sub-image.

[0094] The cropping effect is as follows Figure 4 As shown, by cutting out a square area with the short side of the sample video frame as the side length of the sub-image, a square sub-image of uniform size can be obtained, avoiding the destruction of the characteristic size of the original surgical object due to the need for subsequent resizing. In addition, in LC surgery, the doctor generally pulls the cystic duct to the left side of the field of view to separate the front of the cystic duct and the cystic artery from the surrounding tissues. At this time, the cystic duct and the cystic artery appear in the left and middle areas of the sample video frame. Subsequently, the doctor will flip the lens so that the cystic duct and the cystic artery appear in the middle and right areas of the sample video frame to separate the back of the cystic duct and the cystic artery from the surrounding tissues. This image cropping processing method can make the cystic duct and the cystic artery appear completely in at least two areas of the left, middle and right, that is, appear in two of the three sub-images, thereby helping the surgical field of view evaluation model to learn the details of the cystic duct and the cystic artery presented in the sub-images.

[0095] In some embodiments of the present application, a first number of sample video frames belonging to negative samples may be compared with a second number of sample video frames belonging to positive samples, and if the ratio between the first number and the second number is greater than a ratio threshold, an upsampling operation is performed on the sample video frames belonging to positive samples, and / or a downsampling operation is performed on the sample video frames belonging to negative samples. The ratio threshold may be set according to actual conditions, for example, set to 2.

[0096] The upsampling operation is to keep the first number unchanged and expand the sample video frames of the positive sample. The downsampling operation is to keep the second number unchanged and reduce the sample video frames of the negative sample. Through the upsampling operation and / or the downsampling operation, the number of sample video frames of the positive and negative samples can be made similar or equal, so as to avoid the model from focusing too much on the features presented by the positive or negative samples.

[0097] In some embodiments of the present application, a sample video frame and a plurality of sub-images corresponding to the sample video frame are used as model inputs, and the model is trained with the goal of minimizing the error between the model output and the target annotation information to obtain a surgical field evaluation model. The obtained surgical field evaluation model can be used to determine the surgical field evaluation result of the input video frame, that is, the surgical field operation presented by the video frame input to the surgical field evaluation model meets the surgical requirements.

[0098] Please refer to Figure 5 , Figure 5The present invention provides a schematic diagram of the implementation process of a surgical visual field assessment method provided in an embodiment of the present invention, which can be applied to electronic devices. The electronic devices can refer to medical devices such as endoscopes, or can refer to smart devices such as computer devices and tablet computers, which are not limited in the present invention.

[0099] Specifically, the above surgical field evaluation method may include the following steps S501 to S503.

[0100] Step S501, obtaining a surgical video to be evaluated.

[0101] The surgical video to be evaluated refers to a video that requires surgical field evaluation, and may refer to a LC surgical video.

[0102] Step S502, determining video segments in the surgical video whose surgical fields meet the surgical conditions.

[0103] Specifically, by analyzing each video frame to be evaluated in the surgical video, a video segment in which the surgical field meets the surgical conditions can be determined. A video segment in which the surgical field meets the surgical conditions means that the morphology of the surgical object presented in the surgical field in the video segment meets the surgical requirements. Taking LC surgery as an example, a video segment in which the surgical field meets the surgical conditions in the surgical video can mean that the video segment has CVS.

[0104] Step S503: determining a first evaluation result of the surgical field of the surgical video based on the time distribution of the video clips in the surgical video.

[0105] The first evaluation result of the surgical field of view refers to the prediction result of whether the surgical field of view of the surgical video meets the surgical requirements. Based on the time distribution of the video clips in the surgical video, the distribution of the video clips on the timeline of the entire surgical video and the duration of the video clips can be analyzed to determine the first evaluation result of the surgical field of view of the surgical video.

[0106] In an embodiment of the present application, by determining the video segments in the surgical video to be evaluated whose surgical field of view meets the surgical conditions, and determining a first evaluation result of the surgical field of view of the surgical video based on the time distribution of the video segments in the surgical video, the surgical video can be classified in combination with the time domain information of the video segments in which the surgical field of view meets the surgical conditions, which helps to improve the classification accuracy of the surgical video. When applied to cholecystectomy, it can evaluate whether the doctor has the awareness of key safety field operation, thereby improving the safety of the operation and the convenience of teaching.

[0107] Specifically, Figure 6 As shown, in some embodiments of the present application, determining the video segment in the surgical video whose surgical field meets the surgical conditions may include steps S601 to S603.

[0108] Step S601: input each video frame to be evaluated in the surgical video into the surgical field evaluation model.

[0109] The surgical field evaluation model is obtained by training the model based on the sample video frame, multiple sub-images corresponding to the sample video frame, and target annotation information of the sample video frame. The multiple sub-images corresponding to the sample video frame are obtained by cropping the sample video frame, and when the sample video frame belongs to a positive sample, the surgical object is completely presented in at least part of the corresponding multiple sub-images.

[0110] In some embodiments of the present application, the sample video frame is extracted from a sample video segment of a sample video, wherein the surgical object first appears and is separated from surrounding tissues in the starting video frame of the sample video segment, and the surgical object is clamped in the ending video frame of the sample video segment.

[0111] In some embodiments of the present application, the sample video frame is extracted from the sample video segment based on a preset frame rate.

[0112] In some embodiments of the present application, the picture similarity between adjacent sample video frames extracted from the sample video segment is lower than a similarity threshold.

[0113] In some implementations of the present application, the target annotation information of the sample video frame belonging to the positive sample is obtained by correcting the initial annotation information, wherein the initial annotation information is the annotation information of the sample video segment to which the sample video frame belongs.

[0114] In some embodiments of the present application, the multiple sub-images corresponding to the sample video frame include any of the following: a sub-image obtained by cutting out a square area in the sample video frame with the short side of the sample video frame as the sub-image side length and the upper left corner pixel of the sample video frame as the upper left corner pixel of the sub-image; a sub-image obtained by cutting out a square area in the sample video frame with the short side of the sample video frame as the sub-image side length and the upper right corner pixel of the sample video frame as the upper right corner pixel of the sub-image; a sub-image obtained by cutting out a square area in the sample video frame with the short side of the sample video frame as the sub-image side length and the midline of the sample video frame as the midline of the sub-image.

[0115] In some implementations of the present application, before performing image cropping processing on the sample video frame, the sample video frame undergoes black edge cropping processing, and the black edge cropping processing is used to crop the black edge area in the sample video frame.

[0116] In some embodiments of the present application, the boundary of the black border area is a pixel column whose pixel mean is greater than a preset value, wherein the pixel mean corresponding to each column of pixels is the average of the pixel values ​​of a preset number of pixel points located in the middle of the column of pixels.

[0117] In some other embodiments of the present application, the boundary of the black border area is determined based on contour information obtained by performing contour recognition on the sample video frame.

[0118] In some embodiments of the present application, the boundary of the black border area is an average of boundaries of the black border areas respectively corresponding to a plurality of sampled video frames, and the plurality of sampled video frames originate from a sample video segment to which the sample video frame belongs.

[0119] In some embodiments of the present application, the plurality of sampled video frames are the first N (N>1) frames of the sample video segment. In other embodiments of the present application, the frame spacing between the plurality of sampled video frames and the middle frame of the sample video segment is less than a frame number threshold.

[0120] It should be noted that the specific training process of the surgical field evaluation model can be referred to Figures 1 to 5 This application does not elaborate on the description of the training method of the surgical field assessment model.

[0121] Step S602, obtaining the surgical field evaluation result of each video frame to be evaluated output by the surgical field evaluation model.

[0122] Step S603: determining a video segment according to the surgical field evaluation result of each video frame to be evaluated.

[0123] That is to say, each video frame to be evaluated in the surgical video can be used as the input of the surgical field evaluation model to obtain the surgical field evaluation results of each video frame to be evaluated by the surgical field evaluation model. Based on the surgical field evaluation results of each video frame to be evaluated, adjacent video frames to be evaluated whose surgical fields meet the surgical requirements can be combined to obtain video clips in the surgical video whose surgical fields meet the surgical conditions.

[0124] Specifically, the black-bordered areas in each video frame to be evaluated can be cropped, and image cropping processing can be performed to obtain multiple sub-images corresponding to the video frame to be evaluated. The methods of black-border cropping processing and image cropping processing can refer to the processing methods of the sample video frames in the previous article, and this application will not go into details. Subsequently, each video frame to be evaluated and its corresponding multiple sub-images can be input into the surgical field of view evaluation model to obtain the surgical field of view evaluation results of each video frame to be evaluated output by the surgical field of view evaluation model.

[0125] In some embodiments of the present application, the surgical field evaluation model can be specifically used to: extract features from each video frame to be evaluated and its corresponding multiple sub-images, respectively, to obtain a first feature image of the video frame to be evaluated and multiple sub-images corresponding to the first feature image of the video frame to be evaluated. Figure 1A corresponding plurality of second feature maps. The first feature map and the plurality of second feature maps are concatenated, and the concatenated feature map is classified to obtain the surgical field evaluation result of the video frame to be evaluated.

[0126] Specifically, as Figure 7 shown in the schematic architecture diagram of the surgical field evaluation model, the surgical field evaluation model is a convolutional neural network. The surgical field evaluation model may include a feature extractor and a fully connected layer. The first feature extractor accepts the video frame to be evaluated as input, while the second feature extractor contains three sub-extractors, which respectively accept the input of 3 sub-images. The 4 feature outputs output by the feature extractor will be input into the fully connected layer after a concatenation operation to obtain the surgical field evaluation result of the video frame to be evaluated.

[0127] In this way, using a dual feature extractor to obtain multiple feature maps and realizing the classification of the video frame to be evaluated can effectively improve the classification accuracy.

[0128] In some embodiments of the present application, the surgical field evaluation results of each video frame to be evaluated can be smoothed, and then according to the surgical field evaluation results of each video frame to be evaluated, the video segments in the surgical video whose surgical fields meet the surgical conditions are determined.

[0129] In some embodiments of the present application, as Figure 8 shown, based on the time distribution of the video segments in the surgical video, determining the first evaluation result of the surgical field of the surgical video may include: step S801 to step S802.

[0130] Step S801, comparing the video segments in the surgical video pairwise to determine the video segment pairs whose total number of frames is greater than the frame number threshold and the frame number ratio is greater than the ratio threshold.

[0131] Wherein, the total number of frames is the number of frames included in the two video segments and between the two video segments, and the frame number ratio is the ratio of the total number of all video frames that meet the surgical conditions between the two video segments to the total number of frames.

[0132] Let start i , end i be the head frame and the tail frame of the i-th video segment, where i = 1, 2... N, and N is the total number of video segments in the surgical video. For example, the 5th video segment is between frame 1000 and frame 1023, then start 5 = 1000, end 5 = 1023.

[0133] For any two video segments j and k, j, k = 1, 2... N. Assuming j < k, then the total number of frames end k - startj Greater than the frame number threshold and all video frames that meet the surgical conditions At the end of the total number of frames k -start j The percentage of frames Video segment pairs with a percentage greater than the threshold.

[0134] The frame number threshold and the percentage threshold can be set according to actual conditions, for example, 100 frames and 5% respectively.

[0135] Step S802: If the number of video segment pairs is greater than the number threshold, it is determined that the first evaluation result of the surgical field is that the surgical video meets the surgical conditions.

[0136] That is to say, if the total number of frames is greater than the frame number threshold, and the number of video clip pairs whose frame number ratio is greater than the ratio threshold is greater than the number threshold, it means that most of the video frames to be evaluated in the surgical video are those whose surgical field meets the surgical requirements, and the video frames to be evaluated whose surgical field meets the surgical requirements are evenly distributed in the time domain. At this time, it can be determined that the first evaluation result of the surgical field is that the surgical video meets the surgical conditions. If the number of video clip pairs is less than or equal to the number threshold, it can be determined that the first evaluation result of the surgical field is that the surgical video does not meet the surgical conditions.

[0137] by Fig. 9 In the surgical video shown, there are three video clips a, b, and c on the progress bar of the surgical video whose surgical fields meet the surgical conditions. Assume that:

[0138]

[0139] If the frame number threshold and percentage threshold are 100 frames and 5% respectively, the following calculation can be performed:

[0140] The total number of frames between video segment a and video segment b is 95 frames (less than 100).

[0141] The total number of frames between video segment a and video segment c is 170 frames (greater than 100), and the frame ratio between video segment a and video segment c is

[0142] The distance between video segment b and video segment c is 90 frames (less than 100).

[0143] Finally, in this surgical video, only video clip a and video clip c are video clip pairs whose total frame number is greater than the frame number threshold and whose frame number ratio is greater than the ratio threshold. The number of video clip pairs is 1. If the number threshold is set to 10, the first evaluation result of the surgical field is that the surgical video does not meet the surgical conditions.

[0144] In some embodiments of the present application, after determining a first evaluation result of the surgical field of the surgical video based on the time distribution of the video clips in the surgical video, it may also include: in response to the user's input operation on the second evaluation result of the surgical field of the surgical video, respectively displaying the first evaluation result of the surgical field and the second evaluation result of the surgical field on the same interface, and / or, based on the first evaluation result of the surgical field and the second evaluation result of the surgical field, displaying the user's evaluation accuracy.

[0145] Specifically, the second evaluation result of the surgical field of view is the result of the user's input of whether the surgical field of view of the surgical video meets the surgical requirements. The user can input the second evaluation result of the surgical field of view of the surgical video through the control of the software interface.

[0146] By displaying the first assessment result of the surgical field and the second assessment result of the surgical field on the same interface, the predicted result of the surgical video and the user's input result can be visually compared, making it easier for the user to compare and review.

[0147] Based on the first evaluation result and the second evaluation result of the surgical field of view, the evaluation accuracy of the user's input results of each surgical video in the video library can be counted and displayed, thereby improving the user's ability to judge whether the surgical field of view meets the surgical requirements.

[0148] Applying the method provided in this application to LC surgery videos can automatically classify LC surgery videos for young doctors and medical students to learn from, thereby improving doctors' overall CVS operation awareness and fundamentally reducing the incidence of serious complications in LC surgery.

[0149] It should be noted that, for the sake of simplicity of description, the aforementioned method embodiments are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited to the described order of actions, because according to the present application, certain steps can be performed in other orders.

[0150] like Fig.10 FIG. 1 is a schematic structural diagram of a training device 1000 for a surgical visual field assessment model provided in an embodiment of the present application. The training device 1000 for a surgical visual field assessment model is configured on an electronic device.

[0151] Specifically, the training device 1000 for the surgical field evaluation model may include:

[0152] The sample acquisition unit 1001 is used to acquire a sample video frame and target annotation information of the sample video frame;

[0153] An image cropping unit 1002 is used to perform image cropping processing on the sample video frame to obtain a plurality of sub-images corresponding to the sample video frame, wherein when the sample video frame belongs to a positive sample, the surgical object is completely presented in at least a part of the corresponding plurality of sub-images;

[0154] The model training unit 1003 is used to perform model training based on the sample video frame, multiple sub-images corresponding to the sample video frame, and target annotation information of the sample video frame to obtain a surgical field of view evaluation model, and the surgical field of view evaluation model is used to determine the surgical field of view evaluation result of the input video frame.

[0155] In some embodiments of the present application, the image cropping unit 1002 may be specifically used to: use the short side of the sample video frame as the side length of the sub-image, use the upper left corner pixel of the sample video frame as the upper left corner pixel of the sub-image, and crop a square area in the sample video frame to obtain a sub-image; use the short side of the sample video frame as the side length of the sub-image, use the upper right corner pixel of the sample video frame as the upper right corner pixel of the sub-image, and crop a square area in the sample video frame to obtain a sub-image; use the short side of the sample video frame as the side length of the sub-image, use the midline of the sample video frame as the midline of the sub-image, and crop a square area in the sample video frame to obtain a sub-image.

[0156] In some implementations of the present application, the image cropping unit 1002 may also be specifically configured to: determine a black border region of the sample video frame; and crop the black border region in the sample video frame.

[0157] In some embodiments of the present application, the image cropping unit 1002 can also be specifically used to: determine the pixel mean corresponding to each column of pixels in the sample video frame, the pixel mean corresponding to each column of pixels is the mean of the pixel values ​​of a preset number of pixels located in the middle of the column of pixels; use the pixel column whose pixel mean is greater than the preset value as the boundary of the black border area, and determine the black border area based on the boundary.

[0158] In some embodiments of the present application, the image cropping unit 1002 can also be specifically used to: obtain the boundaries of the black-border area corresponding to multiple sampled video frames in the sample video segment to which the sample video frame belongs; and determine the black-border area based on the average of the boundaries of the black-border area corresponding to the multiple sampled video frames.

[0159] It should be noted that, for the convenience and simplicity of description, the specific working process of the training device 1000 of the surgical visual field assessment model can be referred to Figures 1 to 4 The corresponding process of the method will not be repeated here.

[0160] like Fig.11FIG. 1 is a schematic diagram of the structure of a surgical visual field evaluation device 1100 provided in an embodiment of the present application. The surgical visual field evaluation device 1100 is configured on an electronic device.

[0161] Specifically, the surgical field evaluation device 1100 may include:

[0162] The video acquisition unit 1101 is used to acquire the surgical video to be evaluated;

[0163] A segment analysis unit 1102 is used to determine the video segments in the surgical video whose surgical field meets the surgical conditions;

[0164] The surgical field evaluation unit 1103 is configured to determine a first evaluation result of the surgical field of the surgical video based on the time distribution of the video clips in the surgical video.

[0165] In some embodiments of the present application, the segment analysis unit 1102 can be specifically used to: input each video frame to be evaluated in the surgical video into a surgical field of view evaluation model, wherein the surgical field of view evaluation model is obtained by model training based on sample video frames, multiple sub-images corresponding to the sample video frames, and target annotation information of the sample video frames, and the multiple sub-images corresponding to the sample video frames are obtained by image cropping the sample video frames, and when the sample video frames belong to positive samples, the surgical objects are completely presented in at least part of the corresponding multiple sub-images; obtain the surgical field of view evaluation results of each video frame to be evaluated output by the surgical field of view evaluation model; and determine the video segment according to the surgical field of view evaluation results of each video frame to be evaluated.

[0166] In some embodiments of the present application, the multiple sub-images corresponding to the sample video frame include any of the following: a sub-image obtained by cutting out a square area in the sample video frame with the short side of the sample video frame as the sub-image side length and the upper left corner pixel of the sample video frame as the upper left corner pixel of the sub-image; a sub-image obtained by cutting out a square area in the sample video frame with the short side of the sample video frame as the sub-image side length and the upper right corner pixel of the sample video frame as the upper right corner pixel of the sub-image; a sub-image obtained by cutting out a square area in the sample video frame with the short side of the sample video frame as the sub-image side length and the midline of the sample video frame as the midline of the sub-image.

[0167] In some implementations of the present application, the sample video frame undergoes a black border cropping process, and the black border cropping process is used to crop the black border area in the sample video frame.

[0168] In some embodiments of the present application, the boundary of the black border area is a pixel column whose pixel mean is greater than a preset value, wherein the pixel mean corresponding to each column of pixels is the average of the pixel values ​​of a preset number of pixel points located in the middle of the column of pixels.

[0169] In some embodiments of the present application, the boundary of the black border area is an average of boundaries of the black border areas respectively corresponding to a plurality of sampled video frames, and the plurality of sampled video frames all originate from a sample video segment to which the sample video frame belongs.

[0170] In some embodiments of the present application, the surgical field of view assessment unit 1103 may be specifically used to: compare the video segments in the surgical video in pairs, and determine video segment pairs whose total frame number is greater than a frame number threshold and whose frame number ratio is greater than a ratio threshold, wherein the total frame number is the number of frames contained in the two video segments and between the two video segments, and the frame number ratio is the ratio of the total number of video frames that meet the surgical conditions between the two video segments to the total number of frames; if the number of video segment pairs is greater than the number threshold, it is determined that the first assessment result of the surgical field of view is that the surgical video meets the surgical conditions.

[0171] In some embodiments of the present application, the surgical field of view evaluation device 1100 also includes a teaching unit, which is specifically used to: respond to the user's input operation on the second evaluation result of the surgical field of view of the surgical video, respectively display the first evaluation result of the surgical field of view and the second evaluation result of the surgical field of view on the same interface, and / or, based on the first evaluation result of the surgical field of view and the second evaluation result of the surgical field of view, display the evaluation accuracy of the surgical user.

[0172] It should be noted that, for the convenience and simplicity of description, the specific working process of the above-mentioned surgical field evaluation device 1100 can be referred to Figures 5 to 9 The corresponding process of the method will not be repeated here.

[0173] like Fig.12 , which is a schematic diagram of an electronic device 12 provided in an embodiment of the present application. Specifically, the electronic device 12 may include: a processor 120, a memory 121, and a computer program 122 stored in the memory 121 and executable on the processor 120, such as a training program for a surgical visual field assessment model. When the processor 120 executes the computer program 122, the steps in the above-mentioned training method embodiments of the surgical visual field assessment model are implemented, such as Figure 1 Alternatively, when the processor 120 executes the computer program 122, the steps in the above-mentioned various surgical field evaluation method embodiments are implemented, for example Figure 5 Steps S501 to S503 are shown.

[0174] Alternatively, when the processor 120 executes the computer program 122, the functions of the modules / units in the above-mentioned device embodiments are realized, for example Fig.10 The functions of the sample acquisition unit 1001, the image cropping unit 1002 and the model training unit 1003 are shown. Alternatively, when the processor 120 executes the computer program 122, the functions of each module / unit in the above-mentioned device embodiments are realized, for example Fig.11 The functions of the video acquisition unit 1101 , the segment analysis unit 1102 and the surgical field evaluation unit 1103 are shown.

[0175] The computer program may be cut into one or more modules / units, which are stored in the memory 121 and executed by the processor 120 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program in the electronic device 12.

[0176] For example, the computer program can be cut into a sample acquisition unit, an image cropping unit, and a model training unit. The specific functions of each unit are as follows: a sample acquisition unit, used to acquire a sample video frame and target annotation information of the sample video frame; an image cropping unit, used to perform image cropping processing on the sample video frame to obtain multiple sub-images corresponding to the sample video frame, wherein when the sample video frame belongs to a positive sample, the surgical object is completely presented in at least part of the corresponding multiple sub-images; a model training unit, used to perform model training based on the sample video frame, the multiple sub-images corresponding to the sample video frame, and the target annotation information of the sample video frame to obtain a surgical field evaluation model, wherein the surgical field evaluation model is used to determine the surgical field evaluation result of the input video frame.

[0177] For another example, the computer program can be cut into: a video acquisition unit, a segment analysis unit, and a surgical field evaluation unit. The specific functions of each unit are as follows: a video acquisition unit, used to acquire the surgical video to be evaluated; a segment analysis unit, used to determine the video segments in the surgical video whose surgical field meets the surgical conditions; a surgical field evaluation unit, used to determine the first evaluation result of the surgical field of the surgical video based on the time distribution of the video segments in the surgical video.

[0178] The electronic device 12 may include, but is not limited to, a processor 120 and a memory 121. Those skilled in the art will appreciate that Fig.12It is only an example of the electronic device 12 and does not constitute a limitation of the electronic device 12. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device 12 may also include input and output devices, network access devices, buses, etc.

[0179] The processor 120 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), off-the-shelf programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0180] The memory 121 may be an internal storage unit of the electronic device 12, such as a hard disk or memory of the electronic device 12. The memory 121 may also be an external storage device of the electronic device 12, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 12. Further, the memory 121 may also include both an internal storage unit of the electronic device 12 and an external storage device. The memory 121 is used to store the computer program and other programs and data required by the electronic device 12. The memory 121 may also be used to temporarily store data that has been output or is to be output.

[0181] It should be noted that, for the convenience and brevity of description, the structure of the electronic device 12 can also refer to the specific description of the structure in the method embodiment, which will not be repeated here.

[0182] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0183] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0184] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0185] In the embodiments provided in the present application, it should be understood that the disclosed devices / equipment and methods can be implemented in other ways. For example, the device / equipment embodiments described above are merely schematic, for example, the division of the modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0186] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0187] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0188] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0189] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A training method for a surgical field assessment model, characterized in that: include: Acquire sample video frames and target annotation information of the sample video frames; Performing image cropping processing on the sample video frame to obtain a plurality of sub-images corresponding to the sample video frame, wherein when the sample video frame belongs to a positive sample, the surgical object is completely presented in at least a part of the corresponding plurality of sub-images; Model training is performed based on the sample video frame, multiple sub-images corresponding to the sample video frame, and target annotation information of the sample video frame to obtain a surgical field evaluation model, which is used to determine a surgical field evaluation result of an input video frame.

2. The training method of the surgical field assessment model according to claim 1, characterized in that: The performing image cropping processing on the sample video frame to obtain a plurality of sub-images corresponding to the sample video frame may include any of the following: Using the short side of the sample video frame as the side length of the sub-image, using the upper left corner pixel point of the sample video frame as the upper left corner pixel point of the sub-image, and cutting out a square area in the sample video frame to obtain a sub-image; Using the short side of the sample video frame as the side length of the sub-image, using the upper right corner pixel point of the sample video frame as the upper right corner pixel point of the sub-image, and cutting out a square area in the sample video frame to obtain a sub-image; The short side of the sample video frame is used as the side length of the sub-image, the midline of the sample video frame is used as the midline of the sub-image, and a square area is cropped from the sample video frame to obtain a sub-image.

3. The training method of the surgical field assessment model according to claim 1 or 2, characterized in that: Before performing image cropping processing on the sample video frame to obtain a plurality of sub-images corresponding to the sample video frame, the method further includes: Determining a black border area of ​​the sample video frame; The black border area in the sample video frame is cropped.

4. The training method of the surgical field assessment model according to claim 3, characterized in that: The determining of the black border area of ​​the sample video frame includes: Determine a pixel mean value corresponding to each column of pixels in the sample video frame, wherein the pixel mean value corresponding to each column of pixels is the mean value of pixel values ​​of a preset number of pixel points located in the middle of the column of pixels; The pixel column whose pixel mean value is greater than a preset value is used as a boundary of the black border area, and the black border area is determined based on the boundary.

5. A method for evaluating surgical visual field, characterized in that: include: Obtain surgical videos to be evaluated; Determine a video segment in the surgical video in which the surgical field meets the surgical conditions; Based on the time distribution of the video clips in the surgical video, a first evaluation result of the surgical field of the surgical video is determined.

6. The surgical field assessment method according to claim 5, characterized in that: The video clip of determining that the surgical field in the surgical video meets the surgical conditions includes: Inputting each video frame to be evaluated in the surgical video into a surgical field evaluation model, wherein the surgical field evaluation model is obtained by model training based on a sample video frame, a plurality of sub-images corresponding to the sample video frame, and target annotation information of the sample video frame, wherein the plurality of sub-images corresponding to the sample video frame are obtained by performing image cropping processing on the sample video frame, and when the sample video frame belongs to a positive sample, the surgical object is completely presented in at least part of the corresponding plurality of sub-images; Obtaining surgical field evaluation results of each video frame to be evaluated output by the surgical field evaluation model; The video segment is determined according to the surgical field evaluation result of each video frame to be evaluated.

7. The surgical field assessment method according to claim 6, characterized in that: The multiple sub-images corresponding to the sample video frame include any of the following: The sub-image is obtained by cutting out a square area in the sample video frame with the short side of the sample video frame as the side length of the sub-image and the upper left corner pixel point of the sample video frame as the upper left corner pixel point of the sub-image; The sub-image is obtained by cutting out a square area in the sample video frame with the short side of the sample video frame as the side length of the sub-image and the upper right corner pixel point of the sample video frame as the upper right corner pixel point of the sub-image; The sub-image is obtained by cutting out a square area in the sample video frame with the short side of the sample video frame as the side length of the sub-image and the midline of the sample video frame as the midline of the sub-image.

8. The surgical field assessment method according to claim 5, characterized in that: The determining, based on the time distribution of the video clips in the surgical video, a first evaluation result of the surgical field of view of the surgical video includes: Compare the video clips in the surgical video in pairs to determine video clip pairs whose total number of frames is greater than a frame number threshold and whose frame number ratio is greater than a ratio threshold, wherein the total number of frames is the number of frames included in the two video clips and between the two video clips, and the frame number ratio is the ratio of the total number of all video frames meeting the surgical conditions between the two video clips to the total number of frames; If the number of the video segment pairs is greater than the number threshold, it is determined that the first evaluation result of the surgical field is that the surgical video meets the surgical conditions.

9. The surgical visual field assessment method according to any one of claims 5 to 8, characterized in that: After determining the first evaluation result of the surgical field of the surgical video based on the time distribution of the video clips in the surgical video, the method further includes: In response to the user's input operation on the second evaluation result of the surgical field of view of the surgical video, the first evaluation result of the surgical field of view and the second evaluation result of the surgical field of view are respectively displayed on the same interface, and / or, based on the first evaluation result of the surgical field of view and the second evaluation result of the surgical field of view, the user's evaluation accuracy is displayed.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the training method of the surgical field assessment model as described in any one of claims 1 to 4 are implemented, or when the processor executes the computer program, the steps of the surgical field assessment method as described in any one of claims 5 to 9 are implemented.