Image data processing method, image platform, computer device and storage medium
Through the image data processing method that identifies the target object in the image and conveys three-dimensional depth information, the problem that existing medical devices cannot automatically identify and provide three-dimensional depth information is solved, and the accuracy and safety of surgical operations are improved.
Patent Information
- Application Number
- CN202210182065.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-25
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-02-25
AI Technical Summary
Existing medical devices and equipment cannot automatically identify organs and surgical instruments in the surgical environment, and cannot provide three-dimensional and three-dimensional depth information, affecting the operation of medical workers.
Through the image data processing method, the target object is identified in the image using the target tag, and combined with the depth information, a target tag containing object category information is generated to realize the communication of three-dimensional depth information.
It realizes automatic recognition and accurately conveys three-dimensional depth information in two-dimensional images, improving the operational efficiency and safety of medical workers.
Smart Images

Figure CN114565719B_ABST
Abstract
Description
Technical Field
[0001] This specification belongs to the technical field of medical devices, and particularly relates to an image data processing method, an image platform, a computer device, and a storage medium. Background Art
[0002] When medical staff perform surgical operations (such as abdominal surgery, etc.), they usually use devices such as endoscopes (such as abdominal endoscopes) to assist in observing organs, surgical instruments, etc. in the surgical environment for specific surgical operations.
[0003] However, based on existing medical device equipment, medical staff can often only observe images containing organs and surgical instruments, and medical staff still need to identify specific organs and surgical instruments based on the displayed images. In addition, based on existing medical device equipment, medical staff cannot perceive the three-dimensional depth information of organs and surgical instruments. This will thus affect the surgical operations of medical staff.
[0004] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] This specification provides an image data processing method, an image platform, a computer device, and a storage medium, which can automatically identify and use target labels to mark relevant information of a target object concerned by a user in an image, and at the same time, can accurately and intuitively convey the three-dimensional depth information of the target object to the user through the displayed target labels, enabling the user to obtain a better interaction experience.
[0006] An embodiment of this specification provides an image data processing method, including: obtaining a target image; identifying a target object in the target image to obtain a corresponding identification result; and obtaining the depth information of the target object; and displaying a target label of the target object in the target image according to the identification result and the depth information of the target object; where the target label at least includes object category information of the target object.
[0007] An embodiment of this specification further provides an image platform, including a binocular endoscope and an image data processing device, where the image data processing device is configured to process the target image in the following manner: identifying a target object in the target image to obtain a corresponding identification result; and obtaining the depth information of the target object; and displaying a target label of the target object in the target image according to the identification result and the depth information of the target object; where the target label at least includes object category information of the target object.
[0008] An embodiment of this specification also provides a computer device, including a processor and a memory for storing instructions executable by the processor. When the processor executes the instructions, relevant steps of the image data processing method are implemented.
[0009] An embodiment of this specification also provides a computer-readable storage medium, on which computer instructions are stored. When the instructions are executed, relevant steps of the image data processing method are implemented.
[0010] Based on the image data processing method, image platform, computer device, and storage medium provided in this specification, after obtaining a target image and identifying a target object in the target image to obtain a corresponding recognition result, depth information of the target object can also be obtained; and at the same time, using the recognition result and depth information of the target object, a target label including at least the object category information of the target object is generated and displayed on the target image. Thus, it is possible to automatically identify and use the target label to mark relevant information of the target object concerned by the user in the target image, and through the displayed target label, the three-dimensional depth information of the target object can be accurately and intuitively conveyed to the user in the two-dimensional target image, enabling the user to obtain a better interaction experience.
[0011] Specifically, according to the depth information of the target object, based on the principle of perspective, the label size parameter matching the target object can be determined; then, according to the label size parameter, a target label capable of conveying depth information such as the distance of the target object from the camera is generated; and by displaying the target label in the target image, the user can more conveniently and intuitively know the distance of the target object from the camera based on the target label in the image.
[0012] Furthermore, according to the depth information of the target object, by specifically adjusting the display effect parameters of the target label in the target image, such as shadow parameters, deformation parameters, brightness parameters, etc., the three-dimensional depth information of the target object can be more effectively conveyed to the user through the target image and the target label.
[0013] In addition, based on the augmented reality (AR) technology, first, the difference value between the depth information of different target objects in the target image is obtained and used to construct an effect image layer; then, the target image and the effect image layer are superimposed to obtain a superimposed target image with a stereoscopic visual effect, and then the superimposed target image is displayed to the user, so that the user can obtain a relatively better interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] To more clearly illustrate the embodiments of this specification, the following will briefly introduce the accompanying drawings required for the embodiments. The accompanying drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0015] Figure 1 It is a schematic flowchart of an image data processing method provided by an embodiment of this specification;
[0016] Figure 2 It is a schematic physical diagram of a doctor's console applying the image data processing method provided by an embodiment of this specification;
[0017] Figure 3 It is a schematic physical diagram of a binocular endoscope applying the image data processing method provided by an embodiment of this specification;
[0018] Figure 4 It is a schematic physical diagram of an image platform applying the image data processing method provided by an embodiment of this specification;
[0019] Figure 5 It is a schematic diagram of an embodiment applying the image data processing method provided by an embodiment of this specification in a scenario example;
[0020] Figure 6 It is a schematic diagram of an embodiment applying the image data processing method provided by an embodiment of this specification in a scenario example;
[0021] Figure 7 It is a schematic diagram of an embodiment applying the image data processing method provided by an embodiment of this specification in a scenario example;
[0022] Figure 8 It is a schematic diagram of an embodiment applying the image data processing method provided by an embodiment of this specification in a scenario example;
[0023] Figure 9 It is a schematic diagram of an embodiment applying the image data processing method provided by an embodiment of this specification in a scenario example;
[0024] Figure 10 It is a schematic diagram of an embodiment applying the image data processing method provided by an embodiment of this specification in a scenario example;
[0025] Figure 11 It is a schematic diagram of an embodiment applying the image data processing method provided by an embodiment of this specification in a scenario example;
[0026] Figure 12It is a schematic diagram of the structural composition of an image platform provided by an embodiment of this specification;
[0027] Figure 13 It is a schematic diagram of the structural composition of a computer device provided by an embodiment of this specification;
[0028] Figure 14 It is a schematic diagram of the structural composition of an image data processing device provided by an embodiment of this specification;
[0029] Figure 15 It is a physical diagram of a surgical platform provided by an embodiment of this specification. Detailed implementation manners
[0030] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.
[0031] Refer to Figure 1 As shown, an embodiment of this specification provides an image data processing method. Specifically in implementation, this method may include the following content.
[0032] S101: Obtain a target image.
[0033] S102: Identify the target object in the target image to obtain a corresponding recognition result; and obtain the depth information of the target object.
[0034] S103: According to the recognition result and depth information of the target object, display the target label of the target object in the target image; where the target label at least includes the object category information of the target object.
[0035] In some implementations, the above target image can be specifically understood as an image that at least contains a target object. Among them, the above target object can be specifically an object that the user is concerned about. The number of target objects included in the target image can be one or multiple.
[0036] In some embodiments, the target image can specifically include: a first image, and / or, a second image. Among them, the first image can be specifically an image containing a target object collected by a left camera (or a first camera); the second image can be specifically an image containing a target object collected by a right camera (or a second camera). Among them, the relative positions of the left camera and the right camera are fixed.
[0037] Corresponding to different application scenarios, the above-mentioned target object can be different types of data objects. Specifically, for example, in a surgical scenario, the above-mentioned target object can be the organs of a patient and / or surgical instruments used by medical staff. In a traffic monitoring scenario, the above-mentioned target object can be vehicles traveling on a road, etc.
[0038] Of course, it should be noted that the above-listed application scenarios and target objects are only illustrative. In specific implementation, according to the specific situation and processing requirements, the above image data processing method can also be applied to other types of application scenarios; correspondingly, the above target object can also include other types of data objects in other types of application scenarios. This specification does not make any limitations in this regard.
[0039] Hereinafter, mainly taking the surgical scenario as an example, the image data processing method will be specifically described. For other application scenarios, the embodiments of the following surgical scenario can be referred to. This specification will not elaborate on this.
[0040] Refer to Figure 4 As shown, the image data processing method provided by the embodiments of this specification can be specifically applied to an image platform (or, a surgical robot, a doctor's console, etc.). Specifically, the above image platform is at least further connected to a binocular endoscope facing the patient to be operated on (refer to Figure 3 As shown), and the image platform faces the side of the medical staff (or called the user) performing the surgical operation.
[0041] Specifically, refer to Figure 3 As shown, the above binocular endoscope at least includes two cameras with relatively fixed positions, namely a left camera (or called the first camera) and a right camera (or called the second camera), as well as related components such as an image sensor. The binocular endoscope can be inserted into a specific surgical environment to photograph the target object that the medical staff is concerned about. In specific implementation, the binocular endoscope can image the photographed target object on the image sensor through the cameras, and the image sensor is used to output an electrical signal according to the imaging of the cameras, and the electrical signal is used to generate a target image including the photographed target object.
[0042] Refer to Figure 4As shown, the above image platform may at least include components such as an image data processing device (e.g., a processor supporting image data processing or an image processing host, etc.), and a display device (e.g., a display screen, etc.). In specific implementation, the above image platform may use the image data processing method provided in the embodiments of this specification by the data processing device to process the target image collected by the binocular endoscope; and then display the target image through the display device, and display the target label of the target object in the target image to convey the recognition result of the target object and the three-dimensional depth information of the target object to the medical staff. Among them, the display device may specifically include a two-dimensional display screen, or a three-dimensional display screen, etc.
[0043] Specifically, the above image platform may also be connected to a doctor console (or doctor cart), and reference can be made to Figure 2 as shown. Among them, the doctor console at least includes a monitor (e.g., a stereoscopic monitor, etc.). Correspondingly, the target image carrying the target label generated by the image platform can be transmitted to the doctor console. The doctor console can display the target image containing the target label to the medical staff through the monitor.
[0044] Furthermore, the above doctor console may also be configured with VR devices (e.g., VR glasses, etc.), which support displaying target images based on virtual reality technology so that medical staff can better observe organs and surgical instruments that appear during the operation, etc.
[0045] In addition, the above doctor console may also be connected to the surgical platform, and reference can be made to Figure 15 as shown. Among them, the surgical platform at least includes components such as a robotic arm and a robotic arm control device.
[0046] Correspondingly, medical staff can accurately issue corresponding operation instructions to the surgical platform through the doctor console according to the target image carrying the target label displayed on the monitor in the doctor console; the robotic arm control device of the surgical platform can respond to the operation instructions and control the robotic arm to perform corresponding actions to complete specific surgical operations.
[0047] In some embodiments, the above target image may specifically be an image containing target objects such as organs and / or surgical instruments that are of concern during the surgical operation, which is collected by the binocular endoscope at regular intervals or in real time before or during the surgical operation.
[0048] In some embodiments, the target image may specifically include a first image, and / or a second image collected by the binocular endoscope; correspondingly, the target object may specifically include: tissue organs, and / or surgical instruments.
[0049] In some embodiments, when specifically implementing the above-mentioned acquisition of the target image, it may include: controlling the binocular endoscope to capture the target object to obtain a frame of image as the target image.
[0050] When specifically implementing the above-mentioned acquisition of the target image, it may also include: controlling the binocular endoscope to capture the target object at every preset time interval (for example, every 5 seconds) to obtain multiple frames of images. Among them, the multiple frames of images can be arranged in the order of the acquisition time from front to back. Then, according to the sequence of the acquisition time, the image to be processed currently is obtained frame by frame from the multiple frames of images as the target image.
[0051] In some embodiments, during the process of identifying the target object in the target image to obtain the recognition result, one image of the first image or the second image can be used alone; during the process of obtaining the depth information of the target object, two images of the first image and the second image need to be used simultaneously.
[0052] In some embodiments, when specifically implementing the above-mentioned identification of the target object in the target image to obtain the corresponding recognition result, it may include: using an image processing model to process the target image to determine the object category of the target object in the target image as the recognition result.
[0053] Among them, the above-mentioned image processing model can be specifically understood as a neural network model that can find the target object in the target image and at least can identify information such as the object category of the target object. The establishment method of the above-mentioned image processing model will be described separately later.
[0054] In some embodiments, the image processing model includes an image processing model trained based on the YOLO model.
[0055] Among them, the above-mentioned YOLO model can be specifically understood as a neural network model of a target detection algorithm. Based on this target detection algorithm model, while achieving fast detection, it also has a high accuracy rate.
[0056] In addition, the above-mentioned image processing model can also include an image processing model trained based on other types of models such as Fast YOLO or R-CNN.
[0057] In some embodiments, when specifically implementing the above-mentioned use of the image processing model to process the target image to determine the object category of the target object in the target image, refer to Figure 5As shown, it may include the following: rasterize the target image using an image processing model to obtain the rasterized target image; use a convolutional neural network to process the rasterized target image to detect and identify the object category of the target object in the target image. Among them, the convolutional neural network may specifically be a model structure in the image processing model.
[0058] In some embodiments, for the above-mentioned process of using a convolutional neural network to process the rasterized target image to detect and identify the object category of the target object in the target image, specifically, it may include the following:
[0059] S1: Use a convolutional neural network to use a bounding box to segment a first sub-image region from the rasterized target image; wherein, the first sub-image region is a non-background image region;
[0060] S2: Use a convolutional neural network to detect whether there is a target object in the first sub-image region; and according to the detection result, use a detection box to segment a second sub-image region from the first sub-image region; wherein, the second sub-image region is an image region containing the target object to be recognized;
[0061] S3: Use a convolutional neural network to perform image recognition on the second sub-image region to determine the object category of the target object.
[0062] Through the above embodiments, the background image region that does not contain the target object can be filtered out from the rasterized target image using a bounding box first, and the non-background image region containing the target object is retained as the first sub-image region to be further processed; then, the image region that may contain the target object can be further framed out from the first sub-image region using a detection box as the second sub-image region; furthermore, the convolutional neural network can only perform image recognition on the above second sub-image region to determine the object category of the target object in the target image, reducing the data processing volume and improving the overall recognition efficiency.
[0063] In some embodiments, after using a detection box to segment the second sub-image region from the first sub-image region, the method may further include: using the non-maximum suppression algorithm (NMS) to filter out the second sub-image regions segmented based on redundant detection boxes from the second sub-image region to obtain the corrected second sub-image region; correspondingly, using a convolutional neural network to perform image recognition on the corrected second sub-image region to determine the object category of the target object.
[0064] Through the above embodiments, a second sub-image region segmented by a detection box with a relatively high confidence can be further filtered from the second sub-image region based on the non-maximum suppression algorithm as the corrected second sub-image region. Furthermore, the convolutional neural network can only perform image recognition on the corrected second sub-image region in the second sub-image region, further reducing data processing and further improving the overall recognition efficiency.
[0065] In some embodiments, filtering the second sub-image region segmented based on redundant detection boxes from the second sub-image region by using the non-maximum suppression algorithm (NMS) to obtain the corrected second sub-image region may specifically include the following contents when implemented: S1: Calculate the areas of all detection boxes in the second sub-image region. S2: Calculate the confidence levels for all detection boxes and sort the detection boxes in descending order according to the confidence levels. Among them, when the detection box has a higher matching degree with the sample images in the sample image set, the confidence level is relatively higher. S3: Retain the detection box with the highest confidence level and calculate the intersection between this detection box and the remaining detection boxes. S4: Calculate the intersection over union IoU between the detection box with the highest confidence level and the remaining detection boxes. Reference can be made to Figure 6 as shown. S5: Retain the detection boxes with IoU less than the threshold and eliminate the redundant detection boxes with large IoU. This is because generally, the larger the IoU, the higher the overlap degree with the detection box with the highest confidence level. S6: Detect whether the number of detection boxes with IoU less than the threshold is 0. When it is determined that the number of detection boxes with IoU less than the threshold is 0, the corrected second sub-image region can be determined according to the finally retained detection boxes. When it is determined that the number of detection boxes with IoU less than the threshold is not 0, repeat the above steps S3 to S5 until the number of detection boxes with IoU less than the threshold is 0.
[0066] In some embodiments, when processing the target image by using the image processing model to determine the object category of the target object in the target image, the method may further include when implemented: Extracting the target state features of the target object from the second sub-image region; Determining the state of the target object according to the target state features and a preset state database; Correspondingly, combining the object category of the target object and the state of the target object as the recognition result.
[0067] Among them, the above target state features may specifically include the color feature, texture feature, brightness feature, contour line feature, etc. of the target object in the second sub-image region.
[0068] For different target objects, the states of the above target objects can include multiple different types of states. Specifically, for example, when the target object is a scalpel (a surgical instrument), the states of the target object can include one or more of the following listed usage states: in use, not in use, damaged, undamaged, etc. Another example is that when the target object is the stomach (an organ), the states of the target object can include one or more of the following listed health states: healthy, gastric bleeding, gastric ulcer, etc.
[0069] The above-mentioned preset state database can be specifically understood as a preset state database constructed by pre-learning the image data of the target object in multiple states. Among them, the preset state database includes multiple sub-databases, and each sub-database corresponds to an object category. Further, each sub-database contains multiple preset reference state features respectively corresponding to multiple states of the target object of this object category.
[0070] In some embodiments, during specific implementation, the target state features of the target object can be extracted from the second sub-image area by adopting an adaptive threshold method. Among them, the above-mentioned adaptive threshold method can specifically include: first, calculate the reference value of the local image area in the second sub-image area according to the brightness distribution of different areas in the second sub-image area (for example, the mean value, median value, Gaussian weighted average, etc. of the pixels in the local image area); then, transform the threshold of the corresponding local image area accordingly according to the above reference value; and then extract the target state features of the local image area according to the threshold of the local image area. Thus, the error can be reduced and the target state features of the target object can be extracted more accurately.
[0071] In some embodiments, determining the state of the target object according to the target state features and the preset state database can specifically include the following content during implementation: according to the object category of the target object recognized by the image processing model, determine the sub-database that matches the target object from the preset state database as the target sub-database; match the target state features with the multiple preset reference state features stored in the target sub-database respectively, and find the preset reference state feature with the highest matching degree with the target state features; and determine the state corresponding to this preset reference state feature as the state of this target object.
[0072] In some embodiments, the object category of the target object and the state of the target object can be combined to obtain relatively richer and closer recognition results. Correspondingly, the above recognition results can include relevant information of the target object such as the object category of the target object and the state of the target object. In addition, the above recognition results can also include other relevant information such as the number of the target object.
[0073] In some embodiments, refer to Figure 5 As shown, while processing the target image with the image processing model to obtain the corresponding recognition result, the depth information of the target object in the target image can also be acquired; then, by integrating the recognition result and the depth information, a target label of the target object is generated and displayed in the target image.
[0074] In some embodiments, refer to Figure 7 As shown, the above-mentioned acquisition of the depth information of the target object, in specific implementation, may include the following content:
[0075] S1: Determine a first distance between a first key pixel point and the optical center of the left camera, and a second distance between a second key pixel point and the optical center of the right camera according to the first image and the second image; wherein, the first key pixel point is the pixel point corresponding to the key point of the target object in the first image, and the second key pixel point is the pixel point corresponding to the key point of the target object in the second image;
[0076] S2: Calculate the vertical distance between the key point of the target object and the plane where the left camera and the right camera are located according to the first distance and the second distance, and use it as the depth information of the target object.
[0077] Among them, the above-mentioned key point may specifically be the center point of the target object.
[0078] In some embodiments, the above-mentioned calculation of the vertical distance between the key point of the target object and the plane where the left camera and the right camera are located according to the first distance and the second distance, in specific implementation, may include the following content: Obtain the baseline distance between the left camera and the right camera, and the camera focal length; use the first distance, the second distance, the baseline distance, and the camera focal length to calculate the vertical distance between the key point of the target object and the plane where the left camera and the right camera are located.
[0079] In some embodiments, in specific implementation, the vertical distance between the key point of the target object and the plane of the left camera and the right camera can be calculated according to the following formula:
[0080]
[0081] Among them, z is the vertical distance between the key point of the target object and the plane where the left camera and the right camera are located, f is the camera focal length, u L is the first distance, u R is the second distance.
[0082] In some embodiments, the larger the value of the depth information of the target object, the farther the target object is from the lens; on the contrary, the smaller the value of the depth information of the target object, the closer the target object is to the lens.
[0083] Through the above embodiments, the first image and the second image collected by the binocular endoscope can be utilized simultaneously, and the depth information of the target object can be calculated by computing the parallax between the two images.
[0084] In some embodiments, the above-mentioned displaying the target label of the target object in the target image according to the recognition result and the depth information of the target object may specifically include the following when implemented:
[0085] S1: Generate a label for the target object according to the recognition result of the target object;
[0086] S2: Determine the label size parameter matching the target object by using the depth information of the target object according to a preset matching rule;
[0087] S3: Adjust the character size in the label of the target object according to the label size parameter to obtain the target label of the target object; and display the target label of the target object in the target image.
[0088] Among them, the label of the above-mentioned target object at least includes the object category of the target object. Further, other relevant information such as the state of the target object and the number of the target object may also be included in the label.
[0089] The above-mentioned preset matching rule can be specifically understood as a matching rule constructed based on the perspective theory. Based on the above-mentioned preset matching rule, for the label of a target object relatively close to the lens, a relatively larger size parameter will be matched according to the visual sense of the human eye for this target object; on the contrary, for the label of a target object relatively far from the lens, a relatively smaller size parameter will be matched according to the visual sense of the human eye for this target object. Thus, the distance of this target object relative to the lens can be reflected by the size of the label, and the three-dimensional depth information about this target object can be accurately conveyed to the user.
[0090] In some embodiments, the above-mentioned determining the label size parameter matching the target object by using the depth information of the target object according to a preset matching rule, as shown in Figure 8 When implemented, it may specifically include the following:
[0091] S1: Project the pixel points of the key points of the target object in the target image onto a reference plane to obtain corresponding reference pixel points; wherein, the reference plane is the plane where the pixel point with the maximum depth information in the target image is located;
[0092] S2: Obtain the distance between the reference pixel point and the center of the reference plane, and the maximum value of the depth information in the target image;
[0093] S3: According to the preset matching relationship parameters, use the depth information of the target object, the distance between the reference pixel point and the center of the reference plane, and the maximum value of the depth information in the target image to calculate the label size parameters matching the target object.
[0094] Wherein, the reference plane can be specifically understood as the plane where the data object at the farthest end recognized in the target image is located. The maximum value of the depth information in the target image can be specifically understood as the depth information of this farthest data object.
[0095] According to the preset matching rules, in specific implementation, reference can be made to Figure 8 as shown. First, project the pixel point P of the key point of the target object in the target image onto the reference plane for imaging to obtain the corresponding reference pixel point P'. Then, the maximum value d1 of the depth information in the target image (for example, the vertical distance between the farthest data object and the plane where the left camera and the right camera are located) and the depth information d2 of the key point of the target object can be obtained, and the vertical distance d3 between the reference pixel point and the center of the reference plane can be calculated. Among them, the viewing angle range of the lens can be denoted as θ, and the angle between the lens and the relative central axis of P can be denoted as β.
[0096] In specific calculation, the viewing angle θ of the lens can be obtained first. According to the viewing angle of the lens and the maximum value of the depth information in the target image, the viewing range of the farthest data object can be calculated: Then use D to calculate the vertical distance d3 between the reference pixel point and the center of the reference plane.
[0097] Furthermore, the label size parameters matching the target object can be calculated according to the following formula using the preset matching relationship parameters:
[0098]
[0099] Wherein, DF is the label size parameter, B is the preset matching relationship parameter, d2 is the depth information of the target object, d3 is the vertical distance between the reference pixel point and the center of the reference plane, and d1 is the maximum value of the depth information in the target image.
[0100] In some embodiments, the above label size parameters can specifically be the character size in the target label (for example, the font size of the character, etc.). The specific value of the above preset matching relationship parameter can be determined according to the preset matching rules. Specifically, the preset matching relationship parameter can be determined according to the number of target objects included in the target image and the maximum value of the depth information in the target image.
[0101] In some embodiments, by adjusting the character size in the label of the target object according to the label size parameter, the target label of the target object relatively closer to the lens in the target image can be made relatively larger, and the target label of the target object relatively farther from the lens can be made relatively smaller, so that the user can intuitively feel the layering between different target objects in the target image and effectively convey the three-dimensional depth information of the target object to the user. Reference can be made to Figure 9 as shown.
[0102] In some embodiments, after obtaining the target label of the target object, when the method is specifically implemented, the following content may further be included: adjusting the display effect parameters of the target label of the target object according to the depth information of the target object to obtain the target label with adjusted display effect; wherein, the display effect parameters include at least one of the following: shadow parameter, deformation parameter, brightness parameter; correspondingly, displaying the target label with adjusted display effect in the target image.
[0103] Specifically, for example, for a target object with larger depth information, the shadow parameter and deformation parameter of the target label of this target object can be increased specifically, and at the same time, the brightness parameter of the target label of this target object can be decreased. For a target object with smaller depth information, the shadow parameter and deformation parameter of the target label of this target object can be decreased specifically, and at the same time, the brightness parameter of the target label of this target object can be increased. In this way, the user can more intuitively and comprehensively feel the depth information of the target object through the target label and know the relative position relationship between the target objects.
[0104] It should be noted that the above-listed display effect parameters are only illustrative. When specifically implemented, according to the specific application scenario and processing requirements, other types of display effect parameters may also be introduced. This specification does not make a limitation in this regard.
[0105] In some embodiments, when displaying the target label of the target object in the target image, the following content may further be included when specifically implemented:
[0106] S1: Obtain the position parameter of the target object in the target image;
[0107] S2: Determine the display position coordinates according to the position parameter of the target object in the target image;
[0108] S3: Set and display the target label of the target object at the corresponding position in the target image according to the display position coordinates.
[0109] In some embodiments, the position parameter of the above-mentioned target object in the target image can be specifically understood as the position coordinates of the key points of the target object in the target image, which can be expressed as (x1, y1). Specifically in implementation, the position parameter of the target object can be determined according to the detection frame of the target object.
[0110] In some embodiments, after using the detection frame to segment the second sub-image region from the first sub-image region, when the method is specifically implemented, the following content can also be included: obtaining the position coordinates of the detection frame; determining the position parameter of the target object in the target image according to the position coordinates of the detection frame.
[0111] Further, a matching offset parameter (for example, it can be denoted as A) can be set according to the distance between the target object and other adjacent target objects, the size of the target object, and the position parameters of other adjacent target objects around the target object, etc.; then, according to the position parameter of the target object and the offset parameter, the corresponding display position coordinates can be calculated, which can be denoted as: (x1 + A, y1 + A).
[0112] Furthermore, according to the above-mentioned display position coordinates, the target label of the target object can be set and displayed at the corresponding position in the target image.
[0113] This can ensure that the target label of the target object displayed in the target image does not occlude the target object and other adjacent target objects around the target object, further improving the user's interaction experience.
[0114] In some embodiments, when specifically implemented, based on the target tracking algorithm, according to the position parameter of the target object, the position parameters of the target object in two adjacent frames of images captured at adjacent acquisition time points can be tracked, so that the user can more quickly and accurately understand the changes of the same target object between adjacent acquisition time points.
[0115] Specifically, for example, it can be referred to Figure 10 As shown, based on the target tracking algorithm, according to the position parameter of the surgical instrument used during the surgical process, the same surgical instrument in the Nth frame and the (N + 1)th frame can be tracked.
[0116] In addition, the above-mentioned image platform and / or doctor control also support the user to perform tracking settings on one or more target objects identified in the current frame of the target image. Correspondingly, the operation of the tracking setting for the current frame of the target image can be received and, according to this, the target object indicated by the user for tracking can be determined and marked; during the process of processing the next frame of the target image, the marked target object can be tracked according to the recognition result.
[0117] In some embodiments, during specific implementation, augmented reality technology can also be introduced and used to further process the target image. By presenting the target image processed based on augmented reality technology to the user, the user can more intuitively and vividly perceive the three-dimensional depth information of the target object in the target image, as well as the sense of hierarchy formed due to the depth information difference between different target objects, thereby further enhancing the user's interaction experience.
[0118] In some embodiments, during specific implementation, the method may further include the following: constructing an effect image layer based on the difference value between the depth information of different target objects in the target image; wherein, the effect image layer includes the target object and the three-dimensional effect data of the target label of the target object; superimposing the target image and the effect image layer to obtain a superimposed target image; and presenting the superimposed target image.
[0119] Among them, the above-mentioned effect image layer can specifically be constructed based on augmented reality technology.
[0120] The above-mentioned Augmented Reality (AR) technology can specifically refer to a technology that combines virtual information with real image information, capable of realizing the "augmentation" of the user's perception of the real world.
[0121] By superimposing the target image and the effect image layer constructed based on the above-mentioned augmented reality technology, a superimposed target image with stronger three-dimensional sense can be obtained. Then, by presenting the above-mentioned superimposed target image, the user can more clearly understand the target object in the image and obtain a relatively better interaction experience. For example, by using the above AR technology, it is possible to better display the hierarchical information between organs and surgical instruments in the surgical environment, avoid misperceptions when medical staff use the image platform and the doctor's console, and reduce the risk of surgical operation errors.
[0122] In some embodiments, during specific implementation, the method may further include: constructing a target image with the target label presented based on virtual reality technology according to the target image and the target label of the target object.
[0123] Among them, the above-mentioned Virtual Reality (VR) technology can specifically refer to a three-dimensional simulated reality technology developed relying on technologies such as three-dimensional real-time graphics display, three-dimensional positioning and tracking, tactile and olfactory sensing technologies, artificial intelligence technology, high-speed computing and parallel computing technologies, and human behavior science. It can enable users to be in a virtual environment with three-dimensional vision, hearing, touch, and even smell, and support users to interact with information in this virtual environment.
[0124] In specific implementation, based on the virtual environment modeling algorithm, by processing the target image and the target label of the target object, the above-mentioned target image with the target label based on the virtual reality technology can be constructed.
[0125] Correspondingly, medical workers can wear VR devices configured by the doctor's console to view the above-mentioned target image with the target label based on the virtual reality technology, so as to more clearly understand the hierarchical information between organs and surgical instruments in the surgical environment and obtain a relatively better interaction experience.
[0126] In some embodiments, after the target label of the target object is displayed in the target image, the method further includes: displaying the above-mentioned target image with the target label of the target object to the user, so that the user can efficiently understand relevant information such as the object category of the target object through the displayed target image; and at the same time intuitively perceive the depth information of the distance between the target object and the camera lens, enabling the user to better perform specific data processing based on the above information.
[0127] For example, medical workers can clearly understand the patient's organs and surgical instruments under the current binocular endoscope view according to the target image with the target label of the target object displayed on the display device of the image platform and / or the doctor's console, so as to precisely perform specific surgical operations on the patient by controlling the robotic arm through the doctor's console.
[0128] In some embodiments, before processing the target image using the image processing model, specifically in implementation, the method may further include the following:
[0129] S1: Collect a sample image set; wherein, the sample image set includes multiple sample images containing sample objects arranged according to the collection time;
[0130] S2: Use a bounding box to mark the image region of the sample object in the sample image and mark the object category of the sample object to obtain a labeled sample image set;
[0131] S3: Use the labeled sample image set to train an initial model to obtain an image processing model.
[0132] In some embodiments, specifically in implementation, the binocular endoscope can be used to capture sample objects (such as organs or surgical instruments, etc.) in the corresponding application scenario (such as a surgical scenario) at preset time intervals to obtain multiple sample images arranged according to the collection time, so as to construct a sample image set. It is also possible to record a sample video containing sample objects in the application scenario; then intercept multiple screenshots from the sample video as the multiple sample images to construct a sample image set.
[0133] In some embodiments, after obtaining a plurality of sample images, the plurality of images may be preprocessed to remove error data in the sample images; and then a sample image set may be constructed according to the processed sample images.
[0134] In some embodiments, the specific preprocessing may include: screening the sample images according to the clarity, the size of the blurred area, the degree of image stability, etc. of the sample images, so as to eliminate the sample images that are too bright, too dark, or blurred, and retain the sample images with smaller errors and higher precision.
[0135] In some embodiments, when specifically annotating the sample images, an image area where the sample object is located may be framed in the sample image by using a bounding box, and the object category of the sample object may be annotated to obtain the annotated sample image, so as to obtain the annotated sample image set.
[0136] Furthermore, relevant information such as the state and number of the sample object may also be annotated in the sample image to obtain the annotated sample image containing richer data information. Reference may be made to Figure 11 as shown.
[0137] In some embodiments, during specific training, an initial model based on YOLO may be constructed; the annotated sample image set may be divided into a training set and a test set; the initial model may be continuously trained using the training set, and the trained model may be tested using the test set; until the test result meets the corresponding precision requirement, the training is stopped, and the model at this time is determined as the image processing model.
[0138] As can be seen from the above, based on the image data processing method provided in the embodiments of this specification, after obtaining the target image and identifying the target object in the target image to obtain the corresponding recognition result, the depth information of the target object can also be obtained simultaneously; then, by using the recognition result and depth information of the target object at the same time, a target label including at least the object category information of the target object is generated and displayed on the target image. Thus, it is possible to automatically identify and use the target label to mark the relevant information of the target object concerned by the user in the target image, and through the displayed target label, the three-dimensional depth information of the target object can be accurately and intuitively conveyed to the user in the two-dimensional target image, enabling the user to obtain a better interaction experience. Specifically, according to the depth information of the target object, based on the principle of perspective, the label size parameter matching the target object can be determined; then, according to the label size parameter, a target label capable of expressing the distance of the target object relative to the camera lens is generated; and the target label is displayed in the target image so that the user can directly and intuitively determine the distance of the target object relative to the camera lens based on the label size of the target label. Further, according to the depth information of the target object, by specifically adjusting the display effect parameters of the target label such as shadow parameters, deformation parameters, and brightness parameters, the three-dimensional depth information of the target object can be more effectively conveyed to the user by displaying the target label in the target image. In addition, based on the augmented reality (AR) technology, first, the difference value between the depth information of different target objects in the target image is obtained and used to construct an effect image layer; then, the target image and the effect image layer are superimposed to obtain the superimposed target image, and then the above superimposed target image is displayed to the user, so that the user can obtain a relatively better interaction experience. It is also possible to automatically track the target object concerned by the user, further improving the user's interaction experience.
[0139] Referring to Figure 12 As shown, the embodiments of this specification also provide a doctor console 1200. Specifically, it may include a binocular endoscope 1201 and an image data processing device 1202. Among them, the above binocular endoscope 1201 is specifically used to collect a target image; the above image data processing device 1202 is specifically used to process the above target image in the following manner: identify the target object in the target image to obtain the corresponding recognition result; and obtain the depth information of the target object; according to the recognition result and depth information of the target object, display the target label of the target object in the target image; where the target label includes at least the object category information of the target object.
[0140] An embodiment of this specification also provides a computer device, including a processor and a memory for storing executable instructions of the processor. When specifically implemented, the processor may execute the following steps according to the instructions: obtaining a target image; identifying a target object in the target image to obtain a corresponding identification result; and obtaining depth information of the target object; according to the identification result and depth information of the target object, displaying a target label of the target object in the target image; wherein, the target label at least includes object category information of the target object.
[0141] In order to be able to more accurately complete the above instructions, refer to Figure 13 As shown, an embodiment of this specification also provides another specific computer device 1300. Among them, the computer device includes a network communication port 1301, a processor 1302, and a memory 1303. The above structures are connected by internal cables so that each structure can perform specific data interactions.
[0142] Among them, the network communication port 1301 can specifically be used to obtain a target image.
[0143] The processor 1302 can specifically be used to identify a target object in the target image to obtain a corresponding identification result; and obtain depth information of the target object; according to the identification result and depth information of the target object, display a target label of the target object in the target image; wherein, the target label at least includes object category information of the target object.
[0144] The memory 1303 can specifically be used to store corresponding instruction programs.
[0145] In this embodiment, the network communication port 1301 can be bound to different communication protocols, so as to send or receive different data virtual ports. For example, the network communication port can be a port responsible for web data communication, can also be a port responsible for FTP data communication, can also be a port responsible for mail data communication, and a port for industrial Ethernet fieldbus EtherCAT communication. In addition, the network communication port can also be a physical communication interface or communication chip. For example, it can be a wireless mobile network communication chip, such as GSM, CDMA, etc.; it can also be a Wifi chip; it can also be a Bluetooth chip.
[0146] In this embodiment, the processor 1302 can be implemented in any suitable manner. For example, the processor can take the form of, for example, a microprocessor or a processor, a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an Application Specific Integrated Circuit (ASIC), a programmable logic controller, and an embedded microcontroller, etc. This specification does not make any limitations.
[0147] In this embodiment, the memory 1303 can include multiple levels. In a digital system, anything that can store binary data can be a memory; in an integrated circuit, a circuit with a storage function without a physical form is also called a memory, such as a RAM, FIFO, etc.; in a system, a storage device with a physical form is also called a memory, such as a memory stick, a TF memory card, etc.
[0148] This specification embodiment also provides a computer storage medium based on the above image data processing method. The computer storage medium stores computer program instructions, and when the computer program instructions are executed, the following steps are implemented: obtaining a target image; identifying a target object in the target image to obtain a corresponding identification result; and obtaining depth information of the target object; according to the identification result and depth information of the target object, displaying a target label of the target object in the target image; where the target label at least includes object category information of the target object.
[0149] In this embodiment, the above storage medium includes but is not limited to a Random Access Memory (RAM), a Read-Only Memory (ROM), a Cache, a Hard Disk Drive (HDD), or a Memory Card. The memory can be used to store computer program instructions. The network communication unit can be set according to the standards specified by the communication protocol and is used as an interface for network connection communication.
[0150] In this embodiment, the functions and effects specifically implemented by the program instructions stored in the computer storage medium can be explained by comparison with other embodiments and will not be elaborated here.
[0151] Refer to Figure 14 As shown, at the software level, this specification embodiment also provides an image data processing device 1400, which specifically can include the following structural modules:
[0152] An obtaining module 1401, which specifically can be used to obtain a target image;
[0153] The first processing module 1402 can be specifically configured to identify a target object in a target image to obtain a corresponding recognition result; and acquire depth information of the target object.
[0154] The second processing module 1403 can be specifically configured to display a target label of the target object in the target image according to the recognition result and depth information of the target object; wherein, the target label at least includes object category information of the target object.
[0155] In some embodiments, the target image may specifically include: a first image, and / or, a second image; wherein, the first image is an image captured by a left camera, the second image is an image captured by a right camera, and the relative positions of the left camera and the right camera are fixed.
[0156] In some embodiments, the target image may specifically include a first image, and / or, a second image captured by a binocular endoscope; correspondingly, the target object may specifically include: a tissue organ, and / or, a surgical instrument.
[0157] In some embodiments, when the first processing module 1402 is specifically implemented, the depth information of the target object can be acquired in the following manner: determine a first distance between a first key pixel point and the optical center of the left camera, and a second distance between a second key pixel point and the optical center of the right camera according to the first image and the second image; wherein, the first key pixel point is a pixel point corresponding to a key point of the target object in the first image, and the second key pixel point is a pixel point corresponding to a key point of the target object in the second image; calculate a vertical distance between the key point of the target object and the plane where the left camera and the right camera are located according to the first distance and the second distance, as the depth information of the target object.
[0158] In some embodiments, when the first processing module 1402 is specifically implemented, the vertical distance between the key point of the target object and the plane where the left camera and the right camera are located can be calculated in the following manner according to the first distance and the second distance: acquire a baseline distance between the left camera and the right camera, and a camera focal length; calculate the vertical distance between the key point of the target object and the plane where the left camera and the right camera are located by using the first distance, the second distance, the baseline distance, and the camera focal length.
[0159] In some embodiments, when the second processing module 1403 is specifically implemented, the target label of the target object may be displayed in the target image according to the recognition result and depth information of the target object in the following manner: generate a label for the target object according to the recognition result of the target object; determine a label size parameter matching the target object by using the depth information of the target object according to a preset matching rule; adjust the character size in the label of the target object according to the label size parameter to obtain the target label of the target object; and display the target label of the target object in the target image.
[0160] In some embodiments, when the second processing module 1403 is specifically implemented, the label size parameter matching the target object may be determined by using the depth information of the target object according to a preset matching rule in the following manner: project the pixel points of the key points of the target object in the target image onto a reference plane to obtain corresponding reference pixel points; wherein, the reference plane is the plane where the pixel point with the maximum depth information in the target image is located; obtain the distance between the reference pixel points and the center of the reference plane, and the maximum value of the depth information in the target image; calculate the label size parameter matching the target object by using the depth information of the target object, the distance between the reference pixel points and the center of the reference plane, and the maximum value of the depth information in the target image according to a preset matching relationship parameter.
[0161] In some embodiments, when the second processing module 1403 is specifically implemented, after obtaining the target label of the target object, the display effect parameter of the target label of the target object may be adjusted according to the depth information of the target object to obtain the target label with the adjusted display effect; wherein, the display effect parameter includes at least one of the following: shadow parameter, deformation parameter, brightness parameter; correspondingly, the target label with the adjusted display effect is displayed in the target image.
[0162] In some embodiments, when the second processing module 1403 is specifically implemented, the target label of the target object may be displayed in the target image in the following manner: obtain the position parameter of the target object in the target image; determine the display position coordinate according to the position parameter of the target object in the target image; set and display the target label of the target object at the corresponding position in the target image according to the display position coordinate.
[0163] In some embodiments, when the second processing module 1403 is specifically implemented, it can also be used to construct an effect image layer according to the difference value between the depth information of different target objects in the target image; wherein, the effect image layer includes the target object and the stereoscopic effect data of the target label of the target object; perform superposition processing on the target image and the effect image layer to obtain the superimposed target image; and display the superimposed target image.
[0164] In some embodiments, when the second processing module 1403 is specifically implemented, it can also be used to construct a target image showing the target label based on virtual reality technology according to the target image and the target label of the target object.
[0165] In some embodiments, when the first processing module 1402 is specifically implemented, it can identify the target object in the target image in the following manner to obtain the corresponding recognition result: process the target image using an image processing model to determine the object category of the target object in the target image as the recognition result.
[0166] In some embodiments, the image processing model includes at least one of the following: an image processing model trained based on the YOLO model, an image processing model trained based on the Fast YOLO model, and an image processing model trained based on the R-CNN model.
[0167] In some embodiments, when the first processing module 1402 is specifically implemented, it can process the target image using the image processing model in the following manner to determine the object category of the target object in the target image: perform rasterization processing on the target image using the image processing model to obtain the rasterized target image; use a convolutional neural network to detect and identify the object category of the target object in the target image by processing the rasterized target image; wherein, the convolutional neural network is the model structure in the image processing model.
[0168] In some embodiments, when the first processing module 1402 is specifically implemented, it can detect and identify the object category of the target object in the target image by using the convolutional neural network to process the rasterized target image in the following manner: use the convolutional neural network to segment a first sub-image region from the rasterized target image using a segmentation box; wherein, the first sub-image region is a non-background image region; use the convolutional neural network to detect whether there is a target object in the first sub-image region; and according to the detection result, use a detection box to segment a second sub-image region from the first sub-image region; wherein, the second sub-image region is an image region containing the target object to be recognized; use the convolutional neural network to perform image recognition on the second sub-image region to determine the object category of the target object.
[0169] In some embodiments, after the second sub-image region is segmented from the first sub-image region using the detection frame, when the first processing module 1402 is specifically implemented, it can also be used to filter out the second sub-image region segmented based on redundant detection frames from the second sub-image region by using the non-maximum suppression algorithm, so as to obtain the corrected second sub-image region; correspondingly, the first processing module 1402 can use a convolutional neural network to perform image recognition on the corrected second sub-image region to determine the object category of the target object.
[0170] In some embodiments, while using the image processing model to process the target image to determine the object category of the target object in the target image, the first processing module 1402 can also be used to extract the target state features from the second sub-image region; according to the target state features and the preset state database, determine the state of the target object; correspondingly, the first processing module 1402 can combine the object category of the target object and the state of the target object as the recognition result.
[0171] In some embodiments, before using the image processing model to process the target image, the image data processing device 1400 can also be used to collect a sample image set; wherein, the sample image set includes a plurality of sample images containing sample objects arranged according to the collection time; use annotation frames to annotate the image regions of the sample objects in the sample images and annotate the object categories of the sample objects to obtain the annotated sample image set; use the annotated sample image set to train the initial model to obtain the image processing model.
[0172] It should be noted that the units, devices, or modules etc. described in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. For the convenience of description, when describing the above devices, various modules are described separately according to their functions. Of course, when implementing this specification, the functions of each module can be realized in the same or multiple software and / or hardware, or the modules realizing the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.
[0173] As can be seen from the above, after obtaining the target image, the image data processing device provided by the embodiments of this specification can identify the target object in the target image to obtain the corresponding recognition result, and can also obtain the depth information of the target object at the same time; then, by using the recognition result and the depth information of the target object at the same time, a target label including at least the object category information of the target object is generated and displayed on the target image. Thus, it can automatically identify and use the target label to mark the relevant information of the target object concerned by the user in the target image, and can also accurately and intuitively convey the three-dimensional depth information of the target object to the user through the displayed target label, so that the user can obtain a better interaction experience.
[0174] Although this specification provides method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The step order listed in the embodiments is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual device or client product is executed, it can be executed in the method order shown in the embodiments or the drawings or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing, or even in a distributed data processing environment). The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, product or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, product or device. Without further limitation, it does not exclude the existence of additional identical or equivalent elements in the process, method, product or device including the said elements. The terms "first", "second", etc. are used to denote names and do not denote any particular order.
[0175] Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, the method steps can be logically programmed to enable the controller to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be regarded as a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0176] This specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc. that perform specific tasks or implement specific abstract data types. This specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0177] From the description of the above embodiments, those skilled in the art can clearly understand that this specification can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of this specification can essentially be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this specification.
[0178] The various embodiments in this specification are described in a progressive manner. For the same or similar parts between the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. This specification can be used in many general or specific computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.
[0179] Although this specification is depicted through embodiments, those of ordinary skill in the art know that this specification has many variations and changes without departing from the spirit of this specification. It is hoped that the appended claims will cover these variations and changes without departing from the spirit of this specification.
Claims
1. An image data processing method, characterized in that, Including: Obtain a target image; The target image includes: a first image, and / or, a second image; wherein, the first image and the second image are images acquired by a camera; the target image is an image in a surgical scene; Identify a target object in the target image based on the first image or the second image to obtain a corresponding recognition result; and obtain depth information of the target object according to the first image and the second image; Display a target label of the target object in the target image according to the recognition result and depth information of the target object; including: determining a label size parameter matching the target object by using the depth information of the target object according to a preset matching rule; adjusting the character size in the label of the target object according to the label size parameter to obtain a target label of the target object for conveying the distance of the target object relative to the lens; displaying the target label of the target object in the target image; wherein, the target label at least includes object category information of the target object; Wherein, determining a label size parameter matching the target object by using the depth information of the target object according to a preset matching rule includes: projecting a pixel point of a key point of the target object in the target image onto a reference plane to obtain a corresponding reference pixel point; wherein, the reference plane is a plane where a pixel point with the maximum depth information in the target image is located; obtaining the distance between the reference pixel point and the center of the reference plane, and the maximum value of the depth information in the target image; calculating a label size parameter matching the target object by using the depth information of the target object, the distance between the reference pixel point and the center of the reference plane, and the maximum value of the depth information in the target image according to a preset matching relationship parameter.
2. The image data processing method according to claim 1, wherein The first image is an image acquired by a left camera, the second image is an image acquired by a right camera, and the relative positions of the left camera and the right camera are fixed.
3. The image data processing method according to claim 2, wherein The target image includes a first image, and / or, a second image acquired by a binocular endoscope; Correspondingly, the target object includes: a tissue organ, and / or, a surgical instrument.
4. The image data processing method according to claim 2, wherein Obtaining depth information of the target object includes: Determine a first distance between a first key pixel point and the optical center of the left camera, and a second distance between a second key pixel point and the optical center of the right camera according to the first image and the second image; wherein, the first key pixel point is a pixel point corresponding to a key point of the target object in the first image, and the second key pixel point is a pixel point corresponding to a key point of the target object in the second image; Calculate a vertical distance between the key point of the target object and the plane where the left camera and the right camera are located according to the first distance and the second distance as the depth information of the target object.
5. The image data processing method according to claim 4, characterized in that Calculating a vertical distance between the key point of the target object and the plane where the left camera and the right camera are located according to the first distance and the second distance includes: Obtain a baseline distance between the left camera and the right camera, and a camera focal length; Calculate the vertical distance between the key point of the target object and the plane where the left camera and the right camera are located by using the first distance, the second distance, the baseline distance, and the camera focal length.
6. The image data processing method according to claim 1, wherein After obtaining the target label of the target object, the method further includes: Adjust the display effect parameters of the target label of the target object according to the depth information of the target object to obtain the target label with adjusted display effect; wherein, the display effect parameters include at least one of the following: shadow parameter, deformation parameter, brightness parameter; Correspondingly, display the target label with adjusted display effect in the target image.
7. The image data processing method according to claim 1, characterized in that Displaying the target label of the target object in the target image includes: Obtain the position parameter of the target object in the target image; Determine the display position coordinates according to the position parameter of the target object in the target image; Set and display the target label of the target object at the corresponding position in the target image according to the display position coordinates.
8. The image data processing method according to claim 7, wherein, The method further includes: Construct an effect image layer according to the difference value between the depth information of different target objects in the target image; wherein, the effect image layer includes the target object and the stereoscopic effect data of the target label of the target object; Perform superposition processing on the target image and the effect image layer to obtain the superimposed target image; Display the superimposed target image.
9. The image data processing method according to claim 2, wherein Identifying the target object in the target image to obtain the corresponding identification result includes: Process the target image by using an image processing model to determine the object category of the target object in the target image as the identification result.
10. The image data processing method according to claim 9, wherein, The image processing model includes at least one of the following: an image processing model trained based on the YOLO model, an image processing model trained based on the Fast YOLO model, and an image processing model trained based on the R-CNN model.
11. The image data processing method according to claim 10, characterized in that, Processing the target image by using an image processing model to determine the object category of the target object in the target image includes: Perform rasterization processing on the target image by using an image processing model to obtain the rasterized target image; Use a convolutional neural network to detect and identify the object category of the target object in the target image by processing the rasterized target image; wherein, the convolutional neural network is the model structure in the image processing model.
12. The image data processing method according to claim 11, wherein Using a convolutional neural network to detect and identify the object category of the target object in the target image by processing the rasterized target image includes: Use a convolutional neural network to use a segmentation box to segment a first sub-image region from the rasterized target image; wherein, the first sub-image region is a non-background image region; Use a convolutional neural network to detect whether there is a target object in the first sub-image region; and according to the detection result, use a detection box to segment a second sub-image region from the first sub-image region; wherein, the second sub-image region is an image region containing the target object to be identified; Use a convolutional neural network to perform image recognition on the second sub-image region to determine the object category of the target object.
13. The image data processing method according to claim 12, wherein After using the detection frame to segment the second sub-image region from the first sub-image region, the method further includes: Using the non-maximum suppression algorithm to filter out the second sub-image regions segmented based on redundant detection frames from the second sub-image region, obtaining a corrected second sub-image region; Correspondingly, Using a convolutional neural network to perform image recognition on the corrected second sub-image region to determine the object category of the target object.
14. The image data processing method according to claim 12, wherein While using the image processing model to process the target image to determine the object category of the target object in the target image, the method further includes: Extracting target state features from the second sub-image region; Determining the state of the target object according to the target state features and a preset state database; Correspondingly, Combining the object category of the target object and the state of the target object as the recognition result.
15. The image data processing method according to claim 9, wherein Before using the image processing model to process the target image, the method further includes: Collecting a sample image set; wherein, the sample image set includes a plurality of sample images containing sample objects arranged according to the collection time; Using a marking frame to mark the image region of the sample object in the sample image and marking the object category of the sample object to obtain a marked sample image set; Using the marked sample image set to train an initial model to obtain an image processing model.
16. The image data processing method according to claim 7, wherein The method further includes: Constructing a target image with the target label displayed based on virtual reality technology according to the target image and the target label of the target object.
17. An image platform, characterized in that, Including a binocular endoscope and an image data processing device, the binocular endoscope is used to collect a target image; the image data processing device is used to process the target image by using the image data processing method according to any one of claims 1 to 16.
18. A computer device, characterized in that, Including a processor and a memory for storing processor-executable instructions, when the processor executes the instructions, the relevant steps of the image data processing method according to any one of claims 1 to 16 are implemented.
19. A computer-readable storage medium, characterized in that, Stored thereon are computer instructions, and when the instructions are executed, the relevant steps of the image data processing method according to any one of claims 1 to 16 are implemented.
Citation Information
Patent Citations
Image processing method and device, equipment and storage medium
CN112115913A
Information processing apparatus, information processing method, and recording medium
US20190012799A1