Video analytics methods and related devices, equipment, systems and storage media
By using video analysis methods to detect the location and process points of interest in video frames scanned by the endoscope, and outputting prompt messages, the problem of low efficiency in endoscopy examinations is solved, and more efficient exploration and analysis are achieved.
Patent Information
- Application Number
- CN202210107160.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-28
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-01-28
AI Technical Summary
Endoscopic examinations inside objects are inefficient, increasing the risk of patient discomfort or leakage of hazardous substances. Improving examination efficiency has become an urgent problem to be solved.
By acquiring video frame images scanned by the endoscope, video analysis methods are used to perform analysis and processing such as location detection, point of interest detection and classification, and cleanliness detection. Prompt messages are output to guide the exploration process, including the detection and classification of unscanned areas and points of interest, as well as cleanliness prompts.
It reduces the need for repeated endoscope probes inside the target object, improves probe efficiency, enhances work efficiency and analysis accuracy, and reduces the complexity of system setup.
Smart Images

Figure CN114445380B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a video analysis method and related apparatus, devices, systems and storage media. Background Technology
[0002] Endoscopes play an important role in many fields such as medicine and industry. For example, in the medical field, endoscopes can be used to examine parts of the digestive tract (stomach, intestines), while in the industrial field, endoscopes can be used for inspections in oil and gas chemical industries.
[0003] However, in many scenarios, endoscopes are not suitable for prolonged placement inside the target object, thus posing a significant challenge in terms of examination efficiency. For example, in gastrointestinal examinations, the insertion of the endoscope into the patient's body causes discomfort, and the longer the endoscope remains inside, the more intense the discomfort generally becomes. Therefore, it is crucial to detect lesions quickly. Similarly, in pipeline inspections, since harmful or toxic substances may exist inside the pipeline, and the longer the endoscope remains inside, the higher the likelihood of leakage, it is essential to detect injuries quickly. Therefore, improving the efficiency of target object examination has become a pressing issue. Summary of the Invention
[0004] This application provides a video analysis method and related apparatus, equipment, system and storage medium.
[0005] The first aspect of this application provides a video analysis method, comprising: acquiring video frame images scanned by an endoscope inside a target object; performing analysis and processing on the video frame images to obtain analysis results; wherein the analysis and processing includes position detection, and the analysis results include the position of the endoscope lens inside the target object; and outputting a prompt message based on the analysis results; wherein the prompt message includes a first message, the first message being used to indicate unscanned areas inside the target object.
[0006] Therefore, by acquiring video frame images scanned inside the target object by the endoscope and analyzing and processing them, an analysis result is obtained. This analysis and processing includes position detection, and the analysis result includes the position of the endoscope lens inside the target object. Based on this, a prompt message is output, including a first message, which is used to indicate the unscanned area inside the target object. Since the position of the endoscope lens can be continuously detected during the endoscope's exploration of the target object, the unscanned area inside the target object can be identified, thereby guiding the exploration of the target object. This greatly reduces the possibility of repeatedly exploring the same position, which is beneficial to improving the exploration efficiency of the target object.
[0007] The first message is presented in a preset manner, and the preset includes presentation through the internal structure of the target object, with unscanned areas marked in the internal structure.
[0008] Therefore, the first message is presented in a preset manner, and the preset manner is presented through the internal structure of the target object, which also marks the unscanned areas, thus improving the intuitiveness of displaying the unscanned areas.
[0009] The analysis and processing includes point of interest detection, the analysis results include the detected regions of points of interest in the video frame image, and the prompt message includes a second message, which is used to prompt the detected regions.
[0010] Therefore, the analysis and processing further includes point of interest detection, and the analysis results correspond to the detection area of points of interest in the video frame image. The prompt message corresponds to a second message, which is used to prompt the detection area. Thus, the endoscope can be used to realize the automatic detection of points of interest, which is beneficial to greatly improve work efficiency.
[0011] The analysis and processing includes point of interest (POI) classification, the analysis results include the predicted category of POI in the video frame image, and the prompt message includes a third message, which is used to prompt the predicted category.
[0012] Therefore, the analysis and processing further includes point of interest classification. The analysis results correspond to the predicted categories of points of interest in the video frame image, and the prompt message corresponds to a third message, which is used to prompt the predicted category. Thus, the endoscope can be used to achieve automatic classification of points of interest, which greatly improves work efficiency.
[0013] The analysis and processing includes point of interest retrieval, and the analysis results include several reference images related to the video frame image. The similarity between the reference images and the video frame image meets the preset conditions. The reference images are marked with the predicted category of the point of interest in the reference images. The prompt message includes a fourth message, which is presented in several reference images.
[0014] Therefore, the analysis and processing further includes point of interest retrieval. The analysis results correspond to several reference images related to the video frame image. The similarity between the reference images and the video frame image meets the preset conditions. The reference images are also marked with the predicted category of the point of interest in the reference images. The prompt message includes a fourth message, which is presented as several reference images. Therefore, it is possible to retrieve reference images related to the current video frame image during the endoscopic examination process for reference points of interest in the video frame image, which greatly improves work efficiency.
[0015] The analysis and processing includes cleanliness detection, and the analysis results include the cleanliness of the target object within the field of view of the video frame image. The prompt message includes a fifth message, which is used to indicate the cleanliness.
[0016] Therefore, the analysis and processing further includes cleanliness detection. The analysis results correspond to the cleanliness of the target object's interior within the video frame image's field of view. The prompt message corresponds to a fifth message, which is used to indicate the cleanliness. Since the cleanliness of the target object's interior has a certain impact on the accuracy of image analysis, indicating the cleanliness within the current field of view helps the user understand the credibility of the current analysis results.
[0017] The analysis and processing includes speed detection, the analysis results include the current speed of the lens inside the target object, and the prompt message includes a sixth message, which is used to prompt whether to maintain, increase or decrease the current speed.
[0018] Therefore, the analysis and processing further includes speed detection. The analysis result includes the current speed of the lens inside the target object. The prompt message includes a sixth message, which is used to prompt whether to maintain, increase, or decrease the current speed. This ensures that the lens movement speed is maintained within a reasonable range, preventing the lens movement speed from being too fast and causing image blurring, which would affect the accuracy of the analysis and processing. It also prevents the lens movement speed from being too slow and causing the endoscope to remain inside the target object for too long, which would affect the efficiency of the analysis and processing.
[0019] The analysis and processing includes point of interest segmentation, the analysis results include the contour regions of points of interest in the video frame image, and the prompt message includes a seventh message, which is used to prompt the contour regions.
[0020] Therefore, the analysis and processing further includes point of interest segmentation, and the analysis results correspond to the contour region of the point of interest in the video frame image. The prompt message corresponds to the seventh message, which is used to prompt the contour region. Thus, the endoscope can be used to achieve automatic segmentation of the point of interest, which greatly improves work efficiency.
[0021] The video analysis method is executed by a video analysis device. The input end of the video analysis device is connected to a video acquisition device, which is connected to an endoscope to obtain video frame images by acquiring video signals from the endoscope. The output end of the video analysis device is connected to a display device to output prompt messages through the display device.
[0022] Therefore, the video analysis method is executed by a video analysis device. The input end of the video analysis device is connected to a video acquisition device, which is connected to an endoscope to obtain video frame images by acquiring the video signal from the endoscope. The output end of the video analysis device is connected to a display device to output prompt messages. In other words, video analysis can be achieved simply by setting up a video analysis device and a video acquisition device between the display device and the endoscope, which helps to reduce the complexity of building a video analysis system.
[0023] A second aspect of this application provides a video analysis device, comprising: an image acquisition module, an analysis and processing module, and a prompt output module. The image acquisition module is used to acquire video frame images scanned by an endoscope inside a target object. The analysis and processing module is used to perform analysis and processing based on the video frame images to obtain analysis results. The analysis and processing includes position detection, and the analysis results include the position of the endoscope lens inside the target object. The prompt output module is used to output a prompt message based on the analysis results. The prompt message includes a first message, which is used to prompt unscanned areas inside the target object.
[0024] A third aspect of this application provides a video analysis device, including a memory and a processor coupled to each other, wherein the processor is used to execute program instructions stored in the memory to implement the video analysis method of the first aspect described above.
[0025] The fourth aspect of this application provides a video analysis system, including a video acquisition device and the video analysis device of the third aspect. The video acquisition device and the video analysis device are connected. The video acquisition device is used to connect to an endoscope to obtain video frame images by acquiring video signals from the endoscope. The video analysis device is used to connect to a display device to output prompt messages through the display device.
[0026] The fifth aspect of this application provides a computer-readable storage medium having program instructions stored thereon, which, when executed by a processor, implement the video analysis method of the first aspect described above.
[0027] The above scheme acquires video frame images scanned inside the target object by the endoscope, and analyzes and processes these images to obtain analysis results. The analysis and processing includes position detection, and the analysis results include the position of the endoscope lens inside the target object. Based on the analysis results, a prompt message is output, which includes a first message used to indicate unscanned areas inside the target object. Because the position of the endoscope lens can be continuously detected during the endoscope's exploration of the target object, the unscanned areas inside the target object can be identified, thereby guiding the exploration of the target object. This greatly reduces the possibility of repeatedly exploring the same position, which is beneficial to improving the exploration efficiency of the target object. Attached Figure Description
[0028] Figure 1 This is a flowchart illustrating an embodiment of the video analysis method of this application;
[0029] Figure 2 This is a schematic diagram of the framework of an embodiment of the video analysis system of this application;
[0030] Figure 3 This is a schematic diagram of an embodiment of an unscanned region;
[0031] Figure 4 This is a schematic diagram of an embodiment of the second message;
[0032] Figure 5 This is a schematic diagram of an embodiment of the fourth message;
[0033] Figure 6 This is a schematic diagram of the framework of an embodiment of the video analysis device of this application;
[0034] Figure 7 This is a schematic diagram of the framework of an embodiment of the video analysis device of this application;
[0035] Figure 8 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. Detailed Implementation
[0036] The following describes the embodiments of the present application in detail with reference to the accompanying drawings.
[0037] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.
[0038] The terms "system" and "network" are often used interchangeably in this document. The term "and / or" is simply a description of an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " generally indicates that the related objects are in an "or" relationship. Furthermore, "multiple" in this document means two or more than two.
[0039] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the video analysis method of this application.
[0040] Specifically, this may include the following steps:
[0041] Step S11: Acquire video frame images scanned inside the target object by the endoscope.
[0042] In an implementation scenario, the target object can be set according to the actual application scenario. For example, the target object can include, but is not limited to: the digestive tract (e.g., stomach, intestines), the respiratory tract, etc., to enable the exploration of the digestive tract, respiratory tract, etc. Or, the target object can also include, but is not limited to: oil and gas pipelines, engines, etc., to enable the exploration of oil and gas pipelines, engines, etc., etc., without limitation.
[0043] In one implementation scenario, it should be noted that the exploration of the inside of a target object using an endoscope can be a continuous process. That is, the endoscope can scan multiple video frame images inside the target object to continuously analyze and process each video frame image. The specific process of analysis and processing can be found in the relevant description below, and will not be elaborated here.
[0044] In one implementation scenario, please refer to the following: Figure 2 , Figure 2 This is a schematic diagram of the framework of an embodiment of the video analysis system of this application. Figure 2 As shown, the video analysis system may specifically include a video acquisition device and a video analysis device. The video analysis device is used to execute the steps in the video analysis method embodiment of this application. The video acquisition device is used to connect to an endoscope to obtain video frame images by acquiring the video signals from the endoscope. The video analysis device is used to connect to a display device to output a prompt message through the display device. The prompt message is obtained based on the analysis results of the video frame images. For details, please refer to the relevant description below, which will not be repeated here.
[0045] In a specific implementation scenario, please refer to the following: Figure 2 The input end of the video analysis device can be connected to the video acquisition device, and the output end of the video analysis device can be connected to the display device.
[0046] In a specific implementation scenario, the video signal from the endoscope can adopt a first protocol, while the video acquisition device can convert the video signal of the first protocol into a second protocol, and extract video frame images based on the video signal of the second protocol. For example, an SDI (Serial Digital Interface) video signal can be acquired through an endoscope, and the video acquisition device can convert the SDI video signal into an HDMI (High Definition Multimedia Interface) video signal, based on which video frame images can be extracted.
[0047] In a specific implementation scenario, the video analysis device can be miniaturized without affecting its analysis efficiency by optimizing the following analysis and processing algorithms. For example, the dimensions of the video analysis device can be 22 cm long, 22 cm wide, and 3 cm high, without limitation, thus making it easier to set up the aforementioned video analysis system in endoscopy rooms with limited space.
[0048] In a specific implementation scenario, the video analysis device can also support multiple video stream inputs / outputs. This allows the device to be coupled to multiple endoscopes and connected to multiple display devices. The resulting prompts, derived from processing video frame images scanned by multiple endoscopes, can be output to the respective display devices. For example, if the video analysis device supports four video streams, it can be coupled to four endoscopes and connected to four display devices. Display device 1 can display prompts derived from the video frame images scanned by endoscope 1, display device 2 can display prompts derived from the video frame images scanned by endoscope 2, and so on. Further examples are not provided here.
[0049] Step S12: Analyze and process the video frame images to obtain the analysis results.
[0050] In this embodiment of the disclosure, the analysis process includes position detection, and the analysis result includes the position of the endoscope lens inside the target object.
[0051] In one implementation scenario, location detection can be achieved based on the monocular SLAM (Simultaneous Localization and Mapping) algorithm. For details on the detection process, please refer to the technical details of monocular SLAM, which will not be elaborated here.
[0052] In one implementation scenario, location detection can also be achieved based on network models such as convolutional neural networks. This network model can specifically include, but is not limited to, convolutional layers, activation layers, pooling layers, etc., and the specific structure of the network model is not limited here. For example, several sample images can be pre-collected, and these sample images are labeled with their sample positions inside the target object. Based on this, the network model can be used to perform location detection on the sample images to obtain the predicted position, and the network parameters of the network model can be adjusted based on the difference between the sample position and the predicted position. It should be noted that the calculation process of the above difference can be found in the technical details of loss functions such as cross-entropy loss, and the adjustment process of the above parameters can be found in the technical details of optimization methods such as gradient descent, which will not be elaborated here. Through the above training process, the network model can learn the image features at various locations inside the target object. Therefore, in the subsequent actual detection process, by inputting video frame images into the network model, the position of the endoscope lens inside the target object when the video frame image was captured can be output.
[0053] Step S13: Based on the analysis results, output a prompt message.
[0054] In this embodiment of the disclosure, the prompt message includes a first message, which is used to prompt the unscanned area inside the target object.
[0055] In one implementation scenario, as mentioned earlier, the exploration of the inside of a target object using an endoscope can be a continuous process. After position detection of each video frame, the position of the endoscope lens inside the target object when the video frame was captured can be obtained. Based on this, the positions corresponding to each video frame can be recorded, and the unrecorded positions inside the target object constitute the unscanned areas.
[0056] In one implementation scenario, the first message can be presented in a preset manner, including presentation based on the internal structure of the target object, where unscanned areas may be marked. For an example in a medical scenario, please refer to [reference needed]. Figure 3 , Figure 3 This is a schematic diagram of an embodiment of an unscanned area. It should be noted that... Figure 3 The diagram shows an unscanned area when the target object is the digestive tract. Figure 3As shown, the diagonally shaded area represents the unscanned region. In practical applications, other methods can also be used to mark unscanned areas, such as using outlines with preset colors (e.g., red, green), preset line types (e.g., solid lines, dashed lines), and preset line thicknesses (e.g., 20 pt, 30 pt), or using fill blocks with preset transparency (e.g., 20%, 50%) and preset fill styles (e.g., solid color fill, dotted fill). These methods are not limited here. In the above approach, the first message is presented in a preset manner, which includes presentation based on the internal structure of the target object. The internal structure also marks the unscanned areas, thus improving the intuitiveness of displaying unscanned areas.
[0057] The above scheme acquires video frame images scanned inside the target object by the endoscope, and analyzes and processes these images to obtain analysis results. The analysis and processing includes position detection, and the analysis results include the position of the endoscope lens inside the target object. Based on the analysis results, a prompt message is output, which includes a first message used to indicate unscanned areas inside the target object. Because the position of the endoscope lens can be continuously detected during the endoscope's exploration of the target object, the unscanned areas inside the target object can be identified, thereby guiding the exploration of the target object. This greatly reduces the possibility of repeatedly exploring the same position, which is beneficial to improving the exploration efficiency of the target object.
[0058] In some disclosed embodiments, please refer to the relevant documents. Figure 2 The processing and analysis can further include point-of-interest (POI) detection. The analysis result can include the detected POI region in the video frame image, and the corresponding prompt message can include a second message, which is used to indicate the detected region. This method enables automatic POI detection using an endoscope, significantly improving work efficiency. It should be noted that the POI can vary depending on the application scenario. For example, in a medical setting, POIs may include, but are not limited to, lesions; or, for example, in an industrial setting, POIs may include, but are not limited to, scars. Other scenarios can be deduced similarly, and will not be listed here.
[0059] In one implementation scenario, network models such as convolutional neural networks can be used to detect points of interest (POIs) in video frames to obtain the regions of interest (ROIs) within those frames. Specific network models can include, but are not limited to, Faster-RCNN, YOLO, etc., and their specific structures are not limited here. Specifically, several sample images can be pre-collected, each labeled with a POI region. Based on this, the network model can be used to detect POIs in the sample images, obtaining predicted POI regions. The network parameters can then be adjusted based on the difference between the sample and predicted regions. It should be noted that the calculation process for this difference can be found in the technical details of loss functions such as cross-entropy loss, and the parameter adjustment process can be found in the technical details of optimization methods such as gradient descent, which will not be elaborated upon here. Through this training process, the network model learns image features related to POI regions. Therefore, in subsequent actual detection processes, video frames can be input into the network model to predict the POI regions within those frames.
[0060] In one implementation scenario, the detection region can be represented by a rectangular bounding box that surrounds the point of interest. Specifically, the analysis result can include the position coordinates of the vertices of this bounding box within the video frame image. For example, in a medical scenario, the detection region could be a rectangular bounding box surrounding a lesion; or, in an industrial scenario, it could be a rectangular bounding box surrounding a wound. Other scenarios can be deduced similarly, and will not be listed here.
[0061] In one implementation scenario, the second message can be presented as a video frame image, and the video frame image is marked with the detected regions identified by point-of-interest detection. For an example in a medical scenario, please refer to [reference needed]. Figure 4 , Figure 4 This is a schematic diagram of an embodiment of the second message. For example... Figure 4 The image shown is a video frame, and the solid-lined rectangle represents the detection area of the lesion.
[0062] In some disclosed embodiments, please refer to the relevant documents. Figure 2 The processing and analysis can further include point-of-interest (POI) classification. The analysis results can include the predicted category of POIs in the video frame image, and the corresponding prompt message can include a third message used to indicate the predicted category. This method enables automatic classification of POIs using an endoscope, significantly improving work efficiency.
[0063] In one implementation scenario, network models such as convolutional neural networks can be used to classify points of interest (POIs) in video frames to obtain the predicted categories of POIs. The network model can specifically include, but is not limited to, convolutional layers, activation layers, pooling layers, fully connected layers, etc., and the specific structure of the network model is not limited here. Specifically, several sample images can be pre-collected, and the sample images are labeled with the categories of POIs. Based on this, the network model can be used to classify the POIs in the sample images to obtain the predicted categories of POIs. The network parameters of the model can then be adjusted based on the difference between the sample categories and the predicted categories. It should be noted that the calculation process of the above differences can be found in the technical details of loss functions such as cross-entropy loss, and the parameter adjustment process can be found in the technical details of optimization methods such as gradient descent, which will not be elaborated here. Through the above training process, the network model can learn image features related to POI categories. Therefore, in subsequent actual detection processes, video frames can be input into the network model to predict the predicted categories of POIs in the video frames.
[0064] In one implementation scenario, the third message can be presented as a video frame image, with the predicted category of the point of interest (POI) marked at the detection region in the video frame image. This allows for a visual representation of the detected POI and its predicted category within the video frame image. For example, in a medical scenario, the predicted category of a lesion (e.g., hematoma, cyst) can be marked at the detection region in the video frame image, thus visually representing the detected lesion and its predicted category. Similarly, in an industrial scenario, the predicted category of a scar (e.g., penetrating hole, scratch, corrosion) can be marked at the detection region in the video frame image, visually representing the detected scar and its predicted category. Other scenarios can be deduced similarly, and will not be listed here.
[0065] In some disclosed embodiments, please refer to the relevant documents. Figure 2 The processing and analysis can further include point-of-interest (POI) retrieval. The analysis results can include several reference images related to the video frame image. The similarity between the reference images and the video frame image meets preset conditions. The reference images are labeled with the predicted category of the POIs within them. The corresponding prompt message can include a fourth message, which can be presented as several reference images. This method allows for the retrieval of reference images related to the current video frame image during endoscopic examinations, enabling doctors to refer to the POIs within the video frame image, thus significantly improving work efficiency.
[0066] In one implementation scenario, network models such as convolutional neural networks can be used to extract features from video frame images and historical scanned images, obtaining first image features of the video frame images and second image features of the historical scanned images. It should be noted that the historical scanned images can be images scanned before the video frame images, and these images can be labeled with the predicted categories of the detected points of interest. Based on this, the similarity between the first image features and each of the second image features can be obtained using similarity metrics such as cosine similarity. Specifically, several sample images can be pre-collected, and reference information representing the correlation between any two sample images can be obtained. Based on this, the network model can be used to extract features from each sample image to obtain sample image features. These sample images can then be used as the current image. Based on the reference information, sample images related to the current image are designated as positive examples, and sample images unrelated to the current image are designated as negative examples. Furthermore, the differences between the sample image features of the current image and those of the positive and negative examples can be measured using a triplet loss function, yielding the sub-loss corresponding to the current image. Based on the sub-losses corresponding to each sample image, the total loss can be obtained, and the network parameters of the network model can be adjusted based on the total loss. It should be noted that the calculation process of the above differences can be found in the technical details of loss functions such as triplet loss, and the parameter adjustment process can be found in the technical details of optimization methods such as gradient descent, which will not be elaborated upon here. Through the above training process, the network model can make the image features of related images as close as possible, and the image features of unrelated images as far apart as possible. Thus, in the subsequent actual detection process, the image features of the video frame can be predicted by inputting the video frame image into the network model.
[0067] In one implementation scenario, the historical scanned images can be sorted in descending order of similarity. The preset condition can be set to be located before a preset position. In other words, historical scanned images located before a preset position (e.g., the first 5, the first 6, etc.) can be used as reference images.
[0068] In one implementation scenario, taking a medical setting as an example, please refer to the relevant documents. Figure 5 , Figure 5 This is a schematic diagram of an embodiment of the fourth message. For Figure 5 The video frame image shown on the left can retrieve 5 reference images related to its content. The upper left corner of the reference image is marked with the predicted category of the lesion, such as T1N0, T1N0+, T1Nx, etc. For the specific meaning, please refer to the relevant details of TMN staging, which will not be repeated here.
[0069] In some disclosed embodiments, please refer to the relevant documents. Figure 2 The processing and analysis can further include cleanliness detection. The analysis results can include the cleanliness of the target object's interior within the video frame's field of view, and the corresponding prompt message can include a fifth message, which is used to indicate the cleanliness level. Since the cleanliness of the target object's interior has a certain impact on the accuracy of image analysis, indicating the cleanliness level within the current field of view helps users understand the reliability of the current analysis results.
[0070] In one implementation scenario, a network model such as a convolutional neural network can be used to detect the cleanliness of video frame images, obtaining the cleanliness of the target object within the field of view of the video frame image. This network model can include, but is not limited to, convolutional layers, activation layers, pooling layers, etc., and the specific structure of the network model is not limited here. Specifically, several sample images can be pre-collected, and the sample images are labeled with sample cleanliness. For example, sample cleanliness can be represented numerically (e.g., from 0 to 10), or it can be represented textually (e.g., "not clean," "clean," "very clean"). Based on this, the network model can be used to detect the cleanliness of the sample images, obtaining a predicted cleanliness, and the network parameters of the network model can be adjusted using the difference between the sample cleanliness and the predicted cleanliness. It should be noted that the calculation process of the above difference can be found in the technical details of loss functions such as cross-entropy loss, and the adjustment process of the above parameters can be found in the technical details of optimization methods such as gradient descent, which will not be elaborated here. Through the above training process, the network model can learn image features related to cleanliness. In subsequent actual detection processes, video frame images can be input into the network model to predict the cleanliness of the target object within the field of view of the video frame image.
[0071] In one implementation scenario, it should be noted that the higher the cleanliness of the target object within the field of view of the video frame image, the higher the reliability of the aforementioned location detection, point of interest detection, point of interest classification, and point of interest retrieval. Conversely, the lower the cleanliness of the target object within the field of view of the video frame image, the lower the reliability of the aforementioned location detection, point of interest detection, point of interest classification, and point of interest retrieval.
[0072] In one implementation scenario, the fifth message can also be presented as a video frame image, and the video frame image can be labeled with the cleanliness level. Based on this, the credibility of the analysis results, such as the detection area and prediction category marked on the video frame image, can be indicated to provide doctors with richer auxiliary information.
[0073] In some disclosed embodiments, please refer to the relevant documents. Figure 2The processing and analysis can further include speed detection. The analysis results can include the current speed of the lens inside the target object, and the corresponding prompt message can include a sixth message, which can be used to prompt maintaining, increasing, or decreasing the current speed. This method ensures that the lens's movement speed remains within a reasonable range, preventing it from moving too fast and causing image blurring, thus affecting the accuracy of the analysis and processing. It also prevents the lens from moving too slowly, causing the endoscope to remain inside the target object for too long, thus affecting the efficiency of the analysis and processing.
[0074] In one implementation scenario, as described in the aforementioned disclosed embodiments, position detection can be performed based on monocular SLAM, neural networks, etc., to obtain the current position of the endoscope lens inside the target object. Furthermore, as described in the aforementioned disclosed embodiments, since the exploration of the target object using the endoscope can be a continuous process, the shooting position corresponding to the video frame image scanned before the current video frame image can be obtained. Therefore, based on the distance difference between the current position and the shooting position, and the time difference between the scanning time of the current video frame image and the scanning time of the previously scanned video frame images, the current velocity of the lens inside the target object can be obtained. For example, to simplify the calculation, the ratio of the distance field to the time difference can be directly used as the current velocity.
[0075] In one implementation scenario, the similarity between the currently scanned video frame and previously scanned video frames can be measured to obtain a similarity score. This similarity score can then be substituted into the mapping relationship between image similarity and movement speed to determine the current speed. It should be noted that this mapping relationship can be linear, with higher image similarity resulting in slower movement speed and lower image similarity indicative of faster movement speed. Furthermore, this mapping relationship can be pre-established. For example, the similarity between adjacent images can be calculated beforehand by moving within the target object at different speeds, and the aforementioned mapping relationship can be established based on this.
[0076] In one implementation scenario, the system can detect whether the current speed is within a preset range. If the current speed is within the preset range, the sixth message can specifically indicate a prompt to maintain the current speed. If the current speed is below the lower limit of the preset range, the sixth message can specifically indicate a prompt to increase the current speed. If the current speed is above the upper limit of the preset range, the sixth message can specifically indicate a prompt to decrease the current speed. It should be noted that the preset range can be set according to the actual application. For example, when the endoscope frame rate is high, the preset range can be set appropriately larger, while when the endoscope frame rate is low, the preset range can be set appropriately smaller. The specific numerical range is not limited here.
[0077] In some disclosed embodiments, the analysis process includes point-of-interest (POI) segmentation, the analysis result includes the contour region of the POI in the video frame image, and the prompt message includes a seventh message used to indicate the contour region. This method enables automatic segmentation of POIs using an endoscope, significantly improving work efficiency.
[0078] In one implementation scenario, network models such as convolutional neural networks can be used to segment points of interest (POIs) in video frame images to obtain the contour regions of POIs within the video frame images. Specific network models can include, but are not limited to, U-Net, etc., and the specific structure of the network model is not limited here. Specifically, several sample images can be pre-collected, and sample contours of POIs can be labeled in these sample images. Based on this, the network model can be used to segment the sample images to obtain predicted contours of the POIs. The network parameters of the network model can then be adjusted based on the difference between the sample contours and the predicted contours. It should be noted that the calculation process of the above differences can be found in the technical details of loss functions such as cross-entropy loss, and the parameter adjustment process can be found in the technical details of optimization methods such as gradient descent, which will not be elaborated here. Through the above training process, the network model can learn image features related to the POI contours. Therefore, in subsequent actual detection processes, video frame images can be input into the network model to predict the contour regions of the POIs in the video frame images.
[0079] In one implementation scenario, the contour region can be represented by a contour line along the point of interest. Specifically, the analysis result can include the position coordinates of each contour point on that contour line within the video frame image. For example, in a medical scenario, the contour region could be the contour of a lesion; or in an industrial scenario, it could be the contour of a scar. Other scenarios can be deduced similarly, and will not be listed here.
[0080] In one implementation scenario, the seventh message can be presented as a video frame image, and the video frame image is marked with the contour region detected by point of interest detection. Specifically, a preset style can be used to mark the contour region. For example, the preset style may include, but is not limited to: line color, line type, line thickness, etc., which are not limited here.
[0081] In some disclosed embodiments, by means of, Figure 2The video analysis system shown can perform detection algorithms such as location detection, point of interest (POI) detection, POI classification, POI retrieval, cleanliness detection, speed detection, and POI segmentation to provide users with as much auxiliary information as possible. Once a POI (e.g., lesion, injury) is detected at a certain location, the system can prompt the user to pull the endoscope lens back to that location for further detailed examination to confirm whether it is indeed a POI (e.g., lesion, injury). Alternatively, depending on the actual situation, the system can prompt the user to use a staining agent or increase the lens magnification, etc., without limitation.
[0082] Please see Figure 6 , Figure 6 This is a schematic diagram of a framework of an embodiment of the video analysis device 60 of this application. The video analysis device 60 includes: an image acquisition module 61, an analysis and processing module 62, and a prompt output module 63. The image acquisition module 61 is used to acquire video frame images scanned by an endoscope inside a target object; the analysis and processing module 62 is used to perform analysis and processing based on the video frame images to obtain analysis results; wherein, the analysis and processing includes position detection, and the analysis results include the position of the endoscope lens inside the target object; the prompt output module 63 is used to output a prompt message based on the analysis results; wherein, the prompt message includes a first message, which is used to prompt the unscanned area inside the target object.
[0083] The above solution, by continuously detecting the position of the endoscope lens during the endoscopic exploration of the target object, can identify unscanned areas inside the target object, thereby guiding the exploration of the target object and greatly reducing the possibility of repeatedly exploring the same location, which helps to improve the exploration efficiency of the target object.
[0084] In some disclosed embodiments, the first message is presented in a preset manner, and the preset manner includes a manner of presentation based on the internal structure of the target object, wherein the internal structure is marked with unscanned areas.
[0085] Therefore, the first message is presented in a preset manner, which includes presentation through the internal structure of the target object, which also marks unscanned areas, thus improving the intuitiveness of displaying unscanned areas.
[0086] In some disclosed embodiments, the analysis process includes point of interest detection, the analysis result includes the detected region of point of interest in the video frame image, and the prompt message includes a second message used to prompt the detected region.
[0087] Therefore, the analysis and processing further includes point of interest detection, and the analysis results correspond to the detection area of points of interest in the video frame image. The prompt message corresponds to a second message, which is used to prompt the detection area. Thus, the endoscope can be used to realize the automatic detection of points of interest, which is beneficial to greatly improve work efficiency.
[0088] In some disclosed embodiments, the analysis process includes point of interest classification, the analysis result includes the predicted category of points of interest in the video frame image, and the prompt message includes a third message used to prompt the predicted category.
[0089] Therefore, the analysis and processing further includes point of interest classification. The analysis results correspond to the predicted categories of points of interest in the video frame image, and the prompt message corresponds to a third message, which is used to prompt the predicted category. Thus, the endoscope can be used to achieve automatic classification of points of interest, which greatly improves work efficiency.
[0090] In some disclosed embodiments, the analysis process includes point of interest retrieval, the analysis result includes several reference images related to the video frame image, the similarity between the reference images and the video frame image meets a preset condition, the reference images are marked with the predicted category of the point of interest in the reference images, and the prompt message includes a fourth message, which is presented in the form of several reference images.
[0091] Therefore, the analysis and processing further includes point of interest retrieval. The analysis results include several reference images related to the video frame image. The similarity between the reference images and the video frame image meets the preset conditions. The reference images are also marked with the predicted category of the points of interest in the reference images. The prompt message includes a fourth message, which is presented as several reference images. Therefore, during the endoscopic examination, reference images related to the current video frame image can be retrieved for doctors to refer to the points of interest in the video frame image, which greatly improves work efficiency.
[0092] In some disclosed embodiments, the analysis process includes cleanliness detection, the analysis result includes the cleanliness of the target object inside the field of view of the video frame image, and the prompt message includes a fifth message, which is used to prompt the cleanliness.
[0093] Therefore, the analysis and processing further includes cleanliness detection. The analysis results correspond to the cleanliness of the target object's interior within the video frame image's field of view. The prompt message corresponds to a fifth message, which is used to indicate the cleanliness. Since the cleanliness of the target object's interior has a certain impact on the accuracy of image analysis, indicating the cleanliness within the current field of view helps the user understand the credibility of the current analysis results.
[0094] In some disclosed embodiments, the analysis process includes speed detection, the analysis result includes the current speed of the lens inside the target object, and the prompt message includes a sixth message, which prompts the user to maintain, increase, or decrease the current speed.
[0095] Therefore, the analysis and processing further includes speed detection. The analysis result includes the current speed of the lens inside the target object. The prompt message includes a sixth message, which is used to prompt whether to maintain, increase, or decrease the current speed. This ensures that the lens movement speed is maintained within a reasonable range, preventing the lens movement speed from being too fast and causing image blurring, which would affect the accuracy of the analysis and processing. It also prevents the lens movement speed from being too slow and causing the endoscope to remain inside the target object for too long, which would affect the efficiency of the analysis and processing.
[0096] In some disclosed embodiments, the analysis process includes point-of-interest (POI) segmentation, the analysis result includes the contour region of the POI in the video frame image, and the prompt message includes a seventh message, which is used to prompt the contour region.
[0097] Therefore, the analysis and processing further includes point of interest segmentation, and the analysis results correspond to the contour region of the point of interest in the video frame image. The prompt message corresponds to the seventh message, which is used to prompt the contour region. Thus, the endoscope can be used to achieve automatic segmentation of the point of interest, which greatly improves work efficiency.
[0098] In some disclosed embodiments, the video analysis method is performed by a video analysis device, the input of which is connected to a video acquisition device, which is connected to an endoscope to obtain video frame images by acquiring video signals from the endoscope, and the output of which is connected to a display device to output prompt messages.
[0099] Therefore, the video analysis method is executed by a video analysis device. The input end of the video analysis device is connected to a video acquisition device, which is connected to an endoscope to obtain video frame images by acquiring the video signal from the endoscope. The output end of the video analysis device is connected to a display device to output prompt messages. In other words, video analysis can be achieved simply by setting up a video analysis device and a video acquisition device between the display device and the endoscope, which helps to reduce the complexity of building a video analysis system.
[0100] Please see Figure 7 , Figure 7This is a schematic diagram of a framework of an embodiment of the video analysis device 70 of this application. The video analysis device 70 includes a memory 71 and a processor 72 coupled to each other. The processor 72 is used to execute program instructions stored in the memory 71 to implement the steps of any of the above-described video analysis method embodiments. In a specific implementation scenario, the video analysis device 70 may include, but is not limited to, a microcomputer or a server. In addition, the video analysis device 70 may also include mobile devices such as laptops and tablets, which are not limited here.
[0101] Specifically, processor 72 controls itself and memory 71 to implement the steps of any of the video analysis method embodiments described above. Processor 72 can also be referred to as a CPU (Central Processing Unit). Processor 72 may be an integrated circuit chip with signal processing capabilities. Processor 72 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 72 can be implemented using integrated circuit chips.
[0102] The above solution, by continuously detecting the position of the endoscope lens during the endoscopic exploration of the target object, can identify unscanned areas inside the target object, thereby guiding the exploration of the target object and greatly reducing the possibility of repeatedly exploring the same location, which helps to improve the exploration efficiency of the target object.
[0103] Please see Figure 8 , Figure 8 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium 80 of this application. The computer-readable storage medium 80 stores program instructions 801 that can be executed by a processor. The program instructions 801 are used to implement the steps of any of the above-described video analysis method embodiments.
[0104] The above solution, by continuously detecting the position of the endoscope lens during the endoscopic exploration of the target object, can identify unscanned areas inside the target object, thereby guiding the exploration of the target object and greatly reducing the possibility of repeatedly exploring the same location, which helps to improve the exploration efficiency of the target object.
[0105] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0106] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0107] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0108] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each embodiment method of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0109] If the technical solution of this application involves personal information, the product that applies the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing personal information. If the technical solution of this application involves sensitive personal information, the product that applies the technical solution of this application has obtained the individual's separate consent before processing sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, a clear and prominent sign is set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that they agree to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload their personal information; among which, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
Claims
1. A video analysis method, characterized in that, include: Acquire video frame images scanned inside the target object by the endoscope; The video frame image is analyzed and processed to obtain analysis results. The analysis and processing includes position detection, and the analysis results include the position of the endoscope lens inside the target object. The analysis and processing also includes point of interest (POI) detection, POI classification, POI retrieval, cleanliness detection, and POI segmentation. The analysis results further include: the detection region of the POI in the video frame image determined by the POI detection; the predicted category of the POI in the video frame image determined by the POI classification; several reference images related to the video frame image determined by the POI retrieval; the cleanliness of the target object within the field of view of the video frame image determined by the cleanliness detection; and the contour region of the POI in the video frame image determined by the POI segmentation. Based on the analysis results, a prompt message is output; wherein, the prompt message includes a first message, which is used to prompt the unscanned area inside the target object; the prompt message also prompts the detection area, the predicted category, the plurality of reference images, and the confidence level of the detection area, the predicted category, and the plurality of reference images, wherein the confidence level is positively correlated with the cleanliness; the analysis processing includes speed detection; the analysis result includes the current speed of the lens inside the target object; the prompt message includes a sixth message, which is used to prompt maintaining, increasing, or decreasing the current speed; and the measurement step of the current speed includes: The object is moved inside the target object at different speeds in advance, and the similarity between adjacent images at different moving speeds is calculated. Based on the similarity between adjacent images at different moving speeds, a mapping relationship between image similarity and moving speed is established. The current speed is obtained by substituting the similarity scores obtained from the similarity detection of the currently scanned video frame image and the previously scanned video frame images into the mapping relationship.
2. The method according to claim 1, characterized in that, The first message is presented in a preset manner, and the preset manner includes a presentation method based on the internal structure of the target object, wherein the internal structure is marked with the unscanned area.
3. The method according to claim 1, characterized in that, The notification message includes a second message, which is used to notify the detection area.
4. The method according to claim 1, characterized in that, The prompt message includes a third message, which is used to prompt the prediction category.
5. The method according to claim 1, characterized in that, The similarity between the reference image and the video frame image meets a preset condition. The reference image is marked with the predicted category of the point of interest in the reference image. The prompt message includes a fourth message, which is presented using the plurality of reference images.
6. The method according to claim 1, characterized in that, The notification message includes a fifth message, which is used to indicate the cleanliness level.
7. The method according to claim 1, characterized in that, The prompt message includes a seventh message, which is used to prompt the outline area.
8. The method according to claim 1, characterized in that, The video analysis method is executed by a video analysis device. The input end of the video analysis device is connected to a video acquisition device, which is connected to the endoscope to obtain the video frame image by acquiring the video signal from the endoscope. The output end of the video analysis device is connected to a display device to output the prompt message through the display device.
9. A video analysis device, characterized in that, include: The image acquisition module is used to acquire video frame images scanned inside the target object by the endoscope. An analysis and processing module is used to perform analysis and processing based on the video frame image to obtain analysis results; wherein, the analysis and processing includes position detection, and the analysis results include the position of the endoscope lens inside the target object; the analysis and processing also includes point of interest detection, point of interest classification, point of interest retrieval, cleanliness detection, and point of interest segmentation; the analysis results also include: the detection region of the point of interest in the video frame image determined by performing point of interest detection, the predicted category of the point of interest in the video frame image determined by performing point of interest classification, several reference images related to the video frame image determined by performing point of interest retrieval, the cleanliness of the target object inside the field of view of the video frame image determined by performing cleanliness detection, and the contour region of the point of interest in the video frame image determined by performing point of interest segmentation; A prompt output module is used to output a prompt message based on the analysis results; wherein the prompt message includes a first message, which is used to prompt the unscanned area inside the target object; the prompt message also prompts the detection area, the predicted category, the plurality of reference images, and the confidence level of the detection area, the predicted category, and the plurality of reference images, wherein the confidence level is positively correlated with the cleanliness; the analysis processing includes speed detection; the analysis result includes the current speed of the lens inside the target object; the prompt message includes a sixth message, which is used to prompt maintaining, increasing, or decreasing the current speed; and the measurement step of the current speed includes: The object is moved inside the target object at different speeds in advance, and the similarity between adjacent images at different moving speeds is calculated. Based on the similarity between adjacent images at different moving speeds, a mapping relationship between image similarity and moving speed is established. The current speed is obtained by substituting the similarity scores obtained from the similarity detection of the currently scanned video frame image and the previously scanned video frame images into the mapping relationship.
10. A video analysis device, characterized in that, The device includes a memory and a processor coupled to each other, the processor being configured to execute program instructions stored in the memory to implement the video analysis method according to any one of claims 1 to 8.
11. A video analysis system, characterized in that, The device includes a video acquisition device and a video analysis device as described in claim 10, wherein the video acquisition device and the video analysis device are connected, the video acquisition device is used to connect to the endoscope to obtain video frame images by acquiring the video signals of the endoscope, and the video analysis device is used to connect to a display device to output prompt messages through the display device.
12. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, they implement the video analysis method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Scanning endoscopic device and control method therefor
CN105744878A
Image processing method based on endoscope system, device, equipment and storage medium
CN110265122A
Capsule endoscopy system
CN112075914A
Medical auxiliary operation method, device and equipment and computer storage medium
CN113143168A
Endoscope image detection method, device, storage medium and electronic equipment
CN113487608A