Lesion distinguishing method and device, storage medium and computer equipment

By obtaining cervical cancer screening information and colposcopy videos, using neural network models to extract and fusion features, and automatically perform lesion identification, solving the problem of insufficient accuracy of doctors' manual identification and achieving higher accuracy of lesion identification.

CN120259707APending Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410009353.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, precancerous lesions screening for cervical cancer mainly relies on manual identification by doctors, which leads to a greater impact on the accuracy of discrimination by doctors’ experience and may lead to misdiagnosis and overtreatment.

Method used

By obtaining the target disease screening information of the target object and the inspection video of the target body surface position, multiple inspection image sequences are extracted, feature extraction and fusion is used for the neural network model, and lesion identification is combined with the classification network to automatically output the lesion identification results.

Benefits of technology

It improves the accuracy of lesion identification, reduces misdiagnosis caused by doctor experience differences or diagnostic errors, and improves the automated accuracy of lesion identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259707A_ABST
    Figure CN120259707A_ABST
Patent Text Reader

Abstract

The invention provides a lesion distinguishing method and device, a storage medium and computer equipment. The method comprises the following steps: acquiring target disease screening information of a target object, wherein the target disease screening information is a screening result obtained by performing cytological screening on the target object for a target disease; acquiring an inspection video acquired from the target body surface position of the target object in a preset time period, wherein the preset time period is a chemical reaction time period after a preset chemical substance is applied to the target body surface position; extracting a plurality of inspection images from the inspection video according to a time sequence to obtain an inspection image sequence with a time sequence relationship; performing feature extraction on the inspection image sequence and the target disease screening information, and performing feature fusion on the extracted image features and the screening result features to obtain fusion features; and inputting the fusion features into a classification network for lesion discrimination result classification to obtain a lesion discrimination result output by the classification network. The method can improve the accuracy of lesion discrimination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of biomedical technologies, and particularly to a method, apparatus, storage medium, and computer device for lesion recognition based on lesion surface images. Background Art

[0002] Cervical cancer is a relatively common gynecological malignant tumor, which is commonly found in female groups of middle and high ages, and in recent years, its incidence has shown a trend of getting younger. However, with the development of medical technologies, early detection and treatment of cervical cancer and precancerous lesions can be achieved, and the incidence and mortality rates of cervical cancer have significantly decreased.

[0003] Currently, for the screening of precancerous lesions of cervical cancer, it mainly relies on doctors for manual discrimination, which leads to a great dependence on doctors' professional levels for the discrimination of precancerous lesions of cervical cancer. When doctors' levels or experience are insufficient, the discrimination accuracy of precancerous lesions of cervical cancer will be relatively low. Summary of the Invention

[0004] Embodiments of the present disclosure provide a lesion discrimination method, apparatus, storage medium, and computer device, and this method can improve the accuracy of lesion discrimination.

[0005] According to one aspect of the present disclosure, a lesion discrimination method is provided, and the method includes:

[0006] Obtain target disease screening information of a target object, where the target disease screening information is a screening result obtained by performing cytological screening on the target object for the target disease;

[0007] Obtain an inspection video collected at a target body surface position of the target object within a preset time period, where the preset time period is a chemical reaction time period after applying a preset chemical substance to the target body surface position;

[0008] Extract multiple inspection images from the inspection video in chronological order to obtain an inspection image sequence with a temporal relationship;

[0009] Extract features from the inspection image sequence and the target disease screening information respectively, and perform feature fusion on the extracted image features and screening result features to obtain fusion features;

[0010] Input the fusion features into a classification network for lesion discrimination result classification to obtain a lesion discrimination result output by the classification network.

[0011] According to one aspect of the present disclosure, a lesion discrimination apparatus is provided, and the apparatus includes:

[0012] A first acquisition unit, configured to acquire target disease screening information of a target object, where the target disease screening information is a screening result obtained by performing a cytological screening on the target object for the target disease;

[0013] A second acquisition unit, configured to acquire an inspection video collected at a target body surface position of the target object within a preset time period, where the preset time period is a chemical reaction time period after applying a preset chemical substance to the target body surface position;

[0014] An extraction unit, configured to extract multiple inspection images from the inspection video in chronological order to obtain an inspection image sequence with a chronological relationship;

[0015] A fusion unit, configured to respectively perform feature extraction on the inspection image sequence and the target disease screening information, and perform feature fusion on the extracted image features and screening result features to obtain fusion features;

[0016] A discrimination unit, configured to input the fusion features into a classification network for classifying a lesion discrimination result, and obtain the lesion discrimination result output by the classification network.

[0017] Optionally, in some embodiments, the fusion unit includes:

[0018] A first extraction subunit, configured to perform feature extraction on each inspection image in the inspection image sequence based on an image feature extraction network to obtain multiple image features;

[0019] A second extraction subunit, configured to perform feature extraction on the target disease screening information based on a screening feature extraction network to obtain screening features;

[0020] A first fusion subunit, configured to perform feature fusion on the multiple image features and the screening features to obtain fusion features;

[0021] Wherein, the image feature extraction network, the screening feature extraction network and the classification network are jointly trained using the same batch of sample data.

[0022] Optionally, in some embodiments, the lesion discrimination device provided by the present disclosure further includes a model training unit, specifically including:

[0023] A first acquisition subunit, configured to acquire training sample data, where the training sample data includes a sample inspection image sequence of a sample object, a sample screening result, and a corresponding classification label, and the sample inspection image sequence is an image sequence obtained by performing image extraction on an inspection video collected at a sample body surface position of the sample object within a sample time period;

[0024] A third extraction subunit, configured to perform feature extraction on each sample inspection image in the sample inspection image sequence based on the image feature extraction network to obtain a plurality of sample image features;

[0025] A fourth extraction subunit, configured to perform feature extraction on the sample screening result based on the screening feature extraction network to obtain a sample screening feature;

[0026] A second fusion subunit, configured to perform feature fusion on the plurality of sample image features and the sample screening feature to obtain a sample fusion feature;

[0027] A classification subunit, configured to perform classification processing on the sample fusion feature based on the classification network to obtain a classification output;

[0028] An adjustment subunit, configured to calculate a gradient value according to the classification output and the classification label, and perform gradient backpropagation processing based on the gradient value to adjust the parameters of the image feature extraction network, the screening feature extraction network, and the classification network.

[0029] Optionally, in some embodiments, the second fusion subunit includes:

[0030] A first calculation module, configured to calculate the mean value of the plurality of sample image features to obtain an image average feature;

[0031] A stacking module, configured to calculate the product of the image average feature and the sample screening feature, and stack the obtained product feature and the plurality of sample image features to obtain a sample fusion feature.

[0032] Optionally, in some embodiments, the sample screening result includes a first sample screening sub-result and a second sample screening sub-result, and the fourth extraction subunit includes:

[0033] A first acquisition module, configured to acquire a first training parameter corresponding to the first sample screening sub-result and a second training parameter corresponding to the second sample screening sub-result, where the values of the first training parameter and the second training parameter are 0 or 1;

[0034] A determination module, configured to determine the input data of the screening feature extraction network during the training process according to the first sample screening sub-result and the corresponding first training parameter, and the second sample screening sub-result and the corresponding second training parameter;

[0035] A first extraction module, configured to perform feature extraction on the input data based on the screening feature extraction network to obtain a sample screening feature.

[0036] Optionally, in some embodiments, the lesion discrimination device provided by the present disclosure further includes:

[0037] A second acquisition subunit, configured to acquire the screening date corresponding to the target disease screening information;

[0038] The second extraction subunit is further configured to:

[0039] When the time interval between the screening date and the acquisition date of the inspection video does not exceed a preset number of days, perform feature extraction on the target disease screening information based on a screening feature extraction network.

[0040] Optionally, in some embodiments, the extraction unit includes:

[0041] A splitting subunit, configured to split the inspection video into frames based on the chronological order to obtain multiple image frames with chronological order;

[0042] A fifth extraction subunit, configured to extract multiple first target image frames from the multiple image frames at a preset frame interval;

[0043] A determination subunit, configured to determine an inspection image sequence with a timing relationship according to the multiple first target image frames and the corresponding time information.

[0044] Optionally, in some embodiments, the lesion discrimination device provided by the present disclosure further includes:

[0045] A display subunit, configured to display the multiple image frames with chronological order on a display interface of a display terminal;

[0046] A third acquisition subunit, configured to acquire a selection instruction for the displayed image frames in the display interface and determine multiple second target image frames that are selected;

[0047] The determination subunit is further configured to:

[0048] Determine an inspection image sequence with a timing relationship according to the multiple first target image frames, the multiple second target image frames, and the corresponding time information.

[0049] Optionally, in some embodiments, the fifth extraction subunit includes:

[0050] A second acquisition module, configured to acquire a preset time interval corresponding to image frame acquisition and the frame rate corresponding to the inspection video;

[0051] A calculation module, configured to calculate the frame interval corresponding to image frame acquisition based on the preset time interval and the frame rate;

[0052] A second extraction module, configured to extract multiple first target image frames from the multiple image frames according to the frame interval.

[0053] Optionally, in some embodiments, the lesion discrimination device provided by the present disclosure further includes:

[0054] A calculation subunit, configured to calculate the differences between adjacent image features according to the temporal relationship between the inspection images, and obtain a plurality of image difference features;

[0055] The first fusion subunit is further configured to:

[0056] Perform feature fusion on the plurality of image difference features and the screening features to obtain fusion features.

[0057] Optionally, in some embodiments, the first extraction subunit is further configured to:

[0058] Based on an image feature extraction network, perform optical flow feature extraction on each inspection image in the inspection sequence to obtain a plurality of image optical flow features;

[0059] The calculation subunit is further configured to:

[0060] Calculate the differences between adjacent image optical flow features according to the temporal relationship between the inspection images, and obtain a plurality of image optical flow difference features;

[0061] The first fusion subunit is further configured to:

[0062] Perform feature fusion on the plurality of image optical flow difference features and the screening features to obtain fusion features.

[0063] According to one aspect of the present disclosure, there is provided a storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned lesion discrimination method is implemented.

[0064] According to one aspect of the present disclosure, there is provided a computer program product including a computer program, and when the computer program is read and executed by a processor of a computer device, the computer device executes the above-mentioned lesion discrimination method.

[0065] The lesion discrimination method provided by the embodiments of the present disclosure includes obtaining the target disease screening information of the target object, where the target disease screening information is the screening result obtained by performing cytological screening on the target object for the target disease; obtaining the inspection video collected at the target body surface position of the target object within a preset time period, where the preset time period is the chemical reaction time period after applying a preset chemical substance to the target body surface position; extracting multiple inspection images from the inspection video in chronological order to obtain an inspection image sequence with a temporal relationship; respectively extracting features from the inspection image sequence and the target disease screening information, and performing feature fusion on the extracted image features and screening result features to obtain fusion features; inputting the fusion features into a classification network for lesion discrimination result classification to obtain the lesion discrimination result output by the classification network.

[0066] In the embodiments of the present disclosure, through the target disease screening information of the target object and multiple sequential images extracted from the inspection video collected at the target body surface position, the neural network model is used to respectively extract and fuse features of the target disease screening information and the multiple sequential images to obtain fusion features, and then the classification network is used to classify the fusion features, so that an accurate lesion discrimination result can be automatically obtained. Compared with the manual lesion diagnosis by doctors, this solution can perform joint automatic diagnosis through multi-modal inspection information, avoiding the problem of incorrect lesion discrimination results caused by differences in experience or diagnostic errors of manual work, and can greatly improve the accuracy of lesion discrimination.

[0067] Other features and advantages of the present disclosure will be described in the following specification, and some of them will become obvious from the specification or be understood by implementing the present disclosure. The objectives and other advantages of the present disclosure can be achieved and obtained through the structures specifically pointed out in the specification, claims, and drawings. Description of the Drawings

[0068] The drawings are used to provide a further understanding of the technical solutions of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure, and do not constitute a limitation to the technical solutions of the present disclosure.

[0069] Figure 1 It is a system architecture diagram applied to the lesion discrimination method of the embodiments of the present disclosure;

[0070] Figure 2 It is a flowchart of the lesion discrimination method provided by the present disclosure;

[0071] Figure 3 It is a schematic diagram of extracting an inspection image sequence from an inspection video in the present disclosure

[0072] Figure 4 It is another schematic diagram of extracting an inspection image sequence from an inspection video in the present disclosure;

[0073] Figure 5 Another schematic diagram for extracting an inspection image sequence from an inspection video in the present disclosure;

[0074] Figure 6 Schematic structural diagram of a lesion discrimination model provided by the present disclosure;

[0075] Figure 7 Another schematic flowchart of a lesion discrimination method provided by the present disclosure;

[0076] Figure 8 Schematic structural diagram of a lesion discrimination device provided by an embodiment of the present disclosure;

[0077] Figure 9 Structural diagram of a terminal for implementing the methods according to an embodiment of the present disclosure;

[0078] Figure 10 Structural diagram of a server for implementing the methods according to an embodiment of the present disclosure. Detailed implementation manners

[0079] In order to make the objectives, technical solutions and advantages of the present disclosure clearer and more understandable, the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.

[0080] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are described. The nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations:

[0081] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, pre-trained models, also known as large models and basic models, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0082] Classification network: A mathematical model obtained after machine learning technology learns the labeled sample data (image-category correspondence). During the learning and training process, the parameters of this mathematical model are obtained. When identifying and predicting, the parameters of this mathematical model are loaded and the probability that the input sample belongs to each category is calculated.

[0083] In related technologies, during the process of screening for precancerous lesions of cervical cancer, it mainly relies on doctors to manually distinguish the taken colposcopy images based on experience to obtain a diagnosis result. However, since the colposcopy images do not very clearly show the lesion situation, doctors need to carefully examine and combine rich medical experience to obtain an accurate diagnosis result. Moreover, if a doctor makes a positive diagnosis, generally the patient needs to further undergo cervical biopsy to determine whether cervical cancer is confirmed. The cervical biopsy process is relatively painful. If the doctor's diagnosis is incorrect, it will cause the patient to bear the pain brought by unnecessary over-treatment. Therefore, it is particularly important to improve the accuracy of lesion discrimination based on colposcopy images. In this regard, the present disclosure provides a lesion discrimination method in order to improve the accuracy of lesion discrimination based on body surface images to a certain extent.

[0084] System architecture and scenario description applied in the embodiments of the present disclosure

[0085] Figure 1 It is a system architecture diagram applied to the lesion discrimination method according to an embodiment of the present disclosure. It includes a terminal 140, the Internet 130, a gateway 120, a server 110, etc.

[0086] The terminal 140 includes various forms such as a desktop computer, a laptop computer, a PDA (Personal Digital Assistant), a mobile phone, a vehicle-mounted terminal, a home theater terminal, a dedicated terminal, an intelligent voice interaction device, an intelligent household appliance, an aircraft, or a control end of a colposcope image acquisition system. Additionally, it can be a single device or a collection composed of multiple devices. The terminal 140 can communicate with the Internet 130 in a wired or wireless manner to exchange data.

[0087] The server 110 refers to a computer system that can provide certain services to the terminal 140. Compared with an ordinary terminal 140, the server 110 has higher requirements in terms of stability, security, performance, etc. The server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part (such as a virtual machine) allocated from a high-performance computer, a combination of parts (such as virtual machines) allocated from multiple high-performance computers, etc.

[0088] The gateway 120 is also called an internetwork connector and a protocol converter. The gateway realizes network interconnection at the transport layer and is a computer system or device that acts as a conversion function. Between two systems using different communication protocols, data formats, or languages, and even with completely different architectures, the gateway is a translator. At the same time, the gateway can also provide filtering and security functions. The message sent by the terminal 140 to the server 110 needs to be sent to the corresponding server 110 through the gateway 120. The message sent by the server 110 to the terminal 140 also needs to be sent to the corresponding terminal 140 through the gateway 120.

[0089] The lesion discrimination method provided by the embodiments of the present disclosure can be implemented independently in the foregoing terminal 140, can be implemented independently in the foregoing server 110, or can be partially implemented in the terminal 140 and partially implemented in the server 110.

[0090] When the lesion discrimination method provided by the embodiments of the present disclosure is implemented in the terminal 140, the neural network model (including the feature extraction model and the classification model) can be directly deployed in the terminal 140. Specifically, the terminal 140 can obtain the target disease screening information of the target object, and the target disease screening information is the screening result obtained by performing cytological screening on the target object for the target disease; the terminal 140 obtains the inspection video collected at the target body surface position of the target object within a preset time period, and the preset time period is the chemical reaction time period after applying a preset chemical substance to the target body surface position; the terminal 140 extracts multiple inspection images from the inspection video in chronological order to obtain an inspection image sequence with a chronological relationship; the terminal 140 respectively extracts features from the inspection image sequence and the target disease screening information, and performs feature fusion on the extracted image features and the screening result features to obtain fusion features; the terminal 140 inputs the fusion features into a classification network for classifying the lesion discrimination result to obtain the lesion discrimination result output by the classification network.

[0091] When the lesion discrimination method provided by the embodiments of the present disclosure is implemented in the server 110, the neural network model (including the feature extraction model and the classification model) can be deployed in the server 110. Specifically, the server 110 can obtain the target disease screening information of the target object, and the target disease screening information is the screening result obtained by performing cytological screening on the target object for the target disease; the server 110 obtains the inspection video collected at the target body surface position of the target object within a preset time period, and the preset time period is the chemical reaction time period after applying a preset chemical substance to the target body surface position; the server 110 extracts multiple inspection images from the inspection video in chronological order to obtain an inspection image sequence with a chronological relationship; the server 110 respectively extracts features from the inspection image sequence and the target disease screening information, and performs feature fusion on the extracted image features and the screening result features to obtain fusion features; the server 110 inputs the fusion features into a classification network for classifying the lesion discrimination result to obtain the lesion discrimination result output by the classification network.

[0092] When a part of the lesion discrimination method provided by the embodiments of the present disclosure is implemented in the server 110 and another part is implemented in the terminal 140, the neural network model (including the feature extraction model and the classification model) can be deployed in the server 110. Specifically, the terminal 140 can obtain the target disease screening information of the target object, and the target disease screening information is the screening result obtained by performing cytological screening on the target object for the target disease; the terminal 140 obtains the inspection video collected at the target body surface position of the target object within a preset time period, and the preset time period is the chemical reaction time period after applying a preset chemical substance to the target body surface position; then, the terminal 140 sends the obtained target disease screening information and the inspection video to the server 110; the server 110 extracts multiple inspection images from the inspection video in chronological order to obtain an inspection image sequence with a temporal relationship; the server 110 respectively extracts features from the inspection image sequence and the target disease screening information, and performs feature fusion on the extracted image features and the screening result features to obtain fusion features; the server 110 inputs the fusion features into the classification network for lesion discrimination result classification to obtain the lesion discrimination result output by the classification network.

[0093] The lesion discrimination method provided by the embodiments of the present disclosure can be specifically applied to various disease diagnosis scenarios for lesion discrimination. For example, it can be used for lesion discrimination in the aforementioned cervical cancer pre-screening scenario, and can also be used for lesion discrimination in the infectious skin disease diagnosis scenario. Of course, the lesion discrimination method provided by the present disclosure can also be applied to other lesion discrimination scenarios, and the above examples do not limit the protection scope of the present case.

[0094] General description of the embodiments of the present disclosure

[0095] According to an embodiment of the present disclosure, a lesion discrimination method is provided. As Figure 2 shown, it is a flowchart of a lesion discrimination method provided by the present disclosure. This method can be applied to a lesion discrimination device, and this lesion discrimination device can be integrated in a computer device, and the computer device can specifically be a terminal or a server. This lesion discrimination method can include:

[0096] Step 210, obtaining the target disease screening information of the target object.

[0097] In the embodiments of the present disclosure, the provided lesion discrimination method can be used in various scenarios for lesion discrimination based on the temporal change images of the organism's epidermis during the test phase and related screening information. Among them, the organism can specifically be a human or other animals, and the animals can specifically be cats, dogs, cows, horses, etc. In this embodiment, in order to further improve the accuracy of lesion discrimination, multi-modal examination or screening information is used for lesion discrimination. Specifically, the target disease screening information of the target object and the examination video collected at the target body surface position of the target object can be used as the reference information for lesion discrimination.

[0098] First, the target disease screening information of the target object can be obtained. As introduced above, the target object here can be a human or a pet cat, pet dog, etc. When the target object is a human, the target disease can be the aforementioned cervical cancer or skin disease; when the target object is an animal, the target disease can be a skin disease. The target disease screening information can be the screening result obtained by cytological screening of the target object for the target disease. Specifically, blood samples can be collected from the target object, and then the collected blood samples can be tested to check whether the cells are infected with the germs related to the target disease.

[0099] In the embodiments of the present disclosure, the lesion discrimination method provided by the present disclosure will be introduced in detail with the target object being a human and the target disease being cervical cancer as an example. In this scenario, the target disease screening information can specifically include one or both of human papilloma virus (HPV) and thinprep cytologic test (TCT). Among them, human papilloma virus is a group of spherical, tiny, non-enveloped circular double-stranded DNA viruses belonging to the genus Papillomavirus of the family Papovaviridae. Thinprep cytologic test is a widely used cervical lesion screening technology at present. In this solution, during the process of cervical lesion screening, HPV and TCT screening information can be used as auxiliary information to improve the accuracy of cervical lesion screening. When these two pieces of information are lacking as assistance, this solution can also directly determine the lesion discrimination result based on the examination video collected at the target body surface position. The following solution will introduce this in detail.

[0100] Among them, obtaining the target disease screening information of the target object can be inputting the target disease screening information into the lesion discrimination terminal by the object, or the lesion discrimination terminal automatically searching for the corresponding target disease screening information from the medical information of the target object. Before the lesion discrimination terminal accesses the medical information of the target object, it is necessary to obtain the prior knowledge and authorization of the target object and obtain and use the relevant information of the target object on the basis of compliance with the corresponding laws and regulations.

[0101] After obtaining the target disease screening information of the target object, it is also necessary to further obtain the screening time point of the target disease screening information, and judge the validity of the target disease screening information accordingly. Generally, as time goes by, the functions and cytological characteristics of the human body will change and can only remain stable for a period of time. Therefore, when the screening time of the target disease screening information exceeds the preset time length from the current lesion discrimination time, it can be determined that the target disease screening information is invalid screening information. In this case, it is necessary to re-find the accurate target disease screening information or directly ignore the target disease screening information.

[0102] Step 220, obtain the inspection video collected at the target body surface position of the target object within a preset time period.

[0103] After obtaining the target disease screening information of the target object, the inspection video collected at the target body surface position of the target object can be further obtained, where the inspection video needs to be collected within a preset time period. In the embodiment of the present disclosure, the preset time period is the chemical reaction time period after a preset chemical substance is applied to the target body surface position of the target object. In some embodiments, after obtaining the target disease screening information of the target object, a preliminary judgment can be made first according to the screening information. If one of the two items in the screening information is abnormal, that is, there are abnormal cells in the liquid-based (blood sample or urine sample, etc.) examination, or there is a high-risk HPV infection, then the inspection video collection of the target body surface position of the target object can be further carried out.

[0104] For example, still taking the scenario of cervical lesion discrimination described above as an example. The target body surface position of the collected inspection video is generally the cervical position, and the inspection video of this position is generally examined by a colposcope. Colposcopy is a digital imaging technology applied in the field of gynecological examinations. This technology mainly observes the minute changes of the stromal blood vessels by magnifying the cervical epithelial structure, and then makes a judgment on whether there is a cervical lesion, as well as provides correct guidance for the biopsy site for the lesion type, nature, scope, and when it is determined that there is a lesion.

[0105] Before collecting the inspection video of the cervical position using a colposcope, it is necessary to first perform an acetic acid white test on the cervical position to be inspected. Among them, the acetic acid white test, also known as the acetic acid leukoplakia test, is a test method mainly used clinically for latent infection of human papillomavirus (HPV) or condyloma acuminata and subclinical manifestations of condyloma acuminata. Because latent infection of HPV or condyloma acuminata and subclinical manifestations of condyloma acuminata are not very typical clinically or cannot be seen with the naked eye, making them turn white after applying acetic acid may make the lesions clearly visible, aiming at the diagnosis and differential diagnosis of condyloma acuminata or latent HPV infection. The sensitivity of the acetic acid white test is very high and is very helpful for diagnosing HPV infection, especially subclinical infection. The experimental steps of the acetic acid white test are generally to first soak gauze or tissue paper with a 3% - 5% concentration of acetic acid solution, and then apply the soaked gauze or tissue paper on the skin surface of the target body surface position to observe whether the skin surface of the target body surface position turns white.

[0106] Since the acetic acid applied on the skin surface will gradually volatilize over time, making the effect of the acetic acid white test gradually become blurred, it is necessary to strictly control the acquisition time of the inspection video. Generally, the acquisition starts at a certain time after the acetic acid white test starts and ends at a certain time. The specific start time and end time can be determined according to clinical experience. During the acetic acid white test, the color of the skin surface at the target body surface position may change, and accurate lesion discrimination needs to be based on this change. Therefore, it is necessary to continuously collect the skin surface images during this period, that is, use video acquisition to collect data.

[0107] Step 230, extract multiple inspection images from the inspection video in chronological order to obtain an inspection image sequence with a temporal sequence relationship.

[0108] After collecting the inspection video of the target body surface position of the target object within the preset time period, multiple inspection images can be further extracted from the inspection video to obtain an inspection image sequence. That is, after collecting the shooting video of the cervix of the object to be diagnosed during the acetic acid white test, multiple inspection images can be extracted from the captured colposcope video to obtain an inspection image sequence. Among them, since the color change of the epidermis at the target body surface position over time during the acetic acid white test in this case is an important index feature to reveal whether there is a lesion, when extracting multiple inspection images from the inspection video, it is necessary to retain the time information corresponding to each frame of the image, and then construct an inspection image sequence with a temporal sequence relationship according to the time information of the images.

[0109] In some embodiments, extracting multiple inspection images from the inspection video in chronological order to obtain an inspection image sequence with a temporal sequence relationship includes:

[0110] The inspection video is frame-split based on the chronological order to obtain multiple image frames with chronological order;

[0111] Multiple first target image frames are extracted from the multiple image frames at preset frame intervals;

[0112] An inspection image sequence with a timing relationship is determined based on the multiple first target image frames and the corresponding time information.

[0113] Among them, in the embodiments of the present disclosure, an inspection image sequence composed of multiple inspection images can be extracted from the inspection video in a manner of uniform interval extraction according to image frames. Specifically, the inspection video can be first frame-split based on the chronological order to obtain multiple image frames with chronological order. Or, the inspection video can also be first frame-split, and each frame obtained by the split has time information, and then the multiple image frames obtained by the split are sorted according to the time information corresponding to each frame, so as to obtain multiple image frames with chronological order.

[0114] Then, multiple first target image frames can be further extracted from the multiple image frames at preset frame intervals. For example, when the preset frame interval is set to 2 frames, it can be determined that the 1st frame, the 4th frame, the 7th frame, etc. sorted in chronological order are the extracted first target image frames. Then, an inspection image sequence with a timing relationship can be determined based on these multiple first target image frames and their corresponding time information, that is, these multiple first target image frames are sorted according to the time information corresponding to each first target image frame, so as to obtain an inspection image sequence with a timing relationship.

[0115] Such as Figure 3 shown, is a schematic diagram of extracting an inspection image sequence from an inspection video in the present disclosure. As shown in the figure, the inspection video can be split into multiple image frames 310, and then these multiple image frames are sorted in chronological order. Further, multiple first target image frames 320 can be selected from the multiple image frames 310 in the arrangement order of the image frames at a frame interval of 2 frames, and among them, the image frames with small white dots marked at the lower left corner of the picture in the initial image frame sequence are the automatically selected first target image frames.

[0116] Since the collected inspection images contain a large number of image frames, and this solution uses artificial intelligence technology, specifically a deep learning model to process the image sequence and then identify lesions, a large number of image frames need to be processed, which will lead to a decrease in the efficiency of lesion identification. Moreover, the reaction degree of the acetic acid white test is not very intense, so there may be little difference between multiple consecutive image frames. Therefore, by using the above solution to uniformly select multiple first target image frames from multiple image frames of the inspection video, the number of image frames to be processed can be reduced on the basis of ensuring the lesion identification effect, thereby improving the efficiency of lesion identification.

[0117] In some embodiments, extracting multiple first target image frames from multiple image frames at a preset frame number interval includes:

[0118] Obtain the preset time interval corresponding to image frame acquisition and the frame rate corresponding to the inspection video;

[0119] Calculate the frame number interval corresponding to image frame acquisition based on the preset time interval and the frame rate;

[0120] Extract multiple first target image frames from multiple image frames according to the frame number interval.

[0121] In some alternative embodiments, automatic image selection can also be performed from the inspection video according to a preset time interval. In this solution, after determining the preset time interval corresponding to image frame acquisition, the frame rate of the inspection video can be further determined. Then, according to the set preset time interval and the frame rate of the inspection video, the frame number interval corresponding to automatic image acquisition can be calculated. After determining the frame number interval, the method introduced in the above embodiments can be used for automatic image acquisition to obtain multiple first target image frames.

[0122] In some embodiments, before determining multiple inspection images according to multiple first target image frames, it further includes:

[0123] Display multiple image frames with time sequence on the display interface of the display terminal;

[0124] Obtain the selection instruction for the displayed image frames on the display interface and determine multiple second target image frames that are selected;

[0125] Determining an inspection image sequence with a time sequence relationship according to multiple first target image frames and corresponding time information includes:

[0126] Determine an inspection image sequence with a time sequence relationship according to multiple first target image frames, multiple second target image frames and corresponding time information.

[0127] In an embodiment of the present disclosure, a method for generating an inspection image sequence by combining automatic image selection and manual image selection is provided to avoid missing important image frames due to an overly large frame interval setting during automatic image selection. In this embodiment, multiple image frames with a time sequence obtained by splitting frames of an inspection video can be displayed on the display interface of a display terminal. Here, the display terminal can be the control terminal of a colposcope, the personal terminal of a doctor, or a server with a display component.

[0128] After the multiple image frames are displayed in chronological order on the display interface of the display terminal, a selection instruction from a terminal operator can be received. The selection instruction can be input into the display terminal in various forms such as mouse clicking, touch screen clicking, voice selection, or gesture selection. When a selection instruction for the displayed image frames in the display interface is obtained, multiple second target image frames selected can be determined according to the selection instruction. As Figure 4 shown, it is another schematic diagram for extracting an inspection image sequence from an inspection video in the present disclosure. As shown in the figure, when manually selecting image frames, an image frame at any moment can be selected as the second target image frame 410. Among them, the second target image frames manually selected in the initial image sequence are marked with small black dots at the lower left corner of the picture.

[0129] After determining multiple second target image frames according to the selection instruction, these multiple second target image frames can be directly used as the inspection image sequence, or these multiple second target image frames can be used as a supplement to multiple first target image frames, and the inspection image sequence is jointly constituted by the multiple first target image frames and the multiple second target image frames. When the inspection image sequence is jointly constituted by the multiple first target image frames and the multiple second target image frames, the time information corresponding to each first target image frame and each second target image frame can be determined respectively, and then the multiple first target image frames and the multiple second target image frames are sorted according to the time information, and then the inspection image sequence is formed by combining the sorting results.

[0130] In the embodiment of the present disclosure, by collecting an inspection video of the target body surface position of a target object, then performing sequential image acquisition according to the collected inspection video to obtain multiple inspection image sequences with a sequential relationship, and further performing lesion discrimination according to the inspection image sequences. Compared with the related art that uses a single image for lesion discrimination, the sequential image sequence can more clearly reveal the process state of the test, and thus can more accurately reflect the lesion situation, which can improve the accuracy of lesion discrimination.

[0131] In some embodiments, after extracting multiple first target image frames and multiple second target image frames from an inspection video by means of automatic acquisition and manual acquisition, the image quality of the first target image frames and the second target image frames can be further detected to determine whether there are abnormal unqualified image frames with poor clarity, being blocked by foreign objects, etc. among the multiple first target image frames and the multiple second target image frames. If any, the unqualified image frames among the multiple first target image frames and the multiple second target image frames are deleted, thereby obtaining an inspection image sequence. Among them, the quality detection can be automatically performed by the terminal or manually selected by a human.

[0132] As Figure 5 shown, it is another schematic diagram of extracting an inspection image sequence from an inspection video in the present disclosure. As shown in the figure, after collecting multiple first target image frames and multiple second target image frames from multiple image frames obtained by splitting an inspection image by means of automatic acquisition or manual acquisition, further manual image selection can be performed on these collected first target image frames and second target image frames, so as to remove the image frames with poor quality, obtain the image frames with good quality, and then sort these image frames with good quality according to time information to obtain an inspection image sequence.

[0133] Step 240: Extract features from the inspection image sequence and the target disease screening information respectively, and perform feature fusion on the extracted image features and the screening result features to obtain fusion features.

[0134] Among them, the lesion discrimination method provided by the embodiments of the present disclosure is specifically a method for automatically discriminating lesions based on artificial intelligence technology. Compared with the method of a doctor manually discriminating lesions according to inspection images, this method can avoid misdiagnosis caused by insufficient doctor's professional experience or poor state, and thus can improve the accuracy of lesion discrimination.

[0135] Specifically, a lesion discrimination model based on deep learning can be used to discriminate the target disease of a target object based on the inspection image sequence and the target disease screening information. The lesion discrimination model can specifically be composed of four parts, namely an image feature extraction network, a screening feature extraction network, a feature fusion network, and a classification network. Among them, the image feature extraction network and the screening feature extraction network can be convolutional neural networks, the classification network can specifically be a fully connected network, and the image feature extraction network can specifically be a residual network, such as ResNet50.

[0136] As Figure 6As shown, it is a schematic structural diagram of a lesion discrimination model provided by the present disclosure. As shown in the figure, the lesion discrimination model includes an image feature extraction network 610 for extracting image features from a sequence of examination images; a screening feature extraction network 620 for extracting features from the screening information of the target disease; a feature fusion network 630 for fusing the extracted image features and screening features to obtain fused features; and a classification network 640 for performing lesion discrimination classification based on the fused features to obtain the probability value corresponding to each category, and further obtaining the lesion discrimination result.

[0137] In some embodiments, feature extraction is respectively performed on the sequence of examination images and the screening information of the target disease, and the extracted image features and screening result features are fused to obtain fused features, including:

[0138] Based on the image feature extraction network, feature extraction is performed on each examination image in the sequence of examination images to obtain a plurality of image features;

[0139] Based on the screening feature extraction network, feature extraction is performed on the screening information of the target disease to obtain screening features;

[0140] Feature fusion is performed on the plurality of image features and the screening features to obtain fused features;

[0141] Among them, the image feature extraction network, the screening feature extraction network, and the classification network are jointly trained using the same batch of sample data.

[0142] In some embodiments, the functions required by multiple networks in the lesion discrimination model can be trained independently or jointly. In the embodiments of the present disclosure, in order to avoid the accumulation of model accuracy differences after independent training of multiple networks resulting in poor model effects of the lesion discrimination model, the same training sample data can be used for joint training of multiple networks.

[0143] In some embodiments, the specific process of training the above-mentioned lesion discrimination model may include the following steps:

[0144] Obtain training sample data, where the training sample data includes the sample examination image sequence of the sample object, the sample screening result, and the corresponding classification label. The sample examination image sequence is an image sequence obtained by extracting images from the examination video collected at the sample body surface position of the sample object within the sample time period;

[0145] Based on the image feature extraction network, feature extraction is performed on each sample examination image in the sample examination image sequence to obtain a plurality of sample image features;

[0146] Extract features from the sample screening results based on the screening feature extraction network to obtain sample screening features;

[0147] Fuse the features of multiple sample images and the sample screening features to obtain sample fusion features;

[0148] Perform classification processing on the sample fusion features based on the classification network to obtain a classification output;

[0149] Calculate the gradient value according to the classification output and the classification label, and perform gradient backpropagation processing based on the gradient value to adjust the parameters of the image feature extraction network, the screening feature extraction network, and the classification network.

[0150] First, before training the lesion discrimination model, it is necessary to determine the specific model structure of the lesion discrimination model. Please continue to refer to Figure 6 , which is the schematic diagram of the main structure of the lesion discrimination model provided by the present disclosure. The image feature extraction network therein can specifically adopt a convolutional neural network (CNN), and specifically ResNet50 can be used as the backbone network. The screening feature extraction network can specifically be the mapping relationship between the target disease screening information and the feature value. For example, when HPV is negative, the corresponding feature value is 1. This mapping relationship can be determined in advance, and when performing feature extraction, the screening features corresponding to the target disease screening information can be directly mapped according to this mapping relationship. Therefore, when training the lesion discrimination model, the training of the screening feature extraction network can be not considered. As shown in Table 1 below, it is the screening feature extraction schematic table. As shown in the figure, according to several indicators of HPV and TCT in the target disease screening information, the corresponding feature values of each indicator are determined respectively, and then the screening features corresponding to the target disease screening information are obtained by combining these feature values [0,1,0,0,0,0,0,1,0].

[0151]

[0152] Table 1 Screening Feature Extraction Schematic Table

[0153] The feature fusion network can specifically set an algorithm to fuse the image features extracted by the image feature extraction layer and the screening result features. The algorithm can include fixed parameters and variable parameters. When there are variable parameters, they can participate in the training of the lesion discrimination model. The classification network can specifically be a full connect (FC) network.

[0154] After determining the model structure of the lesion discrimination model, the training sample data for training the lesion discrimination model can be obtained. The training sample data includes the sample inspection image sequence of the sample object, the sample screening result, and the corresponding classification label. Among them, in order to ensure the accuracy of the lesion discrimination model, in different usage scenarios, the lesion discrimination model should be trained using the training sample data corresponding to that scenario. For example, when the lesion discrimination model is used for the early screening of cervical cancer, the sample object in the training sample data should also be a person, and the sample inspection image sequence of the sample object should also be the image obtained by extracting image frames from the colposcopy video collected during the acetic acid white test on the cervical surface of the sample object. The sample screening result should be the TCT and HPV screening results of the sample object. The classification label can be that the sample object is diagnosed with pre-cancerous lesions of cervical cancer or has no pre-cancerous lesions.

[0155] After determining the model structure of the lesion discrimination model and obtaining the training sample data for training the model, the lesion discrimination model can be trained based on the training sample data. Among them, in order to improve the training efficiency of the lesion discrimination model, in some embodiments, the training sample data can also be divided into multiple batches, and then the lesion discrimination model can be trained batch by batch. It can be understood that the process of training the lesion discrimination model is a process of multiple loops. A model parameter convergence detection can be performed after each loop is executed. When the model parameters are detected to converge, the loop can be terminated. The following introduces a loop process.

[0156] In the embodiments of the present disclosure, after obtaining the training sample data, the features of each sample inspection image in the sample inspection image sequence can be extracted based on the image feature extraction network to obtain multiple sample image features; and the features of the sample screening result can be extracted based on the screening feature extraction network. Then, the multiple sample image features and the sample screening features are fused to obtain sample fusion features. Further, the sample fusion features can be classified based on the classification network to obtain a classification output, and the classification output can be the probability corresponding to the lesion and the probability value corresponding to no lesion. Further, the gradient value is calculated according to the classification output and the corresponding classification label. The gradient value is the backpropagation gradient, and then the network parameters of the image feature extraction network and the classification network can be adjusted according to the gradient value.

[0157] In this way, the above loop process is repeatedly executed using multiple batches of training sample data until the network parameters of each network are detected to converge, thereby obtaining the trained lesion discrimination model.

[0158] In some embodiments, fusing the multiple sample image features and the sample screening features to obtain sample fusion features includes:

[0159] Calculate the mean of the features of multiple sample images to obtain the average image feature;

[0160] Calculate the product of the average image feature and the sample screening feature, and stack the obtained product feature with the features of multiple sample images to obtain the sample fusion feature.

[0161] In the embodiments of the present disclosure, after obtaining the sample image features by performing feature extraction on the sample image sequence, the mean of the features of multiple sample images can be calculated to obtain the average image feature. Then, multiply the average image feature by the sample screening feature to obtain a new screening feature whose value range is not much different from that of the sample image feature. Further, multiple sample image features and the new screening feature can be stacked to obtain the fusion feature. By adopting this method, the screening feature can be converted to be close to the value range of the image feature, so that the screening feature can be effectively fused with the sample image feature.

[0162] In some embodiments, the sample screening result includes a first sample screening sub-result and a second sample screening sub-result. Feature extraction is performed on the sample screening result based on the screening feature extraction network to obtain the sample screening feature, including:

[0163] Obtain the first training parameter corresponding to the first sample screening sub-result and the second training parameter corresponding to the second sample screening sub-result, and the values of the first training parameter and the second training parameter are 0 or 1;

[0164] Determine the input data of the screening feature extraction network during the training process according to the first sample screening sub-result and the corresponding first training parameter, and the second sample screening sub-result and the corresponding second training parameter;

[0165] Perform feature extraction on the input data based on the screening feature extraction network to obtain the sample screening feature.

[0166] Since in the actual lesion discrimination scenario, there may be a lack of effective screening information. If during the training process, the screening information is used to train the lesion discrimination model in each training cycle, then in the actual lesion discrimination scenario where there is a lack of effective screening information, it may lead to inaccurate model output results. Therefore, in the embodiments of the present disclosure, when training the lesion discrimination model, a certain degree of random discarding of the screening information can be performed. When the lesion discrimination scenario is a specific cervical lesion discrimination scenario, the two corresponding screening information, HPV and TCT, can be independently and randomly discarded.

[0167] That is, in the embodiments of the present disclosure, the sample screening result may include a first sample screening sub-result and a second sample screening sub-result. When extracting features from the sample screening result based on the screening feature extraction network, the first training parameter corresponding to the first sample screening sub-result and the second training parameter corresponding to the second sample screening sub-result may be obtained first. Among them, the values of the first training parameter and the second training parameter are 0 or 1. That is, when the value of the training parameter is 0, the corresponding sample screening sub-result may be discarded; when the value of the training parameter is 1, the corresponding sample screening sub-result is then adopted. Alternatively, in some embodiments, data augmentation may be performed by randomly discarding, that is, new sample data is generated by discarding any one or two sample screening sub-results in the sample data to enrich the training sample data. Then, the trained sample data after data augmentation may be used to train the lesion discrimination model, thereby improving the generalization ability of the model.

[0168] In some embodiments, before extracting features from the target disease screening information based on the screening feature extraction network to obtain screening features, it further includes:

[0169] Obtain the screening date corresponding to the target disease screening information;

[0170] Extracting features from the target disease screening information based on the screening feature extraction network includes:

[0171] When the time interval between the screening date and the acquisition date of the inspection video does not exceed the preset number of days, extract features from the target disease screening information based on the screening feature extraction network.

[0172] Among them, in the embodiments of the present disclosure, in order to ensure the effectiveness of the obtained target disease screening information, before extracting features from the target disease screening information based on the screening feature extraction network, the effectiveness of the target disease screening information may be checked first. Specifically, the screening date corresponding to the target disease screening information may be obtained first, and then the time interval between the screening date and the acquisition date of the inspection video is calculated. When the time interval exceeds the preset number of days, it is determined that the target disease screening information is invalid information; on the contrary, when the time interval does not exceed the preset number of days, it may be determined that the target disease screening information is valid information. When the target disease screening information is valid information, features are then extracted from the target disease screening information based on the screening feature extraction network.

[0173] In some embodiments, before fusing multiple image features and screening features to obtain fused features, it further includes:

[0174] Calculate the differences between adjacent image features according to the temporal relationship between the inspection images to obtain multiple image difference features;

[0175] Perform feature fusion on multiple image features and screening features to obtain fused features, including:

[0176] Perform feature fusion on multiple image difference features and screening features to obtain fused features.

[0177] In the embodiments of the present disclosure, the fused features can be obtained by fusing based on the differences between image features and screening features. In the scenario of cervical lesion discrimination, the color of the cervical epidermis changes over time during the acetic acid white test, and this change varies due to differences in skin color among different people, making it difficult for doctors to make manual discrimination, which may lead to misjudgment and over-treatment, and the possibility that patients suffer unnecessary medical pain. However, with the characteristics of the neural network model in this solution that can accurately grasp and identify these subtle changes, a video of the changes in the cervical epidermis during the acetic acid white test is collected through a colposcope, and a sequence of temporal images is extracted from it, and then accurate discrimination is given by integrating screening information. During the process of image feature extraction, the changes in the image are more important information compared to the content of the image. Therefore, the differences between adjacent image features can be calculated according to the temporal relationship between the inspection images to obtain multiple image difference features.

[0178] Then, multiple image difference features can be further fused with screening features to obtain fused features. In this embodiment, since the fusion of image difference features and screening features is used as the input of the classification network, the model can pay more attention to the changes in the images over time in the image sequence, and the changes in the images over time can more accurately reveal the lesion results. Therefore, the embodiments of the present disclosure can further improve the accuracy of lesion discrimination. Correspondingly, during the training process of the lesion discrimination model, the differences between adjacent sample image features in the sample temporal image sequence can also be extracted to obtain sample image difference features, and then sample fused features are obtained by fusing the sample image difference features and sample screening features, and the lesion discrimination model is further trained based on the sample fused features.

[0179] In some embodiments, based on an image feature extraction network, feature extraction is performed on each inspection image in the inspection image sequence to obtain multiple image features, including:

[0180] Based on an image feature extraction network, optical flow feature extraction is performed on each inspection image in the inspection sequence to obtain multiple image optical flow features;

[0181] According to the temporal relationship between the inspection images, the differences between adjacent image features are calculated to obtain multiple image difference features, including:

[0182] According to the temporal relationship between the inspection images, the differences between adjacent image optical flow features are calculated to obtain multiple image optical flow difference features;

[0183] Fuse multiple image difference features and screening features to obtain fused features, including:

[0184] Fuse multiple image optical flow difference features and screening features to obtain fused features.

[0185] In the embodiments of the present disclosure, when extracting features of inspection images in an inspection image sequence, the optical flow features of the inspection images can be specifically extracted. In this way, when calculating the differences between image features, the differences between the optical flow features of adjacent images can be calculated, and then the fused features can be obtained by fusing the differences between the optical flow features and the screening features.

[0186] Step 250: Input the fused features into a classification network for classifying the lesion discrimination results to obtain the lesion discrimination results output by the classification network.

[0187] After obtaining the fused features, the fused features can be further input into a classification network for classifying the lesion discrimination results to obtain the lesion discrimination results output by the classification network. Herein, the classification network can specifically be the classification network in the lesion discrimination model trained by using the aforementioned training method of the lesion discrimination model. The lesion discrimination results output by the classification network include the probabilities of the presence and absence of lesions at the target body surface position of the target object, and the presence or absence of lesions at the target body surface position of the target object can be further determined according to the probability values as the final result of lesion discrimination.

[0188] In the embodiments of the present disclosure, by obtaining the target disease screening information of the target object, where the target disease screening information is the screening result obtained by performing cytological screening on the target object for the target disease; obtaining an inspection video collected for the target body surface position of the target object within a preset time period, where the preset time period is the chemical reaction time period after applying a preset chemical substance to the target body surface position; extracting multiple inspection images from the inspection video in chronological order to obtain an inspection image sequence with a temporal relationship; respectively extracting features from the inspection image sequence and the target disease screening information, and fusing the extracted image features and the screening result features to obtain fused features; inputting the fused features into a classification network for classifying the lesion discrimination results to obtain the lesion discrimination results output by the classification network.

[0189] In the embodiments of the present disclosure, multiple sequential images are extracted from the target disease screening information of the target object and the examination video collected at the target body surface position. The neural network model is used to perform feature extraction and fusion processing on the target disease screening information and the multiple sequential images respectively to obtain the fusion features, and then the classification network is used to perform classification processing on the fusion features, so that an accurate lesion discrimination result can be automatically obtained. Compared with the manual lesion diagnosis by doctors, this solution can perform joint automatic diagnosis through multi-modal examination information, avoiding the problem of incorrect lesion discrimination results caused by manual experience differences or diagnostic errors, and can greatly improve the accuracy of lesion discrimination.

[0190] Detailed description of the embodiments of the present disclosure in combination with specific application scenarios

[0191] As Figure 7 shown, it is another schematic flowchart of the lesion discrimination method provided by the present disclosure. In this embodiment, the lesion discrimination method will be introduced in detail in combination with the execution subject of each step. The method specifically includes the following steps:

[0192] Step 701, the computer device obtains the training sample data for training the lesion discrimination model.

[0193] In the embodiments of the present disclosure, the lesion discrimination method provided by the present disclosure can be specifically applied to the discrimination of cervical lesions of a target object (any patient) in a colposcope acquisition system. Before deploying the lesion discrimination model to the control terminal of the colposcope acquisition system, it is necessary to first train the lesion discrimination model. The training of the lesion discrimination model can be specifically carried out in a computer device, and the computer device can be a terminal, a server, or also the control terminal of the colposcope acquisition system. In the embodiments of the present disclosure, in order to use a larger amount of data to train the lesion discrimination model more fully to improve the generalization ability and robustness of the lesion discrimination model, and to improve the training efficiency of the lesion discrimination model, the training of the lesion discrimination model can be carried out in the server, that is, the computer device here can specifically be a server.

[0194] When training the lesion discrimination model in the server, the server can first obtain the training sample data for training the lesion discrimination model. Among them, the training sample data can include the colposcope sample image sequence, sample screening information, and corresponding cervical lesion labels of multiple sample objects. Among them, the colposcope sample image sequence is a sample image sequence selected from the video collected during the acetic acid white laboratory examination of the target object at the cervix, and the sample screening information is the HPV and TCT screening information of the target object.

[0195] Step 702, the computer device trains the lesion discrimination model according to the training sample data.

[0196] After obtaining the training sample data, the server can train the lesion discrimination model based on the training sample data. Specifically, the image feature extraction network of the lesion discrimination model can be first used, which can be the ResNet50 network specifically, to extract image features from each sample image in the sample image sequence, obtaining multiple sample image features. For example, when the sample image sequence contains 5 images, then 5 sample image features F1, F2, F3, F4, and F5 are extracted.

[0197] Then, sample screening features can be constructed according to the HPV and TCT screening information. According to inductive statistics, the possible diagnostic results of HPV include: none, negative, non-16 / 18hr-HPVs, and HPV16 / 18+; the possible diagnostic results of TCT include: none, negative, ≥ASC-US, LSIL, and HSIL. Thus, 1 can be used as the feature construction for the screening information, including 9 elements, which successively represent none, negative, non-16 / 18hr-HPVs, and HPV16 / 18+ of HPV and none, negative, ≥ASC-US, LSIL, and HSIL of TCT. Correspondingly, the corresponding result is set to 1 according to the specific result, and the others are 0. For example, if the HPV result is negative and the TCL result is LSIL, then the sample screening feature F6 = [0, 1, 0, 0, 0, 0, 0, 1, 0].

[0198] Then, the sample image features and the sample screening features are fused. Specifically, the mean M of F1 to F5 can be first calculated F , and then F6 is multiplied by M F to obtain a new sample screening feature F7 = F6 * M whose value range is not much different from that of F1 to F5 F . Then, multiple sample image features and the new sample screening feature are stacked to obtain the sample fusion feature F = [F1, F2, F3, F4, F5, F7].

[0199] Furthermore, the sample fusion feature can be input into the classification model, and the parameters of the classification model and the image feature extraction model can be adjusted according to the difference between the output result of the classification model and the cervical lesion label. Then, the above training steps are cyclically executed until the network parameters of the lesion discrimination model converge, obtaining the trained lesion discrimination model.

[0200] Step 703, the computer device sends the model parameters of the lesion discrimination model to the control terminal of the colposcope acquisition system.

[0201] In the embodiments of the present disclosure, the automatic cervical lesion discrimination ability can be deployed in the colposcope acquisition system, so that the colposcope images can complete cervical lesion discrimination locally, avoiding the possible transmission and leakage of medical information, and greatly improving the information security and privacy protection of the patients.

[0202] Specifically, after the computer device completes the training of the lesion discrimination model, it can package and store the model parameters, and then send the model parameters to the control terminal of the colposcope acquisition system.

[0203] Step 704, the control terminal of the colposcope acquisition system deploys the lesion discrimination model.

[0204] After receiving the model parameters, the control terminal of the colposcope acquisition system performs local deployment of the lesion discrimination model. Specifically, the control terminal of the colposcope acquisition system can obtain the corresponding resources and files in advance according to the backbone network required by the lesion discrimination model, and then deploy the lesion discrimination model based on the obtained corresponding resources and files combined with the model parameters received from the computer device.

[0205] Step 705, the control terminal of the colposcope acquisition system obtains the colposcope video of the target object during the acetic acid white test and the HPV and TCT screening information of the target object.

[0206] After the lesion discrimination model is deployed on the control terminal of the colposcope acquisition system, the cervical lesions of the target object can be discriminated during the acetic acid white test and colposcopy of the target object.

[0207] Specifically, the colposcope acquisition system can collect videos of multiple positions of the cervical epidermis of the target object using a colposcope during the acetic acid white test stage of the target object to obtain a colposcope video. Also, the control terminal of the colposcope acquisition system can obtain the HPV and TCT screening information of the target object.

[0208] Among them, in some cases, the HPV and TCT screening information of the target object may not be obtained, or only one of them can be obtained, or the obtained HPV and TCT screening information of the target object may be invalid due to a long interval. In these cases, the cervical lesions can be discriminated only based on the captured colposcope video.

[0209] Step 706, the control terminal of the colposcope acquisition system splits the colposcope video into frames and selects multiple first image frames from the split image frames.

[0210] After collecting the colposcope video, the control terminal of the colposcope acquisition system can split the colposcope video into frames to obtain multiple colposcope image frames.

[0211] Then, the control terminal of the colposcope acquisition system can automatically select image frames from the multiple split colposcope images. This selection can be made by uniformly selecting at a fixed number of image frame intervals or by uniformly selecting at a fixed time interval to obtain multiple first image frames.

[0212] Step 707, the control terminal of the colposcope acquisition system displays the multiple image frames obtained by splitting, and determines multiple second image frames according to the received first selection instruction.

[0213] After automatic image frame selection, the control terminal of the colposcope acquisition system can further display the multiple colposcope images obtained by splitting in the control terminal of the colposcope acquisition system, and then receive the first selection instruction input by the operator of the colposcope acquisition system. Among them, the first selection instruction is used to select the colposcope images that the operator deems relatively important from the displayed multiple colposcope images, so as to obtain multiple second image frames.

[0214] Step 708, the control terminal of the colposcope acquisition system displays the first image frame and the second image frame, and determines the target image frame based on the received second selection instruction.

[0215] Furthermore, the control terminal of the colposcope acquisition system can display the automatically selected first image frame and the multiple manually selected second image frames in the display interface of the control terminal of the colposcope acquisition system in chronological order.

[0216] Then, the second selection instruction for target image frame selection can be received. The second selection instruction is used to select multiple target image frames with higher image quality from the initially selected first image frame and second image frames.

[0217] Step 709, the control terminal of the colposcope acquisition system inputs the target image frame and the HPV and TCT screening information of the target object into the lesion discrimination model to obtain the cervical lesion discrimination result of the target object.

[0218] After determining the target image frame, the target image frame and the HPV and TCT screening information of the target object can be input into the lesion discrimination model deployed in the control terminal of the colposcope acquisition system for cervical lesion discrimination to obtain an accurate cervical lesion discrimination result.

[0219] Among them, when the cervical lesion discrimination result is confirmed as a lesion, the lesion location can also be determined according to the specific cervical position of the video acquisition to provide a basis for cervical biopsy.

[0220] Device and equipment description of the embodiments of the present disclosure

[0221] It can be understood that although the steps in each of the above flowcharts are sequentially shown according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this embodiment, the execution of these steps has no strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowchart may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0222] It should be noted that in each specific embodiment of the present disclosure, when it comes to performing relevant processing based on data related to the characteristics of the target object, such as target object attribute information or attribute information set, etc., the permission or consent of the target object will be obtained first. Moreover, the collection, use, and processing of these data will comply with the relevant laws, regulations, and standards of the relevant region. In addition, when the embodiment of the present application needs to obtain the target object attribute information, it will obtain the separate permission or separate consent of the target object through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the necessary target object-related data for the normal operation of the embodiment of the present application will be obtained.

[0223] Figure 8 FIG. 800 is a schematic structural diagram of a lesion discrimination device provided by an embodiment of the present disclosure. The device includes:

[0224] A first acquisition unit 810, configured to acquire target disease screening information of a target object, where the target disease screening information is a screening result obtained by performing cytological screening on the target object for a target disease;

[0225] A second acquisition unit 820, configured to acquire an inspection video collected at a target body surface position of the target object within a preset time period, where the preset time period is a chemical reaction time period after applying a preset chemical substance to the target body surface position;

[0226] An extraction unit 830, configured to extract multiple inspection images from the inspection video in chronological order to obtain an inspection image sequence with a temporal relationship;

[0227] A fusion unit 840, configured to respectively perform feature extraction on the inspection image sequence and the target disease screening information, and perform feature fusion on the extracted image features and screening result features to obtain fusion features;

[0228] A discrimination unit 850, configured to input the fusion features into a classification network for lesion discrimination result classification to obtain a lesion discrimination result output by the classification network.

[0229] Optionally, in some embodiments, the fusion unit includes:

[0230] A first extraction subunit, configured to extract features from each inspection image in the inspection image sequence based on an image feature extraction network to obtain a plurality of image features;

[0231] A second extraction subunit, configured to extract features from the target disease screening information based on a screening feature extraction network to obtain screening features;

[0232] A first fusion subunit, configured to perform feature fusion on the plurality of image features and the screening features to obtain fusion features;

[0233] Wherein, the image feature extraction network, the screening feature extraction network, and the classification network are jointly trained using the same batch of sample data.

[0234] Optionally, in some embodiments, the lesion discrimination device provided by the present disclosure further includes a model training unit, specifically including:

[0235] A first acquisition subunit, configured to acquire training sample data, where the training sample data includes a sample inspection image sequence of a sample object, a sample screening result, and a corresponding classification label, and the sample inspection image sequence is an image sequence obtained by extracting images from an inspection video collected at the sample body surface position of the sample object within a sample time period;

[0236] A third extraction subunit, configured to extract features from each sample inspection image in the sample inspection image sequence based on an image feature extraction network to obtain a plurality of sample image features;

[0237] A fourth extraction subunit, configured to extract features from the sample screening result based on a screening feature extraction network to obtain sample screening features;

[0238] A second fusion subunit, configured to perform feature fusion on the plurality of sample image features and the sample screening features to obtain sample fusion features;

[0239] A classification subunit, configured to perform classification processing on the sample fusion features based on a classification network to obtain a classification output;

[0240] An adjustment subunit, configured to calculate a gradient value according to the classification output and the classification label, and perform gradient backpropagation processing based on the gradient value to adjust the parameters of the image feature extraction network, the screening feature extraction network, and the classification network.

[0241] Optionally, in some embodiments, the second fusion subunit includes:

[0242] The first calculation module is used to calculate the mean value of multiple sample image features to obtain the image average feature;

[0243] The stacking module is used to calculate the product of the image average feature and the sample screening feature, and stack the obtained product feature with multiple sample image features to obtain the sample fusion feature.

[0244] Optionally, in some embodiments, the sample screening result includes a first sample screening sub-result and a second sample screening sub-result. The fourth extraction subunit includes:

[0245] The first acquisition module is used to acquire the first training parameter corresponding to the first sample screening sub-result and the second training parameter corresponding to the second sample screening sub-result. The values of the first training parameter and the second training parameter are 0 or 1;

[0246] The determination module is used to determine the input data of the screening feature extraction network during the training process according to the first sample screening sub-result and the corresponding first training parameter, and the second sample screening sub-result and the corresponding second training parameter;

[0247] The first extraction module is used to perform feature extraction on the input data based on the screening feature extraction network to obtain the sample screening feature.

[0248] Optionally, in some embodiments, the lesion discrimination device provided by the present disclosure further includes:

[0249] The second acquisition subunit is used to acquire the screening date corresponding to the target disease screening information;

[0250] The second extraction subunit is further used for:

[0251] When the time interval between the screening date and the acquisition date of the inspection video does not exceed the preset number of days, perform feature extraction on the target disease screening information based on the screening feature extraction network.

[0252] Optionally, in some embodiments, the extraction unit includes:

[0253] The splitting subunit is used to split the inspection video into frames based on the chronological order to obtain multiple image frames with chronological order;

[0254] The fifth extraction subunit is used to extract multiple first target image frames from the multiple image frames at a preset frame interval;

[0255] The determination subunit is used to determine the inspection image sequence with a timing relationship according to the multiple first target image frames and the corresponding time information.

[0256] Optionally, in some embodiments, the lesion discrimination device provided by the present disclosure further includes:

[0257] A display subunit, configured to display multiple image frames with a time sequence on a display interface of a display terminal;

[0258] A third acquisition subunit, configured to acquire a selection instruction for the displayed image frames in the display interface and determine multiple selected second target image frames;

[0259] The determination subunit is further configured to:

[0260] Determine a sequence of inspection images with a timing relationship based on multiple first target image frames, multiple second target image frames, and corresponding time information.

[0261] Optionally, in some embodiments, the fifth extraction subunit includes:

[0262] A second acquisition module, configured to acquire a preset time interval corresponding to image frame acquisition and a frame rate corresponding to an inspection video;

[0263] A calculation module, configured to calculate a frame interval corresponding to image frame acquisition based on the preset time interval and the frame rate;

[0264] A second extraction module, configured to extract multiple first target image frames from multiple image frames according to the frame interval.

[0265] Optionally, in some embodiments, the lesion discrimination device provided by the present disclosure further includes:

[0266] A calculation subunit, configured to calculate the difference between adjacent image features according to the timing relationship between inspection images to obtain multiple image difference features;

[0267] The first fusion subunit is further configured to:

[0268] Perform feature fusion on multiple image difference features and screening features to obtain a fusion feature.

[0269] Optionally, in some embodiments, the first extraction subunit is further configured to:

[0270] Extract optical flow features for each inspection image in the inspection sequence based on an image feature extraction network to obtain multiple image optical flow features;

[0271] The calculation subunit is further configured to:

[0272] Calculate the difference between adjacent image optical flow features according to the timing relationship between inspection images to obtain multiple image optical flow difference features;

[0273] The first fusion subunit is further configured to:

[0274] Perform feature fusion on multiple image optical flow difference features and screening features to obtain a fusion feature.

[0275] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of the module or unit.

[0276] Referring to Figure 9 , Figure 9 , the structural block diagram of a part of the terminal 140 for implementing the lesion discrimination method of the embodiments of the present disclosure. The terminal 140 includes components such as a Radio Frequency (RF) circuit 1410, a memory 915, an input unit 930, a display unit 940, a sensor 950, an audio circuit 960, a wireless fidelity (WiFi) module 970, a processor 980, and a power supply 990. Those skilled in the art can understand that Figure 9 the shown structure of the terminal 140 does not constitute a limitation on a mobile phone or a computer, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0277] The RF circuit 910 can be used for receiving and sending signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is given to the processor 980 for processing; in addition, the designed uplink data is sent to the base station.

[0278] The memory 915 can be used to store software programs and modules. The processor 980 executes various functional applications and document editing of the terminal by running the software programs and modules stored in the memory 915.

[0279] The input unit 930 can be used to receive input digital or character information, and generate key signal inputs related to the settings and function controls of the terminal. Specifically, the input unit 930 may include a touch panel 931 and other input devices 932.

[0280] The display unit 940 can be used to display the input information or the provided information and various menus of the terminal. The display unit 940 may include a display panel 941.

[0281] The audio circuit 960, the speaker 961, and the microphone 962 can provide an audio interface.

[0282] In this embodiment, the processor 980 included in the terminal 140 can execute the lesion discrimination method of the previous embodiments.

[0283] The terminal 140 in the embodiments of the present disclosure includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc.

[0284] Figure 10 The structural block diagram of a part of the server 110 for implementing the lesion discrimination method in the embodiments of the present disclosure. The server 110 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 1022 (for example, one or more processors) and a storage device 1032, and one or more storage media 1030 (for example, one or more mass storage devices) for storing application programs 1042 or data 1044. Among them, the storage device 1032 and the storage media 1030 may be transient storage or persistent storage. The program stored in the storage media 1030 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server 110. Further, the central processor 1022 may be configured to communicate with the storage media 1030 and execute a series of instruction operations in the storage media 1030 on the server 110.

[0285] The server 110 may further include one or more power supplies 1026, one or more wired or wireless network interfaces 1050, one or more input / output interfaces 1058, and / or one or more operating systems 1041, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0286] The central processor 1022 in the server 110 may be used to execute the lesion discrimination method in the embodiments of the present disclosure.

[0287] The embodiments of the present disclosure further provide a storage medium for storing program codes for executing the lesion discrimination methods in the foregoing respective embodiments.

[0288] The embodiments of the present disclosure further provide a computer program product, which includes a computer program. The processor of the computer device reads and executes the computer program, so that the computer device executes to implement the above-mentioned lesion discrimination method.

[0289] The terms "first", "second", "third", "fourth", etc. (if any) in the description of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0290] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B may be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c may be single or plural.

[0291] It should be understood that in the description of the embodiments of the present disclosure, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as "greater than", "less than", "exceeding", etc. do not include the present number, and understandings such as "above", "below", "within", etc. include the present number.

[0292] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0293] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0294] In addition, each functional unit in various embodiments of the present disclosure may be integrated in a processing unit, may exist as individual physical units, or two or more units may be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0295] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present disclosure. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0296] It should also be understood that the various embodiments provided in the present disclosure can be combined arbitrarily to achieve different technical effects.

[0297] The above is a specific description of the embodiments of the present disclosure, but the present disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present disclosure.

Claims

1. A method for lesion discrimination, characterized in that, The method includes: Obtaining target disease screening information of a target object, where the target disease screening information is a screening result obtained by performing cytological screening on the target object for the target disease; Obtaining an inspection video collected for a target body surface position of the target object within a preset time period, where the preset time period is a chemical reaction time period after applying a preset chemical substance to the target body surface position; Extracting multiple inspection images from the inspection video in chronological order to obtain an inspection image sequence with a temporal relationship; Respectively extracting features from the inspection image sequence and the target disease screening information, and performing feature fusion on the extracted image features and screening result features to obtain fusion features; Inputting the fusion features into a classification network for classifying lesion discrimination results to obtain the lesion discrimination results output by the classification network.

2. The method according to claim 1, wherein The respectively extracting features from the inspection image sequence and the target disease screening information, and performing feature fusion on the extracted image features and screening result features to obtain fusion features includes: Extracting features from each inspection image in the inspection image sequence based on an image feature extraction network to obtain multiple image features; Extracting features from the target disease screening information based on a screening feature extraction network to obtain screening features; Performing feature fusion on the multiple image features and the screening features to obtain fusion features; Among them, the image feature extraction network, the screening feature extraction network, and the classification network are jointly trained using the same batch of sample data.

3. The method according to claim 2, characterized in that, The training process of the image feature extraction network, the screening feature extraction network, and the classification network includes the following steps: Obtaining training sample data, where the training sample data includes a sample inspection image sequence of a sample object, a sample screening result, and a corresponding classification label, and the sample inspection image sequence is an image sequence obtained by extracting images from an inspection video collected for a sample body surface position of the sample object within a sample time period; Extracting features from each sample inspection image in the sample inspection image sequence based on the image feature extraction network to obtain multiple sample image features; Extracting features from the sample screening result based on the screening feature extraction network to obtain sample screening features; Performing feature fusion on the multiple sample image features and the sample screening features to obtain sample fusion features; Performing classification processing on the sample fusion features based on the classification network to obtain a classification output; Calculating a gradient value according to the classification output and the classification label, and performing gradient backpropagation processing based on the gradient value to adjust the parameters of the image feature extraction network, the screening feature extraction network, and the classification network.

4. The method according to claim 3, characterized in that, The performing feature fusion on the multiple sample image features and the sample screening features to obtain sample fusion features includes: Calculating the mean of the multiple sample image features to obtain an image average feature; Calculating the product of the image average feature and the sample screening feature, and stacking the obtained product feature with the multiple sample image features to obtain sample fusion features.

5. The method according to claim 3, wherein The sample screening result includes a first sample screening sub-result and a second sample screening sub-result. Feature extraction of the sample screening result by the screening feature extraction network to obtain sample screening features includes: Obtaining a first training parameter corresponding to the first sample screening sub-result and a second training parameter corresponding to the second sample screening sub-result, where the values of the first training parameter and the second training parameter are 0 or 1; Determining the input data of the screening feature extraction network during training according to the first sample screening sub-result and the corresponding first training parameter, and the second sample screening sub-result and the corresponding second training parameter; Performing feature extraction on the input data based on the screening feature extraction network to obtain sample screening features.

6. The method according to claim 2, wherein Before performing feature extraction on the target disease screening information by the screening feature extraction network to obtain screening features, it further includes: Obtaining the screening date corresponding to the target disease screening information; Performing feature extraction on the target disease screening information by the screening feature extraction network, including: When the time interval between the screening date and the acquisition date of the examination video does not exceed a preset number of days, performing feature extraction on the target disease screening information by the screening feature extraction network.

7. The method according to claim 1, wherein Extracting multiple examination images from the examination video in chronological order to obtain a sequence of examination images with a temporal relationship, including: Splitting the examination video into frames based on chronological order to obtain multiple image frames with a time sequence; Extracting multiple first target image frames from the multiple image frames at a preset frame interval; Determining a sequence of examination images with a temporal relationship according to the multiple first target image frames and the corresponding time information.

8. The method according to claim 7, wherein Before determining multiple examination images according to the multiple first target image frames, it further includes: Displaying the multiple image frames with a time sequence on the display interface of the display terminal; Obtaining a selection instruction for the displayed image frames on the display interface to determine multiple selected second target image frames; Determining a sequence of examination images with a temporal relationship according to the multiple first target image frames and the corresponding time information, including: Determining a sequence of examination images with a temporal relationship according to the multiple first target image frames, the multiple second target image frames, and the corresponding time information.

9. The method according to claim 7, characterized in that Extracting multiple first target image frames from the multiple image frames at a preset frame interval, including: Obtaining a preset time interval corresponding to image frame acquisition and the frame rate corresponding to the examination video; Calculating the frame interval corresponding to image frame acquisition based on the preset time interval and the frame rate; Extracting multiple first target image frames from the multiple image frames according to the frame interval.

10. The method according to claim 2, characterized in that, Before performing feature fusion on the multiple image features and the screening features to obtain fusion features, it further includes: Calculating the differences between adjacent image features according to the temporal relationship between the examination images to obtain multiple image difference features; Performing feature fusion on the multiple image features and the screening features to obtain fusion features, including: Fuse the multiple image difference features and the screening features to obtain fused features.

11. The method according to claim 11, wherein Based on the image feature extraction network, extract features from each of the inspection images in the inspection image sequence to obtain multiple image features, including: Based on the image feature extraction network, extract optical flow features from each of the inspection images in the inspection sequence to obtain multiple image optical flow features; Calculate the differences between adjacent image features according to the temporal relationship between the inspection images to obtain multiple image difference features, including: Calculate the differences between adjacent image optical flow features according to the temporal relationship between the inspection images to obtain multiple image optical flow difference features; Fuse the multiple image difference features and the screening features to obtain fused features, including: Fuse the multiple image optical flow difference features and the screening features to obtain fused features.

12. A lesion discrimination device, characterized in that, The device includes: A first acquisition unit for acquiring the target disease screening information of the target object, where the target disease screening information is the screening result obtained by performing cytological screening on the target object for the target disease; A second acquisition unit for acquiring the inspection video collected at the target body surface position of the target object within a preset time period, where the preset time period is the chemical reaction time period after applying a preset chemical substance to the target body surface position; An extraction unit for extracting multiple inspection images from the inspection video in chronological order to obtain an inspection image sequence with a temporal relationship; A fusion unit for respectively extracting features from the inspection image sequence and the target disease screening information, and fusing the extracted image features and the screening result features to obtain fused features; A discrimination unit for inputting the fused features into a classification network for lesion discrimination result classification to obtain the lesion discrimination result output by the classification network.

13. A storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the lesion discrimination method according to any one of claims 1 to 11.

14. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the lesion discrimination method according to any one of claims 1 to 11.

15. A computer program product, which includes a computer program that is read and executed by a processor of a computer device, so that the computer device executes the lesion discrimination method according to any one of claims 1 to 11.