Artificial intelligence-based image processing method, apparatus, and electronic device

By acquiring and expanding image samples to train a size prediction model, the problem of low accuracy in image target recognition was solved, and high-precision target recognition was achieved in different scenarios.

CN113538228BActive Publication Date: 2025-10-24TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011407277.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-04
Publication Date
2025-10-24
Estimated Expiration
2041-01-04

AI Technical Summary

Technical Problem

Existing technologies suffer from low recognition accuracy in image target recognition due to image differences.

Method used

By acquiring multiple first samples including images and pixel sizes, expanding them to generate second samples covering a range of pixel sizes, training a size prediction model, and using the trained model for size prediction and target recognition.

Benefits of technology

It improves the accuracy of target recognition, ensures the effective use of computing resources, and is suitable for scenarios such as face recognition, satellite monitoring, and clinical medicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113538228B_ABST
    Figure CN113538228B_ABST
Patent Text Reader

Abstract

The application provides an image processing method and device based on artificial intelligence, electronic equipment and computer readable storage medium, relates to the field of artificial intelligence and big data technology in the field of cloud technology, and the method comprises the following steps: acquiring a plurality of first samples comprising images and pixel sizes of the images; performing expansion processing on the plurality of first samples according to a set pixel size range to obtain a plurality of second samples covering the pixel size range; training a size prediction model according to the plurality of second samples; performing size prediction processing on a to-be-tested image according to the trained size prediction model; and performing target recognition processing on the to-be-tested image according to the pixel size obtained through the size prediction processing. Through the application, the accuracy of target recognition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to artificial intelligence technology and cloud technology, and in particular to an image processing method, device, electronic device and computer-readable storage medium based on artificial intelligence. Background Art

[0002] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. Computer vision (CV) technology is a key branch of AI, focusing on related theories and technologies, aiming to build AI systems capable of extracting information from images or multidimensional data.

[0003] Object recognition is a typical application of computer vision technology, used to identify specific objects in images, such as humans, cats, or dogs. Conventional solutions typically perform object recognition processing directly on the original image, applying the same processing method to all images. However, due to the inherent differences between images, these conventional solutions often suffer from poor object recognition accuracy. Summary of the Invention

[0004] The embodiments of the present application provide an artificial intelligence-based image processing method, device, electronic device, and computer-readable storage medium, which can improve the accuracy of target recognition processing.

[0005] The technical solution of the embodiment of the present application is implemented as follows:

[0006] The present invention provides an artificial intelligence-based image processing method, including:

[0007] Acquire a plurality of first samples including an image and pixel dimensions of the image;

[0008] performing expansion processing on the plurality of first samples according to a set pixel size range to obtain a plurality of second samples covering the pixel size range;

[0009] training a size prediction model based on the plurality of second samples;

[0010] Performing size prediction processing on the image to be tested according to the trained size prediction model;

[0011] According to the pixel size obtained by the size prediction process, target recognition processing is performed on the image to be tested.

[0012] The present invention provides an artificial intelligence-based image processing device, comprising:

[0013] an acquisition module configured to acquire a plurality of first samples each comprising an image and a pixel size of the image;

[0014] an expansion module configured to expand the plurality of first samples according to a preset pixel size range to obtain a plurality of second samples covering the pixel size range;

[0015] a training module configured to train a size prediction model according to the plurality of second samples;

[0016] a size prediction module configured to perform size prediction processing on a to-be-tested image according to the trained size prediction model;

[0017] a target recognition module configured to perform target recognition processing on the to-be-tested image according to a pixel size obtained through the size prediction processing.

[0018] An electronic device is provided in an embodiment of the present application, and the electronic device comprises:

[0019] a memory configured to store executable instructions;

[0020] a processor configured to execute the executable instructions stored in the memory to implement the image processing method based on artificial intelligence provided in the embodiments of the present application.

[0021] A computer readable storage medium is provided in an embodiment of the present application, and the computer readable storage medium stores executable instructions, and the executable instructions are used to cause a processor to execute the image processing method based on artificial intelligence provided in the embodiments of the present application.

[0022] The embodiments of the present application have the following beneficial effects:

[0023] The plurality of first samples acquired are expanded according to a preset pixel size range to obtain a plurality of second samples covering the pixel size range, wherein each first sample comprises an image and a pixel size of the image. Then, a size prediction model is trained according to the plurality of second samples, so that the trained size prediction model can accurately predict the pixel size of a to-be-tested image. The pixel size of the to-be-tested image can be used as reference information for target recognition processing on the to-be-tested image, thereby improving the accuracy of the target recognition processing. BRIEF DESCRIPTION OF DRAWINGS

[0024] Figure 1 is an architecture schematic diagram of the image processing system based on artificial intelligence provided in the embodiments of the present application;

[0025] Figure 2 is an architecture schematic diagram of the terminal device provided in the embodiments of the present application;

[0026] Figure 3Ais a flowchart of an image processing method based on artificial intelligence provided by an embodiment of the present application;

[0027] Figure 3B is a flowchart of an image processing method based on artificial intelligence provided by an embodiment of the present application;

[0028] Figure 3C is a flowchart of an image processing method based on artificial intelligence provided by an embodiment of the present application;

[0029] Figure 3D is a flowchart of an image processing method based on artificial intelligence provided by an embodiment of the present application;

[0030] Figure 3E is a flowchart of an image processing method based on artificial intelligence provided by an embodiment of the present application;

[0031] Figure 3F is a flowchart of an image processing method based on artificial intelligence provided by an embodiment of the present application;

[0032] Figure 4 is a flowchart of an image processing method based on artificial intelligence provided by an embodiment of the present application;

[0033] Figure 5 is a schematic diagram of determining pixel sizes corresponding to different magnifications provided by an embodiment of the present application;

[0034] Figure 6 is a schematic diagram of a pathological image including a virtual ruler provided by an embodiment of the present application;

[0035] Figure 7 is a schematic diagram of predicting a pixel size provided by an embodiment of the present application;

[0036] Figure 8 is a schematic diagram of predicting a pixel size provided by an embodiment of the present application;

[0037] Figure 9 is a schematic diagram of predicting a pixel size provided by an embodiment of the present application. DETAILED DESCRIPTION

[0038] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings, and the described embodiments should not be regarded as limiting the present application, and all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0039] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments, but which can be understood as possibly representing a same or different subset of all possible embodiments, and which can be combined with each other, without conflicts, in the following description.

[0040] In the following description, the terms "first\second\third" are merely used to distinguish similar objects, and do not represent a specific order of the objects, and it can be understood that the "first\second\third" can be interchanged in a specific order or sequence as allowed, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein. In the following description, the term "a plurality of" refers to at least two.

[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this application is for the purpose of describing embodiments of this application only, and is not intended to be limiting of this application.

[0042] Before further detailing the embodiments of the application, the terms and phrases involved in the embodiments of the application are explained, and the terms and phrases involved in the embodiments of the application are applicable to the following explanations.

[0043] 1) Pixel Size: the size of one pixel in an image, the unit of the pixel size is not limited in the embodiments of the application, for example, it can be microns per pixel (MPP). The pixel size of different images can be different, for example, the pixel size of an image observed by a microscope at 10 times magnification (such as the magnification of an objective lens) is different from the pixel size of an image observed at 20 times magnification.

[0044] 2) Sample: refers to a training sample used for model training. For different models, the content included in the corresponding sample can also be different, for example, for a size prediction model, the corresponding sample includes an image and the pixel size of the image; for an image classification model, the corresponding sample includes an image and the category of the image; for a target recognition model, the corresponding sample includes an image, the position of the target in the image, and the category of the target in the image.

[0045] 3) Machine Learning (ML): an important branch of artificial intelligence, mainly studying how computers simulate or implement human learning behavior to acquire new knowledge or skills, reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent, and its applications are widespread in various fields of artificial intelligence. In the embodiments of the present application, the model involved can be a model constructed based on the principle of machine learning, such as a deep learning model.

[0046] 4) Backpropagation: for models constructed based on the principle of machine learning, it often involves two processes of forward propagation (also known as forward propagation) and backpropagation. Taking a neural network model including an input layer, a hidden layer and an output layer as an example, forward propagation refers to a series of calculations on the input data in the order of input layer-hidden layer-output layer, and finally obtaining the prediction result output by the output layer; backpropagation refers to the difference between the prediction result and the actual result being propagated to each layer in the order of output layer-hidden layer-input layer, and the weight parameters of each layer are updated in the gradient descent direction during the process.

[0047] 5) Virtual ruler: can refer to a virtual ruler obtained by imaging after shooting an actually existing ruler card. For example, in the scenario of cell observation, a ruler card is fixed on a glass slide, and at the same time, a tissue section (such as a human tissue section) is placed on the glass slide. A microscope is used to observe the glass slide, and a special microscope camera is used to shoot the image observed by the microscope. Then, the image shot includes not only the observed tissue section, but also the observed ruler card (i.e. virtual ruler). In the embodiments of the present application, the virtual ruler can be used to determine the pixel size of the image.

[0048] 6) Database: similar to an electronic file cabinet, i.e. a place to store electronic files, users can perform operations such as adding, querying, updating and deleting data in the files. The database can also be understood as a collection of data stored together in a certain way, shared by multiple users, with as little redundancy as possible, and independent of application programs. In the embodiments of the present application, the database can be used to store samples, and can also be used to store images to be tested.

[0049] 7) Big Data: refers to a collection of data that cannot be captured, managed and processed within a certain time range by conventional software tools, and is a massive, high-growth and diversified information asset that requires new processing mode to have stronger decision-making, insight discovery and process optimization capabilities. The technologies suitable for big data include large-scale parallel processing database, data mining, distributed file system, distributed database, cloud computing platform, Internet and scalable storage system. In the embodiments of the present application, the big data technology can be used to realize image processing, such as storage and expansion processing of samples.

[0050] The embodiments of the present application provide an image processing method and device based on artificial intelligence, electronic equipment and computer readable storage medium, which can improve the accuracy of target identification. The following describes an exemplary application of the electronic equipment provided by the embodiments of the present application. The electronic equipment provided by the embodiments of the present application can be implemented as various types of terminal equipment, or as a server.

[0051] Referring to Figure 1 , Figure 1 is an architecture schematic diagram of an image processing system 100 based on artificial intelligence provided by the embodiments of the present application. The terminal equipment 400 is connected to the server 200 through the network 300, and the server 200 is connected to the database 500. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0052] In some embodiments, taking the electronic equipment as an example, the image processing method based on artificial intelligence provided by the embodiments of the present application can be implemented by the terminal equipment. For example, the terminal equipment 400 runs the client 410, and the client 410 expands the first sample according to the set pixel size range to obtain the second sample covering the pixel size range. The first sample can be pre-stored in the local client 410, or can be obtained from the database 500 through the server 200. Then, the client 410 trains the size prediction model according to the second sample, and performs size prediction processing on the obtained test image according to the trained size prediction model. The test image can be pre-stored in the local client 410, can be obtained by real-time shooting, or can be obtained from the database 500, which is not limited. After the client 410 obtains the pixel size of the test image through the size prediction processing, the client 410 performs target identification processing on the test image according to the pixel size to obtain the identification result

[0053] In some embodiments, the electronic device is a server, and the AI-based image processing method provided by the embodiments of the present application can also be implemented by the server. For example, the server 200 obtains a plurality of first samples from the database 500, and performs expansion processing on the plurality of first samples according to a set pixel size range to obtain a plurality of second samples covering the pixel size range. Then, the server 200 trains the size prediction model based on the plurality of second samples, performs size prediction processing on the test image according to the trained size prediction model, and performs target recognition processing on the test image according to the obtained pixel size. The test image can be an image stored in the database 500.

[0054] In some embodiments, the AI-based image processing method provided by the embodiments of the present application can also be implemented by the server and the terminal device. For example, after the server 200 completes the training of the size prediction model, it performs size prediction processing on the test image obtained from the client 410 according to the trained size prediction model, and then performs target recognition processing according to the obtained pixel size, and returns the obtained recognition result to the client 410.

[0055] It is worth noting that the server 200 can store the trained size prediction model locally, for example, in a distributed file system deployed locally, to invoke the size prediction model to complete the size prediction processing of the test image, or can send the trained size prediction model to the client 410 to perform size prediction processing according to the trained size prediction model. In the case of target recognition processing according to the target recognition model, the server 200 can also store the target recognition model locally or send it to the client 410.

[0056] The electronic device (such as the terminal device 400 or the server 200) can improve the accuracy of target recognition processing of the image, i.e., improve the actual utilization rate of the computing resources consumed by the electronic device itself in the processing process, by using the image processing scheme provided by the embodiments of the present application, which is suitable for various application scenarios. For example, in the face recognition scenario, the accuracy of face recognition of the image can be improved, and the misjudgment rate can be reduced; in the satellite monitoring scenario, specific targets such as human bodies or vehicles in the satellite image can be accurately recognized, and further processing such as population density analysis or vehicle density analysis in a certain area can be performed according to the recognition result; in the clinical medicine scenario, whether the pathological image (such as the pathological image of a human tissue section) includes tumor cells (i.e., the target) can be identified, so as to predict and quantify the pathological condition of the corresponding part of the human body, and provide accurate and effective data support for clinical research. Figure 1 In the face recognition scenario, the recognition result obtained after the target recognition processing of the test image is shown, i.e., three face frames.

[0057] In some embodiments, the terminal device 400 or the server 200 can implement the artificial intelligence-based image processing method provided in the embodiments of the present application by running a computer program. For example, the computer program can be a native program or a software module in an operating system; can be a native application program (APP), i.e., a program that needs to be installed in an operating system to run, such as a camera application program (corresponding to the client 410) and the like; can also be a mini program, i.e., a program that only needs to be downloaded into a browser environment to run; and can also be a mini program that can be embedded into any APP, such as a mini program component embedded into a camera application program, for performing size prediction processing and target identification processing on an image. The mini program component can be controlled by a user to run or be turned off. In summary, the above computer program can be any form of application program, module or plug-in.

[0058] In some embodiments, the server 200 can be a standalone physical server, or can be a server cluster or a distributed system composed of multiple physical servers, or can be a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDNs, and basic cloud computing services such as big data and artificial intelligence platforms, etc. The cloud service can be an image processing service, which is called by the terminal device 400 to perform size prediction processing and target identification processing on an image. The terminal device 400 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart television, a smart watch, etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in the embodiments of the present application.

[0059] For example, the electronic device is a terminal device provided in the embodiments of the present application. It can be understood that for the case where the electronic device is a server, Figure 2 some of the structures shown in the above description (e.g., user interface, presentation module and input processing module) can be omitted. See Figure 2 , Figure 2 is a structural schematic diagram of the terminal device 400 provided in the embodiments of the present application, Figure 2 The terminal device 400 shown in FIG. 4 includes at least one processor 460, a memory 450, at least one network interface 420 and a user interface 430. The various components in the terminal 400 are coupled together through a bus system 440. It can be understood that the bus system 440 is used to realize the connection and communication between the components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the purpose of clear illustration, all kinds of buses are marked as the bus system 440 in Figure 2 .

[0060] The processor 460 can be an integrated circuit chip that has a processing capability of signals, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc., wherein the general-purpose processor can be a microprocessor or any conventional processor.

[0061] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432 that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0062] The memory 450 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 optionally includes one or more storage devices physically located in proximity to the processor 460.

[0063] The memory 450 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. Non-volatile memory can be read only memory (ROM), and volatile memory can be random access memory (RAM). The memory 450 described in embodiments of the present application is intended to include any suitable type of memory.

[0064] In some embodiments, the memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or a subset or superset thereof, which are exemplarily illustrated below.

[0065] The operating system 451 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0066] The network communication module 452 is used to reach other computing devices via one or more (wired or wireless) network interfaces 420, examples of which include Bluetooth, wireless compatibility authentication (WiFi), and universal serial bus (USB), etc.

[0067] a presentation module 453 for enabling presentation of information (e.g., a user interface for operating a peripheral device and displaying content and information) via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430;

[0068] an input processing module 454 for detecting and translating one or more user inputs or interactions from one or more input devices 432.

[0069] In some embodiments, the apparatus provided by the embodiments of the present application can be implemented in software, Figure 2 An artificial intelligence-based image processing apparatus 455 stored in the memory 450 is shown, which can be software in the form of programs and plug-ins, etc., including the following software modules: an acquisition module 4551, an expansion module 4552, a training module 4553, a size prediction module 4554, and a target recognition module 4555. These modules are logical, and thus can be combined or further split according to the functions implemented. The functions of each module will be described below.

[0070] The artificial intelligence-based image processing method provided by the embodiments of the present application will be described in conjunction with an exemplary application and implementation of an electronic device provided by the embodiments of the present application.

[0071] Referring to Figure 3A , Figure 3A is a flowchart of an artificial intelligence-based image processing method provided by the embodiments of the present application, which will be described in conjunction with the steps shown. Figure 3A

[0072] In step 101, a plurality of first samples including images and pixel sizes of the images are acquired.

[0073] In the embodiments of the present application, a model is used to realize automatic prediction of the pixel size. First, a plurality of first samples are acquired, wherein each first sample includes an image and a pixel size of the image. The pixel size here can be obtained through manual labeling or automatically determined through other means.

[0074] The embodiments of the present application do not limit the type of the image in the first sample. For example, in a satellite monitoring scenario, the image in the first sample can be a satellite image; in a clinical medicine scenario, the image in the first sample can be a pathological image. The source of the image in the first sample is also not limited, for example, it can be obtained by shooting with a mobile phone, or by computer screen capture, or it can be an image observed through a microscope (image acquisition can be performed with the aid of a microscope camera).

[0075] ​In some embodiments, the image in the first sample includes a virtual scale; the above-mentioned obtaining a plurality of first samples including images and pixel sizes of the images can also be achieved in the following manner: obtaining a plurality of first samples including images, wherein the image in the first sample includes a virtual scale; performing the following processing on the image in each first sample: determining the length of any one part of the virtual scale included in the image; determining the number of pixels corresponding to the any one part; dividing the length by the number of pixels to obtain the pixel size of the image.

[0076] In the case where the image in the first sample includes a virtual scale, the pixel size of the image can be automatically determined based on the virtual scale, wherein the virtual scale can be obtained by shooting an actually existing scale card. For the image in any one first sample, the length of any one part (such as between any two scales) of the virtual scale included in the image can be determined, and the number of pixels corresponding to the part in the image can also be determined, wherein the number of pixels is counted according to the direction of the virtual scale and the pixels covered by the part.

[0077] For example, the direction of the virtual scale in the image is the horizontal direction, and the determined any one part is the part from the 30th scale to the 40th scale in the virtual scale, wherein the 30th scale represents 3 mm, and the horizontal coordinate (i.e., the abscissa) of the 30th scale in the image is 300 (i.e., the 300th pixel); the 40th scale represents 4 mm, and the horizontal coordinate of the 40th scale in the image is 800. The length of the part from the 30th scale to the 40th scale is 1 mm, the number of pixels corresponding to the part is 800-300=500, and thus the pixel size of the image can be determined as 1 / 500=0.002 mm / pixel=2.00 MPP. Through the above manner, the automatic determination of the pixel size can be achieved based on the virtual scale, and the labor cost can be effectively saved.

[0078] In step 102, the plurality of first samples are expanded according to a set pixel size range to obtain a plurality of second samples covering the pixel size range.

[0079] After obtaining the plurality of first samples, the plurality of first samples are expanded according to a set pixel size range to obtain a plurality of second samples covering each pixel size in the pixel size range, wherein the pixel size range can be set according to the actual application scenario. Here, the accuracy of the pixel size can be obtained to determine the plurality of pixel sizes included in the pixel size range, and the accuracy of the pixel size can also be set in advance. For example, the accuracy of the pixel size is 0.01 MPP, and the pixel size range is 0.10 MPP to 2.00 MPP, i.e., [0.10 MPP, 2.00 MPP], and the pixel size range includes 191 pixel sizes.

[0080] The embodiment of the present application does not limit the manner of expansion processing, which may be scaling processing, for example. For each first sample, a scaling ratio between each pixel size in the pixel size range and the pixel size in the first sample can be determined respectively, the image in the first sample is scaled according to the obtained scaling ratio, and then a second sample is constructed according to the image obtained by scaling and the corresponding pixel size (here, the pixel size in the pixel size range). In this way, for one first sample, multiple second samples covering the pixel size range completely can be obtained. Of course, this is only an example, and expansion processing can also be performed in other manners.

[0081] For example, the image size of the image in a certain first sample is 200 pixels*200 pixels, the pixel size of the image is 1.00MPP, if a certain pixel size in the pixel size range is 2.00MPP, the image in the first sample is scaled according to the scaling ratio 2 (i.e., the image size is reduced to 1 / 2 of the original), and a second sample is constructed according to the image with the image size of 100 pixels*100 pixels obtained by scaling and the pixel size of 2.00MPP. If another pixel size in the pixel size range is 0.50MPP, the image in the first sample is scaled according to the scaling ratio 2 (i.e., the image size is enlarged to 2 times of the original), and a second sample is constructed according to the image with the image size of 400 pixels*400 pixels obtained by scaling and the pixel size of 0.50MPP.

[0082] It is worth noting that the scaling ratio K of the image involved in the embodiment of the present application refers to the image size of the image being enlarged to K times of the original, and the scaling ratio K of the image refers to the image size of the image being reduced to 1 / K of the original. Wherein, K is a number greater than 0.

[0083] In step 103, the size prediction model is trained according to the multiple second samples.

[0084] After obtaining the multiple second samples through expansion processing, all the second samples are used as training samples of the size prediction model, i.e., added to the training sample set, so as to train the size prediction model according to the training sample set. The type of the size prediction model is not limited by the embodiment of the present application, which may be a deep learning model, such as an InceptionV3 model or a model with a full convolution structure.

[0085] It is worth noting that the training sample set of the size prediction model can only include the second sample, or can also include the first sample and the second sample, and the latter case can increase the number of training samples.

[0086] In step 104, the size prediction model is trained according to the training sample set.

[0087] In the training process of the size prediction model, if a stop condition is met, the training is stopped. The stop condition can be set according to actual application scenarios, for example, a set number of iterations or a threshold of an index, where the index can be at least one of precision, recall, and F1 index, and of course can also be other indexes, where the F1 index is the harmonic mean of precision and recall.

[0088] After the training of the size prediction model is completed, the pixel size of the to-be-tested image is obtained by performing size prediction processing on the to-be-tested image according to the trained size prediction model.

[0089] In step 105, target recognition processing is performed on the to-be-tested image according to the pixel size obtained by the size prediction processing.

[0090] For example, the to-be-tested image can be scaled according to a scaling ratio between the set pixel size and the pixel size of the to-be-tested image, and target recognition processing is performed on the scaled to-be-tested image according to a target recognition model (such as a target recognition model trained according to an image conforming to the set pixel size), to obtain a recognition result. In this way, the pixel size of the to-be-tested image can be referred to in the target recognition processing process, and an accurate recognition result can be obtained.

[0091] The target to be recognized by the target recognition processing is not limited by the embodiments of the present application. The target can be a single type, such as a face, or can include multiple types. For example, in a clinical medicine scenario, the purpose of the target recognition processing is to recognize tumor cells (i.e., targets) in a pathological image (to-be-tested image); in a satellite monitoring scenario, the purpose of the target recognition processing is to recognize human bodies, vehicles, and buildings, etc. in a satellite image (to-be-tested image), i.e., human bodies, vehicles, and buildings, etc. are all targets to be recognized.

[0092] As shown in Figure 3A , the embodiments of the present application train the size prediction model according to multiple second samples that can cover a range of pixel sizes, which can improve the model training effect, i.e., improve the accuracy of the size prediction processing performed according to the trained size prediction model; the pixel size of the to-be-tested image is used as reference information in the target recognition processing process, which can improve the accuracy of the final recognition result, so that the computing resources consumed by the electronic device when performing image processing will not be wasted. The embodiments of the present application can be applied to various application scenarios, such as a face recognition scenario, a satellite monitoring scenario, and a clinical medicine scenario.

[0093] In some embodiments, referring to Figure 3B , Figure 3B is a flowchart of an image processing method based on artificial intelligence provided by the embodiments of the present application,Figure 3A The step 102 shown can be implemented by steps 201 to 202, which will be described in combination with the steps.

[0094] In step 201, the image in the first sample is scaled according to the pixel size range to obtain an augmented image.

[0095] For each obtained first sample, the image in the first sample can be scaled according to the pixel size range, and for the sake of distinction, the image obtained by scaling here is named as an augmented image. Through step 201, the augmentation at the image level can be realized.

[0096] In some embodiments, the scaling of the image in the first sample according to the pixel size range to obtain the augmented image can be implemented in the following manner: the pixel size range is divided into a plurality of pixel size sub-ranges; for each pixel size sub-range, random selection processing is performed on the plurality of pixel sizes included in the pixel size sub-range, and the image in the first sample is scaled according to the scaling ratio between the pixel size obtained by the random selection processing and the pixel size in the first sample to obtain the augmented image.

[0097] Here, the pixel size range can be divided into N pixel size sub-ranges, and the division manner is not limited in the embodiments of the present application, which can be average or uneven division, wherein N is an integer greater than 1. It is worth noting that for different first samples, the division manner of the pixel size range is the same.

[0098] For each pixel size sub-range divided, random selection processing is performed in the pixel size sub-range, and the scaling ratio between the pixel size obtained by the random selection processing and the pixel size in the first sample is determined. Then, the image in the first sample is scaled according to the scaling ratio to obtain the augmented image. In this way, for the image in a certain first sample, N corresponding augmented images can be obtained.

[0099] In some embodiments, after the scaling of the image in the first sample according to the scaling ratio between the pixel size obtained by the random selection processing and the pixel size in the first sample, the pixel size obtained by the random selection processing is further used as the pixel size of the augmented image.

[0100] After the pixel size obtained by the random selection processing is determined as the expanded image corresponding to the image in the first sample, the pixel size obtained by the random selection processing is taken as the pixel size of the expanded image. For example, a certain pixel size sub-range is [1.80MPP, 2.00MPP], the pixel size obtained by the random selection processing on the pixel size sub-range is 2.00MPP, the image size of the image in the first sample is 200 pixels*200 pixels, and the pixel size of the image is 1.00MPP, and then the image is processed by the scaling processing according to the scaling ratio 2 to obtain an expanded image with a size of 100 pixels*100 pixels, and the pixel size of the expanded image is the pixel size obtained by the random selection processing, that is, 2.00MPP.

[0101] In the foregoing manner, the distribution of the pixel sizes of all the expanded images (referring to the expanded images corresponding to all the images in the first sample) obtained can approximately satisfy the random distribution in the pixel size range, that is, the number of the expanded images corresponding to each pixel size in the pixel size range is basically consistent. Furthermore, the training effect on the size prediction model can be improved by the obtained expanded images and the pixel sizes of the expanded images.

[0102] In some embodiments, before the random selection processing is performed on the plurality of pixel sizes included in the pixel size sub-range, the method further includes: obtaining the precision of the pixel size to determine the plurality of pixel sizes included in the pixel size sub-range.

[0103] Here, the plurality of pixel sizes included in the pixel size sub-range can be determined according to the precision of the pixel size, and then the random selection processing is performed on the plurality of pixel sizes included in the pixel size sub-range, and then the expanded image is determined according to the pixel size obtained by the random selection processing. The precision of the pixel size can be set according to the actual application scenario, for example, 0.01MPP.

[0104] In step 202, the second sample is constructed according to the expanded image and the pixel size of the expanded image.

[0105] After the expanded image is obtained by the scaling processing, the second sample is constructed according to the expanded image and the pixel size of the expanded image. Finally, all the obtained second samples can completely cover the pixel size range.

[0106] As shown in Figure 3B , the embodiments of the present application realize effective expansion of the sample by the scaling processing, and thus the training effect on the size prediction model can be improved.

[0107] In some embodiments, referring to Figure 3C , Figure 3C is a flowchart of an image processing method based on artificial intelligence provided by the embodiments of the present application, Figure 3AStep 103 shown can be implemented through steps 301 to 303 , which will be described in conjunction with each step.

[0108] In step 301, the image in the second sample is cropped according to the input image size corresponding to the size prediction model to obtain a cropped image.

[0109] Here, the size prediction model is trained based on the training samples in the training sample set. During the training process, the size prediction model performs size prediction processing (i.e., forward propagation) on the images in the training samples to obtain predicted pixel sizes. For ease of distinction, the predicted pixel sizes are referred to as the pixel sizes to be compared. The training sample set may include only the second sample or both the first and second samples. For ease of understanding, the following description assumes only the second sample.

[0110] If the size prediction model requires an input image size, the image in the second sample is cropped according to the input image size to obtain a cropped image. The cropping process can be performed at a set position or a random position, and the set position can be, for example, the center, upper left corner, lower left corner, upper right corner, or lower right corner of the image, but is not limited thereto.

[0111] In some embodiments, after step 301 , the method further includes: when the image size of the cropped image does not conform to the input image size, padding the cropped image so that the image size of the cropped image after padding conforms to the input image size.

[0112] After the image in the second sample is cropped, the image size of the resulting cropped image may be smaller than the input image size, that is, it may not conform to the input image size. For example, the input image size is 299 pixels * 299 pixels, and the image size of the cropped image obtained by the cropping process is 200 pixels * 200 pixels. In this case, new pixels are filled in the cropped image so that the image size of the cropped image after the filling process conforms to (is equal to) the input image size. In order to avoid the new pixels filled in from having an adverse effect on the content of the cropped image itself (such as confusion with the content of the cropped image itself), the pixels can be uniformly filled with black color. In this way, it can be ensured that the image size of the cropped image of the input size prediction model conforms to the input image size.

[0113] In step 302, size prediction processing is performed on the cropped image according to the size prediction model to obtain the pixel size to be compared.

[0114] After the image in the second sample is cropped to obtain a cropped image, a size prediction process is performed on the cropped image according to the size prediction model to obtain a pixel size to be compared.

[0115] In step 303, according to the difference between the pixel size to be compared and the pixel size in the second sample, the back propagation is performed in the size prediction model, and the weight parameters of the size prediction model are updated during the back propagation.

[0116] After obtaining the pixel size to be compared through a certain second sample, according to the loss function of the size prediction model, the difference (loss value) between the pixel size to be compared and the pixel size in the second sample is determined, and the back propagation is performed in the size prediction model according to the difference, wherein the type of the loss function is not limited. During the back propagation, the weight parameters of the size prediction model are updated in the gradient descent direction.

[0117] In Figure 3C , the step 104 shown can be implemented by steps 304 to 306, which will be described in combination with each step. Figure 3A

[0118] In step 304, the test image is cropped multiple times according to the input image size to obtain multiple test cropped images.

[0119] After the training of the size prediction model is completed, the test image can be cropped at a set position or a random position to obtain a test cropped image. Then, the test cropped image is subjected to size prediction processing through the trained size prediction model to obtain the pixel size of the test cropped image, and the pixel size of the test cropped image is taken as the pixel size of the test image.

[0120] In order to further improve the accuracy of size prediction, the test image can be cropped multiple times according to the input image size to obtain multiple test cropped images.

[0121] In some embodiments, after step 304, it further includes: when the image size of the test cropped image does not conform to the input image size, the test cropped image is subjected to padding processing so that the image size of the test cropped image after the padding processing conforms to the input image size.

[0122] Similarly, when the image size of the test cropped image is smaller than the input image size of the size prediction model, the test cropped image can also be subjected to padding processing so that the image size of the test cropped image after the padding processing is equal to the input image size. In this way, the test cropped image after the padding processing can be subjected to size prediction processing according to the trained size prediction model.

[0123] In step 305, the test cropped image is subjected to size prediction processing according to the trained size prediction model to obtain the pixel size of the test cropped image.

[0124] ​For each to-be-tested cropped image, the size prediction model is used to perform size prediction on the to-be-tested cropped image, to obtain the pixel size of the to-be-tested cropped image.

[0125] In step 306, the pixel sizes of the plurality of to-be-tested cropped images are averaged to obtain the pixel size of the to-be-tested image.

[0126] Here, the pixel sizes of all to-be-tested cropped images are averaged, and the result of the average processing is taken as the pixel size of the to-be-tested image. In this way, by fusing the size prediction results (i.e., the predicted pixel sizes) of different parts (i.e., different to-be-tested cropped images) in the to-be-tested image, the accuracy of the pixel size of the to-be-tested image obtained finally can be improved.

[0127] As shown in Figure 3C , the embodiments of the present application can effectively and accurately update the weight parameters of the size prediction model based on the mechanism of back propagation; the plurality of to-be-tested cropped images are obtained by performing multiple cropping processing on the to-be-tested image, and the pixel sizes of the plurality of to-be-tested cropped images are averaged, which can improve the accuracy of the pixel size of the to-be-tested image obtained finally.

[0128] In some embodiments, referring to Figure 3D , Figure 3D is a flowchart of an image processing method based on artificial intelligence provided by the embodiments of the present application, based on Figure 3A , before step 103, a plurality of image classification samples including images and categories of the images can also be obtained in step 401.

[0129] In the embodiments of the present application, in order to improve the model training effect of the training sample set of the size prediction model, the size prediction model can be pre-trained. In the pre-training process, a plurality of image classification samples can be obtained, each image classification sample including an image and a category of the image. In order to improve the pre-training effect, a plurality of image classification samples including different categories can be obtained.

[0130] In step 402, the image classification model is trained according to the plurality of image classification samples; wherein the prediction task of the image classification model is a multi-category prediction task.

[0131] According to the plurality of image classification samples obtained, the image classification model is trained, wherein the prediction task of the image classification model is a multi-category prediction task, i.e., predicting a category to which an image belongs from a plurality of categories (such as human, cat, dog, etc., which are not limited).

[0132] Taking any one image classification sample as an example, the training process of the image classification model is described. First, according to the image classification model, the image in the image classification sample is subjected to image classification processing (that is, a multi-class prediction task is performed), to obtain a confidence corresponding to each of a plurality of categories, and then the category with the highest confidence is determined as the category of the image. In order to facilitate distinction, the category determined here is named as a to-be-compared category. Then, according to the difference between the to-be-compared category and the category in the image classification sample, back propagation is performed in the image classification model, and the weight parameters of the image classification model are updated in the process of back propagation.

[0133] In step 403, the prediction task of the trained image classification model is updated to a single-class prediction task to obtain a size prediction model; wherein the prediction target of the single-class prediction task is a pixel size.

[0134] The training process of the image classification model is a pre-training process of the size prediction model. After the training of the image classification model is completed, the prediction task of the trained image classification model is updated from the multi-class prediction task to the single-class prediction task, so that the trained image classification model can be used as the size prediction model. The prediction target of the single-class prediction task is a pixel size, and the size prediction processing performed by the size prediction model is equivalent to performing a single-class prediction task.

[0135] As shown in Figure 3D , since the processing objects of the image classification model and the size prediction model are both images, there is a correlation, so the trained image classification model can be used as an initialized size prediction model (which involves updating of the prediction task), thereby improving the training effect and training efficiency of the size prediction model, that is, accelerating optimization.

[0136] In some embodiments, referring to Figure 3E , Figure 3E is a flowchart of an image processing method based on artificial intelligence provided by the embodiments of the present application, Figure 3A The step 105 shown in the figure can be implemented by steps 501 to 504, which will be described in conjunction with each step.

[0137] In step 501, a plurality of target recognition samples including an image, a position of a target in the image, and a category of the target in the image are obtained; wherein the pixel size of the image in the target recognition sample is a set pixel size.

[0138] In the embodiment of the present application, the target recognition model can be used to implement the target recognition processing. The type of the target recognition model is not limited herein, for example, it can be a region-based convolutional neural network (R-CNN) model. Before the target recognition model is used to perform the target recognition processing on the to-be-tested image, the target recognition model is trained. First, a plurality of target recognition samples are obtained, wherein each target recognition sample includes an image, the positions of the targets (the number of targets can be one or more) in the image, and the categories of the targets in the image. It should be noted that the pixel size of the image in all target recognition samples is the set pixel size, so that a target recognition model suitable for the set pixel size can be trained.

[0139] Herein, the target recognition samples can be obtained according to the actual application scenario. For example, in the face recognition scenario, the image in the obtained target recognition sample can include a human face; in the clinical medicine scenario, the image (such as a pathological image) in the obtained target recognition sample can include a tumor cell; and in the satellite monitoring scenario, the image (such as a satellite image) in the obtained target recognition sample can include at least one of a human body, a vehicle, and a building.

[0140] In step 502, the target recognition model is trained according to the plurality of target recognition samples.

[0141] Herein, the target recognition model is trained according to the plurality of target recognition samples, i.e., the weight parameters of the target recognition model are updated, in combination with the mechanism of back propagation.

[0142] In step 503, the to-be-tested image is scaled according to the scaling ratio between the set pixel size and the pixel size of the to-be-tested image.

[0143] After the size prediction model is trained, the to-be-tested image is scaled according to the scaling ratio between the set pixel size and the pixel size of the to-be-tested image. For example, the set pixel size is 0.20MPP, the image size of the to-be-tested image is 200 pixels*200 pixels, and the pixel size of the to-be-tested image is 0.10MPP. Then, the to-be-tested image is scaled according to the scaling ratio of 2, i.e., the image size of the to-be-tested image is reduced to 1 / 2 of the original size, to obtain an image with an image size of 100 pixels*100 pixels and a pixel size of 0.20MPP.

[0144] In step 504, the target recognition model is trained, and the target recognition processing is performed on the scaled to-be-tested image to obtain the positions of the targets in the scaled to-be-tested image and the categories of the targets.

[0145] The pixel size of the to-be-measured image after the scaling processing is the set pixel size, and therefore, the target recognition model trained for the set pixel size can be used to perform target recognition processing on the to-be-measured image after the scaling processing, to obtain the position of the target in the to-be-measured image after the scaling processing and the category of the target.

[0146] It should be noted that the target recognition model can also have requirements on the size of the input image, and therefore, in the training process in step 502, the images in the target recognition samples can be cropped, and the cropped images can be input into the target recognition model; similarly, in step 504, the to-be-measured image after the scaling processing can be cropped, and the cropped image can be input into the target recognition model.

[0147] As shown in Figure 3E , the embodiments of the present application train the target recognition model applicable to the set pixel size, to perform target recognition processing on the to-be-measured image with the pixel size conforming to the set pixel size, which can further improve the accuracy of the target recognition processing.

[0148] In some embodiments, referring to Figure 3F , Figure 3F is a flowchart of an image processing method based on artificial intelligence provided by the embodiments of the present application, Figure 3A The step 103 shown in the figure can be implemented by steps 601 to 603, which will be described in combination with the steps.

[0149] In step 601, at least one of flip processing, crop processing, and color jittering processing is performed on the images in the second samples to obtain enhanced images.

[0150] In the embodiments of the present application, in order to strengthen the training effect of the size prediction model, the mechanism of data enhancement can be used to further expand the number of training samples for training the size prediction model. Here, at least one of flip (Flip) processing, crop (Crop) processing, and color jittering (Color Jittering) processing can be performed on the images in the second samples to obtain enhanced images, wherein the flip processing can be random flipping in any direction, the crop processing can be cropping at a random position, and the color jittering processing can be multiplying the pixel value (such as the pixel value of the red channel, the green channel, and the blue channel) of each pixel in the image by a set value. Of course, the above is only an example of flip processing, crop processing, and color jittering processing, and does not constitute a limitation on the embodiments of the present application. In addition, other data enhancement methods, such as rotation processing or adding Gaussian noise, can also be used to obtain enhanced images.

[0151] It is worth mentioning that for an image in a second sample, step 601 can be performed multiple times to obtain multiple enhanced images, and the number of enhanced images to be obtained can be set according to an actual application scenario.

[0152] In step 602, a third sample is constructed according to the enhanced image and the pixel size in the second sample.

[0153] After obtaining the enhanced image according to the image in the second sample, a third sample is constructed according to the enhanced image and the pixel size in the second sample.

[0154] In some embodiments, at least one of the flipping processing, the cropping processing and the color disturbance processing can also be performed on the image in the first sample to obtain an enhanced image, and then a third sample is constructed according to the enhanced image and the pixel size in the first sample.

[0155] In step 603, the size prediction model is trained according to multiple third samples.

[0156] In the embodiments of the present application, the training sample set used to train the size prediction model can include at least one of the first sample, the second sample and the third sample. In the case where the training sample set only includes all the third samples, the size prediction model is trained according to all the third samples, and the training process here is similar to the training process of the second sample described above.

[0157] As shown in FIG. 6, the embodiments of the present application further expand the samples based on the mechanism of data enhancement, which helps to improve the generalization ability of the size prediction model in the training process. Figure 3F

[0158] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described. In order to facilitate understanding, the scenario of clinical medicine (i.e. pathological analysis) is taken as an example. In the embodiments of the present application, an image observed by a microscope can be captured by a microscope camera, and target recognition processing can be performed to obtain a target (such as the position of a diseased area or a diseased cell) in the image, wherein the microscope can be used to observe an ex vivo tissue section (such as a tissue section of a human body or other organisms), and the finally recognized target cannot be directly or indirectly used for disease diagnosis. In the recognition process, the pixel size is very important information. For example, in an image with a pixel size of 0.2MPP, the size of a normal cell is about 10 pixels, and the size of a tumor cell is 30 pixels or more. Therefore, the pixel size is an important reference information for target recognition processing, which can affect the accuracy of target recognition.

[0159] Figure 4 ​​As shown, the embodiments of the present application realize the automatic determination of the pixel size of the image through two stages of training stage and prediction stage, and can guarantee the accuracy of the determined pixel size, which will be described respectively.

[0160] 1) Training stage.

[0161] ① Sample collection.

[0162] The microscope camera is used to shoot the pathological images observed by the microscope at different magnifications, and the pixel size of each pathological image is determined, such as Figure 5 As shown, the pathological images observed by the microscope at 10 times magnification (10 times objective lens), 20 times magnification and 40 times magnification can be obtained respectively, so as to obtain the pixel size corresponding to each magnification.

[0163] It is worth noting that the pathological images obtained here include a virtual ruler. For example, a ruler card can be fixed on the slide, and a pathological tissue section can also be fixed at the same time. In this way, after observing the slide through the microscope, the obtained pathological image includes not only the image corresponding to the pathological tissue section, but also the virtual ruler.

[0164] As an example, the embodiments of the present application provide a schematic diagram of a pathological image as shown in Figure 6 As shown in the pathological image 61 shown in Figure 6 The pixel size of the pathological image 61 can be determined according to the virtual ruler, for example, the length of any part of the virtual ruler included in the pathological image 61 can be determined, the number of pixels corresponding to the part is also determined, and the length is divided by the number of pixels to obtain the pixel size of the pathological image 61. For example, Figure 6 Taking the virtual ruler in the horizontal direction as an example, the pixel position of the 30 scale (representing 3 mm) of the virtual ruler is 650, and the pixel position of the 70 scale is 2600. Then the number of pixels corresponding to the total length of 4 mm is 2600-650=1950, and the pixel size of each pixel in the pathological image 61 can be determined as 4 / 1950≈0.00205 mm / pixel=2.05MPP.

[0165] In order to facilitate understanding, the pathological images collected here are combined with the obtained pixel size, which is named as the first sample, that is, the first sample includes a pathological image and the pixel size of the pathological image. Here, taking 10000 pathological images (i.e. forming 10000 first samples) as an example, the 10000 pathological images include three types, and the pixel sizes are 0.21MPP, 0.42MPP and 0.85MPP respectively.

[0166] ② Randomly generated.

[0167] Here, taking the pixel size range to be predicted as an example, which is 0.10MPP to 2.00MPP with a precision of 0.01MPP, after obtaining a plurality of first samples through step ①, a batch of new samples is generated by a random generation method, so that the batch of new samples can completely cover the above-mentioned pixel size range and be distributed randomly. In order to facilitate understanding, the generated new samples are named as second samples, wherein the random distribution means that the number of second samples corresponding to each pixel size in the pixel size range is basically consistent.

[0168] One way of random generation is to divide the pixel size range into 10 intervals (corresponding to the pixel size sub-range in the above), for each first sample, randomly select in each interval, and determine the scaling ratio between the randomly selected pixel size and the pixel size in the first sample, and then scale the pathological image in the first sample according to the scaling ratio, and construct a second sample according to the scaled pathological image and the randomly selected pixel size. Wherein, when dividing the pixel size range into 10 intervals, it can be uniformly divided, or it can be unevenly divided according to the actual application scene.

[0169] For example, a first sample includes a pathological image with a pixel size of 0.21MPP, and a certain interval divided from the pixel size range is [0.60MPP, 0.80MPP], and the randomly selected pixel size in the interval is 0.63MPP, then according to the scaling ratio of 0.63 / 0.21=3, the pathological image in the first sample is scaled down (i.e. the image size is reduced to 1 / 3 of the original), and a second sample is constructed according to the scaled pathological image and the pixel size 0.63MPP.

[0170] Here, taking 100000 pathological images randomly generated (i.e. forming 100000 second samples) as an example, wherein for each pixel size in the pixel size range of [0.10MPP, 2.00MPP], about 524 pathological images correspond to each pixel size, and 191 is the number of pixel sizes in the pixel size range.

[0171] ③Training the size prediction model.

[0172] After obtaining a plurality of second samples through step ②, at least one of random flipping, random cropping and color disturbance is performed on the pathological image in each second sample, i.e. data augmentation is performed to obtain an augmented image, and then a third sample is constructed according to the augmented image and the pixel size in the second sample. For each second sample, a plurality of corresponding third samples can be constructed through the mechanism of data augmentation.

[0173] The third sample is used to train a size prediction model, the input of the size prediction model is an image in the third sample, and the output is a predicted pixel size (corresponding to the pixel size to be compared in the foregoing). The type of the size prediction model is not limited in the embodiments of the present application. For example, the size prediction model can be a deep learning model, such as an InceptionV3 model, the number of output categories of the InceptionV3 model is 1, that is, the InceptionV3 model is used to perform a single-class prediction task, or the size prediction model can be a model with a full convolution structure, and the output (that is, the predicted pixel size) of the model with the full convolution structure is the average value of the last layer feature map.

[0174] Here, the size prediction model can be pre-trained by using a transfer learning mechanism. For example, an image classification model used to perform a multi-class prediction task is trained by using an open-source ImageNet dataset (corresponding to the image classification sample in the foregoing). Then, the number of output categories of the trained image classification model is adjusted to 1 (that is, the prediction task is updated from a multi-class prediction task to a single-class prediction task), and the size prediction model is obtained.

[0175] When the size prediction model is trained by using the third sample, a root mean square propagation (RMSprop) algorithm can be used, the input image size can be 299 pixels*299 pixels, the batch processing size can be 64, the initial learning rate can be 0.0001, the maximum number of iterations can be 10,000, and the loss function can be a mean squared error (MSE) loss function. Of course, this does not constitute a limitation on the embodiments of the present application. It should be noted that the size prediction model can have a requirement for the input image size, and therefore, an image in the third sample can be cropped to obtain a cropped image, and then the cropped image is input into the size prediction model. The cropping processing involved in the embodiments of the present application can be cropping at a fixed position (such as a center position) or a random position. In addition, when the image size of the obtained cropped image is smaller than the input image size, the cropped image can be padded, for example, black pixels are filled in the cropped image until the image size of the cropped image is equal to the input image size.

[0176] 2) Prediction phase.

[0177] ① Pixel size prediction.

[0178] In the prediction process, for the input to-be-measured pathological image (corresponding to the to-be-measured image in the foregoing), first, according to the input image size of the size prediction model, the to-be-measured pathological image is cropped to obtain a to-be-measured cropped image, and then the to-be-measured cropped image is input into the size prediction model, and the output of the size prediction model is the pixel size of the to-be-measured cropped image predicted. Then, the pixel size of the to-be-measured cropped image is taken as the pixel size of the to-be-measured pathological image.

[0179] It is worth noting that when the to-be-measured pathological image is cropped, the cropping can be performed at fixed positions (such as the center and four corners of the to-be-measured pathological image) or random positions to obtain a plurality of to-be-measured cropped images. After obtaining the pixel size of each to-be-measured cropped image through the size prediction model, the pixel sizes of all to-be-measured cropped images are averaged to obtain the pixel size of the to-be-measured pathological image, so that the accuracy of the predicted pixel size can be further improved.

[0180] The embodiments of the present application can provide services in the form of software interfaces, as shown in Figure 7 , Figure 8 and Figure 9 , the input of the software service is a to-be-measured pathological image, such as the pathological image 71 in Figure 7 , the pathological image 81 in Figure 8 and the pathological image 91 in Figure 9 ; the output of the software service is the pixel size of the to-be-measured pathological image predicted.

[0181] ②Automatic identification of pathological information.

[0182] In the embodiments of the present application, the automatic identification of pathological information can be realized by a target recognition model, wherein the images in the samples used to train the target recognition model are of a set pixel size (such as 0.42MPP). When the pixel size of the to-be-measured pathological image is obtained, the to-be-measured pathological image is scaled according to the scaling ratio between the set pixel size and the pixel size of the to-be-measured pathological image, so that the pixel size of the scaled to-be-measured pathological image is equal to the set pixel size. Then, the scaled to-be-measured pathological image is input into the target recognition model, and the output of the target recognition model is the pathological information in the to-be-measured pathological image, such as the position of tumor cells (such as cancer cells).

[0183] The embodiments of the present application can achieve the following technical effects: 1) strong applicability, not dependent on hardware environment, can process various types of pathological images, such as pathological images obtained by different combinations of microscopes and microscope cameras, such as pathological images obtained by screen capture of a computer screen, such as pathological images obtained by shooting a computer screen with a mobile phone; 2) can improve the accuracy of the predicted pixel size, provide accurate reference information for automatic identification of pathological information, and thus provide effective and reliable research basis for doctors or other researchers; 3) can effectively save labor costs.

[0184] The following continues to illustrate the exemplary structure of the image processing apparatus 455 based on artificial intelligence provided by the embodiments of the present application implemented as a software module. In some embodiments, as shown in Figure 2 The software module stored in the image processing apparatus 455 based on artificial intelligence in the memory 450 can include: an acquisition module 4551 configured to acquire a plurality of first samples including images and pixel sizes of the images; an expansion module 4552 configured to perform expansion processing on the plurality of first samples according to a set pixel size range, to obtain a plurality of second samples covering the pixel size range; a training module 4553 configured to train a size prediction model according to the plurality of second samples; a size prediction module 4554 configured to perform size prediction processing on a to-be-tested image according to the trained size prediction model; and a target identification module 4555 configured to perform target identification processing on the to-be-tested image according to the pixel size obtained by the size prediction processing.

[0185] In some embodiments, the expansion module 4552 is further configured to perform the following processing for each first sample: performing scaling processing on the image in the first sample according to the pixel size range, to obtain an expanded image; and constructing the second sample according to the expanded image and the pixel size of the expanded image.

[0186] In some embodiments, the expansion module 4552 is further configured to: divide the pixel size range into a plurality of pixel size sub-ranges; for each pixel size sub-range, perform random selection processing in the plurality of pixel sizes included in the pixel size sub-range, and perform scaling processing on the image in the first sample according to a scaling ratio between the pixel size obtained by the random selection processing and the pixel size in the first sample, to obtain the expanded image.

[0187] In some embodiments, the expansion module 4552 is further configured to: take the pixel size obtained by the random selection processing as the pixel size of the expanded image.

[0188] In some embodiments, the expansion module 4552 is further configured to: acquire the accuracy of the pixel size, to determine the plurality of pixel sizes included in the pixel size sub-range.

[0189] In some embodiments, the training module 4553 is further configured to: perform size prediction on the image in the second sample according to the size prediction model to obtain a pixel size to be compared; and perform back propagation in the size prediction model according to a difference between the pixel size to be compared and a pixel size in the second sample, and update a weight parameter of the size prediction model in the process of back propagation.

[0190] In some embodiments, the training module 4553 is further configured to: perform cropping processing on the image in the second sample to obtain a cropped image according to an input image size corresponding to the size prediction model; and perform size prediction on the cropped image according to the size prediction model.

[0191] In some embodiments, the size prediction module 4554 is further configured to: perform multiple times of cropping processing on the to-be-tested image according to the input image size to obtain multiple to-be-tested cropped images; perform size prediction on the to-be-tested cropped images according to the trained size prediction model to obtain pixel sizes of the to-be-tested cropped images; and perform average processing on the pixel sizes of the multiple to-be-tested cropped images to obtain a pixel size of the to-be-tested image.

[0192] In some embodiments, the training module 4553 is further configured to: when the image size of the cropped image does not conform to the input image size, perform padding processing on the cropped image, so that the image size of the cropped image after the padding processing conforms to the input image size.

[0193] In some embodiments, the image processing apparatus based on artificial intelligence 455 further includes a pre-training module configured to: obtain multiple image classification samples including images and categories of the images; train an image classification model according to the multiple image classification samples; wherein a prediction task of the image classification model is a multi-category prediction task; and update the prediction task of the trained image classification model to a single-category prediction task to obtain a size prediction model; wherein a prediction target of the single-category prediction task is a pixel size.

[0194] In some embodiments, the target identification module 4555 is further configured to: obtain multiple target identification samples including images, positions of targets in the images, and categories of the targets in the images; wherein a pixel size of the image in the target identification sample is a set pixel size; train a target identification model according to the multiple target identification samples; perform scaling processing on the to-be-tested image according to a scaling ratio between the set pixel size and a pixel size of the to-be-tested image; and perform target identification processing on the to-be-tested image after the scaling processing according to the trained target identification model to obtain the position of the target and the category of the target in the to-be-tested image after the scaling processing.

[0195] In some embodiments, the training module 4553 is further configured to perform the following processing for each second sample: performing at least one of a flipping processing, a cropping processing, and a color perturbation processing on the image in the second sample to obtain an enhanced image; constructing a third sample according to the enhanced image and a pixel size in the second sample; and training the size prediction model according to the third sample.

[0196] In some embodiments, the image in the first sample includes a virtual ruler; the artificial intelligence-based image processing apparatus 455 further includes: a length determination module configured to determine a length of any one part in the virtual ruler; a pixel number determination module configured to determine a pixel number corresponding to the any one part; and a pixel size determination module configured to divide the length by the pixel number to obtain the pixel size of the image.

[0197] The embodiments of the present application provide a computer program product or a computer program, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to enable the computer device to perform the artificial intelligence-based image processing method provided by the embodiments of the present application.

[0198] The embodiments of the present application provide a computer readable storage medium storing executable instructions, wherein the executable instructions, when executed by a processor, cause the processor to perform the method provided by the embodiments of the present application, for example, the artificial intelligence-based image processing method shown in Figure 3A 、 Figure 3B 、 Figure 3C 、 Figure 3D 、 Figure 3E and Figure 3F .

[0199] In some embodiments, the computer readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or various devices including one or any combination of the above memories.

[0200] In some embodiments, the executable instructions can be in the form of a program, software, software module, script, or code, written in any form of programming language (including a compiled or interpreted language, or a declarative or procedural language), and can be deployed in any form, including being deployed as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0201] By way of example, executable instructions can correspond to a file in a file system, but in many cases can not be so confined. For example, executable instructions can be stored in one or more files that are used jointly by one or more programs, stored in a single file that is part of a larger file, stored in a file that is common to both programs and data, stored in a file that is common to multiple programs and data, or stored in one or more executable files, such as one or more scripts or other executable files.

[0202] By way of example, executable instructions can be deployed across one computer system, or across multiple computer systems that are located at one site, or distributed across multiple sites and interconnected by a communication network.

[0203] The above merely provides example embodiments of the present application, but is not intended to limit the protection scope of the present application. Any modification, equivalent replacement and improvement made within the spirit and scope of the present application shall fall within the protection scope of the present application.

Claims

1. An artificial intelligence-based image processing method, characterized by, The method comprises: acquiring a plurality of first samples comprising images and pixel sizes of the images; performing expansion processing on the plurality of first samples according to a set pixel size range to obtain a plurality of second samples covering the pixel size range; training a size prediction model according to the plurality of second samples; performing size prediction processing on a plurality of to-be-tested cropped images according to the trained size prediction model to obtain a pixel size of each to-be-tested cropped image, wherein the plurality of to-be-tested cropped images are obtained by performing multiple cropping processing on a to-be-tested image according to an input image size corresponding to the size prediction model; performing average processing on the pixel sizes corresponding to the plurality of to-be-tested cropped images respectively to obtain a pixel size of the to-be-tested image; performing target identification processing on the to-be-tested image according to the pixel size.

2. The method of claim 1, wherein, The method comprises: performing the following processing for each first sample: performing scaling processing on the image in the first sample according to the pixel size range to obtain an expanded image; constructing a second sample according to the expanded image and a pixel size of the expanded image.

3. The method of claim 2, wherein, The method comprises: dividing the pixel size range into a plurality of pixel size sub-ranges; performing random selection processing in the pixel sizes included in each pixel size sub-range, and performing scaling processing on the image in the first sample according to a scaling ratio between the pixel size obtained by the random selection processing and the pixel size in the first sample to obtain an expanded image.

4. The method of claim 3, wherein, After performing scaling processing on the image in the first sample according to the scaling ratio between the pixel size obtained by the random selection processing and the pixel size in the first sample, the method further comprises: taking the pixel size obtained by the random selection processing as the pixel size of the expanded image. Before performing random selection processing in the pixel sizes included in the pixel size sub-ranges, the method further comprises: acquiring the precision of the pixel size to determine the pixel sizes included in the pixel size sub-ranges.

5. The method according to any one of claims 1 to 4, characterized in that, The method comprises: performing size prediction processing on the image in the second sample according to the size prediction model to obtain a to-be-compared pixel size; performing back propagation in the size prediction model according to the difference between the to-be-compared pixel size and the pixel size in the second sample, and updating the weight parameters of the size prediction model in the process of back propagation.

6. The method of claim 5, wherein, The method comprises: performing cropping processing on the image in the second sample according to an input image size corresponding to the size prediction model to obtain a cropped image; performing size prediction processing on the cropped image according to the size prediction model.

7. The method of claim 6, wherein, After the cropping image in the second sample is cropped, the method further comprises: When the image size of the cropped image does not conform to the input image size, the cropped image is padded to make the image size of the padded cropped image conform to the input image size.

8. The method according to any one of claims 1 to 4, characterized in that, Before the size prediction model is trained according to the plurality of second samples, the method further comprises: Obtaining a plurality of image classification samples comprising images and categories of the images; Training an image classification model according to a plurality of the image classification samples; wherein the prediction task of the image classification model is a multi-category prediction task; Updating the prediction task of the trained image classification model to a single-category prediction task to obtain a size prediction model; Wherein the prediction target of the single-category prediction task is a pixel size.

9. The method according to any one of claims 1 to 4, characterized in that, The target recognition processing of the to-be-tested image according to the pixel size comprises: Obtaining a plurality of target recognition samples comprising images, positions of targets in the images, and categories of the targets in the images; wherein the pixel size of the image in the target recognition sample is a set pixel size; Training a target recognition model according to a plurality of the target recognition samples; Scaling the to-be-tested image according to a scaling ratio between the set pixel size and the pixel size of the to-be-tested image; According to the trained target recognition model, the target recognition processing is performed on the to-be-tested image after scaling to obtain the position of the target and the category of the target in the to-be-tested image after scaling.

10. The method according to any one of claims 1 to 4, characterized in that, The training of the size prediction model according to the plurality of second samples comprises: For each of the second samples, the following processing is performed: At least one of the following processing is performed on the image in the second sample to obtain an enhanced image: flipping processing, cropping processing, and color perturbation processing; According to the enhanced image and the pixel size in the second sample, a third sample is constructed; The size prediction model is trained according to the third sample.

11. The method according to any one of claims 1 to 4, characterized in that, The image in the first sample comprises a virtual ruler; The method further comprises: Determining the length of any one part of the virtual ruler; Determining the number of pixels corresponding to the any one part; Dividing the length by the number of pixels to obtain the pixel size of the image.

12. An artificial intelligence-based image processing apparatus, characterized by comprising: The device comprises: An acquisition module configured to acquire a plurality of first samples comprising images and pixel sizes of the images; An expansion module configured to expand a plurality of the first samples according to a set pixel size range to obtain a plurality of second samples covering the pixel size range; A training module configured to train a size prediction model according to the plurality of second samples; A size prediction module configured to perform size prediction processing on a plurality of to-be-tested cropped images according to the trained size prediction model to obtain the pixel size of each of the to-be-tested cropped images, wherein the plurality of to-be-tested cropped images are obtained by cropping a to-be-tested image multiple times according to the input image size corresponding to the size prediction model. averaging the pixel sizes corresponding to the plurality of to-be-tested clipping images respectively to obtain a pixel size of the to-be-tested image; a target recognition module, configured to perform target recognition processing on the to-be-tested image according to the pixel size.

13. The apparatus of claim 12, wherein, the expansion module is further configured to perform, for each first sample, the following processing: performing scaling processing on an image in the first sample according to the pixel size range to obtain an expanded image; and constructing a second sample according to the expanded image and a pixel size of the expanded image.

14. The apparatus of claim 13, wherein, the expansion module is further configured to divide the pixel size range into a plurality of pixel size sub-ranges; perform, for each pixel size sub-range, random selection processing on the plurality of pixel sizes included in the pixel size sub-range; and perform scaling processing on an image in the first sample according to a scaling ratio between the pixel size obtained through the random selection processing and a pixel size in the first sample to obtain an expanded image.

15. The apparatus of claim 14, wherein, the expansion module is further configured to take the pixel size obtained through the random selection processing as a pixel size of the expanded image; and obtain a precision of the pixel size to determine the plurality of pixel sizes included in the pixel size sub-range.

16. An electronic device, comprising: including: a memory configured to store executable instructions; a processor configured to execute the executable instructions stored in the memory to implement the artificial intelligence-based image processing method of any one of claims 1 to 11.

17. A computer-readable storage medium, characterized in that, executable instructions stored in the memory, and configured to be executed by a processor to implement the artificial intelligence-based image processing method of any one of claims 1 to 11.

18. A computer program product comprising computer instructions, characterized in that, the computer instructions are executed by a processor to implement the artificial intelligence-based image processing method of any one of claims 1 to 11.

Citation Information

Patent Citations

  • Image processing method and device

    CN107578375A

  • Training sample image generation method and device

    CN110598785A