Patient assistance intelligent auditing system based on deep learning

Through the deep learning-based intelligent patient assistance audit system, the intelligent detection and identification of image and text information is used by YOLO3 and CRNN networks, the problem of low qualification audit efficiency in the patient assistance system is solved, and an efficient audit process and cost-reducing effect is achieved.

CN120388386APending Publication Date: 2025-07-29WUHAN WANMU HEALTH TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311800748.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The qualification review of the patient assistance system in the prior art is inefficient, the audit workload is large, and the efficiency cannot be improved by increasing manpower or simplifying procedures.

Method used

The intelligent patient assistance audit system based on deep learning is adopted, including image preprocessing module, intelligent detection module, intelligent recognition module and natural language analysis and processing module. The intelligent detection and recognition of image and text information is carried out through the YOLO3 algorithm and CRNN network, reducing the requirements for material format and clarity.

Benefits of technology

It improves the efficiency of patient assistance materials review, reduces the threshold for patient use and the work pressure of auditors, and reduces human resources costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388386A_ABST
    Figure CN120388386A_ABST
Patent Text Reader

Abstract

The invention provides a patient assistance intelligent auditing system based on deep learning, which comprises an image preprocessing module, an intelligent detection module, an intelligent identification module and a natural language analysis processing module, and is characterized in that the image preprocessing module is used for preprocessing an original image material uploaded by a patient; the intelligent detection module is used for performing intelligent detection on the corrected image through an intelligent detection model based on a YOLO3 algorithm to obtain a to-be-recognized area; the intelligent identification module is used for acquiring original text information of all rows in a to-be-identified area; and the natural language analysis processing module is used for verifying and correcting the original text information through a pattern matching algorithm. The auditing system performs intelligent auditing and key information detection and identification on the patient materials based on a deep learning method, so that the auditing efficiency of the patient assistance materials can be greatly improved, the dispensing process of patient assistance is accelerated, the workload of auditing personnel is reduced, and the human resource cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning technology, and in particular to a patient assistance intelligent review system based on deep learning. Background Art

[0002] Patient assistance refers to the provision of charitable medicines donated by charitable foundations and caring businesses. Through collaborative efforts, expensive medications are provided to eligible low-income patients, improving their quality of life and alleviating their financial burden. A crucial step in patient assistance is the qualification review of applicants.

[0003] The existing solution uses an internet-based patient-assisted online review system to complete the qualification review process. In this system, patients submit various paper documents in the form of scanned copies and photos. Online reviewers review these documents and, if any documents are missing or inconsistent with the requirements, notify the patient promptly to supplement or modify them, significantly reducing delays caused by mailing documents.

[0004] However, online reviewers in this existing technology face a massive amount of non-standard image data daily, creating an enormous workload. The average review cycle for each patient is long. This is partly because the cost of reviewing each patient is limited, making it impossible to simply increase staffing to improve review efficiency. Furthermore, given the seriousness and importance of patient assistance, simplifying review procedures is not an option. Deep learning technology can be used to address the inefficiencies in qualification review. Image recognition is a universal requirement, but its effectiveness in recognizing medical documents and non-standard certificates falls far short of practical requirements. Summary of the Invention

[0005] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a patient assistance intelligent review system based on deep learning to solve the problems raised in the above background technology. The present invention can adapt to the actual needs of the patient assistance system and reduce the requirements on the format and clarity of the materials uploaded by patients, thereby effectively lowering the threshold for patients to use.

[0006] To achieve the above object, the present invention is implemented by the following technical solutions: A patient assistance intelligent review system based on deep learning, the review system includes an image preprocessing module, an intelligent detection module, an intelligent recognition module, and a natural language analysis and processing module. The image preprocessing module is used to preprocess the original image materials uploaded by the patient. The preprocessing includes determining the type of the image materials and correcting the image and then sending it to the intelligent detection module. The intelligent detection module is used to perform intelligent detection on the corrected image through an intelligent detection model based on the YOLO3 algorithm to obtain the area to be recognized, and then send it to the intelligent recognition module. The intelligent recognition module is used to obtain the original text information of all rows in the area to be recognized through an intelligent recognition model based on the CRNN network algorithm, and send it to the natural language analysis and processing module. The natural language analysis and processing module is used to verify and correct the original text information through a pattern matching algorithm to obtain the final image text information. The natural language analysis technology is mainly based on linguistic principles and statistical principles.

[0007] Further, the steps of the image preprocessing include: First, scale the image in the UMat coordinate system. Then, detect the elements in the image through SIFT, and use the fast nearest neighbor approximation search algorithm to achieve fast matching of two-dimensional feature points. Next, make a judgment on the matched features, and define the ones that meet the conditions as a qualified matching point.

[0008] Further, calculate the mapping relationship between images using HomoGraphy; then use matrix inversion and perspective transformation to calculate the corresponding shape of the template image preset in the system on the image uploaded by the patient, and finally determine the type of the image uploaded by the patient through the feature matching similarity.

[0009] Further, the training method of the intelligent detection model is: Prepare training data, use the application materials of the existing patients in the project, and generate a marked file in xml format through a marking software.

[0010] Further, the marked file contains the coordinates of each text area and the included text. Input the marked file into the intelligent detection model to train the intelligent detection model.

[0011] Further, the training method of the intelligent recognition model is: Prepare training data, use the application materials of the existing patients in the project, use a marking software to generate a marked file in xml format. The marked file contains the coordinates of each text area and the included text. Cut out the marked picture areas according to the coordinates of each area in the marked file to generate a dataset of corresponding small areas and text labels.

[0012] Further, during training, the data set is randomly divided into a training data set and a test data set. The intelligent recognition model is trained with the training set data and tested with the test set data until the model loss function is less than the set threshold, at which point the training ends.

[0013] Further, the natural language analysis and processing module analyzes and understands the structure and meaning of human language through linguistic principles; and the natural language analysis and processing module mines rules and patterns from a large amount of language data through statistical principles, trains and optimizes the natural language processing model.

[0014] Further, the image preprocessing module compares the material images uploaded by the user with various image templates preset in the system, confirms the type of image materials, and obtains corrected images that can be used for intelligent detection.

[0015] Further, in the intelligent recognition model, the network structure of the CRNN network algorithm includes three parts. This network structure is, from bottom to top, a convolutional layer, a recurrent layer, and a transcription layer. The convolutional layer uses the CNN algorithm, the recurrent layer uses the RNN algorithm, and the transcription layer uses the CTC algorithm.

[0016] Advantages of the present invention:

[0017] 1. The intelligent review system for patient assistance based on deep learning is based on the deep learning method, conducts intelligent review, key information detection and recognition on patient materials, can greatly improve the review efficiency of patient assistance materials, and speed up the dispensing process of patient assistance; it can adapt to the actual needs of the patient assistance system and lower the threshold for patients to use.

[0018] 2. The intelligent review system for patient assistance based on deep learning reduces the requirements for the format, clarity, etc. of the materials uploaded by patients, while reducing the work pressure and workload of reviewers, and reducing the input of human resource costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is the network structure diagram of the YOLO algorithm adopted in the embodiment of the present invention;

[0020] Figure 2 It is the tensor value interpretation diagram of the YOLO algorithm adopted in the embodiment of the present invention;

[0021] Figure 3 It is another tensor value interpretation diagram of the YOLO algorithm adopted in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] To make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the present invention will be further described below in conjunction with specific embodiments.

[0023] See also Figures 1 to 3 The present invention provides a technical solution: a patient assistance intelligent review system based on deep learning, including an image preprocessing module, an intelligent detection module, an intelligent recognition module, and a natural language analysis and processing module.

[0024] The image preprocessing module is used to preprocess the original image materials uploaded by the patient. The preprocessing includes determining the type of image material and correcting the image before sending it to the intelligent detection module.

[0025] The function of the image preprocessing module is to preprocess the original materials uploaded by the patient that are of unknown type, arbitrary format, and do not meet the required clarity, in preparation for the subsequent intelligent detection, and to improve the accuracy of intelligent detection and intelligent recognition. In this embodiment, the image preprocessing module uses the computer vision library OpenCV to perform corresponding image processing, compares the data image uploaded by the user with various image templates preset by the system, confirms the type of image material, and obtains a corrected image that can be used for intelligent detection. First, the image is scaled in the unified matrix (UMat) coordinate system; then the elements in the image are detected by the feature point matching algorithm (SIFT), and the fast nearest neighbor similarity search algorithm is used to achieve fast matching of two-bit feature points; then the image is projected and mapped using the homography matrix transformation (HomoGraphy); then the corresponding shape of the system preset template image on the image uploaded by the patient is calculated using matrix inversion and perspective transformation, and finally the type of the image uploaded by the patient is determined by feature matching similarity.

[0026] The intelligent detection module is used to perform intelligent detection on the corrected image through an intelligent detection model based on the YOLO3 algorithm to obtain the area to be identified, and then send it to the intelligent recognition module.

[0027] The function of the intelligent detection module is to detect the area to be identified in the image and then send it to the recognition module for intelligent recognition. There are many algorithms for intelligent detection. In this embodiment, the YOLO3 algorithm is used. The YOLO algorithm is a solution proposed by Ross Girshick to address the speed problem of target detection based on deep learning, following RCNN, Fast-RCNN, and Faster-RCNN. It is generally used to detect objects in a scene. As the latest algorithm in the YOLO series, YOLO3 retains and improves on previous algorithms. Therefore, this embodiment uses the YOLO3 algorithm to detect image areas, and the effect is also relatively ideal.

[0028] As attached Figure 1As shown, it is the network structure diagram of the YOLO algorithm. The core idea of YOLO is to use the entire image as the input of the network and directly regress the position of the bounding box and its affiliated category at the output layer.

[0029] The network structure of this embodiment draws on GoogLeNet. There are 24 convolutional layers and 2 fully connected layers. The convolutional layers are mainly used to extract features, and the fully connected layers are mainly used to predict the category probability and coordinates. The alternating 1×1 convolutional layers reduce the feature space of the previous layer. In the ImageNet classification task, the convolutional layers are pre-trained at half the resolution (224×224 input images), and then the resolution is doubled in detection.

[0030] As shown in the appendix Figure 2 and Figure 3 As shown, the implementation method of YOLO is to divide the input image into S x S grids. If the center of an object falls within this grid, then this grid is responsible for predicting this object. Each grid needs to predict B bounding boxes. In addition to regressing its own position x, y, w, h, each bounding box also predicts a confidence value. This confidence represents two kinds of information: the confidence that the predicted bounding box contains an object and how accurate this bounding box prediction is. Its value calculation formula is: Pr(Object)*IOU(truth)(pred). If there is a target in the current grid, that is, the center coordinates of the actual marked box (ground truth) of the target are in the current grid, the first item takes 1, otherwise it takes 0. The second item represents the intersection-over-union (IOU: Intersection-over-Union value) between the predicted bounding box and the actual bounding box. Among them, the intersection-over-union is a concept used in object detection, which is the overlap rate of the generated candidate box (candidate bound) and the original marked box (ground truth bound), that is, the ratio of their intersection to their union. The most ideal situation is complete overlap, that is, the ratio is 1. Each bounding box needs to predict a total of 5 values: (x, y, w, h) and confidence. x and y are the center coordinates of the bounding box, which are the offset values relative to the current grid, and the value range is from 0 to 1. w and h are the width and height of the bounding box, which are normalized by dividing the width and height of the image respectively, and finally the range of w and h is also within 0 to 1. In addition, each grid also needs to predict the probabilities of C assumed categories. Then for S x S grids, each grid needs to predict B bounding boxes and C categories, and the output is a tensor of S x S x (5*B + C). There are 20 categories in the paper. For a 7x7 grid, each grid needs to predict 2 bounding boxes and 20 category probabilities, and the output is a tensor of 7x7x(5x2 + 20).

[0031] During testing, the class information predicted for each grid is multiplied by the confidence information predicted for the bounding box to obtain the class-specific confidence score for each bounding box.

[0032] This embodiment provides the following formula:

[0033]

[0034] In this formula:

[0035]

[0036] Finally, by setting a threshold, the bounding boxes with low scores are filtered out, and the remaining bounding boxes are processed by non-maximum suppression (NMS) to obtain the final detection result.

[0037] The main framework of the YOLO3 algorithm is Darknet53. Darknet53 borrows the idea of resnet and adds residual modules to the network, which is beneficial to solving the gradient problem of deep networks. Each residual module consists of two convolutional layers and a shortcut connection. 1, 2, 8, 8, 4 represent the number of repeated residual modules. In the entire v3 structure, there is no pooling layer and fully connected layer. The downsampling of the network is achieved by setting the stride of the convolution to 2. Whenever passing through this convolutional layer, the size of the image is reduced to half. The implementation of each convolutional layer includes convolution + BN + Leakyrelu, and a zero-padding module is added after each residual module.

[0038] The YOLO3 algorithm adopts multi-scale detection. For multi-scale detection, multiple scales are used for prediction. Specifically, it is achieved by performing upsampling and concatenation operations on some of the last layers of network prediction. The influence of resolution on prediction is explained as follows: The resolution information directly reflects the number of pixels that make up the target. For a target, the more pixels it has, the richer and more specific the details of the target are represented, that is, the richer the resolution information is. This is why the large-scale feature map provides resolution information. Semantic information in object detection refers to the information that distinguishes the target from the background, that is, semantic information is what makes people know that this is the target and the rest is the background. In different categories, semantic information does not require a lot of detailed information. A large resolution information will instead reduce the semantic information. Therefore, the small-scale feature map will provide better semantic information under the condition of providing necessary resolution information. (For small targets, the small-scale feature map cannot provide the necessary resolution information, so the large-scale feature map still needs to be combined.) YOLO3 further adopts three different-scale feature maps for object detection, and can detect more fine-grained features. The final output of the network has three scales, namely 1 / 32, 1 / 16, and 1 / 8. After several convolutional operations after the 79th layer, the prediction result of 1 / 32 (13×13) is obtained. The downsampling multiple is high, and the receptive field of the feature map here is relatively large. Therefore, it is suitable for detecting larger objects in the image. Then this result is upsampled and concatenated with the result of the 61st layer through tensor concatenation (concat), and after several convolutional operations, the prediction result of 1 / 16 is obtained; it has a medium-scale receptive field and is suitable for detecting medium-scale objects. The result of the 91st layer is upsampled and then concatenated with the result of the 36th layer through tensor concatenation. After several convolutional operations, the result of 1 / 8 is obtained. Its receptive field is the smallest and is suitable for detecting small-sized objects. The upsampling of a certain layer in the middle of darknet and the subsequent layer is concatenated. The concatenation operation is different from the addition operation of the residual layer. Concatenation will expand the dimension of the tensor, while add only directly adds without causing a change in the tensor dimension.

[0039] YOLO3 uses the Kmeans clustering method to determine the sizes of the prior boxes (anchor priors). The sizes of the prior boxes are obtained by K-means clustering. Three prior boxes are set for each downsampling scale, and a total of nine sizes of prior boxes are clustered. In the COCO dataset, these nine prior boxes are: (10x13), (16x30), (33x23), (30x61), (62x45), (59x119), (116x90), (156x198), (373x326). In terms of assignment, larger prior boxes (116x90), (156x198), (373x326) are applied on the smallest 13*13 feature map (with the largest receptive field), which is suitable for detecting larger objects. Medium-sized prior boxes (30x61), (62x45), (59x119) are applied on the medium-sized 26*26 feature map (with medium receptive field), which is suitable for detecting medium-sized objects. Smaller prior boxes (10x13), (16x30), (33x23) are applied on the larger 52*52 feature map (with smaller receptive field), which is suitable for detecting smaller objects. When YOLO3 predicts the bounding box (bbox), it uses logistic regression. Each time YOLO3 makes a prediction for a b-box, the output is the same as that of v2, which is (tx,ty,tw,th,to), and then the absolute (x,y,w,h,c) is calculated through formula 1. Logistic regression is used to perform an objectness score for the part surrounded by the anchor (for NMS), that is, how likely it is that there is an object at this position. YOLO3 only operates on one prior, that is, the best prior. And logistic regression is used to find the one with the highest objectness score from the nine prior boxes.

[0040] The object classification softmax of YOLO3 is changed to logistic. When predicting the object category, softmax is not used, but the output of logistic is used for prediction. This can support multi-label objects (for example, a person has two labels, Woman and Person).

[0041] The intelligent recognition module is used to obtain the original text information of all lines in the image material through an intelligent recognition model based on the CRNN network (Convolutional Recurrent Neural Network) algorithm and send it to the natural language analysis and processing module.

[0042] In the intelligent recognition model, the network structure of the CRNN network algorithm includes three parts, which are, from bottom to top, the convolutional layer, the recurrent layer, and the transcription layer. The convolutional layer uses the CNN algorithm, the recurrent layer uses the RNN algorithm, and the transcription layer uses the CTC algorithm. Specifically, the blank mechanism is inserted into the CTC algorithm of the transcription layer.

[0043] In this embodiment, the convolutional layer uses CNN, and its function is to extract the feature sequence from the input image; the recurrent layer uses RNN, and its function is to predict the label (true value) distribution of the feature sequence obtained from the convolutional layer; the transcription layer uses CTC (Connectionist Temporal Classification, connectionist temporal classification algorithm), and its function is to convert the label distribution obtained from the recurrent layer into the final recognition result through operations such as deduplication and integration.

[0044] The difficulty of end-to-end OCR lies in how to handle the problem of aligning variable-length sequences. The CRNN OCR draws on the idea of solving variable-length speech sequences in speech recognition. The algorithm of training RNN based on the connectionist temporal classification algorithm (CTC) significantly outperforms traditional speech recognition algorithms in the field of speech recognition. CRNN borrows the CTC loss function into OCR recognition. The CRNN algorithm inputs a normalized height word image of 100*32, extracts the feature map based on 7 layers of CNN (usually using GG16), splits the feature map by column (Map-to-Sequence), and inputs the 512-dimensional features of each column into a two-layer bidirectional LSTM with 256 units each for classification. During the training process, under the guidance of the CTC loss function, an approximate soft alignment of character positions and class labels is achieved. CRNN draws on the LSTM+CTC modeling method in speech recognition. The difference is that the features input into the LSTM are replaced from the acoustic features (MFCC, etc.) in the speech field with the image feature vectors extracted by the CNN network. The CRNN algorithm combines the potential of CNN for image feature engineering with the potential of LSTM for sequential recognition. It not only extracts robust features but also avoids the extremely difficult single-character segmentation and single-character recognition in traditional algorithms through sequential recognition. At the same time, sequential recognition also embeds temporal dependencies (implicitly using the corpus). In the training stage, CRNN uniformly scales the training images to 100×32 (w×h); in the testing stage, to address the problem of reduced recognition rate caused by character stretching, CRNN maintains the aspect ratio of the input image size, but the image height must still be uniformly 32 pixels, and the size of the convolutional feature map dynamically determines the LSTM temporal length.

[0045] The problem to be solved in CRNN is that the length of the image text is variable, so there is an alignment and decoding problem. Therefore, RNN needs an additional partner to solve this problem, and this partner is the famous CTC decoding. The architecture adopted by CRNN is CNN+RNN+CTC. CNN extracts image pixel features, RNN extracts image temporal features, and CTC generalizes the connection characteristics between characters. The RNN layer in CRNN outputs a variable-length sequence. For example, if the width of the original image is W, the number of sequences output after passing through CNN and RNN may be S. At this time, we need to translate this sequence into the final recognition result. When RNN performs temporal classification, there will inevitably be a lot of redundant information. For example, a letter is recognized twice in a row. This requires a redundancy removal mechanism. However, simply removing redundancy when seeing two consecutive letters also has problems. For example, words like "cook" and "geek", so CTC has a blank mechanism to solve this problem.

[0046] To solve the ambiguity, CTC proposes an inserted blank mechanism. For example, if we use the "-" symbol to represent blank, then if the label is "aaa-aaaabb", it will be mapped to "aab", and "aaaaaaabb" will be mapped to "ab". Introducing the blank mechanism can well handle the problem of repeated characters. The output of the RNN layer is the probability matrix in the sequence. Among them, represents the probability of the first sequence outputting "-". Then, the probability of outputting a certain path is the product of the probabilities of each sequence. Therefore, there can be multiple paths to obtain a label. Intuitively, when an image text is output to the network, it is necessary to maximize the probability that the output is the label L. Since the paths are mutually exclusive, for the labeled sequence, its conditional probability is the sum of the probabilities of all paths that map to it: where π∈B1(l) means all path sets that can be merged into l. This way of making the sum of the mapping B and all candidate path probabilities enables CTC not to require an accurate segmentation of the original input sequence, which makes it possible to translate tasks where the sequence length output by the RNN layer > label length. CTC can be combined with any RNN model. However, considering that the labeling probability is related to the entire input string rather than only related to the fragments in the previous small window range, the bidirectional RNN / LSTM model is more suitable.

[0047] The natural language analysis and processing module is used to verify and correct the original text information through a pattern matching algorithm to obtain the final image text information.

[0048] Natural language analysis and processing can mark the position information of each line of text in an image through a text detection network based on YOLO3. After being recognized by the CRNN network, the text of each line is obtained. Since the original results contain information in units of lines, and in the actual materials to be recognized, there may be different information fields in the same horizontal line, such as name and gender, natural language analysis is required. Natural language analysis further processes the results of network recognition. For different categories of the specific materials to be recognized, by analyzing the relative position relationships of each text area in the real materials and the key characters contained therein, methods such as regular expressions are used to extract key and valuable text information from the recognized result characters. Through natural language analysis and processing, the key information we need, such as name and gender, is extracted from the original possibly chaotic multi-line text.

[0049] Specifically, the system in this embodiment uses Keras (Keras is an open-source artificial neural network library written in Python) to implement YOLO3 for text detection, and implements text recognition based on the CRNN algorithm of TensorFlow (TensorFlow is an open-source software library for numerical computing using dataflow graphs) and Pytouch (Pytorch is the Python version of torch, an open-source neural network framework by Facebook, specifically for programming GPU-accelerated deep neural networks (DNNs). Torch is a classic tensor library for operating on multi-dimensional matrix data and has wide applications in machine learning and other math-intensive applications).

[0050] In this embodiment, the training method of the intelligent detection model is as follows: Prepare training data: Use the existing application materials of patients in the project to generate a marked file in xml format through a marking software. The marked file contains the coordinates of each text area and the text contained therein, and input the marked file into the intelligent detection model to train the intelligent detection model.

[0051] In this embodiment, the training method of the intelligent recognition model is as follows: Prepare training data, use the existing application materials of patients in the project, and use a marking software to generate a marked file in xml format. The marked file contains the coordinates of each text area and the text contained therein; Cut out the marked picture areas according to the coordinates of each area in the marked file to generate a dataset of corresponding small areas and text labels; During training, randomly divide the dataset into a training dataset and a test dataset, train the intelligent recognition model with the training set data, and test the intelligent recognition model with the test set data until the model loss function is less than the set threshold, and the training ends.

[0052] The foregoing has shown and described the basic principles, main features and advantages of the present invention. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, in any regard, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to embrace all changes that fall within the meaning and scope of the equivalent elements of the claims in the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

[0053] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. An intelligent audit system for patient assistance based on deep learning, characterized in that: The audit system includes an image preprocessing module, an intelligent detection module, an intelligent recognition module, and a natural language analysis and processing module. The image preprocessing module is used to preprocess the original image materials uploaded by patients. The preprocessing includes determining the type of the image materials and sending the corrected image to the intelligent detection module after image correction. The intelligent detection module is used to perform intelligent detection on the corrected image through an intelligent detection model based on the YOLO3 algorithm to obtain the area to be recognized, and then send it to the intelligent recognition module. The intelligent recognition module is used to obtain the original text information of all lines in the area to be recognized through an intelligent recognition model based on the CRNN network algorithm, and send it to the natural language analysis and processing module. The natural language analysis and processing module is used to verify and correct the original text information through a pattern matching algorithm to obtain the final image text information. The natural language analysis technology is based on linguistic principles and statistical principles.

2. The intelligent audit system for patient assistance based on deep learning according to claim 1, characterized in that, The steps of the image preprocessing include: first, scaling the image in the UMat coordinate system; then detecting the elements in the image through SIFT and using the fast nearest neighbor approximation search algorithm to achieve fast matching of two-dimensional feature points; next, making a judgment on the matched features, and the ones that meet the conditions are defined as qualified matching points.

3. The intelligent audit system for patient assistance based on deep learning according to claim 2, characterized in that: Calculate the mapping relationship between images using HomoGraphy; then use matrix inversion and perspective transformation to calculate the corresponding shape of the template image preset in the system on the image uploaded by the patient, and finally determine the type of the image uploaded by the patient through the feature matching similarity.

4. The intelligent audit system for patient assistance based on deep learning according to claim 1, characterized in that, The training method of the intelligent detection model is: prepare training data, use the application materials of existing patients in the project, and generate a marked file in xml format through a marking software.

5. The intelligent audit system for patient assistance based on deep learning according to claim 4, wherein: The marked file contains the coordinates of each text area and the contained text. Input the marked file into the intelligent detection model to train the intelligent detection model.

6. The intelligent review system for patient assistance based on deep learning according to claim 1, characterized in that, The training method of the intelligent recognition model is: prepare training data, use the application materials of existing patients in the project, use a marking software to generate a marked file in xml format. The marked file contains the coordinates of each text area and the contained text; cut out the marked picture areas according to the coordinates of each area in the marked file to generate a dataset of corresponding small areas and text labels.

7. The intelligent audit system for patient assistance based on deep learning according to claim 6, characterized in that: During training, randomly divide the dataset into a training dataset and a test dataset. Train the intelligent recognition model through the training set data and test the intelligent recognition model through the test set data until the model loss function is less than the set threshold, and the training ends.

8. An intelligent audit system for patient assistance based on deep learning according to claim 1, characterized in that: The natural language analysis and processing module analyzes and understands the structure and meaning of human language through linguistic principles; and the natural language analysis and processing module mines rules and patterns from a large amount of language data through statistical principles, trains and optimizes the natural language processing model.

9. The intelligent review system for patient assistance based on deep learning according to claim 3, characterized in that: The image preprocessing module compares the material image uploaded by the user with various image templates preset in the system, confirms the type of the image materials, and obtains a corrected image that can be used for intelligent detection.

10. The intelligent audit system for patient assistance based on deep learning according to claim 7, wherein: In the intelligent recognition model, the network structure of the CRNN network algorithm consists of three parts. This network structure is, from bottom to top, a convolutional layer, a recurrent layer, and a transcription layer. The convolutional layer uses the CNN algorithm, the recurrent layer uses the RNN algorithm, and the transcription layer uses the CTC algorithm.