AI-based system for automatically auditing medical first-marketing data
By designing an AI-based automatic audit system, using deep learning models and blockchain technology, the problem of difficult to ensure the efficiency and accuracy of the first-time data review of traditional medicines is solved, and efficient automated audits and data security traceability are achieved.
Patent Information
- Application Number
- CN202510128531.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-27
AI Technical Summary
The review of the first-time data of traditional medicine relies on manual labor, which makes it difficult to guarantee audit efficiency and accuracy, and faces challenges such as semantic understanding, compliance inspection, data security and regulatory differences.
Design an automatic audit system based on AI, including AI preprocessing module, intelligent classification module, content audit module, decision support module and blockchain evidence storage module, and use deep learning models and blockchain technology to achieve efficient automated audit of the first medical data and secure data storage.
It realizes efficient and automated audits, improves audit speed and accuracy, ensures multi-level audit guarantee and data security traceability, and adapts to changing audit needs through continuous optimization and user feedback mechanisms.
Smart Images

Figure CN120047164A_ABST
Abstract
Description
Technical Field
[0001] The present invention mainly relates to the technical field of pharmaceutical data management, and specifically to a system for automatically auditing pharmaceutical first business data based on AI. Background Art
[0002] In the pharmaceutical industry, the audit of first business data is a key link to ensure the safety and compliance of drug sales. Traditionally, this audit process mainly relies on manual operation, which is not only time-consuming and laborious, but also easily affected by human factors, resulting in difficult-to-guarantee audit efficiency and accuracy. With the rapid development of the pharmaceutical industry and the continuous improvement of regulatory requirements, the traditional audit method has been difficult to meet the current needs.
[0003] In recent years, the rapid development of artificial intelligence (AI) technology has provided a new solution for the audit of pharmaceutical first business data. AI technology, especially deep learning models, has powerful data processing and analysis capabilities, and can automatically extract key information from text and perform accurate classification and auditing. However, applying AI technology to the audit of pharmaceutical first business data still faces many challenges.
[0004] Firstly, pharmaceutical first business data usually contains a large number of professional terms and complex regulatory requirements, and the model needs to have strong semantic understanding and compliance checking capabilities. Secondly, due to the particularity of the pharmaceutical industry, the audit process needs to ensure the security and traceability of data to prevent data leakage or tampering. In addition, the regulations and regulatory requirements in different regions may vary, which also increases the complexity and uncertainty of the audit.
[0005] According to a pharmaceutical electronic first business data management system provided by application number CN202211417539.8, it includes: a matching module for matching the seller user and the buyer user to establish a data exchange relationship; a data entry module for the seller user and the buyer user to enter electronic first business data respectively; a data transmission module for transmitting the electronic first business data entered by the seller user and the buyer user to the other user with whom the data exchange relationship is determined; a data audit module for auditing whether the electronic first business data entered by the seller user and the buyer user meets the requirements, and generating corresponding audit results and sending them to the seller user and the buyer user respectively; a data storage module for storing all audited electronic first business data and data exchange process information according to the data exchange relationship involved by the user.
[0006] In the prior art, it focuses on building a complete electronic first business data management system, including various links such as user matching, data entry, transmission, audit, and storage, but there are deficiencies in terms of safety, reliability, stability, accuracy, and certain traceability. Summary of the Invention
[0007] Based on this, the objective of the present invention is to provide a system for automatically auditing pharmaceutical first business materials based on AI, so as to solve the technical problems raised in the above-mentioned background art.
[0008] To achieve the above objective, the present invention provides the following technical solutions:
[0009] A system for automatically auditing pharmaceutical first business materials based on AI, comprising an AI preprocessing module, an intelligent classification module, a content auditing module, a decision support module, and a blockchain evidence storage module;
[0010] The AI preprocessing module is used to receive and preprocess pharmaceutical first business materials, and convert the pharmaceutical first business materials into an analyzable text format;
[0011] The intelligent classification module is used to automatically classify the preprocessed text by using a deep learning model, and identify the types of pharmaceutical first business materials;
[0012] The content auditing module includes a compliance check sub-module, a semantic analysis sub-module, and an anomaly detection sub-module; the content auditing module is used to perform multi-level auditing and analysis on the content of the classified pharmaceutical first business materials;
[0013] The decision support module automatically generates an audit report based on the analysis results of the content auditing module, and provides audit decision support;
[0014] The blockchain evidence storage module is used to securely store the audited pharmaceutical first business materials and their audit records on the blockchain.
[0015] Preferably, the AI preprocessing module further includes:
[0016] The OCR conversion sub-module is used to convert paper materials into electronic text;
[0017] The text preprocessing sub-module is used to perform preprocessing operations on the electronic text, and the preprocessing operations include denoising, word segmentation, and part-of-speech tagging.
[0018] Preferably, the intelligent classification module adopts a convolutional neural network CNN or a long short-term memory network LSTM model, and realizes automatic classification of newly uploaded materials by training a large number of pharmaceutical first business material samples.
[0019] Preferably, the compliance check sub-module is used to compare the content of the materials item by item in combination with a rule library and a regulation database specific to the pharmaceutical industry to ensure information compliance;
[0020] The semantic analysis sub-module performs in-depth semantic understanding on the text based on a BERT or GPT series model to identify potential risks;
[0021] The abnormal detection sub-module uses clustering analysis or abnormal detection algorithms to identify abnormal data patterns.
[0022] Preferably, the decision support module further includes a reinforcement learning component, which is used to continuously optimize the audit strategy according to historical audit data and feedback.
[0023] Preferably, the blockchain evidence storage module uses encryption algorithms to encrypt the approved materials and their audit records to ensure data security, and stores the encrypted data on the blockchain.
[0024] In summary, the present invention mainly has the following beneficial effects:
[0025] The present invention has the characteristics of efficient and automated auditing. By integrating AI technologies, especially deep learning models, the present invention realizes the efficient and automated auditing of pharmaceutical first business materials. This not only significantly improves the auditing speed but also reduces the burden of manual auditing, making the auditing process more efficient and accurate.
[0026] And it has multi-level auditing guarantees. The system designs a multi-level auditing process including intelligent classification, content auditing, and decision support. This design ensures the comprehensive auditing of materials and effectively identifies potential risks such as false propaganda and omission of key information, thus improving the comprehensiveness and accuracy of auditing.
[0027] At the same time, it has continuous optimization and adaptability. By introducing reinforcement learning and user feedback mechanisms, the present invention can continuously optimize the audit strategy according to historical data and user feedback. This continuous optimization mechanism enables the system to adapt to the changing audit requirements and improve the adaptability and accuracy of audit decisions.
[0028] And it includes data security and traceability. Using blockchain technology for data evidence storage ensures the security and traceability of the approved materials and their audit records. The immutability of the blockchain provides additional security guarantees for users, ensuring the authenticity and reliability of audit results.
[0029] Among them, the innovative application of OCR technology. In the AI preprocessing module, an OCR algorithm based on deep learning is adopted, which can accurately convert paper materials into electronic texts. This innovative application not only improves the efficiency of material processing but also lays a solid foundation for subsequent intelligent classification and content auditing.
[0030] In addition, the deep learning model is optimized. Both the intelligent classification module and the content auditing module adopt deep learning models, and by introducing attention mechanisms, using pre-trained models and other optimization strategies, the performance and accuracy of the models are improved. These optimization strategies enable the system to better adapt to complex and changeable audit tasks.
[0031] The system provides an intuitive front - end user interface and a decision - support module, enabling users to conveniently upload materials, view audit reports, and provide feedback. This user - friendly design improves the usability of the system and user satisfaction. Brief Description of the Drawings
[0032] Figure 1 It is the overall structural framework diagram of the system of the present invention;
[0033] Figure 2 It is the specific structural framework diagram of the content audit module of the present invention;
[0034] Figure 3 It is the specific structural framework diagram of the AI pre - processing module of the present invention.
[0035] Brief Description of the Drawings: 10. AI pre - processing module; 20. Intelligent classification module; 30. Content audit module; 40. Decision - support module; 50. Blockchain evidence - storage module; 11. OCR conversion sub - module; 12. Text pre - processing sub - module; 31. Compliance check sub - module; 32. Semantic analysis sub - module; 33. Anomaly detection sub - module. Detailed Embodiment
[0036] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0037] As Figures 1 to 3 shown, a system for automatically auditing pharmaceutical first - business materials based on AI includes an AI pre - processing module 10, an intelligent classification module 20, a content audit module 30, a decision - support module 40, and a blockchain evidence - storage module 50;
[0038] The AI pre - processing module 10 is used to receive and pre - process pharmaceutical first - business materials, and convert the pharmaceutical first - business materials into an analyzable text format;
[0039] The intelligent classification module 20 is used to automatically classify the pre - processed text using a deep - learning model to identify the types of pharmaceutical first - business materials;
[0040] The content audit module 30 includes a compliance check sub - module 31, a semantic analysis sub - module 32, and an anomaly detection sub - module 33; the content audit module 30 is used to perform multi - level audit and analysis on the content of the classified pharmaceutical first - business materials;
[0041] The decision - support module 40 automatically generates an audit report based on the analysis results of the content audit module 30 and provides audit decision support;
[0042] The blockchain evidence - storage module 50 is used to securely store the audited - passed pharmaceutical first - business materials and their audit records on the blockchain.
[0043] The AI preprocessing module 10 further includes:
[0044] The OCR conversion sub-module 11 is used to convert paper materials into electronic texts;
[0045] The text preprocessing sub-module 12 is used to perform preprocessing operations on the electronic texts, and the preprocessing operations include denoising, word segmentation, and part-of-speech tagging.
[0046] The intelligent classification module 20 adopts a convolutional neural network CNN or a long short-term memory network LSTM model, and realizes the automatic classification of newly uploaded materials by training a large number of samples of pharmaceutical first business materials.
[0047] The compliance check sub-module 31 is used to compare the content of the materials item by item in combination with a rule library and a regulation database specific to the pharmaceutical industry to ensure information compliance;
[0048] The semantic analysis sub-module 32 performs in-depth semantic understanding on the text based on the BERT or GPT series models to identify potential risks;
[0049] The anomaly detection sub-module 33 adopts clustering analysis or anomaly detection algorithms to identify abnormal data patterns.
[0050] The decision support module 40 further includes a reinforcement learning component 41, and the reinforcement learning component 41 is used to continuously optimize the review strategy according to historical review data and feedback.
[0051] The blockchain evidence storage module 50 uses encryption algorithms to encrypt the reviewed materials and their review records to ensure data security, and stores the encrypted data on the blockchain.
[0052] Example 1: System Architecture and Implementation Process
[0053] For a system for automatically reviewing pharmaceutical first business materials based on AI of the present invention, specifically in implementation, the system architecture includes a front-end user interface, a back-end server, and a blockchain network. Users upload pharmaceutical first business materials through the front-end interface, and the back-end server receives the materials and starts the AI preprocessing module 10 to perform material preprocessing.
[0054] The preprocessed pharmaceutical first business materials are sent to the intelligent classification module 20 for automatic classification, and the classification results guide the content review module 30 to perform multi-level review and analysis.
[0055] The decision support module 40 generates a review report according to the review analysis results and provides it for users to view.
[0056] At the same time, the blockchain evidence storage module 50 encrypts the reviewed materials and their review records and stores them on the blockchain to ensure the security and traceability of the data.
[0057] Example 2: Training and Optimization of the Intelligent Classification Module 20
[0058] It should be noted that in this embodiment, the intelligent classification module 20 is trained using a deep learning model, such as a convolutional neural network (CNN) or a long short-term memory network (LSTM).
[0059] During the training process, a large number of classified pharmaceutical first business materials are collected as training samples, and the types of materials are labeled with tags. The model learns to extract key information from text features during the training process to achieve automatic classification of newly uploaded materials. To improve the classification accuracy, the training samples are updated regularly, and the transfer learning method is used to transfer the knowledge of new samples to the model, continuously optimizing the classification performance.
[0060] Example 3: Application and Effect of the Content Review Module 30
[0061] In actual application, the compliance check sub-module 31 of the content review module 30 compares each piece of information in the materials with a specific rule library and regulation database in the pharmaceutical industry to ensure the compliance of the information. The semantic analysis sub-module 32 uses a deep learning model to perform in-depth semantic understanding of the text to identify potential risks such as possible false propaganda and omission of key information. The anomaly detection sub-module 33 uses clustering analysis or anomaly detection algorithms to identify abnormal data patterns, such as frequently occurring error types or abnormally high similarities, to assist in identifying possible fraud behaviors or incorrect entries. Through multi-level review and analysis, the accuracy and comprehensiveness of the review are greatly improved.
[0062] Example 4: Optimization and Feedback of the Decision Support Module 40
[0063] While generating the review report, the decision support module 40 also uses the reinforcement learning component 41 to continuously optimize the review strategy based on historical review data and user feedback. The reinforcement learning component simulates the review process and continuously adjusts the review rules and parameters to improve the accuracy and adaptability of the review decision.
[0064] At the same time, the system also provides a user feedback mechanism that allows users to comment on and score the review results, and these feedback data are used for further training and optimization of the reinforcement learning model.
[0065] Example 5: Data Security and Traceability of the Blockchain Evidence Storage Module 50
[0066] The blockchain evidence storage module 50 uses advanced encryption algorithms to encrypt the approved materials and their review records, ensuring the security of data during transmission and storage. The encrypted data is stored on the blockchain. Due to the immutability of the blockchain, any modification to the data will be recorded, thus achieving data traceability. This provides additional security for users and ensures the authenticity and reliability of the review results.
[0067] Example Six: The OCR conversion algorithm in the AI preprocessing module 10
[0068] In the AI preprocessing module 10, the OCR conversion algorithm is a key link for converting paper-based pharmaceutical first business materials into electronic texts.
[0069] In this embodiment, we adopt an OCR algorithm based on deep learning, namely a model that combines a convolutional neural network CNN and connectionist temporal classification CTC.
[0070] Model architecture:
[0071] Input layer: Receives the scanned image data and performs normalization processing.
[0072] CNN layer: Uses multiple layers of convolution, pooling, and activation functions to extract character features in the image.
[0073] RNN layer: To capture the sequential relationship between characters, a long short-term memory network LSTM or a gated recurrent unit GRU can be added, which is an optional solution.
[0074] CTC layer: Uses the connectionist temporal classification algorithm to map the output of CNN (or CNN + RNN) to a character sequence and solve the character alignment problem.
[0075] Training process:
[0076] Collect a large number of labeled pharmaceutical first business material images and their corresponding text data.
[0077] Use this data to train the model and optimize the model parameters through the backpropagation algorithm.
[0078] Evaluate the model performance and adjust the hyperparameters, such as the learning rate, batch size, etc.
[0079] Application and effect:
[0080] The trained OCR model can accurately convert paper-based materials into electronic texts, laying a foundation for subsequent intelligent classification and content review.
[0081] Compared with traditional OCR methods, OCR algorithms based on deep learning have significant advantages in recognizing complex layouts, handwritten texts, special symbols, etc.
[0082] Example 7: Deep learning model in the intelligent classification module 20
[0083] The intelligent classification module 20 uses a deep learning model for automatic classification. In this example, a combination of convolutional neural network CNN and long short-term memory network LSTM, namely CNN-LSTM, is specifically used.
[0084] Model architecture:
[0085] The CNN part: used to extract local features in the text, such as characters, words, etc.
[0086] The LSTM part: used to capture global sequence features in the text, that is, context information.
[0087] Classification layer: uses a fully connected layer and the softmax function to map the feature vector to the class label.
[0088] Training process:
[0089] Collect and annotate a large number of samples of pharmaceutical first business materials to ensure the diversity and representativeness of the samples.
[0090] Preprocess the text data, such as word segmentation, stop word removal, stemming, etc.
[0091] Convert the preprocessed text data into an input format acceptable to the model, such as word vectors or character vectors.
[0092] Use strategies such as cross-validation and early stopping to prevent overfitting.
[0093] Optimization strategies:
[0094] Introduce an attention mechanism to improve the model's attention to key information.
[0095] Use pre-trained models (such as BERT, GPT, etc.) to initialize the model parameters, accelerate the training process and improve the performance.
[0096] Regularly evaluate the model and adjust the model architecture and hyperparameters according to the evaluation results.
[0097] Example Eight: Semantic analysis sub-module 32 in the content review module 30;
[0098] The semantic analysis sub-module 32 in the content review module 30 is used to deeply understand the text content and identify potential risks. In this example, we use a semantic analysis model based on BERT.
[0099] Model architecture:
[0100] Input layer: Receives the preprocessed text data and performs word embedding processing.
[0101] Encoder layer: Uses multiple layers of Transformer encoders to deeply encode the input text and generate context-related word vectors.
[0102] Classification / regression layer: According to the task requirements, uses a fully connected layer or a softmax function for classification or regression prediction.
[0103] Training process:
[0104] Collect and annotate samples of pharmaceutical first business materials containing potential risk information.
[0105] Preprocess the samples to ensure data quality.
[0106] Use the annotated data to fine-tune the BERT model to adapt to semantic analysis tasks in specific domains.
[0107] Application and effect:
[0108] The fine-tuned BERT model can accurately identify potential risk information in the text, such as false propaganda, omission of key information, etc.
[0109] Compared with traditional rule-based methods, the BERT-based semantic analysis algorithm has stronger generalization ability and higher accuracy.
[0110] Example 9: Data encryption algorithm in the blockchain evidence storage module 50
[0111] The data encryption algorithm in the blockchain evidence storage module 50 is used to protect the security of the approved materials and their audit records.
[0112] In this embodiment, we adopt an encryption algorithm based on Elliptic Curve Cryptography (ECC).
[0113] Encryption algorithm:
[0114] Use the ECC algorithm to generate a public-private key pair.
[0115] Perform a hash process on the data to be stored to generate a hash value of a fixed length.
[0116] Use the private key to sign the hash value to ensure the integrity and authenticity of the data.
[0117] Store the signed hash value and the original data together on the blockchain.
[0118] Decryption and verification:
[0119] When it is necessary to verify the authenticity of the data, the public key is used to verify the signature.
[0120] If the verification passes, it proves that the data has not been tampered with during storage.
[0121] The original data can be indexed and retrieved through the hash value, but the original plaintext data cannot be directly obtained, protecting user privacy.
[0122] Security analysis:
[0123] The ECC algorithm has higher security and a smaller key length compared to traditional algorithms such as RSA.
[0124] The use of the hash function ensures the immutability and uniqueness of the data.
[0125] Proper custody of the private key is the key to ensuring data security.
[0126] The above embodiments are only used to illustrate the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any changes made on the basis of the technical solution according to the technical idea proposed by the present invention fall within the protection scope of the present invention.
Claims
1. A system for automatically reviewing pharmaceutical first-time business information based on AI, characterized in that: It includes an AI pre-processing module (10), an intelligent classification module (20), a content review module (30), a decision support module (40), and a blockchain evidence storage module (50); The AI preprocessing module (10) is used to receive and preprocess the medical first-time marketing data, and convert the medical first-time marketing data into an analyzable text format; The intelligent classification module (20) is used to automatically classify the pre-processed text using a deep learning model to identify the type of medical first-time marketing information; The content review module (30) comprises a compliance check submodule (31), a semantic analysis submodule (32) and an anomaly detection submodule (33); the content review module (30) is used to perform multi-level review and analysis on the classified content of the first-time medical marketing materials; The decision support module (40) automatically generates an audit report based on the analysis results of the content audit module (30) and provides audit decision support; The blockchain evidence storage module (50) is used to securely store the approved pharmaceutical first-time marketing information and its audit records on the blockchain.
2. According to claim 1, the system for automatically reviewing the first-time medical business data based on AI is characterized in that: The AI preprocessing module (10) further comprises: The OCR conversion submodule (11) is used to convert paper documents into electronic texts; The text preprocessing submodule (12) is used to perform preprocessing operations on the electronic text, and the preprocessing operations include denoising, word segmentation, and part-of-speech tagging.
3. According to claim 1, the system for automatically reviewing the first-time medical business data based on AI is characterized in that: The intelligent classification module (20) adopts a convolutional neural network (CNN) or a long short-term memory (LSTM) model, and realizes automatic classification of newly uploaded data by training a large number of medical first-time data samples.
4. According to claim 1, the system for automatically reviewing the first-time medical business data based on AI is characterized in that: The compliance check submodule (31) is used to compare the content of the data item by item in combination with the rule base and regulatory database specific to the pharmaceutical industry to ensure that the information is compliant; The semantic analysis submodule (32) performs deep semantic understanding of the text based on the BERT or GPT series model to identify potential risks; The anomaly detection submodule (33) uses cluster analysis or anomaly detection algorithm to identify abnormal data patterns.
5. According to claim 1, the system for automatically reviewing the first-time medical business data based on AI is characterized in that: The decision support module (40) further comprises a reinforcement learning component (41), wherein the reinforcement learning component (41) is used to continuously optimize the audit strategy based on historical audit data and feedback.
6. According to claim 1, the system for automatically reviewing the first-time medical business data based on AI is characterized in that: The blockchain evidence storage module (50) uses an encryption algorithm to encrypt the audited information and its audit records to ensure data security, and stores the encrypted data on the blockchain.
Citation Information
Patent Citations
Medical electronic first-marketing data management system
CN115662569A
Cited By
AI model-based auxiliary data auditing method and system
CN120806948A