Intelligent computer data processing method and system based on artificial intelligence
By constructing an adaptive data preprocessing mechanism and intelligent analysis path planning through reinforcement learning algorithms, the problem of poor adaptability of static rules in traditional computer data processing is solved, realizing intelligent and personalized data processing and improving the accuracy and efficiency of data processing.
Patent Information
- Application Number
- CN202511476083.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-16
- Publication Date
- 2025-11-11
AI Technical Summary
In traditional computer data processing, the data preprocessing stage relies on static rules, which makes it unable to adapt to complex dynamic data. The preprocessed data does not match the analysis model, the analysis process has a low degree of intelligence, and the fixed analysis process cannot respond to personalized needs.
An adaptive dynamic data preprocessing mechanism based on reinforcement learning algorithms is constructed. The preprocessing strategy is adjusted through Q-learning algorithm, and the intelligent data analysis path is planned by combining DQN algorithm to achieve dynamic matching between data characteristics and user goals.
It improves the timeliness and accuracy of data processing, realizes a high degree of intelligence and personalization in the analysis process, forms a closed loop of data processing, and supports efficient business decision-making.
Smart Images

Figure CN120930832A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer data processing technology, specifically to a computer data intelligent processing method and system based on artificial intelligence. Background Technology
[0002] Computer data refers to a collection of information stored in binary form on computer storage media (such as hard drives, memory, USB flash drives, etc.). It encompasses various types, including text, images, audio, video, and numerical signals collected by sensors. This data can be either raw records reflecting the attributes and states of objective things or structured information that has been preliminarily organized and encoded. It serves as the foundation for computer systems to perform calculations, analyses, and decisions, and is widely found in various scenarios such as financial transaction records, industrial production monitoring, medical diagnostic images, and internet user behavior trajectories. With the rapid development of information technology, computer data is experiencing explosive growth, and the types of data are becoming increasingly complex and diverse. Traditional manual processing and simple programmatic processing methods are no longer sufficient to cope with this. Therefore, conducting intelligent processing of computer data is a key step in transforming massive amounts of data into decision-making basis, improving industry operating efficiency, and promoting technological innovation.
[0003] Computer data intelligent processing can break through the bottleneck of data value mining and realize the efficient utilization of data resources. From the perspective of industry applications, it can help enterprises accurately grasp market demands, optimize resource allocation, and enhance core competitiveness. From the perspective of social development, it can promote the construction of smart cities, smart healthcare, intelligent transportation and other fields, and improve the quality of public services. From the perspective of technological innovation, it can promote the deep integration of artificial intelligence, big data, cloud computing and other technologies, provide support for technological iteration and upgrading, and ultimately realize data-driven intelligent development, injecting new momentum into the high-quality development of the social economy.
[0004] However, in the current computer data processing workflow, both the core stages of data preprocessing and data analysis have significant limitations. In the data preprocessing stage, traditional methods generally rely on manually preset static rules to complete data cleaning and transformation. When faced with complex and dynamically changing data environments, static rules cannot be adjusted in time to adapt to real-time data characteristics, which can easily lead to a mismatch between the preprocessed data and the subsequent analysis model. This not only reduces the accuracy of data processing but also delays the processing process due to the need for manual readjustment of rules. In the data analysis stage, traditional methods use fixed analysis processes to process all types of data. They cannot adjust the processing methods according to the differences in data characteristics, nor can they respond to users' personalized analysis goals. It is difficult to plan a targeted optimal analysis path, resulting in a low level of intelligence in the analysis process and a poor match between the analysis results and actual needs. Therefore, developing computer data intelligent processing methods and systems based on artificial intelligence is of great significance. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a computer data intelligent processing method and system based on artificial intelligence. It can build an adaptive dynamic data preprocessing mechanism, use reinforcement learning algorithms to automatically adjust the preprocessing strategy in real time based on data characteristics, achieve accurate matching between preprocessed data and subsequent analysis models, improve the timeliness and accuracy of data processing, and introduce a reinforcement learning framework to realize intelligent data analysis path planning. It can dynamically generate the optimal analysis path according to data characteristics and user analysis goals, and realize a high degree of intelligence and personalization in the analysis process.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a computer data intelligent processing method based on artificial intelligence, the method comprising the following steps: S1. The feature information of the data to be processed is acquired in real time through the data acquisition module, and the feature information is converted into a state vector that can be recognized by the reinforcement learning agent. S2. Construct a reinforcement learning agent based on the Q-learning algorithm, define the agent's action space as a set of data preprocessing operations, and define the reward function as the fit between the preprocessed data and the subsequent analysis model. S3. Input the state vector into the agent, the agent selects a preprocessing action to execute and obtains preprocessed data, calculates the fitness degree as a reward and feeds it back to the agent, the agent updates the Q-value table to adjust the strategy, repeats the iteration until the reward value is stable, and determines the optimal preprocessing strategy. S4. Use the optimal preprocessing strategy to process real-time data and output standard data; S5. Obtain the user's analysis target, extract standard data characteristic parameters, and combine them into agent state information; S6. Construct an intelligent agent for analysis path planning based on DQN, define the action space of the agent as a set of data analysis operations, and define the reward function as a weighted sum of the matching degree between the analysis result and the target and the analysis efficiency; S7. Input the state information into the agent, the agent selects the action sequence to be analyzed and executes it, calculates the reward value and updates the parameters, iterates until the reward value is stable, and determines the optimal analysis path; S8. Perform the analysis according to the optimal analysis path and output the analysis results.
[0007] Furthermore, the feature information of the data to be processed in step S1 includes data format, noise distribution ratio, missing value ratio, and data dimension. The state vector transformation process adopts standardized coding rules. The standardized coding rules set corresponding coding mapping relationships based on the data type of the feature information. The state vector is calculated through weighted coding of feature information, and the calculation formula is as follows: ,in, For state vectors, The number of types of feature information. For the first Encoding weights of class feature information For the first Class feature information The encoding function, The importance of features was determined through a feature importance assessment. Specifically, the random forest algorithm was used to calculate the contribution of various feature information to the selection of subsequent preprocessing strategies, and the contribution normalization result was used as the basis for the assessment. The value of .
[0008] Furthermore, the data preprocessing operation set in step S2 includes cleaning rules and transformation rules. The cleaning rules include outlier removal threshold adjustment and duplicate data filtering methods, while the transformation rules include data standardization methods and classification data encoding methods. By clarifying the specific operation types of the cleaning rules and transformation rules, a complete set of data preprocessing operations is formed, providing a clear range of action choices for the reinforcement learning agent.
[0009] Furthermore, in step S3, the iterative process needs to monitor the fluctuation range of reward values over multiple consecutive rounds. The fluctuation range of reward values is obtained by calculating the percentage deviation between the reward value of each round and the average of the reward values over multiple rounds. When the fluctuation range of reward values over multiple consecutive rounds reaches the stability judgment standard, the reward value is determined to be stable. The fit is calculated by weighted summation of the compliance rate and the data deviation of the model input data. The compliance rate is the proportion of data that meets the input requirements of the subsequent analysis model to the total amount of preprocessed data. The data deviation is the degree of deviation between the preprocessed data and the sample mean.
[0010] Furthermore, in step S4, during the processing of real-time data, the feature information of the real-time data is continuously collected and compared with the feature information of the data to be processed obtained in step S1. When there is a significant difference between the feature information of the real-time data and the feature information of the data to be processed in step S1, steps S1 to S3 are re-executed. After updating the optimal preprocessing strategy, the updated optimal preprocessing strategy is used to continue processing the real-time data.
[0011] Furthermore, in step S6, the data analysis operation set includes a data feature extraction method, an analysis model type, and a combination order of analysis steps. The data feature extraction method is used to extract key features from standard data, the analysis model type is used to perform analysis operations on the extracted key features, and the combination order of analysis steps clarifies the sequential execution logic of data feature extraction and analysis model operation. In the reward function, the weighted sum of the matching degree between the analysis result and the target and the analysis efficiency is calculated using the following formula: ,in, For the reward function value, The weighting coefficients for the matching degree. To analyze the degree of matching between the results and the target, The processing time for the analysis process, The computational resource utilization rate and weighting coefficients in the analysis process. Based on the priority of the user's analysis goals, the system receives the user's priority settings for "result accuracy" and "processing efficiency" through the user interaction module, and converts the priorities into corresponding weight coefficients, eliminating the need for manual setting of specific values.
[0012] Furthermore, in step S7, the iterative process needs to monitor the reward values for multiple consecutive rounds. When the reward values for multiple consecutive rounds reach the stability judgment criterion, the reward value is determined to be stable and the optimal analysis path is determined. During the analysis, the agent's state information, the selected analysis action sequence, the calculated reward value, and the next state information after executing the analysis action sequence are combined into experience data and stored in the experience replay pool. Each time the agent updates its parameters, it samples experience data from the experience replay pool and updates the neural network parameters based on the sampled experience data. During the neural network parameter update process, the parameter adjustment amount is calculated using the following formula: ,in, This represents the adjustment amount for the current neural network parameters. This is the learning rate coefficient. For the current loss function with respect to parameters gradient, The momentum coefficient, The previous parameter adjustment amount, learning rate coefficient With momentum coefficient Based on the distribution characteristics of empirical data in the empirical replay pool, the parameters are adaptively determined by calculating the variance of the empirical data and dynamically adjusting the coefficients according to the magnitude of the variance, so that the parameter update speed is adapted to the changes in data distribution.
[0013] The computer data intelligent processing system based on artificial intelligence is applicable to any of the above-mentioned computer data intelligent processing methods based on artificial intelligence. The system includes: a data acquisition module, an adaptive preprocessing module, a user interaction module, an intelligent analysis module, a model storage module, a result output module, and a log management module. The data acquisition module is used to acquire the data to be processed and its feature information in real time, and convert the feature information into a state vector. The adaptive preprocessing module has a built-in reinforcement learning agent based on the Q-learning algorithm, which is used to perform preprocessing strategy optimization and dynamic preprocessing of real-time data, and output standard data. The user interaction module is used to receive user input of analysis objectives and parameter settings; The intelligent analysis module integrates a DQN-based intelligent agent for analysis path planning, which is used to perform analysis path planning and personalized analysis of standard data; The model storage module is used to store analysis models, preprocessing rules, and reinforcement learning agent parameters; The results output module is used to output analysis results, and the log management module is used to store data during system operation.
[0014] Furthermore, the data acquisition module incorporates statistical analysis algorithms and format recognition algorithms. The statistical analysis algorithm calculates the noise distribution ratio and missing value ratio of the data to be processed, while the format recognition algorithm identifies the data format of the data to be processed. The result output module supports converting the analysis results into various types of output files and provides feedback on the analysis results through interface display, email push, and API interface calls. The result output module is also used to receive user evaluation information on the analysis results and transmit the evaluation information to the intelligent analysis module, providing a reference for the intelligent analysis module to optimize the analysis path.
[0015] Compared with existing technologies, this AI-based computer data intelligent processing method and system has the following beneficial effects: This invention constructs an adaptive dynamic data preprocessing mechanism and uses reinforcement learning algorithms to enable the system to automatically adjust the preprocessing strategy based on real-time data characteristics. This breaks through the limitations of traditional static preprocessing that relies on fixed rules, achieving precise adaptation between preprocessed data and subsequent analysis models. This effectively improves the timeliness and accuracy of data processing. By introducing a reinforcement learning framework to achieve intelligent data analysis path planning, the system can dynamically generate the optimal analysis path based on data characteristics and user analysis goals. This solves the problem that traditional fixed analysis processes cannot respond to personalized needs, achieving a high degree of intelligence and personalization in the analysis process. Through a complete system design, data acquisition, preprocessing, analysis, and result output form a closed loop, ensuring efficient collaboration among all links. This provides more accurate and efficient data analysis support for subsequent business decisions, further unlocking data value and promoting the development of data-driven applications.
[0016] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0018] Figure 1A flowchart of a computer data intelligent processing method based on artificial intelligence; Figure 2 This is a flowchart of a computer data intelligent processing method based on artificial intelligence. Figure 3 This is a schematic diagram of the structure of a computer data intelligent processing system based on artificial intelligence. Detailed Implementation
[0019] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0020] The computer data intelligent processing method and system based on artificial intelligence provided by this invention are described in [reference needed]. Figure 1 , Figure 2 and Figure 3 The details are as follows: Methodologically, the system first acquires real-time feature information of the data to be processed through a data acquisition module. These features include data format and noise distribution ratio, and are then converted into state vectors recognizable by the reinforcement learning agent. Next, a reinforcement learning agent based on the Q-learning algorithm is constructed. The action space is set as a set of data preprocessing operations containing cleaning and transformation rules, and the reward function is defined as the fit between the preprocessed data and the subsequent analysis model. The state vector is input into the agent, which selects and executes preprocessing actions to obtain preprocessed data. The fit feedback is calculated to update the Q-value table and adjust the strategy. This process iterates until the reward value stabilizes to determine the optimal preprocessing strategy. This strategy is then used to process real-time data and output standard data. If the real-time data features differ significantly from the initial data, the strategy needs to be re-optimized.
[0021] Next, the user's analysis target is acquired, and standard data characteristic parameters are extracted and combined to form the agent's state information. A DQN-based analysis path planning agent is then constructed. The action space is a set of data analysis operations including feature extraction and model type. The reward function is a weighted sum of the matching degree between the analysis result and the target, and the analysis efficiency. The state information is input into the agent, an analysis action sequence is selected for execution, a reward value is calculated and fed back, and parameters are updated. The process iterates until the reward value stabilizes to determine the optimal analysis path. The analysis is then executed according to the path, and the results are output. During the iteration, empirical data is also stored for parameter updates.
[0022] The system comprises seven modules, including data acquisition and adaptive preprocessing. Each module has a clear division of labor: the data acquisition module acquires and transforms data, the adaptive preprocessing module optimizes preprocessing strategies, the user interaction module receives user requests, the intelligent analysis module plans analysis paths, the model storage module stores relevant data, the result output module outputs results and receives evaluations, and the log management module stores system operation data. All modules work together to ensure efficient data processing.
[0023] Example 1
[0024] This embodiment applies to the transaction data processing scenario of financial institutions. Financial institutions generate massive amounts of transaction data daily, covering customer transfer records, consumption details, account balance changes, etc. This data comes in various formats and suffers from issues such as noisy data and missing values. Traditional static preprocessing rules are difficult to adapt to dynamic data changes, and fixed analysis processes cannot meet the analytical needs of different business scenarios. See [link to relevant documentation]. Figure 1 , Figure 2 and Figure 3 This invention can achieve efficient processing and accurate analysis of financial transaction data through adaptive preprocessing and intelligent analysis path planning, providing support for risk management and business decision-making.
[0025] In this embodiment, the data acquisition module is first activated. This module incorporates statistical analysis and format recognition algorithms. After acquiring financial transaction data to be processed in real time, the format recognition algorithm identifies the data format (e.g., determining whether a batch of data is in CSV format or JSON format). The statistical analysis algorithm calculates the noise distribution ratio (e.g., the proportion of abnormal transaction records to total transaction records) and the proportion of missing values (e.g., the proportion of transaction records lacking merchant information). Simultaneously, it calculates the statistical data dimensions (e.g., the number of fields contained in each transaction record) to obtain the feature information of the data to be processed. Subsequently, according to standardized coding rules, an encoding mapping relationship is set based on the data type of the feature information, and a weighted encoding is used to calculate the state vector. The calculation formula is as follows: ,in, For state vectors, The types and quantities of characteristic information in financial transaction data (such as four categories of characteristics: data format, noise distribution ratio, etc.) ), For the first The encoding weights of class feature information are determined by calculating the contribution of each type of feature to the selection of the preprocessing strategy using the random forest algorithm and then normalizing them. For the first Class feature information Encoding functions (e.g., encoding CSV as 1 and JSON as 2 for data format characteristics).
[0026] Next, a reinforcement learning agent based on the Q-learning algorithm is constructed. The agent's action space is defined as a set of preprocessing operations for financial transaction data, which includes cleaning rules and transformation rules. Cleaning rules cover outlier removal threshold adjustment (e.g., adjusting the outlier judgment threshold according to a reasonable range of transaction amounts) and duplicate data filtering methods (e.g., filtering duplicate records based on transaction serial numbers). Transformation rules include data standardization methods (e.g., converting transaction amounts in different currencies to a unified currency unit) and categorized data encoding methods (e.g., encoding merchant types "restaurant" as 001 and "retail" as 002). Simultaneously, a reward function is defined as the fit between the preprocessed data and subsequent analysis models (e.g., risk identification models).
[0027] The calculated state vector is input into the agent, which then selects preprocessing actions from the action space. For example, it might choose the action combination of "adjusting the threshold for removing abnormal transaction amounts + filtering duplicate records by transaction serial number + unifying the currency unit" to preprocess the financial transaction data, resulting in preprocessed data. The fit is then calculated and fed back to the agent as a reward. The fit is obtained by weighted summing the compliance rate and data deviation of the model input data. The compliance rate is the proportion of data that meets the input requirements of the risk identification model to the total amount of preprocessed data. The data deviation is the degree of deviation between the preprocessed data and the mean of the transaction data samples (such as the average historical normal transaction amount). The agent updates the Q-value table based on the reward and adjusts its strategy, while simultaneously monitoring the fluctuation range of the reward values over multiple rounds (calculating the percentage deviation of each round's reward value from the average of multiple rounds). When the fluctuation range reaches the stability criterion, iteration stops, and the optimal preprocessing strategy is determined.
[0028] The system employs an optimal preprocessing strategy to process real-time financial transaction data and outputs standard data. During processing, the data acquisition module continuously collects feature information from the real-time data and compares it with the initially acquired feature information of the data to be processed. If significant differences are found in the real-time data features (e.g., the noise distribution ratio suddenly increases from 1% to 5%), the above steps of data acquisition, state vector transformation, Q-learning agent construction, and iterative optimization are repeated. After updating the optimal preprocessing strategy, the system continues to process the real-time data.
[0029] Subsequently, the analysis objectives from financial institutions are received through the user interaction module. If the objective is "identifying high-risk transactions," characteristic parameters of standard data (such as transaction amount fluctuation range, transaction location change frequency, etc.) are extracted and combined with the analysis objective to form the agent's state information. An analysis path planning agent based on DQN is constructed, defining the action space as a set of data analysis operations. This set includes data feature extraction methods (such as extracting transaction amount fluctuation features and transaction time interval features), analysis model types (such as selecting a support vector machine model or a neural network model), and the order of analysis steps (such as extracting features first and then inputting them into the model for computation). The reward function is set as the weighted sum of the matching degree between the analysis result and the objective, and the analysis efficiency. The calculation formula is: ,in, For the reward function value, The weighting coefficient for the matching degree is determined based on the user's priority settings for "result accuracy" and "processing efficiency." If the user is more concerned about the accuracy of risk identification, The value is biased towards 0.7). To analyze the degree of matching between the results and the goal of "identifying high-risk transactions" (such as the proportion of high-risk transactions identified by the model that are confirmed to be genuine high-risk transactions by manual verification). The processing time for the analysis process, This refers to the computational resource utilization rate during the analysis process.
[0030] The state information is input into the DQN analysis path planning agent. The agent selects and executes an analysis action sequence (e.g., "extract transaction amount fluctuation features + select neural network model + feature extraction followed by model calculation"). The reward value is calculated and fed back to the agent. Simultaneously, the agent's state information, selected action sequence, reward value, and next state information after the action are combined to form experience data, which is stored in the experience replay pool. When updating parameters, the agent samples experience data from the experience replay pool and updates the neural network parameters based on the sampled data. The formula for calculating the parameter adjustment is: in, This represents the adjustment amount for the current neural network parameters. This is the learning rate coefficient. For the current loss function with respect to parameters gradient, The momentum coefficient, The amount of the previous parameter adjustment and (Dynamically adjusted based on the variance of empirical data).
[0031] During the iteration process, the reward values are monitored for multiple consecutive rounds. When a stable judgment standard is reached, the optimal analysis path is determined, and the analysis is executed according to this path. The high-risk transaction identification results are output through the result output module. The result output module supports converting the results into Excel reports, PDF reports, and other formats. The results are displayed through the system interface, pushed to business personnel's email addresses, and accessed via API calls (for the risk control system to obtain data). Simultaneously, the module receives evaluation information from business personnel regarding the analysis results (such as "high accuracy" or "missed risk transactions"), and transmits this evaluation information to the intelligent analysis module for reference in subsequent optimization of the analysis path. Furthermore, the model storage module stores the risk identification analysis model, preprocessing rules, and reinforcement learning agent parameters, while the log management module records data during system operation (such as data acquisition time, preprocessing actions, and analysis time).
[0032] This embodiment, by applying the present invention, achieves dynamic adjustment of the preprocessing strategy for financial transaction data, solving the problem of poor adaptability of traditional static rules and improving the adaptability of preprocessed data to the analysis model. Simultaneously, it dynamically generates the optimal analysis path based on the goal of "identifying high-risk transactions," improving the accuracy and efficiency of risk identification. The collaborative work of each module in the system forms a closed loop, ensuring not only the efficiency and accuracy of financial transaction data processing but also providing reliable data support for financial institutions' risk management, helping to reduce business risks and enhance the scientific nature of decision-making.
[0033] Example 2
[0034] This embodiment applies to the medical image data processing scenario in medical institutions. Medical institutions generate a large amount of medical image data daily, including CT, MRI, and ultrasound images. Traditional medical image processing relies on manually set fixed preprocessing workflows, which are difficult to adapt to the characteristics of image data from different devices and locations. Furthermore, fixed analysis paths cannot meet the diverse analytical needs in clinical diagnosis (such as determining the benignity or malignancy of tumors and measuring lesion size). See also Figure 1 , Figure 2 and Figure 3 This invention enables automated and precise processing of medical image data through an adaptive preprocessing mechanism and intelligent analysis path planning, providing efficient and reliable data analysis support for clinical diagnosis.
[0035] In this embodiment, the data acquisition module is first activated. This module incorporates statistical analysis and format recognition algorithms. After acquiring the medical image data to be processed in real time, the format recognition algorithm identifies the image data format (e.g., determining that a batch of CT images is in DICOM format and a batch of ultrasound images is in JPEG format). The statistical analysis algorithm calculates the noise distribution ratio (e.g., the proportion of image frames containing artifacts to the total number of image frames) and the proportion of missing values (e.g., the proportion of images with missing scan slice thickness records). Simultaneously, it analyzes the statistical dimensions (e.g., the pixel resolution and frame count of the images) to form the feature information of the data to be processed. Subsequently, according to standardized coding rules, an encoding mapping relationship is established based on the data type of the feature information. A state vector is calculated through weighted coding, using the following formula: ,in, For state vectors, The types and quantities of feature information in medical image data (such as five categories of features including data format and noise distribution ratio) ), For the first The encoding weights of class feature information (the contribution of each type of feature to the selection of preprocessing strategy is calculated by the random forest algorithm and determined after normalization. For example, if the noise distribution ratio has a greater impact on the preprocessing strategy, its weight value is higher). For the first Class feature information The encoding functions (such as encoding DICOM format as 1, JPEG format as 2, and encoding "low" artifact percentage as 0, "medium" as 1, and "high" as 2).
[0036] Next, a reinforcement learning agent based on the Q-learning algorithm is constructed. The agent's action space is defined as a set of medical image data preprocessing operations, which includes cleaning rules and transformation rules. Cleaning rules cover outlier removal threshold adjustment (e.g., adjusting the artifact removal threshold according to a reasonable range of image grayscale values) and duplicate data filtering methods (e.g., filtering duplicate uploaded image data based on image scan time and patient ID). Transformation rules include data standardization methods (e.g., mapping image grayscale values generated by different devices to the same range) and classification data encoding methods (e.g., encoding the image examination site "head" as 001 and "chest" as 002). Simultaneously, a reward function is defined as the fit between the preprocessed data and the subsequent analysis model (e.g., a tumor detection model).
[0037] The calculated state vector is input into the agent, which selects a combination of preprocessing actions from the action space to execute. For example, it might select the action sequence "adjusting artifact removal threshold + filtering duplicate images by scan time and patient ID + unifying image grayscale range" to preprocess the medical image data, resulting in preprocessed data. The fitness score is then calculated and fed back to the agent as a reward. The fitness score is obtained by weighted summing the compliance rate and data deviation of the model input data. The compliance rate is the proportion of data that meets the input requirements of the tumor detection model (e.g., image resolution, grayscale range) to the total amount of preprocessed data. The data deviation is the degree of deviation between the preprocessed image data and the mean of standard image samples (e.g., the mean grayscale value of normal tissue). The agent updates the Q-value table and adjusts the strategy based on the reward value, while simultaneously monitoring the fluctuation range of the reward value over multiple rounds (achieved by calculating the percentage deviation between each round's reward value and the average of multiple rounds' reward values). When the fluctuation range reaches the stability criterion, iteration stops, and the optimal preprocessing strategy is determined.
[0038] The system employs an optimal preprocessing strategy to process real-time medical image data and outputs standard image data. During processing, the data acquisition module continuously collects feature information from the real-time image data and compares it with the feature information of the initially acquired data to be processed. If significant differences are found in the real-time data features (e.g., the proportion of artifacts in the received MRI images suddenly increases from "low" to "high" at a certain time period), the data acquisition, state vector transformation, Q-learning agent construction, and iterative optimization steps are repeated. After updating the optimal preprocessing strategy, the system continues to process the real-time image data.
[0039] Subsequently, the analysis target from the clinician is received through the user interaction module. If the target is "to determine whether a lesion in a lung image is a malignant tumor," characteristic parameters of the standard image data (such as lesion area, edge smoothness, density uniformity, etc.) are extracted and combined with the analysis target to form the agent's state information. An analysis path planning agent based on DQN is constructed, defining the action space as a set of medical image data analysis operations. This set includes data feature extraction methods (such as extracting texture and morphological features of the lesion area), analysis model types (such as selecting a convolutional neural network model or a support vector machine model), and the order of analysis steps (such as first extracting lesion features and then inputting them into the model for classification). The reward function is set as the weighted sum of the matching degree between the analysis result and the target, and the analysis efficiency. The calculation formula is: ,in, For the reward function value, The weighting coefficient for the matching degree is determined based on the clinician's priority settings for "diagnostic accuracy" and "analysis speed." If the clinician prioritizes diagnostic accuracy, The value is biased towards 0.8). To analyze the degree of matching between the results and the goal of "determining whether the lesion is a malignant tumor" (e.g., the proportion of results in which the model determines that the lesion is a malignant tumor, but which are confirmed by pathological examination). The processing time for the analysis process (the total time from inputting standard image data to outputting analysis results). This refers to the computational resource utilization during the analysis process (such as the ratio of CPU to GPU resources used during the analysis).
[0040] The state information is input into the DQN analysis path planning agent. The agent selects an analysis action sequence to execute, such as the action combination of "extracting texture and morphological features of lesion areas + using a convolutional neural network model + feature extraction followed by model classification," and performs the analysis operation. The reward value is calculated and fed back to the agent. Simultaneously, the agent's state information, the selected action sequence, the reward value, and the next state information after the action execution are combined into experience data and stored in the experience replay pool. When updating parameters, the agent randomly samples experience data from the experience replay pool and updates the neural network parameters based on the sampled data. The formula for calculating the parameter adjustment is: ,in, This represents the adjustment amount for the current neural network parameters. This is the learning rate coefficient. For the current loss function with respect to parameters gradient, The momentum coefficient, The amount of the previous parameter adjustment and The variance of the empirical data in the experience replay pool is dynamically adjusted; for example, if the distribution of the empirical data fluctuates greatly, the variance is reduced. (To slow down the parameter update speed).
[0041] During the iteration process, the reward value is continuously monitored for multiple rounds. When the reward value reaches a stable judgment standard, the optimal analysis path is determined, the analysis is executed according to this path, and the benign or malignant lesion judgment result is output through the result output module. The result output module supports converting the analysis results into diagnostic report documents (such as Word format and PDF format), and provides feedback through system interface display (for doctors to view in real time), push to doctor workstations, and API interface calls (for electronic medical record system to obtain data). At the same time, it receives doctors' evaluation information on the analysis results (such as "accurate judgment" or "misdiagnosis exists"), and transmits the evaluation information to the intelligent analysis module to provide a reference for subsequent optimization of the analysis path. In addition, the model storage module stores the tumor detection analysis model, image preprocessing rules, and reinforcement learning agent parameters, and the log management module records key data during system operation (such as data acquisition time, preprocessing action type, analysis time, doctor evaluation results, etc.).
[0042] This embodiment, by applying the present invention, achieves dynamic adaptation of medical image data preprocessing strategies, solving the problem that traditional fixed preprocessing workflows are unable to handle image data from different devices and of different qualities, and improving the adaptability of preprocessed data to clinical analysis models. Simultaneously, it dynamically generates the optimal analysis path based on the doctor's diagnostic goals, improving the accuracy and efficiency of medical image analysis. The various modules of the system work together to form a complete closed loop, not only reducing the manual processing burden on doctors but also providing scientific data support for clinical diagnosis, helping to reduce misdiagnosis rates and improve the diagnostic and treatment levels of medical institutions.
[0043] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A computer data intelligent processing method based on artificial intelligence, characterized in that, The method includes the following steps: S1. The feature information of the data to be processed is acquired in real time through the data acquisition module, and the feature information is converted into a state vector that can be recognized by the reinforcement learning agent. S2. Construct a reinforcement learning agent based on the Q-learning algorithm, define the agent's action space as a set of data preprocessing operations, and define the reward function as the fit between the preprocessed data and the subsequent analysis model. S3. Input the state vector into the agent, the agent selects a preprocessing action to execute and obtains preprocessed data, calculates the fitness degree as a reward and feeds it back to the agent, the agent updates the Q-value table to adjust the strategy, repeats the iteration until the reward value is stable, and determines the optimal preprocessing strategy. S4. Use the optimal preprocessing strategy to process real-time data and output standard data; S5. Obtain the user's analysis target, extract standard data characteristic parameters, and combine them into agent state information; S6. Construct an intelligent agent for analysis path planning based on DQN, define the action space of the agent as a set of data analysis operations, and define the reward function as a weighted sum of the matching degree between the analysis result and the target and the analysis efficiency; S7. Input the state information into the agent, the agent selects the action sequence to be analyzed and executes it, calculates the reward value and updates the parameters, iterates until the reward value is stable, and determines the optimal analysis path; S8. Perform the analysis according to the optimal analysis path and output the analysis results.
2. The computer data intelligent processing method based on artificial intelligence according to claim 1, characterized in that, The feature information of the data to be processed in step S1 includes data format, noise distribution ratio, missing value ratio, and data dimension. The state vector transformation process adopts standardized coding rules. The standardized coding rules set corresponding coding mapping relationships based on the data type of the feature information. The state vector is calculated through weighted coding of feature information. The calculation formula is as follows: ,in, For state vectors, The number of types of feature information. For the first Encoding weights of class feature information For the first Class feature information The encoding function.
3. The computer data intelligent processing method based on artificial intelligence according to claim 1, characterized in that, The data preprocessing operation set in step S2 includes cleaning rules and transformation rules. The cleaning rules include outlier removal threshold adjustment and duplicate data filtering methods. The transformation rules include data standardization methods and classification data encoding methods. By clarifying the specific operation types of the cleaning rules and transformation rules, a complete set of data preprocessing operations is formed, providing a clear range of action choices for the reinforcement learning agent.
4. The computer data intelligent processing method based on artificial intelligence according to claim 1, characterized in that, In step S3, the iterative process needs to monitor the fluctuation range of reward values over multiple consecutive rounds. The fluctuation range of reward values is obtained by calculating the percentage deviation between the reward value of each round and the average of the reward values over multiple rounds. When the fluctuation range of reward values over multiple consecutive rounds reaches the stability judgment standard, the reward value is determined to be stable. The fit is calculated by weighted summation of the compliance rate and the data deviation of the model input data. The compliance rate is the proportion of data that meets the input requirements of the subsequent analysis model to the total amount of preprocessed data. The data deviation is the degree of deviation between the preprocessed data and the sample mean.
5. The computer data intelligent processing method based on artificial intelligence according to claim 1, characterized in that, In step S4, during the processing of real-time data, the feature information of the real-time data is continuously collected and compared with the feature information of the data to be processed obtained in step S1. When there is a significant difference between the feature information of the real-time data and the feature information of the data to be processed in step S1, steps S1 to S3 are re-executed. After updating the optimal preprocessing strategy, the updated optimal preprocessing strategy is used to continue processing the real-time data.
6. The computer data intelligent processing method based on artificial intelligence according to claim 1, characterized in that, The data analysis operation set in step S6 includes data feature extraction methods, analysis model types, and the order of analysis steps. The data feature extraction methods are used to extract key features from standard data, the analysis model types are used to perform analysis and calculation on the extracted key features, and the order of analysis steps clarifies the execution logic of data feature extraction and analysis model calculation. In the reward function, the weighted sum of the matching degree between the analysis result and the target and the analysis efficiency is calculated using the following formula: ,in, For the reward function value, The weighting coefficients for the matching degree. To analyze the degree of matching between the results and the target, The processing time for the analysis process, This refers to the computational resource utilization rate during the analysis process.
7. The computer data intelligent processing method based on artificial intelligence according to claim 1, characterized in that, In step S7, the iterative process requires monitoring the reward values for multiple consecutive rounds. When the reward values for multiple consecutive rounds reach the stability criterion, the reward value is determined to be stable, and the optimal analysis path is identified. During the analysis, the agent's state information, the selected analysis action sequence, the calculated reward value, and the next state information after executing the analysis action sequence are combined into empirical data and stored in the experience replay pool. Each time the agent updates its parameters, it samples empirical data from the experience replay pool and updates the neural network parameters based on the sampled empirical data. During the neural network parameter update process, the parameter adjustment amount is calculated using the following formula: ,in, This represents the adjustment amount for the current neural network parameters. This is the learning rate coefficient. For the current loss function with respect to parameters gradient, The momentum coefficient, This is the amount adjusted in the previous parameter adjustment.
8. A computer data intelligent processing system based on artificial intelligence, applicable to the computer data intelligent processing method based on artificial intelligence as described in any one of claims 1-7, characterized in that, The system includes: a data acquisition module, an adaptive preprocessing module, a user interaction module, an intelligent analysis module, a model storage module, a result output module, and a log management module; The data acquisition module is used to acquire the data to be processed and its feature information in real time, and convert the feature information into a state vector. The adaptive preprocessing module has a built-in reinforcement learning agent based on the Q-learning algorithm, which is used to perform preprocessing strategy optimization and dynamic preprocessing of real-time data, and output standard data. The user interaction module is used to receive user input of analysis objectives and parameter settings; The intelligent analysis module integrates a DQN-based intelligent agent for analysis path planning, which is used to perform analysis path planning and personalized analysis of standard data; The model storage module is used to store analysis models, preprocessing rules, and reinforcement learning agent parameters; The results output module is used to output analysis results, and the log management module is used to store data during system operation.
9. The computer data intelligent processing system based on artificial intelligence according to claim 8, characterized in that, The data acquisition module incorporates statistical analysis algorithms and format recognition algorithms. The statistical analysis algorithm calculates the noise distribution ratio and missing value ratio of the data to be processed, while the format recognition algorithm identifies the data format of the data to be processed. The result output module supports converting the analysis results into various types of output files and provides feedback on the analysis results through interface display, email push, and API interface calls. The result output module is also used to receive user evaluation information on the analysis results and transmit the evaluation information to the intelligent analysis module, providing a reference for the intelligent analysis module to optimize the analysis path.