A Database Information Acquisition and Analysis Method for Pharmacological Analysis of Heatstroke
By determining the data status based on the diversity and stability of rat input data, and selecting appropriate data processing methods and search strategies, the problem of low database information acquisition efficiency in heatstroke pharmacological analysis was solved, and data accuracy and search efficiency were improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2026-03-13
AI Technical Summary
In existing technologies, the database information acquisition efficiency for heatstroke pharmacological analysis is low, and different data acquisition methods cannot be selected according to the actual performance of the data, resulting in poor information acquisition efficiency.
The data state is determined based on the data diversity and data instability of the rat input data. Data processing methods such as simulated data supplementation or keyword analysis are selected. Search and adjustment methods are determined by factors such as the proportion of effective standard index, the distribution of unstable paragraphs, the number of times effective keywords are combined and the interval distance, so as to optimize the data acquisition process.
The improved selection of data processing methods aligns with actual work scenarios, enhancing data accuracy and search efficiency, and providing better data support for subsequent pharmacological analysis.
Smart Images

Figure CN120336615B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis, and more particularly to a database information acquisition and analysis method for pharmacological analysis of heatstroke. Background Technology
[0002] Exertional heatstroke is a type of severe heatstroke caused by strenuous physical activity in hot and humid environments. This leads to an imbalance in the body's regulatory functions, resulting in heat production exceeding heat dissipation, a rapid rise in core temperature, and a severe, potentially fatal illness accompanied by burning skin, altered consciousness, and multiple organ dysfunction. Currently, researchers often construct rat models and conduct pharmacological analyses of relevant therapeutic drugs. However, due to time and resource limitations, experimental evidence and insights into the mechanisms by which these drugs play a role in heatstroke prevention are limited. With advancements in science and technology, data sharing has contributed to overcoming the limitations of experiments conducted by individuals or institutions. A key challenge for those skilled in the art is how to quickly and effectively match the necessary experimental and pharmacological data from a large volume of data.
[0003] Chinese Patent Publication No. CN116701353A discloses a method for constructing an extranodal lymphoma pathology database, including: Step 1: Determining a list of lymphoma subtypes to be included in the database based on the classification standards for tumors of the lymphatic and hematopoietic systems; Step 2: Generating a lymphoma subtype data package by obtaining at least a preset number of target cases for each lymphoma subtype according to the lymphoma subtype query list, wherein the target cases contain various clinical information; Step 3: Storing the lymphoma subtype data package to establish an extranodal lymphoma pathology database. This invention establishes a digital slide library for extranodal lymphoma pathology. Medical students can search digital pathology knowledge graphs, retrieve digital pathology images related to pathology teaching text content, and the images are accompanied by detailed clinical information. The library can automatically display the differences between normal images and images with detailed annotations of lesions. Students can highlight key points and save notes, facilitating their pathology learning and improving learning efficiency. However, the above technical solution has the following problems: it only determines the query list based on classification standards, and cannot select different data acquisition methods according to the actual data performance, resulting in poor database information acquisition efficiency. Summary of the Invention
[0004] Therefore, the present invention provides a database information acquisition and analysis method for pharmacological analysis of heatstroke, which overcomes the problem in the prior art that it is impossible to select different data acquisition methods according to the actual performance of the data, resulting in poor information acquisition efficiency.
[0005] To achieve the above objectives, the present invention provides a database information acquisition and analysis method for pharmacological analysis of heatstroke, comprising:
[0006] The data status is determined based on the data diversity and data instability of the rat input data, and the data processing method is determined based on the data status, namely, simulated data supplementation or keyword analysis.
[0007] In the process of supplementing the simulated data, the method of supplementation is determined by comparing the effective standard index ratio with the preset effective standard index ratio. The method is to supplement the simulated data based on the difference in the effective standard index ratio or the distribution status of the unstable segments.
[0008] The distribution of unstable segments is determined based on the number of unstable segments and the degree of distribution of unstable segments.
[0009] In keyword analysis, the usage status is determined based on the number of times effective keywords are combined and the interval distance, and the corresponding search method, either combined search or single search, is determined based on the usage status.
[0010] During the search for effective keywords, the page distribution status is determined based on the page similarity and information relevance of the target website, and the acquisition adjustment method is determined based on the page distribution status, namely, trajectory simulation adjustment or search simulation adjustment.
[0011] Under the preset completion conditions, pharmacological analysis was performed.
[0012] Furthermore, the data state of the rat input data is that the data diversity is less than the preset data diversity or the data instability is greater than or equal to the preset data instability, and the data processing method is simulated data supplementation;
[0013] In the supplementation of simulated data, the supplementation method is determined based on the proportion of effective standard indices;
[0014] If the effective standard index percentage is less than the preset effective standard index percentage, the supplementation method is to supplement the data by simulating the difference in the effective standard index percentage.
[0015] If the effective standard index is greater than or equal to the preset effective standard index percentage, the supplementation method is to supplement the data by simulating the distribution of unstable segments.
[0016] Furthermore, when supplementing simulated data based on the difference in the proportion of effective standardized indexes, the difference in the proportion of effective standardized indexes is detected, and the adjustment range of rat training parameters is determined based on the difference in the proportion of effective standardized indexes.
[0017] The difference between the adjustment range of the rat training parameters and the percentage of the effective standard index is positively correlated.
[0018] Supplementary data refers to the stored datasets within the database whose corresponding rat training parameters fall within the adjustment range of the rat input data's training parameters.
[0019] Furthermore, when supplementing the simulation data based on the distribution state of the unstable segments, the distribution state of the unstable segments is determined based on the number of unstable segments and the degree of distribution of the unstable segments.
[0020] If the distribution of unstable segments is such that the number of unstable segments is within the first preset range of unstable segment numbers, then the range of rat training parameter adjustment is determined based on the number of unstable segments. The difference between the rat training parameter adjustment range and the effective standard index ratio is positively correlated. The stored datasets in the search database whose corresponding rat training parameters are within the range of rat training parameter adjustment of the rat input data are recorded as supplementary data.
[0021] If the unstable segment distribution state is such that the number of unstable segments is within the second preset range of the number of unstable segments and the unstable segment distribution degree is within the second preset range of the unstable segment distribution degree, then the preset experimental environment difference degree is determined based on the unstable segment distribution degree, and the stored dataset that satisfies the experimental environment difference degree being less than the preset experimental environment difference degree is recorded as supplementary data.
[0022] Furthermore, when the rat input data is in a state where the data diversity is greater than or equal to the preset data diversity and the data stability is less than the preset data stability, the data processing method is keyword analysis.
[0023] In keyword analysis, both the rat input data and the supplementary data are recorded as data to be analyzed. Keywords in the experimental text corresponding to each data to be analyzed are detected, and keywords with a valid frequency greater than the preset valid frequency are recorded as valid keywords.
[0024] Furthermore, for a single effective keyword, the corresponding search method is determined based on the usage status of the effective keyword;
[0025] If the usage status of a valid keyword is that the number of times it is used is greater than or equal to the preset number of times it is used and the interval distance is less than the preset interval distance, then the search method for the valid keyword is a combination search.
[0026] If the usage status of a valid keyword is less than the preset number of combinations or the interval distance is greater than or equal to the preset interval distance, then the search method for that valid keyword is a standalone search.
[0027] Furthermore, the preset interval distance is determined based on the keyword density;
[0028] The preset interval distance is negatively correlated with keyword density.
[0029] Furthermore, during the search for effective keywords, the acquisition and adjustment methods are determined based on the page distribution of the target website;
[0030] If the page distribution status of the target website is such that the page similarity is greater than the preset page similarity or the information relevance is greater than the preset information relevance, the adjustment method is trajectory simulation adjustment;
[0031] If the page distribution status of the target website is such that the page similarity is less than or equal to the preset page similarity and the information relevance is less than or equal to the preset information relevance, the adjustment method is obtained through search simulation adjustment.
[0032] Furthermore, in the trajectory simulation adjustment, the number of trajectory change points and the trajectory change frequency are increased based on the maximum effective difference;
[0033] The maximum effective difference is positively correlated with the increase in the number of trajectory change points and the trajectory change frequency.
[0034] Furthermore, in the search simulation adjustment, the dwell time variation frequency and IP rotation frequency are determined based on the search depth;
[0035] Search depth is positively correlated with the frequency of dwell time changes and IP rotation frequency.
[0036] Compared with the prior art, the technical solution of the present invention determines the data status based on the data diversity and data stability of the rat input data, and determines the data processing method (simulated data supplementation or keyword analysis) based on the data status. The data status effectively reflects whether the current data can meet the needs of pharmacological analysis, and selects different data processing methods accordingly. This makes the selection of data processing methods more in line with the actual working scenario, improves data accuracy, and provides data support for subsequent pharmacological analysis.
[0037] Furthermore, in the technical solution of the present invention, in the process of supplementing simulated data, the method of supplementation is determined by comparing the effective standard index ratio with the preset effective standard index ratio. The method of supplementing simulated data is based on the difference in the effective standard index ratio or the distribution state of unstable segments. The difference in the effective standard index ratio reflects whether the data stability of the reference data required for pharmacological analysis meets the requirements. The distribution state of unstable segments characterizes the degree of difference between data. Different supplementation methods are selected accordingly to make the supplemented data more in line with the current data requirements, thereby improving the accuracy of subsequent pharmacological analysis.
[0038] Furthermore, in the technical solution of this invention, the keyword analysis determines the usage status based on the number of times effective keywords are combined and the interval distance, and determines the corresponding search method as a combination search or a single search based on the usage status. The usage status reflects the correlation between the current keywords, thereby enabling the use of effective search methods in subsequent search processes to improve search efficiency.
[0039] Furthermore, in the technical solution of the present invention, during the search for effective keywords, the page distribution status is determined based on the page similarity and information relevance of the target website, and the acquisition adjustment method is determined to be trajectory simulation adjustment or search simulation adjustment based on the page distribution status, so as to avoid the website's restrictions on the search program during the search process and improve the data search capability. Attached Figure Description
[0040] Figure 1 This is a schematic diagram of the database information acquisition and analysis method for the pharmacological analysis of heatstroke according to the present invention;
[0041] Figure 2 This is a flowchart illustrating how the present invention determines the data processing method based on the data status;
[0042] Figure 3 This is a flowchart illustrating how the present invention determines the corresponding search method based on the usage status of effective keywords. Detailed Implementation
[0043] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0044] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0045] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.
[0046] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0047] Please see Figures 1 to 3 As shown, this invention provides a database information acquisition and analysis method for pharmacological analysis of heatstroke, comprising:
[0048] The data status is determined based on the data diversity and data instability of the rat input data, and the data processing method is determined based on the data status, namely, simulated data supplementation or keyword analysis.
[0049] In the process of supplementing the simulated data, the method of supplementation is determined by comparing the effective standard index ratio with the preset effective standard index ratio. The method is to supplement the simulated data based on the difference in the effective standard index ratio or the distribution status of the unstable segments.
[0050] The distribution of unstable segments is determined based on the number of unstable segments and the degree of distribution of unstable segments.
[0051] In keyword analysis, the usage status is determined based on the number of times effective keywords are combined and the interval distance, and the corresponding search method, either combined search or single search, is determined based on the usage status.
[0052] During the search for effective keywords, the page distribution status is determined based on the page similarity and information relevance of the target website, and the acquisition adjustment method is determined based on the page distribution status, namely, trajectory simulation adjustment or search simulation adjustment.
[0053] Under the preset completion conditions, pharmacological analysis was performed.
[0054] This invention is applied to network pharmacological analysis of heatstroke-related drugs. By constructing and analyzing biological networks, it reveals the interaction between drugs and the body. This invention is applied to the front-end data acquisition and storage stage. The rat input data in this invention consists of several datasets uploaded by the user. Each dataset includes rat training parameters, rat physical parameters, rat experimental data, and experimental environment data. This invention utilizes an information database that stores several stored datasets. The stored datasets also include rat training parameters, rat physical parameters, rat experimental data, and experimental environment data. The rat experimental data includes functional indicators measured daily during the rat training phase, as well as body temperature and functional indicators measured before and after each experiment. Functional indicators include ALT, AST, BUN, CK, and CREA. The units for ALT, AST, and CK are U / L, the unit for BUN is mmol / L, and the unit for CREA is μmol / L.
[0055] In this invention, rats need to undergo relevant training before the experiment, including: training on a treadmill with a treadmill wheel for a preset time every day. Starting from the second day, the preset time each day is increased by the preset time difference compared to the preset time of the previous day, until the preset number of training days is reached. During each training session, the initial speed of the treadmill wheel is the preset initial speed, and the preset speed difference is increased every 2 minutes.
[0056] The rat physical parameters include the rat's age, weight, and blood pressure;
[0057] The experimental text is the textual description of the rat input data, including but not limited to descriptions of the training process, rat training results, and rat symptoms.
[0058] The rat training parameters include preset time, preset time difference, preset training days, preset initial speed, and preset speed difference. One embodiment is provided in which the preset time is 20 minutes, the preset time difference is 10 minutes, the preset number of days is 5 days, the preset initial speed is 5 meters / minute, and the preset speed difference is 1 meter / minute.
[0059] The rat experiment in this invention includes: rats being randomly divided into 8 groups based on their weight, with a weight difference of less than 400g between groups. In subsequent experiments, body temperature, blood, and tissue samples were collected from each group. A temperature monitoring capsule was implanted into the experimental rats, and core body temperature was measured every 5 minutes. Each rat was anesthetized with isoflurane, and the sterile temperature monitoring capsule was implanted into the abdominal cavity through a small sterile incision. The rats were allowed to rest for 2 days post-surgery. On day 8, a treadmill was placed in an artificial climate chamber with a temperature of 40±1℃ and a relative humidity of 60±5%. The rats exercised under these high temperature and high humidity conditions. Core body temperature and physical condition were continuously monitored in real-time. When the rats' core body temperature exceeded 42℃ and they showed signs of unconsciousness, they were removed from the room and placed in a room temperature environment. Three hours later, serum, plasma, organ tissue, and fecal matter were collected from the rats and analyzed to obtain experimental data.
[0060] The preset completion conditions are the completion of simulated data supplementation or the end of the keyword search process in keyword analysis. Pharmacological analysis includes:
[0061] Users collect relevant drug active components and targets from rat input data, supplementary data, and keyword search data;
[0062] Drug composition analysis;
[0063] Prediction of potential targets: Utilizing online compound target databases to predict potential protein targets of active ingredients;
[0064] Collect disease targets, and combine disease-related gene or protein databases to collect potential targets for heatstroke;
[0065] Construct a "drug-target-disease" network, screen common targets, construct an interaction network, and build a PPI network;
[0066] Enrichment analysis was performed on potential targets to identify biological processes, cellular components, molecular functions, and signaling pathways significantly associated with heatstroke.
[0067] Molecular docking verification uses molecular docking software to verify the binding activity of predicted key compounds to targets, and evaluates the binding affinity between compounds and targets based on docking scores.
[0068] Data integration and interpretation;
[0069] It is worth noting that if no simulation data supplementation or keyword analysis is performed, then there is no need to collect relevant effective components and targets of the drug in the supplementary data or keyword search data. The above is content that is easy for those skilled in the art to understand, and will not be elaborated here.
[0070] In this invention, the preset threshold value can be set by the user based on historical experience and actual application scenarios. A method for setting the value is provided, which determines the preset threshold based on the detection values of past historical records. For example, for the preset effective standard index ratio, the detection values corresponding to the historical records that meet the user's needs are extracted, and the detection values are the effective standard index ratio. Data cleaning is performed on the effective standard index ratio to remove outliers. The outlier identification method can be Z-Score method or IQR method. The average value of the effective standard index ratio after removing outliers is recorded as the preset effective standard index ratio value. The preset threshold can also be optimized by deep learning model. This is easy for those skilled in the art to understand and will not be elaborated here.
[0071] This invention utilizes several historical records. Each historical record documents at least one past information acquisition and analysis process and the corresponding detection data. The detection data includes data diversity, data instability, data stability, number of unstable segments, distribution of unstable segments, number of effective times, and number of combinations. Each historical record is also marked with a pass / fail flag. The pass / fail flag indicates whether the information acquisition and analysis process corresponding to the historical record meets the user's needs. Whether the user's needs are met is determined by the user based on the actual information analysis results.
[0072] Specifically, the data state of the rat input data is that the data diversity is less than the preset data diversity or the data deviation is greater than or equal to the preset data deviation, and the data processing method is simulated data supplementation;
[0073] In the supplementation of simulated data, the supplementation method is determined based on the proportion of effective standard indices;
[0074] If the effective standard index percentage is less than the preset effective standard index percentage, the supplementation method is to supplement the data by simulating the difference in the effective standard index percentage.
[0075] If the effective standard index is greater than or equal to the preset effective standard index percentage, the supplementation method is to supplement the data by simulating the distribution of unstable segments.
[0076] Wherein, data diversity = (H / H0)×α1+(C / C0)×α2, H is the absolute value of the difference between the maximum and minimum values of rat physical parameters in the rat input data, C is the absolute value of the difference between the maximum and minimum values of rat training parameters in the rat input data, α1 is the first diversity coefficient, and α2 is the second diversity coefficient. The values of α1 and α2 are determined by the user based on historical experience and deep learning networks to determine the degree of influence of rat training parameters and rat physical parameters on the accuracy of subsequent pharmacological analysis, and the corresponding values of α1 and α2 are determined. For example, the greater the influence of rat training parameters on the accuracy of subsequent pharmacological analysis, the larger the value of α1. One possible value of α1 and α2 is provided: α1 = 0.5, α2 = 0.5.
[0077] The method for confirming data stability is as follows: extract each dataset of rat input data; for a single dataset, extract ALT, AST, BUN, CK, and CREA from the rat experimental data; calculate the reference area for each functional index; then calculate the maximum difference in reference area between the same functional indexes in different datasets and record it as the data stability. For a single functional index, the reference area is calculated by presenting the values measured in each experiment as data points on a two-dimensional coordinate system. The horizontal axis of the coordinate system represents the number of experiments when the values were measured, such as at the beginning of the first experiment, at the end of the first experiment, at the beginning of the second experiment, etc.; the vertical axis represents the numerical unit; connect adjacent points with straight lines to obtain a reference data line; and connect the data points with the minimum and maximum horizontal coordinates with perpendicular lines to the horizontal axis. The closed figure formed is recorded as the reference area.
[0078] In this embodiment of the invention, the preset data diversity is 1, the preset data instability is 20% of the maximum value of the reference area corresponding to the data instability, the effective standard index ratio is the number of effective standard indices / 5, and the effective standard index is a functional indicator whose data stability is greater than the preset data stability. For a single functional indicator, the calculation formula for its corresponding data stability M is as follows:
[0079]
[0080] Where Mi is the average of all values of this functional index measured in the i-th dataset. n is the total number of datasets corresponding to the rat input data, i = 1, 2, 3, ..., n. In this embodiment of the invention, the preset data stability value is 5% of M0.
[0081] Specifically, when supplementing simulated data based on the difference in the proportion of effective standardized indices, the difference in the proportion of effective standardized indices is detected and the adjustment range of rat training parameters is determined based on the difference in the proportion of effective standardized indices.
[0082] The difference between the adjustment range of the rat training parameters and the percentage of the effective standard index is positively correlated.
[0083] Supplementary data refers to the stored datasets within the database whose corresponding rat training parameters fall within the adjustment range of the rat input data's training parameters.
[0084] The effective standard index ratio difference is the absolute value of the difference between the effective standard index ratio and the preset effective standard index ratio.
[0085] In this invention, the range of rat training parameters is the range of values with the average value of the rat training parameters of the current rat input data as the median point, and the absolute value of the difference between the maximum and minimum values of the range and the median point is the same.
[0086] Specifically, when supplementing simulation data based on the distribution of unstable segments, the distribution of unstable segments is determined based on the number of unstable segments and the degree of distribution of unstable segments.
[0087] If the unstable segment distribution state is such that the number of unstable segments is within the first preset range of the number of unstable segments or the distribution degree of unstable segments is within the first preset range of the distribution degree of unstable segments, then the rat training parameter adjustment range is determined according to the number of unstable segments. The difference between the rat training parameter adjustment range and the effective standard index ratio is positively correlated. The stored datasets in the search database whose corresponding rat training parameters are within the rat input data rat training parameter adjustment range are recorded as supplementary data.
[0088] If the unstable segment distribution state is such that the number of unstable segments is within the second preset range of the number of unstable segments and the unstable segment distribution degree is within the second preset range of the unstable segment distribution degree, then the preset experimental environment difference degree is determined according to the unstable segment distribution degree, and the stored dataset that satisfies the experimental environment difference degree being less than the preset experimental environment difference degree is recorded as supplementary data.
[0089] The experimental environment variability was assumed to be positively correlated with the distribution of unstable segments.
[0090] The experimental environment data includes the experimental environment temperature and humidity. The temperature and humidity are each taken as the average of several measurements. The experimental environment variability includes the absolute value of the difference in experimental environment temperature and the absolute value of the difference in experimental environment humidity. The corresponding preset experimental environment variability is different. If the absolute value of the difference in experimental environment temperature and the absolute value of the difference in experimental environment humidity are both less than the corresponding preset experimental environment variability, the corresponding stored dataset is recorded as supplementary data. In the specific implementation of this invention, the preset experimental environment variability corresponding to the absolute value of the difference in experimental environment temperature is 5℃, and the preset experimental environment variability corresponding to the absolute value of the difference in experimental environment humidity is 4%.
[0091] The values within the first preset range of unstable segment quantity are all less than the preset unstable segment quantity. The values within the second preset range of unstable segment quantity are all greater than or equal to the preset unstable segment quantity. The values within the first preset range of unstable segment distribution degree are all less than the preset unstable segment distribution degree. The values within the second preset range of unstable segment distribution degree are all greater than or equal to the preset unstable segment distribution degree. The preset unstable segment quantity and preset unstable segment distribution degree are determined by extracting the unstable segment quantity and unstable segment distribution degree corresponding to the historical records that meet the user's needs. Data cleaning is performed on the unstable segment quantity and unstable segment distribution degree to remove outliers. The average values of the unstable segment quantity and unstable segment distribution degree after removing outliers are recorded as the preset unstable segment quantity and preset unstable segment distribution degree, respectively.
[0092] The method for confirming the number of unstable segments is as follows: extract the reference data lines for each rat's experimental data from the rat input data, divide each reference data line into several segments with the horizontal axis as the reference, and calculate the similarity between segments at the same position, including:
[0093] Randomly extract the same number of points from each paragraph;
[0094] For each point, calculate its relative position with all other points and quantize these relative positions into a set of discrete bins;
[0095] For any two paragraphs S1 and S2, the similarity Z is... D(C1, C2) is Earth Mover's Di stance between the two reference data lines, and Dmax is the maximum bin value of C1 and C2.
[0096] If a given position corresponds to a number of paragraphs that have a similarity to other paragraphs with a similarity less than the given number of paragraphs, then that position is recorded as an unstable paragraph position, and the number of unstable paragraph positions is recorded as the number of unstable paragraphs.
[0097] If two paragraphs have the same minimum x-coordinate and the same maximum x-coordinate, then the two paragraphs are in the same position.
[0098] Users can set the number of points extracted from a paragraph and the preset number of points according to the actual scenario. It is understandable that the greater the user's requirement for the accuracy of the determination of the number of unstable paragraphs, the greater the value of the number of points extracted from the paragraph and the preset number of points. One possible value is that the number of points extracted from each paragraph is 10, and the preset number is 70% of the total number of paragraphs at the corresponding position.
[0099] The method for confirming the distribution degree of unstable segments is to calculate the absolute value of the difference between the maximum and minimum values of the horizontal coordinates of the unstable segment positions, and record this absolute value as the distribution degree of unstable segments.
[0100] Specifically, when the rat input data is in a state where the data diversity is greater than or equal to the preset data diversity and the data stability is less than the preset data stability, the data processing method is keyword analysis.
[0101] In keyword analysis, both the rat input data and the supplementary data are recorded as data to be analyzed. Keywords in the experimental text corresponding to each data to be analyzed are detected, and keywords with a valid frequency greater than the preset valid frequency are recorded as valid keywords.
[0102] The number of valid occurrences refers to the number of times the keyword appears.
[0103] The keyword detection method in the experimental text can be, but is not limited to, the BERT model, the TextRank algorithm, or the RAKE algorithm. This is easy for those skilled in the art to understand and will not be elaborated here. The number of valid counts is used to determine the importance of the keywords. The greater the user's demand for the accuracy of information acquisition using keywords, the larger the value of the preset number of valid counts. A preset number of valid counts is provided, which is 5.
[0104] Specifically, for a single valid keyword, the corresponding search method is determined based on the usage status of the valid keyword;
[0105] If the usage status of a valid keyword is that the number of times it is used is greater than or equal to the preset number of times it is used and the interval distance is less than the preset interval distance, then the search method for the valid keyword is a combination search.
[0106] If the usage status of a valid keyword is less than the preset number of combinations or the interval distance is greater than or equal to the preset interval distance, then the search method for that valid keyword is a standalone search.
[0107] The method for determining the number of collocations is to detect the number of times each effective keyword appears in the experimental text along with other effective keywords, and record the maximum number of times as the number of collocations. For example, for an effective keyword A, if it appears in the experimental text along with effective keyword B 5 times and along with effective keyword C 8 times, then the number of collocations for the effective keyword is 8. The number of collocations is used to characterize the degree of association between effective keywords and other keywords. The greater the user's demand for the precision of the combined search, the larger the value of the preset number of collocations. One preset value for the number of collocations is provided, which is 5 times.
[0108] Specifically, the preset interval distance is determined based on the keyword density;
[0109] The preset interval distance is negatively correlated with keyword density.
[0110] Keyword density is calculated as the number of words in a paragraph that includes all valid keywords divided by the total number of words in the text.
[0111] Specifically, during the search for effective keywords, the acquisition and adjustment methods are determined based on the page distribution of the target website;
[0112] If the page distribution status of the target website is such that the page similarity is greater than the preset page similarity or the information relevance is greater than the preset information relevance, the adjustment method is trajectory simulation adjustment;
[0113] If the page distribution status of the target website is such that the page similarity is less than or equal to the preset page similarity and the information relevance is less than or equal to the preset information relevance, the adjustment method is obtained through search simulation adjustment.
[0114] The method for confirming page similarity is as follows: obtain several relevant pages from the target website, crop images from each relevant page (the specific area is not limited, but should be smaller than the area of the relevant page), and detect the number of hyperlinks within each image. Page similarity = 1 / the absolute value of the difference between the maximum and minimum number of links. The method for confirming information relevance is as follows: identify the keywords in the hyperlinks within each image, using the same method as for the keywords in the experimental text, and detect the number of images corresponding to each keyword. The maximum number of images is recorded as the information relevance. The method for obtaining relevant web pages is as follows: obtain the linked pages corresponding to several hyperlinks within the initial page of the target website and record them as relevant pages. The number of relevant pages obtained is not specifically limited. It can be understood that the greater the user's demand for the accuracy of the page distribution determination, the greater the number of relevant pages.
[0115] Among them, page similarity and information relevance are used to characterize the similarity of page link distribution and information content in the target website, and different acquisition adjustment methods are determined accordingly. It can be understood that the greater the requirement for stability in the user information acquisition process, the smaller the preset page similarity and preset information relevance values will be.
[0116] Specifically, in trajectory simulation adjustment, the number of trajectory change points and the trajectory change frequency are increased based on the maximum effective difference;
[0117] The maximum effective difference is positively correlated with the increase in the number of trajectory change points and the trajectory change frequency.
[0118] In this invention, searches are performed separately for each valid keyword. If a combined search is performed, both the currently searched valid keyword and any valid keywords whose number of combinations with the currently searched valid keyword exceeds a preset number of combinations are recorded as combined keywords. These combined keywords are used simultaneously as search content. If a single search is performed, the currently searched valid keyword is used as the search content. During the search process, pages containing valid keywords are retained. During the crawler search, the hyperlink positions on the main page are recorded as movement points. The mouse movement trajectory is a straight line movement between these movement points. The trajectory change points are randomly generated non-hyperlink positions. The mouse will randomly execute a straight line movement trajectory between a movement point, a trajectory change point, and another movement point. The trajectory change frequency is the generation speed of the trajectory change points, measured in points per minute.
[0119] Specifically, in the search simulation adjustment, the dwell time variation frequency and IP rotation frequency are determined based on the search depth;
[0120] Search depth is positively correlated with the frequency of dwell time changes and IP rotation frequency.
[0121] The method for confirming the number of search pages is as follows: detect the main page currently used for keyword search, detect the number of indirect links from the target website's initial page to the main page, and record this number of indirect links as the number of search pages. The number of indirect links is the number of links clicked from the target website's initial page to the main page. The main page is the currently used page, the initial page is the target website's initial page, and the target website is the website from which information needs to be obtained. In this invention, the number of search pages is recorded as the search depth.
[0122] The search process uses a web crawler and simulates human control by operating the mouse. The mouse hovers at a preset frequency, that is, once every 2 minutes. The mouse hovers after 2 minutes of operation. The hover duration is initially set by the user and varies randomly. The hover duration variation frequency is the frequency at which the hover duration is modified. The IP rotation frequency is the frequency at which the web crawler changes its IP address.
[0123] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
[0124] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for acquiring and analyzing database information for pharmacological analysis of heatstroke, characterized in that, include: The data status is determined based on the data diversity and data instability of the rat input data, and the data processing method is determined based on the data status, namely, simulated data supplementation or keyword analysis. In the process of supplementing the simulated data, the method of supplementation is determined by comparing the effective standard index ratio with the preset effective standard index ratio. The method is to supplement the simulated data based on the difference in the effective standard index ratio or the distribution status of the unstable segments. The distribution of unstable segments is determined based on the number of unstable segments and the degree of distribution of unstable segments. In keyword analysis, the usage status is determined based on the frequency of effective keyword combinations and the interval distance, and the search method corresponding to the effective keywords is determined as either a combined search or a single search based on the usage status; During the search for effective keywords, the page distribution status is determined based on the page similarity and information relevance of the target website, and the acquisition adjustment method is determined based on the page distribution status, namely, trajectory simulation adjustment or search simulation adjustment. Under the preset completion conditions, pharmacological analysis was performed.
2. The database information acquisition and analysis method for pharmacological analysis of heatstroke according to claim 1, characterized in that, The rat input data is in the following states: data diversity is less than the preset data diversity or data instability is greater than or equal to the preset data instability. The data processing method is simulated data supplementation. In the supplementation of simulated data, the supplementation method is determined based on the proportion of effective standard indices; If the effective standard index percentage is less than the preset effective standard index percentage, the supplementation method is to supplement the data by simulating the difference in the effective standard index percentage. If the effective standard index is greater than or equal to the preset effective standard index percentage, the supplementation method is to supplement the data by simulating the distribution of unstable segments.
3. The database information acquisition and analysis method for pharmacological analysis of heatstroke according to claim 2, characterized in that, When supplementing simulated data based on the difference in the proportion of effective standardized indexes, the difference in the proportion of effective standardized indexes is detected and the adjustment range of rat training parameters is determined based on the difference in the proportion of effective standardized indexes. The difference between the adjustment range of the rat training parameters and the percentage of the effective standard index is positively correlated. Supplementary data refers to the stored datasets within the database whose corresponding rat training parameters fall within the adjustment range of the rat input data's training parameters.
4. The database information acquisition and analysis method for pharmacological analysis of heatstroke according to claim 3, characterized in that, When supplementing simulation data based on the distribution of unstable segments, the distribution of unstable segments is determined based on the number of unstable segments and the degree of distribution of unstable segments. If the distribution of unstable segments is such that the number of unstable segments is within the first preset range of unstable segment numbers, then the range of rat training parameter adjustment is determined based on the number of unstable segments. The difference between the rat training parameter adjustment range and the effective standard index ratio is positively correlated. The stored datasets in the search database whose corresponding rat training parameters are within the range of rat training parameter adjustment of the rat input data are recorded as supplementary data. If the unstable segment distribution state is such that the number of unstable segments is within the second preset range of the number of unstable segments and the unstable segment distribution degree is within the second preset range of the unstable segment distribution degree, then the preset experimental environment difference degree is determined based on the unstable segment distribution degree, and the stored dataset that satisfies the experimental environment difference degree being less than the preset experimental environment difference degree is recorded as supplementary data.
5. The database information acquisition and analysis method for pharmacological analysis of heatstroke according to claim 4, characterized in that, When the rat input data is in a state where the data diversity is greater than or equal to the preset data diversity and the data stability is less than the preset data stability, the data processing method is keyword analysis. In keyword analysis, both the rat input data and the supplementary data are recorded as data to be analyzed. Keywords in the experimental text corresponding to each data to be analyzed are detected, and keywords with a valid frequency greater than the preset valid frequency are recorded as valid keywords.
6. The database information acquisition and analysis method for pharmacological analysis of heatstroke according to claim 5, characterized in that, For a single valid keyword, determine the corresponding search method based on the usage status of the valid keyword; If the usage status of a valid keyword is that the number of times it is used is greater than or equal to the preset number of times it is used and the interval distance is less than the preset interval distance, then the search method for the valid keyword is a combination search. If the usage status of a valid keyword is less than the preset number of combinations or the interval distance is greater than or equal to the preset interval distance, then the search method for that valid keyword is a standalone search.
7. The database information acquisition and analysis method for pharmacological analysis of heatstroke according to claim 6, characterized in that, The preset interval distance is determined based on the keyword density; The preset interval distance is negatively correlated with keyword density.
8. The database information acquisition and analysis method for pharmacological analysis of heatstroke according to claim 7, characterized in that, During the search for effective keywords, the acquisition and adjustment methods are determined based on the page distribution of the target website; If the page distribution status of the target website is such that the page similarity is greater than the preset page similarity or the information relevance is greater than the preset information relevance, the adjustment method is trajectory simulation adjustment; If the page distribution status of the target website is such that the page similarity is less than or equal to the preset page similarity and the information relevance is less than or equal to the preset information relevance, the adjustment method is obtained through search simulation adjustment.
9. The database information acquisition and analysis method for pharmacological analysis of heatstroke according to claim 8, characterized in that, In trajectory simulation adjustment, the number of trajectory change points and the trajectory change frequency are increased based on the maximum effective difference. The maximum effective difference is positively correlated with the increase in the number of trajectory change points and the trajectory change frequency.
10. The database information acquisition and analysis method for pharmacological analysis of heatstroke according to claim 9, characterized in that, In the search simulation adjustment, the dwell time variation frequency and IP rotation frequency are determined based on the search depth. Search depth is positively correlated with the frequency of dwell time changes and IP rotation frequency.
Citation Information
Patent Citations
Construction method of extrajunctional lymphoma pathological database
CN116701353A
Content extraction method based on keyword matching
CN107229668A
Data recommendation method based on large model
CN119719353A