Database information acquisition and analysis method for pharmacological analysis of heat stroke
By selecting the data processing method based on the diversity and destability of the input data of rats, combining effective standard index and keyword analysis, the data acquisition process is optimized, and the problem of low information acquisition efficiency in pharmacological analysis of heat stroke pathology is solved, and more efficient and accurate data acquisition and analysis are achieved.
Patent Information
- Application Number
- CN202510419177.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-03
AI Technical Summary
In the prior art, the database information acquisition efficiency of heat stroke pharmacological analysis is low, and different data acquisition methods cannot be selected based on the actual data performance status, resulting in inaccurate information acquisition.
Determine the data state based on the data diversity and data destabilization of the input data of rats, select simulated data supplement or keyword analysis, supplement data through the proportion of effective standard index and the distribution status of instable paragraphs, determine the search method based on the number of effective keywords and the interval distance, and use page similarity and information relevance for acquisition and adjustment to improve data search capabilities.
It improves the accuracy and efficiency of data acquisition of pharmacological analysis of heat stroke, ensures that the data processing method meets the needs of actual work scenarios, and enhances the accuracy and search efficiency of subsequent pharmacological analysis.
Smart Images

Figure CN120336615A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data analysis, and in particular to a method for obtaining and analyzing database information for pharmacological analysis of heat stroke. Background Art
[0002] Exertional heat stroke is one of the subtypes of severe heat stroke. It is a serious and fatal disease caused by exposure to a high-temperature and high-humidity environment during high-intensity physical activities, resulting in an imbalance in the body's regulatory functions, more heat production than heat dissipation, a rapid increase in core temperature, accompanied by skin burning, disturbance of consciousness, and multiple organ dysfunction. In the prior art, technicians often construct rat models and conduct pharmacological analysis on related therapeutic drugs. However, due to time and resource limitations, the experimental evidence and insights into the mechanisms of the role of related therapeutic drugs in heat stroke prevention are limited. With the development of science and technology, data sharing has contributed to breaking the experimental limitations of independent individuals or institutions. How to quickly and effectively match the relevant experiments and pharmacological data needed from a large amount of data is an urgent problem for those skilled in the art.
[0003] Chinese Patent Publication No. CN116701353A discloses a method for constructing a pathological database of extranodal lymphoma, including: Step 1: Based on the classification criteria of lymphoid and hematopoietic system tumors, determine the query list of lymphoma subtypes to be included in the database; Step 2: According to the query list of lymphoma subtypes, obtain at least a preset number of target cases for each lymphoma subtype to generate a lymphoma subtype data packet, where the target cases contain various clinical information; Step 3: Store the lymphoma subtype data packet to establish a pathological database of extranodal lymphoma. The present invention establishes a digital slice library of extranodal lymphoma pathology. Medical students can search the digital pathology knowledge graph and retrieve digital pathology pictures related to the text content of pathology teaching. The pictures are attached with detailed clinical information and can automatically pop up the differences between normal images and images of lesions with detailed annotations. Students can mark the key points and save notes, which is convenient for students' pathology learning and improves the efficiency of pathology learning. It can be seen that the above technical solution has the following problems: Only determining the query list based on the classification criteria cannot select different data acquisition methods according to the performance status of the actual data, resulting in poor efficiency of obtaining database information. Summary of the Invention
[0004] Therefore, the present invention provides a method for obtaining and analyzing database information for pharmacological analysis of heat stroke to overcome the problem in the prior art that different data acquisition methods cannot be selected according to the performance status of the actual data, resulting in poor information acquisition efficiency.
[0005] To achieve the above object, the present invention provides a method for obtaining and analyzing database information for pharmacological analysis of heat stroke, including:
[0006] Determine the data status based on the data diversity and data deviation stability of the rat input data, and determine the data processing method as simulated data supplementation or keyword analysis according to the data status;
[0007] In simulated data supplementation, determine the supplementation method according to the comparison result of the proportion of the effective standard index and the preset proportion of the effective standard index, that is, supplement the simulated data according to the difference in the proportion of the effective standard index or the distribution state of the unstable paragraphs;
[0008] The distribution state of the unstable paragraphs is determined according to the number of unstable paragraphs and the distribution degree of the unstable paragraphs;
[0009] In keyword analysis, determine the usage status according to the collocation times and interval distances of the effective keywords, and determine the corresponding search method as combined search or individual search according to the usage status;
[0010] During the search process of the effective keywords, determine the page distribution state according to the page similarity and information relevance of the target website, and determine the acquisition adjustment method as trajectory simulation adjustment or search simulation adjustment according to the page distribution state;
[0011] Under the preset completion conditions, perform pharmacological analysis.
[0012] Furthermore, when the data status of the rat input data is that the data diversity is less than the preset data diversity or the data deviation stability is greater than or equal to the preset data deviation stability, the data processing method is simulated data supplementation;
[0013] In simulated data supplementation, determine the supplementation method according to the proportion of the effective standard index;
[0014] If the proportion of the effective standard index is less than the preset proportion of the effective standard index, the supplementation method is to supplement the simulated data according to the difference in the proportion of the effective standard index;
[0015] If the effective standard index is greater than or equal to the preset proportion of the effective standard index, the supplementation method is to supplement the simulated data according to the distribution state of the unstable paragraphs.
[0016] Furthermore, when supplementing the simulated data according to the difference in the proportion of the effective standard index, detect the difference in the proportion of the effective standard index and determine the adjustment range of the rat training parameters according to the difference in the proportion of the effective standard index;
[0017] The adjustment range of the rat training parameters has a positive correlation with the difference in the proportion of the effective standard index;
[0018] The storage data set in the search database whose corresponding rat training parameters are within the adjustment range of the rat training parameters of the rat input data is recorded as supplementary data.
[0019] Further, when supplementing simulation data according to the distribution state of the instability segments, the distribution state of the instability segments is determined according to the number of instability segments and the distribution degree of the instability segments;
[0020] If the distribution state of the instability segments is that the number of instability segments is within the first preset range of the number of instability segments, the adjustment range of the rat training parameters is determined according to the number of instability segments. The adjustment range of the rat training parameters has a positive correlation with the difference in the proportion of the effective standard index. The stored data set in the database whose corresponding rat training parameters are within the adjustment range of the rat training parameters of the rat input data is recorded as supplementary data;
[0021] If the distribution state of the instability segments is that the number of instability segments is within the second preset range of the number of instability segments and the distribution degree of the instability segments is within the second preset range of the distribution degree of the instability segments, the preset experimental environment difference degree is determined according to the distribution degree of the instability segments, and the stored data set that satisfies the experimental environment difference degree less than the preset experimental environment difference degree is recorded as supplementary data.
[0022] Further, when the data state of the rat input data is that the data diversity is greater than or equal to the preset data diversity and the data instability degree is less than the preset data instability degree, the data processing method is keyword analysis;
[0023] In keyword analysis, the rat input data and the supplementary data are respectively recorded as a data to be analyzed, the keywords in the corresponding experimental text of each data to be analyzed are detected, and the keywords with the effective times greater than the preset effective times are recorded as effective keywords.
[0024] Further, for a single effective keyword, the corresponding search method is determined according to the usage state of the effective keyword;
[0025] If the usage state of the effective keyword is that the collocation times are greater than or equal to the preset collocation times and the interval distance is less than the preset interval distance, the search method of the effective keyword is combined search;
[0026] If the usage state of the effective keyword is that the collocation times are less than the preset collocation times or the interval distance is greater than or equal to the preset interval distance, the search method of the effective keyword is individual search.
[0027] Further, the preset interval distance is determined according to the keyword density;
[0028] The preset interval distance has a negative correlation with the keyword density.
[0029] Further, during the search process of the effective keyword, the acquisition adjustment method is determined according to the page distribution state of the target website;
[0030] If the page distribution status of the target website is that the page similarity is greater than the preset page similarity or the information relevance is greater than the preset information relevance, the obtained adjustment method is trajectory simulation adjustment;
[0031] If the page distribution status of the target website is that the page similarity is less than or equal to the preset page similarity and the information relevance is less than or equal to the preset information relevance, the obtained adjustment method is search simulation adjustment.
[0032] Further, in the trajectory simulation adjustment, the number of trajectory transformation points and the trajectory transformation frequency are increased according to the maximum effective difference;
[0033] The maximum effective difference has a positive correlation with both the number of trajectory transformation points and the increased amount of the trajectory transformation frequency.
[0034] Further, in the search simulation adjustment, the residence time transformation frequency and the IP rotation frequency are determined according to the search depth;
[0035] The search depth has a positive correlation with both the residence time transformation frequency and the IP rotation frequency.
[0036] Compared with the prior art, in the technical solution of the present invention, the data state is determined according to the data diversity and data deviation stability of the input data of the rats, and the data processing method is determined as simulated data supplementation or keyword analysis according to the data state. Whether the current data can meet the requirements of pharmacological analysis is effectively reflected by the data state, and different data processing methods are correspondingly selected, so that the selection of the data processing method can more conform to the actual working scenario, improve the data accuracy, and provide data support for subsequent pharmacological analysis.
[0037] Further, in the technical solution of the present invention, in the simulated data supplementation, the supplementation method is determined according to the comparison result of the effective standard index ratio and the preset effective standard index ratio to perform simulated data supplementation according to the effective standard index ratio difference or the unstable paragraph distribution state. Whether the data stability of the reference data required for pharmacological analysis meets the requirements is reflected by the effective standard index ratio difference, and the difference degree between data is characterized by the unstable paragraph distribution state, and different supplementation methods are correspondingly selected, so that the supplemented data more conforms to the current data requirements, and further improves the accuracy of subsequent pharmacological analysis.
[0038] Further, in the technical solution of the present invention, in the keyword analysis, the usage status is determined according to the collocation times and interval distances of effective keywords, and the corresponding search method is determined as combined search or individual search according to the usage status. The association relationship between current keywords is reflected by the usage status, and thus an effective search method is used in the subsequent search process to improve the search efficiency.
[0039] Furthermore, in the technical solution of the present invention, during the search process of effective keywords, the page distribution status is determined according to the page similarity and information relevance of the target website, and the acquisition adjustment method is determined as trajectory simulation adjustment or search simulation adjustment according to the page distribution status, so as to avoid the restrictions of the website on the search program during the search process and improve the data search ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a schematic diagram of the method for obtaining and analyzing database information for pharmacological analysis of heat stroke in the present invention;
[0041] Figure 2 It is a flowchart of the method for determining the data processing method according to the data status in the present invention;
[0042] Figure 3 It is a flowchart of the method for determining the corresponding search method according to the usage status of effective keywords in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] In order to make the objectives and advantages of the present invention clearer, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0044] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.
[0045] It should be noted that in the description of the present invention, the terms indicating directions or positional relationships such as "upper", "lower", "left", "right", "inner", "outer", etc. are based on the directions or positional relationships shown in the drawings. This is only for convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.
[0046] In addition, it should be noted that in the description of the present invention, unless otherwise clearly defined and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0047] Please refer to Figures 1 to 3 As shown, the present invention provides a method for obtaining and analyzing database information for pharmacological analysis of heat stroke, including:
[0048] Determine the data status based on the data diversity and data deviation stability of the rat input data, and determine the data processing method as simulated data supplementation or keyword analysis according to the data status;
[0049] In simulated data supplementation, determine the supplementation method according to the comparison result of the proportion of the effective standard index and the preset proportion of the effective standard index, and supplement the simulated data according to the difference in the proportion of the effective standard index or the distribution state of the unstable paragraphs;
[0050] The distribution state of the unstable paragraphs is determined according to the number of unstable paragraphs and the distribution degree of the unstable paragraphs;
[0051] In keyword analysis, determine the usage status according to the collocation times and interval distances of the effective keywords, and determine the corresponding search method as combined search or individual search according to the usage status;
[0052] During the search process of the effective keywords, determine the page distribution state according to the page similarity and information relevance of the target website, and determine the acquisition adjustment method as trajectory simulation adjustment or search simulation adjustment according to the page distribution state;
[0053] Under the preset completion conditions, perform pharmacological analysis.
[0054] The present invention is applied to the network pharmacological analysis of heat stroke-related drugs. By constructing and analyzing biological networks, the interaction between drugs and the body is revealed. The present invention is applied to the front-end data collection and storage stage. The rat input data in the present invention is a number of data sets uploaded by users themselves. Each data set includes rat training parameters, rat physical parameters, rat experimental data, and experimental environment data. The present invention applies an information database, and there are a number of stored data sets stored in the database. The stored data sets also include rat training parameters, rat physical parameters, rat experimental data, and experimental environment data. The rat experimental data includes the functional indicators measured every day during the rat training stage, and the body temperature and functional indicators measured before and after each experiment. The functional indicators include ALT, AST, BUN, CK, and CREA. Among them, the units of ALT, AST, and CK are U / L, the unit of BUN is mmol / L, and the unit of CREA is μmol / L.
[0055] In the present invention, the rats need to undergo relevant training before the experiment, including: training on a running wheel treadmill for a preset time every day. Starting from the second day, the preset time every day is increased by a preset time difference compared to the previous day until the number of training days reaches the preset number of training days. Among them, during each training, the initial speed of the running wheel treadmill is the preset initial speed, and the preset speed difference is increased every 2 minutes;
[0056] The rat physical parameters include the age, weight, and blood pressure of the rats;
[0057] The experimental text is the text description related to the input data of rats, including but not limited to the description of the training process, the description of the training results of rats, and the description of the symptoms of rats;
[0058] The rat training parameters include a preset time, a preset time difference, a preset number of training days, a preset initial speed, and a preset speed difference. An embodiment is provided. In this embodiment, the preset time is 20 minutes, the preset time difference is 10 minutes, the preset number of days is 5 days, the preset initial speed is 5 m / min, and the preset speed difference is 1 m / min;
[0059] The rat experiment in the present invention includes: randomly dividing rats into 8 groups according to their body weights, and the difference in body weights between rats within each group should be less than 400 g. In the subsequent experiment, body temperature, blood, and tissue samples are collected from each group of rats. A temperature monitoring capsule is implanted into the experimental rats, and the core body temperature is measured every 5 minutes. Among them, each rat is anesthetized with isoflurane, and the sterile temperature monitoring capsule is implanted into the abdominal cavity through a small sterile incision. After the operation, the rats are allowed to rest for 2 days. On the 8th day, a treadmill is placed in an artificial climate chamber with a temperature of 40 ± 1 °C and a relative humidity of 60 ± 5%. The rats exercise under the above high temperature and high humidity conditions. The core body temperature and physical condition are continuously monitored in real time. When the core body temperature of the rats exceeds 42 °C and shows signs of unconsciousness, they are removed from the room and placed in a room temperature environment. After 3 hours, the serum, plasma, organ tissues, and fecal matter of the rats are collected and data detection is performed to obtain experimental data.
[0060] The preset completion condition is that the simulation data supplementation is completed or the keyword search process in the keyword analysis ends. The pharmacological analysis includes:
[0061] The user collects the effective components and targets of relevant drugs in the rat input data, supplementary data, and keyword search data;
[0062] Analysis of the composition of the drug;
[0063] Prediction of potential action targets, using an online compound target database to predict the potential protein targets of active ingredients;
[0064] Collect disease targets, and collect potential targets of heat stroke by combining disease-related gene or protein databases;
[0065] Construct a "drug-target-disease" network, screen common targets, construct an interaction network and construct a PPI network;
[0066] Enrichment analysis, perform GO function and KEGG pathway enrichment analysis on potential targets, and identify biological processes, cell components, molecular functions, and signal pathways significantly related to heat stroke;
[0067] Molecular docking verification is carried out to verify the binding activity of the predicted key compounds with the targets by using molecular docking software, and the binding affinity between the compounds and the targets is evaluated according to the docking scores;
[0068] Data integration and interpretation;
[0069] It should be noted that if no simulated data supplementation or keyword analysis is carried out, then there is no need to collect the effective components and targets of the relevant drugs in the supplementary data or keyword search data. The above is easy to understand for those skilled in the art and will not be elaborated here.
[0070] In the present invention, for the value of the preset threshold, the user can set it according to historical experience and actual application scenarios. A method for obtaining the value is provided. The preset threshold is determined according to the detection values in the past historical records. For example, for the value of the preset effective standard index ratio, the detection values corresponding to the historical records that meet the user's requirements are extracted, and the detection value is the effective standard index ratio. Data cleaning is performed on the effective standard index ratio to remove the outliers therein. The method for identifying outliers can be the Z-Score method or the IQR method. The average value of the effective standard index ratio after removing the outliers is recorded as the value of the preset effective standard index ratio. Subsequently, the preset threshold can also be optimized through a deep learning model. The above is easy to understand for those skilled in the art and will not be elaborated here.
[0071] The present invention applies a number of historical records. Each individual historical record records at least one past information acquisition and analysis process and the corresponding detection data. The detection data includes data diversity, data deviation stability, data stability, the number of unstable paragraphs, the distribution degree of unstable paragraphs, the number of effective times, and the number of matching times. And each historical record is correspondingly set with a qualified mark, and the qualified mark records whether the information acquisition and analysis process corresponding to the historical record meets the user's requirements. Among them, whether it meets the user's requirements is determined by the user according to the actual information analysis results.
[0072] Specifically, when the data state of the input data of the rat is that the data diversity is less than the preset data diversity or the data deviation stability is greater than or equal to the preset data deviation stability, the data processing method is simulated data supplementation;
[0073] In the simulated data supplementation, the supplementation method is determined according to the effective standard index ratio;
[0074] If the effective standard index ratio is less than the preset effective standard index ratio, the supplementation method is to supplement the simulated data according to the difference in the effective standard index ratio;
[0075] If the effective standard index is greater than or equal to the preset effective standard index ratio, the supplementation method is to supplement the simulated data according to the distribution state of the unstable paragraphs.
[0076] Among them, data diversity = (H / H0)×α1+(C / C0)×α2, where H is the absolute value of the difference between the maximum value and the minimum value of the rat physical parameter in the rat input data, C is the absolute value of the difference between the maximum value and the minimum value of the rat training parameter in the rat input data, α1 is the first diversity coefficient, α2 is the second diversity coefficient. For the values of α1 and α2, the user can determine the influence degree of the rat training parameter and the rat physical parameter on the accuracy of subsequent pharmacological analysis according to historical experience and the deep learning network, and correspondingly determine the values of α1 and α2. For example, the greater the influence degree of the rat training parameter on the accuracy of subsequent pharmacological analysis, the greater the value of α1. Provide a set of values for α1 and α2, α1 = 0.5, α2 = 0.5.
[0077] The confirmation method of data deviation stability is as follows: extract each data set of the rat input data. For a single data set, extract the ALT, AST, BUN, CK, and CREA of its rat experimental data, calculate the reference area of each functional index respectively, then calculate the maximum difference of the reference areas between the same functional indexes of each data set, and record it as the data deviation stability. For a single functional index, the calculation method of its reference area is to present the measured values of each experiment in the form of data points on a two-dimensional coordinate system. The horizontal axis of the coordinate system of the data points is the number of experiments when the value is measured, such as the start of the first experiment, the completion of the first experiment, the start of the second experiment, etc., and the vertical axis is the value unit. Connect the adjacent points with straight lines to obtain the reference data line, and connect the data points with the smallest and largest abscissas perpendicular to the horizontal axis. The enclosed figure formed is recorded as the reference area.
[0078] In the embodiment of the present invention, the preset data diversity is 1, the value of the preset data deviation stability is 20% of the maximum value of the reference area corresponding to the data deviation stability, the effective standard index ratio is the number of effective standard indexes / 5, and the effective standard index is the functional index with a data stability greater than the preset data stability. For a single functional index, the calculation formula of its corresponding data stability M is:
[0079]
[0080] Among them, M i is the average value of all the measured values of this functional index in the i-th data set in the experiment, n is the total number of data sets corresponding to the rat input data, i = 1, 2, 3,..., n. In the embodiment of the present invention, the value of the preset data stability is 5% of M0.
[0081] Specifically, when supplementing simulation data according to the difference ratio of the effective standard index, detect the difference ratio of the effective standard index and determine the adjustment range of the rat training parameters according to the difference ratio of the effective standard index;
[0082] The adjustment range of the rat training parameters is positively correlated with the difference ratio of the effective standard index;
[0083] The stored data set in the database whose corresponding rat training parameters are within the adjustment range of the rat training parameters of the rat input data is recorded as supplementary data.
[0084] The difference ratio of the effective standard index is the absolute value of the difference between the ratio of the effective standard index and the preset ratio of the effective standard index.
[0085] In the present invention, the adjustment range of the rat training parameters is a numerical range with the average value of the rat training parameters of the current rat input data as the median point, and the absolute values of the differences between the maximum and minimum values of the numerical range and the median point are the same.
[0086] Specifically, when supplementing simulation data according to the instability paragraph distribution state, determine the instability paragraph distribution state according to the number of instability paragraphs and the instability paragraph distribution degree;
[0087] If the instability paragraph distribution state is that the number of instability paragraphs is within the first preset range of the number of instability paragraphs or the instability paragraph distribution degree is within the first preset range of the instability paragraph distribution degree, then determine the adjustment range of the rat training parameters according to the number of instability paragraphs. The adjustment range of the rat training parameters is positively correlated with the difference ratio of the effective standard index. The stored data set in the database whose corresponding rat training parameters are within the adjustment range of the rat training parameters of the rat input data is recorded as supplementary data;
[0088] If the instability paragraph distribution state is that the number of instability paragraphs is within the second preset range of the number of instability paragraphs and the instability paragraph distribution degree is within the second preset range of the instability paragraph distribution degree, then determine the preset experimental environment difference degree according to the instability paragraph distribution degree, and record the stored data set that satisfies the experimental environment difference degree less than the preset experimental environment difference degree as supplementary data;
[0089] The preset experimental environment difference degree is positively correlated with the instability paragraph distribution degree.
[0090] The experimental environment data includes the experimental environment temperature and the experimental environment humidity. The temperature and humidity are respectively the average values of several detections. The experimental environment difference degree includes the absolute value of the difference in the experimental environment temperature and the absolute value of the difference in the experimental environment humidity. Their corresponding preset experimental environment difference degrees are different. If the absolute value of the difference in the experimental environment temperature and the absolute value of the difference in the experimental environment humidity are both less than the corresponding preset experimental environment difference degree, the corresponding storage data set is recorded as supplementary data. In the specific implementation of the present invention, the preset experimental environment difference degree corresponding to the absolute value of the difference in the experimental environment temperature is 5°C, and the preset experimental environment difference degree corresponding to the absolute value of the difference in the experimental environment humidity is 4%.
[0091] The values within the first preset number of unstable segments range are all less than the preset number of unstable segments, the values within the second preset number of unstable segments range are all greater than or equal to the preset number of unstable segments, the values within the first preset unstable segment distribution degree range are all less than the preset unstable segment distribution degree, and the values within the second preset unstable segment distribution degree range are all greater than or equal to the preset unstable segment distribution degree. The value-taking methods of the preset number of unstable segments and the preset unstable segment distribution degree are to extract the number of unstable segments and the unstable segment distribution degree corresponding to the historical records that meet the user requirements, perform data cleaning on the number of unstable segments and the unstable segment distribution degree respectively to remove the outliers therein, and record the respective average values of the number of unstable segments and the unstable segment distribution degree after removing the outliers as the preset number of unstable segments and the preset unstable segment distribution degree respectively.
[0092] The method for confirming the number of unstable segments is to extract the reference data lines of each rat's experimental data in the rat input data, divide each reference data line into several segments based on the horizontal axis, and calculate the similarity between the segments at the same position, including:
[0093] Randomly extract the same number of several points on each segment;
[0094] For each point, calculate its relative position with all other points, and quantify these relative positions into a set of discrete bins;
[0095] Regarding the similarity Z between any two segments S1 and S2, D(C1, C2) is the Earth Mover's Distance between the two reference data lines, and Dmax is the maximum bin value of C1 and C2.
[0096] If in the segment corresponding to a position, there are a preset number of segments corresponding to other segments with a similarity less than the preset similarity, then this position is recorded as an unstable segment position, and the number of unstable segment positions is recorded as the number of unstable segments.
[0097] If two paragraphs have the same minimum abscissa and the same maximum abscissa, then the two paragraphs are in the same position.
[0098] Regarding the number of points extracted from a paragraph and the value of a preset number, the user can set them according to the actual scenario. It can be understood that the greater the user's demand for the determination accuracy of the number of buckling paragraphs, the greater the number of points extracted from the paragraph and the value of the preset number. A value is provided where the number of points extracted from each paragraph is 10, and the preset number is 70% of the total number of paragraphs at the corresponding position.
[0099] The method for confirming the distribution degree of buckling paragraphs is to calculate the absolute value of the difference between the maximum abscissa and the minimum abscissa of the buckling paragraph position, and record this absolute value as the buckling paragraph distribution degree.
[0100] Specifically, when the data state of the input data of the rat is that the data diversity is greater than or equal to the preset data diversity and the data stability deviation is less than the preset data stability deviation, the data processing method is keyword analysis;
[0101] In keyword analysis, both the input data of the rat and the supplementary data are respectively recorded as a data to be analyzed, the keywords in the experimental text corresponding to each data to be analyzed are detected, and the keywords with the effective times greater than the preset effective times are recorded as effective keywords.
[0102] The effective times is the number of times the keyword appears.
[0103] Among them, the detection method of keywords in the experimental text can be, but is not limited to, the BERT model, the TextRank algorithm or the RAKE algorithm. This is easy to understand for those skilled in the art and will not be elaborated here; the effective times is used to determine the importance degree of the keyword. The greater the user's demand for the information acquisition accuracy using keywords, the greater the value of the preset effective times. A preset effective times is provided, and the preset effective times is 5.
[0104] Specifically, for a single effective keyword, determine the corresponding search method according to the usage status of the effective keyword;
[0105] If the usage status of the effective keyword is that the collocation times is greater than or equal to the preset collocation times and the interval distance is less than the preset interval distance, then the search method of this effective keyword is combined search;
[0106] If the usage status of the effective keyword is that the collocation times is less than the preset collocation times or the interval distance is greater than or equal to the preset interval distance, then the search method of this effective keyword is individual search.
[0107] The way to confirm the collocation times is to detect the times when valid keywords appear in the experimental text simultaneously with other valid keywords respectively, and record the maximum times as the collocation times. For example, for a valid keyword A, the number of times it appears in the experimental text simultaneously with valid keyword B is 5 times, and the number of times it appears in the experimental text simultaneously with valid keyword C is 8 times. Then the collocation times of the valid keyword is 8. The collocation times are used to characterize the degree of association between valid keywords and other keywords. The greater the user's demand for the accuracy of combined search, the greater the value of the preset collocation times. A value of the preset collocation times is provided, and the preset collocation times is 5 times.
[0108] Specifically, the preset interval distance is determined according to the keyword density;
[0109] The preset interval distance has a negative correlation with the keyword density.
[0110] Among them, keyword density = the number of words in the paragraph that can include all valid keywords / the total number of words in the text.
[0111] Specifically, during the search process of valid keywords, the acquisition adjustment method is determined according to the page distribution status of the target website;
[0112] If the page distribution status of the target website is that the page similarity is greater than the preset page similarity or the information relevance is greater than the preset information relevance, the acquisition adjustment method is trajectory simulation adjustment;
[0113] If the page distribution status of the target website is that the page similarity is less than or equal to the preset page similarity and the information relevance is less than or equal to the preset information relevance, the acquisition adjustment method is search simulation adjustment.
[0114] The way to confirm the page similarity is to obtain several relevant pages of the target website, perform image capture on each relevant page, and the captured image areas are the same. The specific area is not limited, but it should be less than the area of the relevant page. Detect the number of hyperlinks in each image. Page similarity = 1 / the absolute value of the difference between the maximum number of links and the minimum number of links. The way to confirm the information relevance is to identify the keywords of the hyperlinks in each image. The identification method is the same as the keyword method of the experimental text. Detect the number of images corresponding to each keyword, and record the maximum number of images as the information relevance. The way to obtain the relevant web pages is to obtain the link pages corresponding to several hyperlinks in the initial page of the target website and record them as relevant pages. The number of relevant pages obtained is not specifically limited. It can be understood that the greater the user's demand for the determination accuracy of the page distribution status, the greater the number of relevant pages.
[0115] Among them, the similarity degree of page links and the relevance degree of information in the target website are characterized, and different acquisition adjustment methods are determined accordingly. It can be understood that the greater the stability requirement for the user information acquisition process, the smaller the values of the preset page similarity and the preset information relevance degree.
[0116] Specifically, in the trajectory simulation adjustment, the number of trajectory transformation points and the trajectory transformation frequency are increased according to the maximum effective difference.
[0117] The maximum effective difference has a positive correlation with both the increase amount of the number of trajectory transformation points and the trajectory transformation frequency.
[0118] In the present invention, searches are performed separately for each valid keyword. Among them, if a combined search is to be performed, the valid keyword currently to be searched and the valid keywords whose corresponding matching times are greater than the preset matching times with the valid keyword currently to be searched are all recorded as combined keywords, and the above-mentioned combined keywords are used as the search content simultaneously. If a separate search is to be performed, the valid keyword currently to be searched is used as the search content. During the search process, the pages with valid keywords are retained. During the crawler search process, the hyperlink position in the main page is recorded as a moving point, the mouse movement trajectory is a straight-line movement between each moving point, the trajectory transformation point is a randomly generated point position that is not a hyperlink position, and the mouse will randomly execute a mouse movement trajectory of straight-line movement between a moving point - a trajectory transformation point - a moving point. The trajectory transformation frequency is the generation speed of the trajectory transformation point, with the unit of number / min.
[0119] Specifically, in the search simulation adjustment, the residence time transformation frequency and the IP rotation frequency are determined according to the search depth.
[0120] The search depth has a positive correlation with both the residence time transformation frequency and the IP rotation frequency.
[0121] The confirmation method for the number of search pages is to detect the main page where the keyword search is currently being performed, detect the number of indirect links from the initial page of the target website to the main page, and record this number of indirect links as the number of search pages. The number of indirect links is the number of links clicked from the initial page of the target website to the main page. The main page is the page currently being used, the initial page is the initial page of the target website, and the target website is the website for which information needs to be acquired. In the implementation of the present invention, the number of search pages is recorded as the search depth.
[0122] During the search process, a web crawler is used for searching, and the mouse is operated to simulate manual control. The mouse pauses at a preset frequency, that is, the preset frequency is to pause once every 2 minutes. Then, after the mouse runs for 2 minutes, it pauses. The pause duration is initially set by the user and changes randomly. The pause time change frequency is the frequency at which the pause time is modified. The IP rotation frequency is the replacement frequency of the IP address of the web crawler.
[0123] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or replacements to the relevant technical features, and the technical solutions after these changes or replacements will all fall within the protection scope of the present invention.
[0124] The above are only the preferred embodiments of the present invention and are not used to limit the present invention; for those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for obtaining and analyzing database information for pharmacological analysis of heat stroke, characterized in that Including: Determine the data status according to the data diversity and data instability degree of the rat input data, and determine the data processing method as simulated data supplementation or keyword analysis according to the data status; In the simulated data supplementation, determine the supplementation method according to the comparison result of the proportion of the effective standard index and the preset proportion of the effective standard index, and supplement the simulated data according to the difference in the proportion of the effective standard index or the distribution state of the unstable paragraphs; The distribution state of the unstable paragraphs is determined according to the number of unstable paragraphs and the distribution degree of the unstable paragraphs; In the keyword analysis, determine the usage status according to the collocation times and interval distances of the effective keywords, and determine the search method corresponding to the effective keywords as combined search or individual search according to the usage status; During the search process of the effective keywords, determine the page distribution state according to the page similarity and information relevance of the target website, and determine the acquisition adjustment method as trajectory simulation adjustment or search simulation adjustment according to the page distribution state; Under the preset completion conditions, perform pharmacological analysis.
2. The database information acquisition and analysis method for heat stroke pharmacological analysis according to claim 1, characterized in that, When the data status of the rat input data is that the data diversity is less than the preset data diversity or the data instability degree is greater than or equal to the preset data instability degree, the data processing method is simulated data supplementation; In the simulated data supplementation, determine the supplementation method according to the proportion of the effective standard index; If the proportion of the effective standard index is less than the preset proportion of the effective standard index, the supplementation method is to supplement the simulated data according to the difference in the proportion of the effective standard index; If the effective standard index is greater than or equal to the preset proportion of the effective standard index, the supplementation method is to supplement the simulated data according to the distribution state of the unstable paragraphs.
3. The database information acquisition and analysis method for heat stroke pharmacological analysis according to claim 2, characterized in that When supplementing the simulated data according to the difference in the proportion of the effective standard index, detect the difference in the proportion of the effective standard index and determine the adjustment range of the rat training parameters according to the difference in the proportion of the effective standard index; The adjustment range of the rat training parameters has a positive correlation with the difference in the proportion of the effective standard index; The storage data set in which the corresponding rat training parameters in the search database are within the adjustment range of the rat training parameters of the rat input data is recorded as supplementary data.
4. The method for obtaining and analyzing database information for pharmacologic analysis of heat stroke according to claim 3, wherein, When supplementing the simulated data according to the distribution state of the unstable paragraphs, determine the distribution state of the unstable paragraphs according to the number of unstable paragraphs and the distribution degree of the unstable paragraphs; If the distribution state of the unstable paragraphs is that the number of unstable paragraphs is within the first preset range of the number of unstable paragraphs, determine the adjustment range of the rat training parameters according to the number of unstable paragraphs. The adjustment range of the rat training parameters has a positive correlation with the difference in the proportion of the effective standard index. The storage data set in which the corresponding rat training parameters in the search database are within the adjustment range of the rat training parameters of the rat input data is recorded as supplementary data; If the distribution state of the unstable paragraphs is that the number of unstable paragraphs is within the second preset range of the number of unstable paragraphs and the distribution degree of the unstable paragraphs is within the second preset range of the distribution degree of the unstable paragraphs, determine the preset experimental environment difference degree according to the distribution degree of the unstable paragraphs, and record the storage data set that satisfies the experimental environment difference degree less than the preset experimental environment difference degree as supplementary data.
5. The database information acquisition and analysis method for heat stroke pharmacological analysis according to claim 4, characterized in that When the data status of the rat input data is that the data diversity is greater than or equal to the preset data diversity and the data instability degree is less than the preset data instability degree, the data processing method is keyword analysis; In keyword analysis, the input data of rats and the supplementary data are each recorded as a data to be analyzed. The keywords in the experimental text corresponding to each data to be analyzed are detected, and the keywords with the effective times greater than the preset effective times are recorded as effective keywords.
6. The database information acquisition and analysis method for heat stroke pharmacological analysis according to claim 5, characterized in that, For a single effective keyword, the corresponding search method is determined according to the usage status of the effective keyword; If the usage status of the effective keyword is that the collocation times are greater than or equal to the preset collocation times and the interval distance is less than the preset interval distance, the search method of this effective keyword is combined search; If the usage status of the effective keyword is that the collocation times are less than the preset collocation times or the interval distance is greater than or equal to the preset interval distance, the search method of this effective keyword is individual search.
7. The database information acquisition and analysis method for heat stroke pharmacological analysis according to claim 6, characterized in that, The preset interval distance is determined according to the keyword density; The preset interval distance and the keyword density have a negative correlation.
8. The method for obtaining and analyzing database information for pharmacological analysis of heat stroke according to claim 7, characterized in that, During the search process of effective keywords, the acquisition adjustment method is determined according to the page distribution status of the target website; If the page distribution status of the target website is that the page similarity is greater than the preset page similarity or the information relevance is greater than the preset information relevance, the acquisition adjustment method is trajectory simulation adjustment; If the page distribution status of the target website is that the page similarity is less than or equal to the preset page similarity and the information relevance is less than or equal to the preset information relevance, the acquisition adjustment method is search simulation adjustment.
9. The database information acquisition and analysis method for heat stroke pharmacological analysis according to claim 8, characterized in that, In trajectory simulation adjustment, the number of trajectory transformation points and the trajectory transformation frequency are increased according to the maximum effective difference; The maximum effective difference and the increase amounts of the number of trajectory transformation points and the trajectory transformation frequency are all positively correlated.
10. The database information acquisition and analysis method for heat stroke pharmacological analysis according to claim 9, characterized in that In search simulation adjustment, the residence time transformation frequency and the IP rotation frequency are determined according to the search depth; The search depth and the residence time transformation frequency and the IP rotation frequency are all positively correlated.
Citation Information
Patent Citations
Construction method of extrajunctional lymphoma pathological database
CN116701353A
Content extraction method based on keyword matching
CN107229668A
Data recommendation method based on large model
CN119719353A
Cited By
Heat stroke medical record text classification method based on reinforcement learning
CN121278096A