Information collecting and processing method for heat stroke pathological data integration based on deep learning
Through deep learning technology, the keyword generation method and search optimization method are selected according to the data status, the problem of inefficient search of pathological data of thermal radiation pathology is solved, and efficient and accurate information collection and processing is achieved.
Patent Information
- Application Number
- CN202510353506.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-03-25
AI Technical Summary
In the prior art, the search efficiency of thermal radiation pathological data is low, and it is impossible to effectively screen and optimize the search method, resulting in poor search efficiency.
Through deep learning technology, the data state is determined based on the data abundance and image conversion completion of the user's analysis data, the associated phrases or image conversion supplements are selected to generate keywords, and the segmentation method is determined based on image disorder and distribution trend, the density of repeated keywords and the use reference deviation value is detected, and the combination or sequential search optimization method is selected.
It improves the efficiency and accuracy of integrated search of thermal radiation pathology data, avoids the problem of inefficient search caused by single keyword selection method, and realizes intelligent information collection and processing.
Smart Images

Figure CN120413073A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data analysis, and particularly to an information collection and processing method for heat stroke pathological data integration based on deep learning. Background Art
[0002] Heat stroke is a life-threatening disease caused by high-temperature environments and has significant health risks. With the development of information technology and the medical field, the sharing degree of heat stroke pathological information has been significantly improved. However, due to the dispersion of heat stroke pathological information and the lack of detailed classification, in the face of a large amount of pathological data, traditional manual search and processing methods are inefficient and difficult to meet actual needs. In recent years, deep learning technology has shown great potential in the field of medical data processing, but still faces many challenges, such as high data annotation costs and difficulties in multi-dimensional search. Therefore, how to collect information from a large amount of data through deep learning technology to improve the search efficiency of heat stroke pathological data integration is a technical problem that needs to be solved urgently by those skilled in the art.
[0003] Chinese Patent Publication No. CN118643128A discloses a pathological literature search and dialogue system based on large language models and RAG technology. The system includes a pathological literature collection and processing module for collecting, organizing, and processing pathological literature data; a real-time interaction and dialogue module for receiving user input and processing it through a large language model to extract user intentions and semantic embeddings; a multi-modal vector representation module for converting pathological literature into semantic vector representations; a multi-way recall module for retrieving the most relevant pathological literature to the user's query; and a content ranking module including a re-ranking model based on importance, which re-ranks the pathological literature recalled from multiple parties using a unified metric dimension to generate an optimal ranking result. It can be seen that the above technical solution has the following problems: during the search process, the search method cannot be effectively screened and optimized, resulting in poor search efficiency. Summary of the Invention
[0004] To this end, the present invention provides an information collection and processing method for heat stroke pathological data integration based on deep learning to overcome the problem in the prior art that the search method cannot be effectively screened and optimized, resulting in poor search efficiency.
[0005] To achieve the above object, the present invention provides an information collection and processing method for heat stroke pathological data integration based on deep learning, including:
[0006] Obtain user analysis data, and determine the data state of the user analysis data according to the data richness and image conversion completion degree of the user analysis data;
[0007] Determine the keyword selection method according to the data state, and the keyword selection method is generated by associated phrase selection or supplemented by image conversion.
[0008] During the generation of associated phrase selection, the number of selected keywords is determined according to the phrase influence coefficient, and keywords are selected in descending order of the sensitivity coefficient.
[0009] During the generation of image conversion supplement, the segmentation method is determined according to the image disorder degree and distribution uniformity degree of the analyzed image in the user analysis data to generate divided paragraphs, and the keyword generation method is determined according to the characteristic representativeness of the divided paragraphs to generate keywords based on the median point or outlier.
[0010] Under the condition that the initial search is completed, optimization analysis is carried out, the density of repeated keywords is detected and the analysis status of repeated keywords is determined using the reference deviation value, and the search optimization method is determined as combined search or sequential search according to the analysis status of repeated keywords.
[0011] Furthermore, the keyword selection method is determined according to the data status of the user analysis data, including:
[0012] When the data status is that the data abundance degree is greater than the preset data abundance degree, the keyword selection method is associated phrase selection generation;
[0013] When the data status is that the data abundance degree is less than or equal to the preset data abundance degree and the image conversion completion degree is greater than the preset image conversion completion degree, the keyword selection method is image conversion supplement generation.
[0014] Furthermore, the confirmation method of the data abundance degree includes:
[0015] Analyze the user analysis data to obtain the text representation coefficient;
[0016] If the text representation coefficient is less than or equal to the preset text representation coefficient, the data abundance degree is determined according to the search frequency reference value and the keyword effective search reference value;
[0017] If the text representation coefficient is greater than the preset text representation coefficient, the data abundance degree is determined according to the keyword quantity reference value and the keyword distribution reference value.
[0018] Furthermore, the confirmation method of the image conversion completion degree includes:
[0019] Determine the preset image associated text length according to the complexity reference value ratio of the image complexity reference value and the preset image complexity reference value;
[0020] Determine the image conversion completion degree according to the difference value of the image associated text length;
[0021] The complexity reference value ratio and the preset image associated text length have a positive correlation.
[0022] Further, the generation of associated word groups includes:
[0023] For a single associated word group,
[0024] Determine the number of keywords to be selected according to the word group influence coefficient, and select keywords in the order of decreasing sensitivity coefficient;
[0025] The number of keywords selected in a single associated word group is positively correlated with the word group influence coefficient corresponding to the associated word group;
[0026] The associated word groups are determined according to the character interval reference value and the similarity threshold.
[0027] Further, the generation of image conversion and supplementation includes:
[0028] For a single analysis image in the user analysis data, determine the segmentation method according to the image disorder degree and the distribution trend degree of the analysis image;
[0029] If the image disorder degree is greater than the preset image disorder degree or the distribution trend degree is less than or equal to the preset distribution trend degree, the segmentation method is associated division;
[0030] If the image disorder degree is less than or equal to the preset image disorder degree and the distribution trend degree is greater than the preset distribution trend degree, the segmentation method is uniform division.
[0031] Further, for a single segmented paragraph, determine the keyword generation method according to the feature representativeness of the segmented paragraph;
[0032] If the feature representativeness is greater than or equal to the preset feature representativeness, the keyword generation method is to generate keywords according to the median point;
[0033] If the feature representativeness is less than the preset feature representativeness, the keyword generation method is to generate keywords according to the outlier.
[0034] Further, the optimization analysis includes:
[0035] Determine the usage reference value corresponding to the repeated keyword according to the influence range of the repeated keyword;
[0036] Detect the density of the repeated keyword and the usage reference deviation value to determine the analysis status of the repeated keyword;
[0037] Determine the search optimization method according to the analysis status of the repeated keyword.
[0038] Further, for a single repeated keyword, the confirmation method for its corresponding influence range includes:
[0039] Extract the text segments where the repeated keyword exists, and conduct sub-analysis on each text segment;
[0040] In the sub - analysis, the initial analysis range of the repeated keywords is determined according to the segmentation coefficient of the text paragraph, the effective data volume within the initial analysis range is detected, and the initial analysis range for extracting the repeated keywords is adjusted to increase according to the difference between the effective data volume and the preset effective data volume;
[0041] The initial analysis range after the adjustment to increase is denoted as the influence range.
[0042] Furthermore, the search optimization method is determined according to the analysis status of the repeated keywords, including:
[0043] If the analysis status of the repeated keywords is that the density of the repeated keywords is greater than the preset density of the repeated keywords or the usage reference deviation value is less than or equal to the preset usage reference deviation value, the search optimization method is combined search;
[0044] If the analysis status of the repeated keywords is that the density of the repeated keywords is less than or equal to the preset density of the repeated keywords and the usage reference deviation value is greater than the preset usage reference deviation value, the search optimization method is sequential search.
[0045] Compared with the prior art, the beneficial effect of the present invention is that in the technical solution of the present invention, the data status of the user - analysis data is determined according to the data abundance of the user - analysis data and the image conversion completion degree. By analyzing the user - analysis data, it is reflected whether the keywords therein can effectively meet the search requirements, and different keyword selection methods are correspondingly selected, so that the keyword selection method can accurately select the keywords required for data search.
[0046] Furthermore, in the present invention, a keyword selection method supplemented by image conversion is applied. By identifying and analyzing the image to generate corresponding keywords, the image - based information is converted into a supplementary basis for keywords, avoiding the problem of poor search efficiency based on keywords caused by the single keyword selection method in the prior art, and thus improving the integration search efficiency of the pathological data of the present invention.
[0047] Furthermore, in the present invention, the segmentation method is determined according to the image disorder degree and the distribution trend degree of the analysis image in the user - analysis data to generate divided paragraphs, making the determination of the divided paragraphs more in line with the actual image state and avoiding the problem of poor data efficiency caused by a single image segmentation method.
[0048] Furthermore, in the present invention, the density of the repeated keywords and the usage reference deviation value are detected to determine the analysis status of the repeated keywords. The importance of the keywords in the current search results for the search contribution is reflected through the keyword analysis status, and different search optimization methods are correspondingly selected, improving the integration search efficiency of the pathological data. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 Schematic diagram of the information collection and processing method for heat stroke pathological data integration based on deep learning according to the present invention;
[0050] Figure 2 Flowchart of determining the keyword selection method according to the data status of the user analysis data in the present invention;
[0051] Figure 3 Flowchart of determining the segmentation method according to the image disorder degree and distribution trend degree of the analysis image in the present invention;
[0052] Figure 4 Flowchart of determining the search optimization method according to the repeated keyword analysis status in the present invention. Detailed implementation manners
[0053] In order to make the objectives and advantages of the present invention clearer and more understandable, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0054] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principle of the present invention and do not limit the protection scope of the present invention.
[0055] It should be noted that in the description of the present invention, the terms indicating directions or positional relationships such as "upper", "lower", "left", "right", "inner", "outer", etc. are based on the directions or positional relationships shown in the drawings. This is only for convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.
[0056] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0057] Please refer to Figures 1 to 4 As shown, the present invention provides an information collection and processing method for heat stroke pathological data integration based on deep learning, including:
[0058] Obtaining user analysis data, and determining the data status of the user analysis data according to the data richness degree and image conversion completion degree of the user analysis data;
[0059] Determine the keyword selection method according to the data status, and the keyword selection method is to generate by associated phrase selection or to supplement and generate by image conversion;
[0060] In the generation by associated phrase selection, determine the number of selected keywords according to the phrase influence coefficient, and select keywords in the order from large to small according to the sensitivity coefficient;
[0061] In the supplement and generation by image conversion, determine the segmentation method according to the image disorder degree and distribution uniformity degree of the analysis image in the user analysis data to generate divided paragraphs, and determine the keyword generation method according to the characteristic representativeness of the divided paragraphs to generate keywords according to the median point or outlier;
[0062] Under the condition that the initial search is completed, perform optimization analysis, detect the density of duplicate keywords and use the reference deviation value to determine the analysis status of duplicate keywords, and determine the search optimization method as combined search or sequential search according to the analysis status of duplicate keywords.
[0063] The application scenario of the present invention is to search for analysis data through deep learning, perform intelligent evaluation of data status, automatic generation of keywords and optimization analysis of search results through deep learning technology, so that the information collection results are more efficient, accurate and intelligent; the analysis data includes detected text and several analysis images, the detected text is text related to heat stroke generated by the user himself, the analysis images include several two-dimensional coordinate images, and the two-dimensional coordinate images include but are not limited to a blood cell analysis graph with red blood cell volume as the abscissa and red blood cell count as the ordinate and a body temperature curve graph with time as the abscissa and body temperature as the ordinate, which is easy for those skilled in the art to understand and will not be elaborated here;
[0064] The user analysis data is a single analysis data currently being analyzed. The keywords of the present invention are keywords related to heat stroke, and the keywords include but are not limited to nuclear condensation, villous damage, intestinal microbiota, and organ damage. In the present invention, the method also needs to apply a pathological information database, and the pathological information database is used to store several data to be selected, and the data to be selected is the analysis data uploaded by the user himself in the historical process. The above are all common technical means in the art and will not be elaborated here;
[0065] In the present invention, after determining the keyword selection method according to the data status, perform an initial search. The initial search includes: recording the keywords selected by the keyword selection method as search keywords, and recording the data to be selected with any one of the search keywords as the initial data; the condition for the completion of the initial search is the completion of the determination of the initial data;
[0066] In the present invention, a number of historical records are correspondingly set. Any one of the historical records records the data abundance, image conversion completion degree, keyword quantity reference value, keyword distribution reference value, change degree, image complexity reference value, etc. in the historical process of at least one data search. And each historical record corresponds to a qualified mark, which records whether the data search process meets the user's needs. The qualified mark can be recorded manually. It can be understood that the user can determine whether the data search process meets the requirements according to the self-set indicators. The self-set indicators can be, but are not limited to, the error search index, which will not be elaborated here. Among them, the error search index is the number of times of data to be selected for error search.
[0067] Specifically, the keyword selection method is determined according to the data state of the user's analyzed data, including:
[0068] When the data state is that the data abundance is greater than the preset data abundance, the keyword selection method is to generate by selecting associated word groups;
[0069] When the data state is that the data abundance is less than or equal to the preset data abundance and the image conversion completion degree is greater than the preset image conversion completion degree, the keyword selection method is to generate by supplementing the image conversion.
[0070] Among them, the data state includes the first data state, the second data state, and the third data state. The first data state is that the data abundance is greater than the preset data abundance. The second data state is that the data abundance is less than or equal to the preset data abundance and the image conversion completion degree is greater than the preset image conversion completion degree. The third data state is that the data abundance is less than or equal to the preset data abundance and the image conversion completion degree is less than or equal to the preset image conversion completion degree;
[0071] It should be noted that in the third data state, the keyword selection method is to generate by selecting associated word groups, and an increase adjustment is made for the initial search volume. The increase value of the initial search volume and the comprehensive evaluation value have a negative correlation;
[0072] The comprehensive evaluation value = data abundance + image conversion completion degree, and the initial search volume is the number of data to be selected for the initial search;
[0073] For the values of the preset data abundance and the preset image conversion completion degree, the user can determine them according to the actual application scenario. The larger the value of the preset data abundance and the smaller the value of the preset image conversion completion degree, the greater the user's need to generate keywords by supplementing the image conversion. Provide a set of values for the preset data abundance and the preset image conversion completion degree, detect the historical records of the user generating keywords by supplementing the image conversion, and record the average value of the data abundance of the historical records that can meet the user's needs as the preset data abundance, and record the average value of the image conversion completion degree of the historical records that can meet the user's needs as the preset image conversion completion degree.
[0074] Specifically, the method for confirming the data abundance includes:
[0075] Analyze the user analysis data to obtain the text representation coefficient;
[0076] If the text representation coefficient is less than or equal to the preset text representation coefficient, determine the data abundance according to the search frequency reference value and the keyword effective search reference value;
[0077] If the text representation coefficient is greater than the preset text representation coefficient, determine the data abundance according to the keyword quantity reference value and the keyword distribution reference value.
[0078] Among them, the text representation coefficient = the sum of the numbers of characters corresponding to each keyword in the user analysis data / the total number of characters in the user analysis data;
[0079] For the value of the preset text representation coefficient, the user can determine it according to the actual application scenario. The larger the value of the preset text representation coefficient, the greater the need for the user to determine the data abundance according to the search frequency reference value and the keyword effective search reference value. Provide a value of the preset text representation coefficient, and the preset text representation coefficient is 70%.
[0080] If the text representation coefficient is less than or equal to the preset text representation coefficient, then the data abundance = the search frequency reference value + the keyword effective search reference value;
[0081] If the text representation coefficient is greater than the preset text representation coefficient, then the data abundance = (the keyword quantity reference value / the preset keyword quantity reference value) + (the keyword distribution reference value / the preset keyword distribution reference value).
[0082] The search frequency reference value is the average value of the sub-search frequencies corresponding to each keyword in the user analysis data; the sub-search frequency corresponding to a single keyword is the number of times of the first search using this keyword in the historical records that can meet the user's needs; the keyword effective search reference value is the average value of the effective coefficients corresponding to each keyword in the user analysis data;
[0083] The method for confirming the effective coefficient is as follows: for a single keyword, record this keyword as the target keyword, record the historical records that can meet the user's needs and use the target keyword for the first search as the reference historical records, record the data to be selected obtained by the first search using the target keyword in the reference historical records as the first reference pathological data, record the data to be selected obtained by the search optimization method and the same as the first reference pathological data in the reference historical records as the second reference pathological data, and the effective coefficient corresponding to the target keyword = the total amount of the second reference pathological data / the total amount of the first reference pathological data;
[0084] The reference value of the number of keywords is the total number of keywords in the user analysis data. The reference value of the keyword distribution is the average of the reference distances corresponding to each keyword in the user analysis data. The reference distance corresponding to a single keyword is the average of the reference values of the character intervals between this keyword and other keywords in the user analysis data. For any two keywords, the reference value of the character interval is the total number of characters between the two keywords;
[0085] For the values of the preset reference value of the number of keywords and the preset reference value of the keyword distribution, the user can determine them according to the historical records. It can be understood that the smaller the values of the preset reference value of the number of keywords and the preset reference value of the keyword distribution, the greater the determination result of the data richness, and the greater the user's demand for generating keywords by selecting associated word groups. Provide a value of the preset reference value of the number of keywords and the preset reference value of the keyword distribution, detect the historical records for determining the data richness according to the reference value of the number of keywords and the reference value of the keyword distribution, and record the average value of the reference values of the number of keywords corresponding to the historical records that can meet the user's needs as the preset reference value of the number of keywords, and record the average value of the reference values of the keyword distribution corresponding to the historical records that can meet the user's needs as the preset reference value of the keyword distribution.
[0086] Specifically, the confirmation method of the image conversion completion degree includes:
[0087] Determine the preset image associated text length according to the complexity reference value ratio of the image complexity reference value to the preset image complexity reference value;
[0088] Determine the image conversion completion degree according to the difference in the image associated text length;
[0089] The complexity reference value ratio and the preset image associated text length have a positive correlation.
[0090] Among them, the complexity reference value ratio = image complexity reference value / preset image complexity reference value;
[0091] The image conversion completion degree has a positive correlation with the difference in the image associated text length. The difference in the image associated text length = image associated text length - preset image associated text length;
[0092] The image complexity reference value is the average of the complexity thresholds corresponding to each analyzed image in the user analysis data. The complexity threshold corresponding to a single analyzed image = the image disorder degree corresponding to this analyzed image + the distribution uniformity degree corresponding to this analyzed image;
[0093] For a single analysis image, denote this analysis image as the target image. Divide the abscissa interval corresponding to the target image into n equal parts, and n equally spaced points can be obtained. The value of n is positively correlated with the total length of the abscissa. The abscissa interval is the range between the abscissa value corresponding to the first point and the abscissa value corresponding to the last point in the target image. The total length of the abscissa = the abscissa corresponding to the last point in the target image - the abscissa corresponding to the first point in the target image. Denote each equally spaced point, as well as the first point and the last point in the target image, as characteristic points;
[0094] The image disorder degree is the standard deviation of the ordinate values corresponding to the characteristic points in the target image;
[0095] The distribution trend degree = 1 - [the number of abnormal points / (n + 2)]. The number of abnormal points is the number of characteristic points with a variation degree greater than the preset variation degree. For a single characteristic point, denote this characteristic point as the target characteristic point, and denote the characteristic points adjacent to the target characteristic point as adjacent points. The variation degree corresponding to the target characteristic point is the larger value among the difference thresholds corresponding to the adjacent points. The difference threshold corresponding to a single adjacent point is the absolute value of the difference between the ordinate value corresponding to the target characteristic point and the ordinate value corresponding to this adjacent point;
[0096] In the present invention, each analysis image corresponds to a figure number. The manifestation form of the figure number can be number 1, number a, and number 1(a), which is easy for those skilled in the art to understand and will not be elaborated herein;
[0097] The length of the image-associated text is the average value of the associated length reference values corresponding to each analysis image in the user analysis data. For a single analysis image, detect the position of the figure number corresponding to this analysis image in the detection text corresponding to the user analysis data, and denote the sum of the number of keywords in each sentence where the figure number is located as the associated length reference value corresponding to this analysis image. The method for confirming the sentence where the figure number is located is as follows: for a single figure number among the figure numbers corresponding to a single analysis image, denote this figure number as the target figure number, and denote the characters between the first full stop before the reading order of the target figure number and the first full stop after the reading order of the target figure number as the sentence where the target figure number is located. The reading order is the order from front to back during text reading;
[0098] The values of the preset degree of change, the preset image complexity reference value, and the preset length of the image-related text can be determined by the user according to the actual application scenario. The larger the value of the preset degree of change, the greater the user's need to determine the feature points as abnormal points. By taking a value of the preset degree of change, the historical records of determining the feature points as abnormal points are detected, and the average value of the degrees of change corresponding to the historical records that can meet the user's needs is recorded as the preset degree of change. The larger the values of the preset image complexity reference value and the preset length of the image-related text, the smaller the determined degree of image conversion completion, that is, the greater the user's need to generate keywords by selecting associated word groups. A value of the preset image complexity reference value and a preset length of the image-related text are provided. The historical records of the user generating keywords by supplementing image conversion are detected, and the average value of the image complexity reference values corresponding to the historical records that can meet the user's needs is recorded as the preset image complexity reference value, and the average value of the lengths of the image-related texts corresponding to the historical records that can meet the user's needs is recorded as the preset length of the image-related text.
[0099] Specifically, the selection of associated word groups to generate keywords includes:
[0100] For a single associated word group,
[0101] The number of keywords to be selected is determined according to the phrase influence coefficient, and the keywords are selected in the order from the largest to the smallest sensitivity coefficient;
[0102] The number of keywords selected from a single associated word group is positively correlated with the phrase influence coefficient corresponding to the associated word group;
[0103] The associated word groups are determined according to the character interval reference value and the similarity threshold.
[0104] Among them, the associated word groups are determined according to the character interval reference value and the similarity threshold. Among them, for each keyword in the user analysis data, correlation analysis is performed in the order from front to back according to the reading order. When performing correlation analysis on a single keyword, the keyword is recorded as the target keyword, and the keywords that are not recorded in the associated word group except the target keyword are recorded as reference keywords. The set of reference keywords whose character interval reference value with the target keyword is less than the preset character interval reference value and whose similarity threshold is greater than the preset similarity threshold is recorded as an associated word group, and the correlation analysis is continued for the keywords that are not recorded in the associated word group until all keywords are recorded in the associated word group;
[0105] The confirmation method of the similarity threshold is as follows: For any two keywords, which are recorded as the analysis keywords, the number of qualified data in which the two analysis keywords exist at the same time is recorded as the similarity threshold. The qualified data are the selected data to be selected by the search optimization method in the historical records that can meet the user's needs;
[0106] The values of the preset character interval reference value and the preset similarity threshold can be determined by the user according to the actual application scenario. The greater the user's demand for improving the search accuracy, the smaller the value of the preset character interval reference value and the greater the value of the preset similarity threshold. A set of values for the preset character interval reference value and the preset similarity threshold is provided, where the preset character interval reference value is 20 characters and the preset similarity threshold is 50% of the number of qualified data;
[0107] For a single associated phrase, denote this associated phrase as the target associated phrase. The phrase influence coefficient corresponding to the target associated phrase = the total number of keywords in the target associated phrase + the radiation coefficient corresponding to the target associated phrase. Denote each keyword in the target associated phrase as the first keyword, and denote the keywords other than the first keyword as the second keyword. Denote the second keyword that is the same as the first keyword and the first keyword as the analysis keyword. Denote the number of characters between the first analysis keyword and the last analysis keyword in the forward reading order as the radiation coefficient corresponding to the target associated phrase;
[0108] The confirmation method of the sensitivity coefficient is as follows: for a single keyword, denote this keyword as the first target keyword. Denote the data to be selected that is selected through the search optimization method in the historical records that can meet the user's needs as the first data, and denote the other data to be selected except the first data as the second data. The sensitivity coefficient corresponding to the first target keyword = (the number of times the first target keyword appears in the first data / the total amount of the first data) - (the number of times the first target keyword appears in the second data / the total amount of the second data).
[0109] Specifically, for a single analysis image in the user analysis data, determine the segmentation method according to the image disorder degree and the distribution trend degree of the analysis image;
[0110] If the image disorder degree is greater than the preset image disorder degree or the distribution trend degree is less than or equal to the preset distribution trend degree, the segmentation method is association division;
[0111] If the image disorder degree is less than or equal to the preset image disorder degree and the distribution trend degree is greater than the preset distribution trend degree, the segmentation method is uniform division.
[0112] Among them, the values of the preset image disorder degree and the preset distribution trend degree can be determined by the user according to the actual application scenario. The greater the value of the preset image disorder degree and the smaller the value of the preset distribution trend degree, the greater the user's demand for uniform division. A set of values for the preset image disorder degree and the preset distribution trend degree is provided. Detect the historical records of the user's uniform division, and denote the average value of the image disorder degrees corresponding to the historical records that can meet the user's needs as the preset image disorder degree, and the preset distribution trend degree is 80%;
[0113] In the associated division, for a single analysis image, the associated paragraph analysis is performed on each feature paragraph in the order of increasing abscissa values of the analysis image. When performing the associated paragraph analysis on a single feature paragraph, the feature paragraph is denoted as the target feature paragraph. The slope differences between each feature paragraph located on the right side of the target feature paragraph are sequentially detected. If the slope difference between a feature paragraph and the target feature paragraph is less than the preset slope difference, it is denoted as the associated feature paragraph of the target feature paragraph. If the slope difference between a feature paragraph and the target feature paragraph is greater than or equal to the preset slope difference, the associated paragraph analysis for the target feature paragraph is stopped, and the target feature paragraph and its corresponding associated feature paragraph are uniformly denoted as a division paragraph, and the associated paragraph analysis is continued for the feature paragraphs not included in the division paragraph;
[0114] For two feature paragraphs, the slope difference is equal to the absolute value of the difference between the paragraph slope of one feature paragraph and the paragraph slope of the other feature paragraph. For a single feature paragraph, the ordinate value corresponding to the leftmost feature point of the feature paragraph is denoted as the first value, and the ordinate value corresponding to the rightmost feature point of the feature paragraph is denoted as the second value. The paragraph slope of the feature paragraph = (second value - first value) / feature paragraph length, where the feature paragraph length is the length of the abscissa corresponding to the feature paragraph; The value of the preset slope difference can be determined by the user according to the actual application scenario. The smaller the value of the preset slope difference, the greater the user's requirement for improving the association degree of the feature paragraphs in the division paragraph. A value of the preset slope difference is provided, and the historical records of the associated division are detected, and the average value of the slope differences corresponding to the historical records that can meet the user's requirements is denoted as the preset slope difference;
[0115] In the uniform division, the analysis image is uniformly divided into several division paragraphs with the same abscissa length, and the number of division paragraphs has a positive correlation with the complexity threshold.
[0116] Specifically, for a single division paragraph, the keyword generation method is determined according to the feature representativeness of the division paragraph;
[0117] If the feature representativeness is greater than or equal to the preset feature representativeness, the keyword generation method is to generate keywords according to the median point;
[0118] If the feature representativeness is less than the preset feature representativeness, the keyword generation method is to generate keywords according to the outlier points.
[0119] The method for confirming the feature representativeness is as follows: for a single segment in a single analysis image, the segment is recorded as the target segment, and the feature representativeness = deviation index + correlation index. The other segments in the analysis image other than the target segment are recorded as reference segments, and the deviation index = the average value of the ordinate values corresponding to each feature point of the target segment - the average value of the ordinate values corresponding to each feature point in the analysis image. The correlation index is the average value of the sub-correlation coefficients corresponding to the target segment and each reference segment. For any two segments, the calculation formula for the sub-correlation coefficient r of the two segments is:
[0120] Detect the number of feature points corresponding to the two divided paragraphs. If the number of feature points is different, the smaller value of the number of feature points corresponding to the two divided paragraphs is recorded as n, and the first n feature points in the two divided paragraphs are recorded as reference points in descending order of the horizontal coordinates; if the number of feature points is the same, all feature points in the two divided paragraphs are recorded as reference points;
[0121]
[0122] Where n is the number of reference points corresponding to a single segment; x i and y i are the ordinate values of the i-th reference point within the monitoring time corresponding to the two divided sections, is x i The average value of the vertical coordinate values corresponding to each reference point of the corresponding divided paragraph, y i The average value of the ordinate values corresponding to the reference points of the corresponding segment division, i = 1, 2, 3, ..., n;
[0123] The value of the preset feature representativeness can be determined by the user according to the actual application scenario. The smaller the value of the preset feature representativeness, the greater the user's demand for generating keywords based on the median point. A value of the preset feature representativeness is provided, and the historical records of users generating keywords based on the median point are detected. The average value of the feature representativeness corresponding to the historical records that can meet the user's needs is recorded as the preset feature representativeness;
[0124] When generating keywords based on the median point, the abscissa value corresponding to the abscissa name of the median point corresponding to the feature paragraph and the ordinate value corresponding to the ordinate name of the median point are used as the generated keywords;
[0125] The median point corresponding to a single divided paragraph is the point whose abscissa is the reference value, where the reference value = (the abscissa corresponding to the leftmost point of the divided paragraph + the abscissa corresponding to the rightmost point of the divided paragraph) / 2;
[0126] When generating keywords based on outliers, the abscissa value corresponding to the abscissa name of the outlier of the characteristic paragraph and the ordinate value corresponding to the ordinate name of the outlier are used as the generated keywords;
[0127] An outlier is the characteristic point with the greatest degree of change in a single divided paragraph.
[0128] Specifically, the optimization analysis includes:
[0129] Determine the usage reference value corresponding to the repeated keyword according to the influence range of the repeated keyword;
[0130] Detect the density of repeated keywords and the usage reference deviation value to determine the analysis status of repeated keywords;
[0131] Determine the search optimization method according to the analysis status of repeated keywords.
[0132] Among them, the repeated keyword is the search keyword that appears in each initial data, and the usage reference value corresponding to the repeated keyword has a positive correlation with the influence range of the repeated keyword;
[0133] The analysis status of repeated keywords includes the first analysis status of repeated keywords and the second analysis status of repeated keywords. The first analysis status of repeated keywords is that the density of repeated keywords is greater than the preset density of repeated keywords or the usage reference deviation value is less than or equal to the preset usage reference deviation value. The second analysis status of repeated keywords is that the density of repeated keywords is less than or equal to the preset density of repeated keywords and the usage reference deviation value is greater than the preset usage reference deviation value.
[0134] Specifically, for a single repeated keyword, the confirmation method of its corresponding influence range includes:
[0135] Extract the text segments where the repeated keyword exists, and perform sub-analysis on each text segment;
[0136] In the sub-analysis, determine the initial analysis range of the repeated keyword according to the segmentation coefficient of the text segment, detect the amount of valid data in the initial analysis range, and increase and adjust the initial analysis range for extracting the repeated keyword according to the difference between the amount of valid data and the preset amount of valid data;
[0137] Record the initial analysis range after the increase and adjustment as the influence range.
[0138] Among them, for a single repeated keyword, the repeated keyword is denoted as the target repeated keyword, and each data to be selected obtained from the initial search is denoted as the initial data. The target repeated keyword corresponds to a text segment in each initial data. The text segment corresponding to the target repeated keyword in a single initial data is the text between the first target repeated keyword and the last target repeated keyword that appear in the forward reading order in the initial data;
[0139] Partition coefficient = 1 - [(the number of target repeated keywords in a single text segment + 2) / (the number of keywords in a single text segment + 2)];
[0140] The initial analysis range of the repeated keyword has a positive correlation with the partition coefficient of the text segment. The initial analysis range of the repeated keyword is the number of characters selected for analysis in a single text segment in the forward reading order;
[0141] The effective data volume is the number of keywords whose co - usage times in the single text segment corresponding to the target repeated keyword are greater than the preset co - usage times. The confirmation method of the co - usage times is as follows: for a single keyword in the text segment, the keyword is denoted as the text segment keyword, and the number of the first data where both the text segment keyword and the target repeated keyword exist is denoted as the co - usage times corresponding to the text segment keyword;
[0142] According to the difference between the effective data volume and the preset effective data volume, the initial analysis range for extracting the repeated keyword is adjusted to increase. The increase value of the initial analysis range for extracting the repeated keyword has a positive correlation with the difference between the effective data volume and the preset effective data volume; the difference between the effective data volume and the preset effective data volume = effective data volume - preset effective data volume;
[0143] For the values of the preset co - usage times and the preset effective data volume, the user can determine them according to the actual application scenario. The greater the user's demand for improving the search efficiency, the greater the values of the preset co - usage times and the preset effective data volume. A set of values for the preset co - usage times and the preset effective data volume is provided. The preset co - usage times is 50% of the total amount of the first data, and the preset effective data volume is 60% of the total amount of keywords existing in a single text segment.
[0144] Specifically, according to the analysis status of the repeated keyword, the search optimization method is determined, including:
[0145] If the analysis status of the repeated keyword is that the repeated keyword density is greater than the preset repeated keyword density or the usage reference deviation value is less than or equal to the preset usage reference deviation value, the search optimization method is combined search;
[0146] If the duplicate keyword analysis status is that the duplicate keyword density is less than or equal to the preset duplicate keyword density and the usage reference deviation value is greater than the preset usage reference deviation value, the search optimization method is sequential search.
[0147] Among them, the duplicate keyword analysis status includes the first duplicate keyword analysis status and the second duplicate keyword analysis status. The first duplicate keyword analysis status is that the duplicate keyword density is greater than the preset duplicate keyword density or the usage reference deviation value is less than or equal to the preset usage reference deviation value. The second duplicate keyword analysis status is that the duplicate keyword density is less than or equal to the preset duplicate keyword density and the usage reference deviation value is greater than the preset usage reference deviation value.
[0148] The method for confirming the duplicate keyword density is to detect the positions of each duplicate keyword in the user analysis data. The duplicate keyword density = the area of the smallest rectangle that can contain each duplicate keyword in the user analysis data / the area of the smallest rectangle that can contain each search keyword in the user analysis data. The usage reference deviation value is the standard deviation of the search quantities corresponding to each initial data, and the search quantity is the number of search keywords existing in a single initial data.
[0149] For the values of the preset duplicate keyword density and the preset usage reference deviation value, the user can determine them according to the actual application scenario. The smaller the value of the preset duplicate keyword density and the larger the value of the preset usage reference deviation value, the greater the user's demand for combined search. Provide a set of values for the preset duplicate keyword density and the preset usage reference deviation value. Detect the historical records of the user's combined search, and record the average value of the duplicate keyword density corresponding to the historical records that can meet the user's needs as the preset duplicate keyword density, and record the average value of the usage reference deviation value corresponding to the historical records that can meet the user's needs as the preset usage reference deviation value.
[0150] When the search optimization method is combined search, combined analysis is performed on each search keyword. When performing combined analysis on a single search keyword, this search keyword is recorded as the target search keyword, and each search keyword other than the target search keyword is recorded as the reference search keyword. The combination of the reference search keyword whose matching coefficient with the target search keyword is greater than the preset matching coefficient and the target search keyword is recorded as a keyword combination, and combined analysis is continued for the search keywords that have not been analyzed for combination until each search keyword is recorded in the keyword combination, and the searches are performed for each keyword combination in the order of the combination influence coefficient from large to small. When performing a search for a single keyword combination, the data to be selected with each search keyword corresponding to this keyword combination is used as the search data.
[0151] When the search optimization method is sequential search, each search keyword is searched in descending order of the emergence coefficient of the search keyword. When searching for a single search keyword, the data to be selected that contains the search keyword is used as the search data;
[0152] The method for confirming the matching coefficient is as follows: for any two search keywords, the matching coefficient = the number of initial data that simultaneously contains the two search keywords / the total amount of initial data; the value of the preset matching coefficient can be determined by the user according to the actual application scenario. The smaller the value of the preset matching coefficient, the greater the user's demand for expanding the search scope. A value of the preset matching coefficient is provided, and the preset matching coefficient is 40%;
[0153] The method for confirming the combined influence coefficient is as follows: for a single keyword combination, the search keywords corresponding to the keyword combination are recorded as the search keywords to be analyzed, and the total amount of the search keywords to be analyzed existing in the user analysis data is recorded as the combined influence coefficient;
[0154] For a single search keyword, the emergence coefficient is the total amount of the search keyword existing in the user analysis data.
[0155] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
[0156] The above are only the preferred embodiments of the present invention and are not used to limit the present invention; for those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An information collection and processing method for integrating heat stroke pathological data based on deep learning, characterized in that, Including: Obtain user analysis data, and determine the data status of the user analysis data according to the data abundance degree and the image conversion completion degree of the user analysis data; Determine the keyword selection method according to the data status, and the keyword selection method is associated phrase selection generation or image conversion supplementary generation; In the associated phrase selection generation, determine the number of selected keywords according to the phrase influence coefficient, and select keywords in the order from large to small according to the sensitivity coefficient; In the image conversion supplementary generation, determine the segmentation method according to the image disorder degree and the distribution trend degree of the analysis image in the user analysis data to generate divided paragraphs, and determine the keyword generation method according to the feature representativeness of the divided paragraphs as generating keywords according to the median point or the outlier point; Under the condition of the initial search completion, perform optimization analysis, detect the density of duplicate keywords and use the reference deviation value to determine the analysis status of duplicate keywords, and determine the search optimization method as combined search or sequential search according to the analysis status of duplicate keywords.
2. The information collection and processing method for heat stroke pathological data integration based on deep learning according to claim 1, characterized in that, Determine the keyword selection method according to the data status of the user analysis data, including: When the data status is that the data abundance degree is greater than the preset data abundance degree, the keyword selection method is associated phrase selection generation; When the data status is that the data abundance degree is less than or equal to the preset data abundance degree and the image conversion completion degree is greater than the preset image conversion completion degree, the keyword selection method is image conversion supplementary generation.
3. The information collection and processing method for heat stroke pathological data integration based on deep learning according to claim 2, characterized in that, The confirmation method of the data abundance degree includes: Analyze the user analysis data to obtain the text representation coefficient; If the text representation coefficient is less than or equal to the preset text representation coefficient, determine the data abundance degree according to the search frequency reference value and the keyword effective search reference value; If the text representation coefficient is greater than the preset text representation coefficient, determine the data abundance degree according to the keyword quantity reference value and the keyword distribution reference value.
4. The information collection and processing method for heat stroke pathological data integration based on deep learning according to claim 3, characterized in that, The confirmation method of the image conversion completion degree includes: Determine the preset image associated text length according to the complex reference value ratio of the image complexity reference value and the preset image complexity reference value; Determine the image conversion completion degree according to the difference of the image associated text length; The complex reference value ratio and the preset image associated text length are in a positive correlation relationship.
5. The information collection and processing method for heat stroke pathological data integration based on deep learning according to claim 4, characterized in that, The associated phrase selection generation includes: For a single associated phrase, Determine the number of selected keywords according to the phrase influence coefficient, and select keywords in the order from large to small according to the sensitivity coefficient; The number of selected keywords in a single associated phrase is in a positive correlation relationship with the phrase influence coefficient corresponding to the associated phrase; The associated phrase is determined according to the character interval reference value and the similarity threshold.
6. The information collection and processing method for heat stroke pathological data integration based on deep learning according to claim 5, wherein For a single analysis image in the user analysis data, determine the segmentation method according to the image disorder degree and the distribution trend degree of the analysis image; If the image disorder degree is greater than the preset image disorder degree or the distribution trend degree is less than or equal to the preset distribution trend degree, the segmentation method is associated segmentation; If the image disorder degree is less than or equal to the preset image disorder degree and the distribution trend degree is greater than the preset distribution trend degree, the segmentation method is uniform segmentation.
7. The information collection and processing method for heat stroke pathological data integration based on deep learning according to claim 6, characterized in that, For a single divided paragraph, determine the keyword generation method according to the feature representativeness of the divided paragraph; If the feature representativeness is greater than or equal to the preset feature representativeness, the keyword generation method is to generate keywords according to the median point; If the feature representation degree is less than the preset feature representation degree, the keyword generation method is to generate keywords based on outliers.
8. The information collection and processing method for heat stroke pathological data integration based on deep learning according to claim 7, characterized in that, The optimization analysis includes: Determining the usage reference value corresponding to the repeated keyword according to the influence range of the repeated keyword; Detecting the density of the repeated keyword and the usage reference deviation value to determine the analysis status of the repeated keyword; Determining the search optimization method according to the analysis status of the repeated keyword.
9. The information collection and processing method for heat stroke pathological data integration based on deep learning according to claim 8, characterized in that, For a single repeated keyword, the method for confirming its corresponding influence range includes: Extracting the text segments where the repeated keyword exists and performing sub-analysis on each text segment; In the sub-analysis, determining the initial analysis range of the repeated keyword according to the segmentation coefficient of the text segment, detecting the amount of valid data in the initial analysis range, and increasing and adjusting the initial analysis range for extracting the repeated keyword according to the difference between the amount of valid data and the preset amount of valid data; Recording the initial analysis range after the increase adjustment as the influence range.
10. The information collection and processing method for heat stroke pathological data integration based on deep learning according to claim 9, wherein, Determining the search optimization method according to the analysis status of the repeated keyword, including: If the analysis status of the repeated keyword is that the density of the repeated keyword is greater than the preset density of the repeated keyword or the usage reference deviation value is less than or equal to the preset usage reference deviation value, the search optimization method is combined search; If the analysis status of the repeated keyword is that the density of the repeated keyword is less than or equal to the preset density of the repeated keyword and the usage reference deviation value is greater than the preset usage reference deviation value, the search optimization method is sequential search.
Citation Information
Patent Citations
Medical data retrieval method and system
CN111223533A
Comprehensive information analysis method and system based on keyword extraction
CN117993392A
High-order medical statistical data processing method and system based on intelligent medical engineering
CN118571498A
Pathological literature searching and dialogue system based on large language model and RAG technology
CN118643128A
Petroleum business data asset retrieval method and system
CN119336809A
Cited By
Heat stroke medical record text classification method based on reinforcement learning
CN121278096A