Information collection processing method for heat stroke pathological data integration based on deep learning

By using deep learning technology to select keywords and optimize search strategies based on data status, the problem of low search efficiency for heatstroke pathological data has been solved, achieving more efficient and accurate information collection.

CN120413073BActive Publication Date: 2026-01-23THE FIRST MEDICAL CENT CHINESE PLA GENERAL HOSPITAL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510353506.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2026-01-23
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

In existing technologies, the search efficiency for pathological data of heatstroke is low, and the search methods cannot be effectively filtered and optimized, resulting in poor search efficiency.

Method used

Using a deep learning-based approach, the data status is determined based on the data richness and image conversion completion of the user's analytical data. Keywords are generated by selecting related phrases or supplementing image conversion, and optimization analysis and search optimization are performed, including detecting the density of duplicate keywords and using reference deviation values ​​to determine the search optimization method.

Benefits of technology

It improves the efficiency and accuracy of integrated search of heatstroke pathology data, avoids the problem of poor search efficiency caused by single keyword selection, and realizes more intelligent information collection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120413073B_ABST
    Figure CN120413073B_ABST
Patent Text Reader

Abstract

The present application relates to the field of data analysis, especially to a heat stroke pathological data integration information collection processing method based on deep learning, comprising: determining the data state of user analysis data according to the data richness and image transformation completion degree of the user analysis data; when the keyword selection mode is the association group selection generation, determining the number of selected keywords according to the group influence coefficient, and selecting the keywords according to the order from large to small of the sensitive coefficient; when the keyword selection mode is the image transformation supplement generation, determining the keyword generation mode according to the feature representative degree of the divided paragraph, which is to generate keywords according to the median point or outlier point; under the initial search completion condition, optimization analysis is carried out, and the search optimization mode is determined to be combination search or sequential search according to the repeated keyword analysis state, which improves the pathological data integration search efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of data analysis, in particular to a heat stroke pathological data integration information collection processing method based on deep learning. BACKGROUND

[0002] Heat stroke is a life-threatening disease caused by high temperature environment, which has a significant health risk. With the development of information technology and medical field, the sharing degree of heat stroke pathological information has been significantly improved. However, due to the dispersion of heat stroke pathological information and the lack of detailed classification, in the face of massive pathological data, the traditional manual search and processing method is low in efficiency and difficult to meet the actual demand. In recent years, deep learning technology has shown great potential in medical data processing field, but still faces many challenges, such as high data annotation cost, multi-dimensional search difficulty and other problems. Therefore, how to collect information from massive data through deep learning technology to improve the search efficiency of heat stroke pathological data integration is a technical problem to be solved by the skilled in the art.

[0003] Chinese patent publication No. CN118643128A discloses a pathological literature search and dialogue system based on large language model and RAG technology. The system includes a pathological literature collection and processing module for collecting, organizing and processing pathological literature data; a real-time interaction and dialogue module for receiving user input and processing through a large language model to extract user intent and semantic embedding; a multi-modal vector representation module for converting pathological literature into semantic vector representation; a multi-path recall module for retrieving the most relevant pathological literature to user query; a content sorting module including an importance-based reordering model that uses a unified metric dimension to reorder the pathological literature recalled by multiple parties to generate the optimal sorting result. As can be seen from the above technical solution, the search process cannot effectively screen and optimize the search method, resulting in poor search efficiency. SUMMARY

[0004] Therefore, the present application provides a heat stroke pathological data integration information collection processing method based on deep learning to overcome the problem of poor search efficiency caused by the inability to effectively screen and optimize the search method in the prior art.

[0005] To achieve the above purpose, the present application provides a heat stroke pathological data integration information collection processing method based on deep learning, comprising:

[0006] Obtain user analysis data, and determine the data state of the user analysis data according to the data richness and image conversion completion degree of the user analysis data;

[0007] Determine the keyword selection method according to the data state, wherein the keyword selection method is association word group selection generation or image conversion supplement generation;

[0008] In the association phrase selection generation, the number of selected keywords is determined according to the phrase influence coefficient, and the keywords are selected according to the order from large to small of the sensitivity coefficient;

[0009] In the image conversion supplement generation, the segmentation manner is determined according to the image disorder degree and the distribution tendency degree of the analyzed image in the user analysis data to generate segmented paragraphs, and the keyword generation manner is determined according to the feature representative degree of the segmented paragraphs, that is, the keywords are generated according to the median point or the outlier point;

[0010] Under the initial search completion condition, optimization analysis is performed, the repetition keyword density is detected, and the repetition keyword analysis state is determined using the reference deviation value, and the search optimization manner is determined according to the repetition keyword analysis state, that is, the combination search or the sequential search.

[0011] Further, the keyword selection manner is determined according to the data state of the user analysis data, including:

[0012] When the data richness is greater than the preset data richness, the keyword selection manner is the association phrase selection generation;

[0013] When the data richness is less than or equal to the preset data richness and the image conversion completion degree is greater than the preset image conversion completion degree, the keyword selection manner is the image conversion supplement generation.

[0014] Further, the confirmation manner of the data richness includes:

[0015] The user analysis data is analyzed to obtain a text representation coefficient;

[0016] If the text representation coefficient is less than or equal to the preset text representation coefficient, the data richness is determined according to the search frequency reference value and the keyword effective search reference value;

[0017] If the text representation coefficient is greater than the preset text representation coefficient, the data richness is determined according to the keyword number reference value and the keyword distribution reference value.

[0018] Further, the confirmation manner of the image conversion completion degree includes:

[0019] The preset image associated text length is determined according to the complexity reference value ratio of the image complexity reference value and the preset image complexity reference value;

[0020] The image conversion completion degree is determined according to the image associated text length difference value;

[0021] The complexity reference value ratio and the preset image associated text length are in a positive correlation relationship.

[0022] Further, the association phrase selection generation comprises:

[0023] For a single association phrase,

[0024] The number of selected keywords is determined according to the phrase influence coefficient, and the keywords are selected according to the order from large to small of the sensitivity coefficient;

[0025] The number of selected keywords in a single association phrase is positively correlated with the phrase influence coefficient corresponding to the association phrase;

[0026] The association phrase is determined according to the character interval reference value and the similarity threshold.

[0027] Further, the image conversion supplement generation comprises:

[0028] For a single analysis image in the user analysis data, the segmentation method is determined according to the image disorder degree and the distribution trend degree of the analysis image;

[0029] If the image disorder degree is greater than the preset image disorder degree or the distribution trend degree is less than or equal to the preset distribution trend degree, the segmentation method is association division;

[0030] If the image disorder degree is less than or equal to the preset image disorder degree and the distribution trend degree is greater than the preset distribution trend degree, the segmentation method is uniform division.

[0031] Further, for a single division paragraph, the keyword generation method is determined according to the feature representative degree of the division paragraph;

[0032] If the feature representative degree is greater than or equal to the preset feature representative degree, the keyword generation method is to generate keywords according to the median point;

[0033] If the feature representative degree is less than the preset feature representative degree, the keyword generation method is to generate keywords according to the outlier point.

[0034] Further, the optimization analysis comprises:

[0035] The use reference value corresponding to the repeated keyword is determined according to the influence range of the repeated keyword;

[0036] The repeated keyword analysis state is determined by detecting the repeated keyword density and the use reference deviation value;

[0037] The search optimization method is determined according to the repeated keyword analysis state.

[0038] Further, for a single repeated keyword, the confirmation method of the corresponding influence range comprises:

[0039] Extract the text segment corresponding to the repeated keyword, and perform sub-analysis on each text segment;

[0040] In the sub-analysis, the initial analysis range of the repeated keyword is determined according to the segmentation coefficient of the text segment, the effective data amount in the initial analysis range is detected, and the initial analysis range for extracting the repeated keyword is adjusted in size according to the difference between the effective data amount and the preset effective data amount;

[0041] The initial analysis range after the size adjustment is recorded as the influence range.

[0042] Further, the search optimization mode is determined according to the repeated keyword analysis state, including:

[0043] If the repeated keyword analysis state is that the repeated keyword density is greater than the preset repeated keyword density or the use reference deviation value is less than or equal to the preset use reference deviation value, the search optimization mode is combined search.

[0044] If the repeated keyword analysis state is that the repeated keyword density is less than or equal to the preset repeated keyword density and the use reference deviation value is greater than the preset use reference deviation value, the search optimization mode is sequential search.

[0045] Compared with the prior art, the beneficial effects of the present application are that in the technical scheme of the present application, the data state of the user analysis data is determined according to the data richness of the user analysis data and the image conversion completion degree, the keyword selection mode is selected according to the analysis of the user analysis data, and the keyword selection mode can accurately select the keywords required for data search.

[0046] Further, in the present application, the keyword selection mode generated by image conversion supplement is applied, the corresponding keywords are generated by recognizing and analyzing the image, the image information is converted into a supplementary basis for keywords, the problem of poor search efficiency caused by the single keyword selection mode in the prior art is avoided, and the pathological data integration search efficiency is improved.

[0047] Further, in the present application, the segmentation mode is determined according to the image disorder degree and the distribution trend degree of the analysis image in the user analysis data to generate the divided paragraph, so that the determination of the divided paragraph is more in line with the actual image state, and the problem of poor data efficiency caused by the single image segmentation mode is avoided.

[0048] Further, in the present application, the repeated keyword analysis state is determined by detecting the repeated keyword density and the use reference deviation value, the importance of the keyword in the current search result is reflected through the keyword analysis state, and different search optimization modes are selected correspondingly, and the pathological data integration search efficiency is improved. BRIEF DESCRIPTION OF DRAWINGS

[0049] Fig. 1 A schematic diagram of the information collection and processing method for heat stroke pathological data integration based on deep learning of the present application is shown in the figure;

[0050] Fig. 2 A flow chart of the keyword selection method according to the data state of the user analysis data of the present application is shown in the figure;

[0051] Fig. 3 A flow chart of the segmentation method according to the image disorder degree and distribution trend degree of the analyzed image of the present application is shown in the figure;

[0052] Fig. 4 A flow chart of the search optimization method according to the repeated keyword analysis state of the present application is shown in the figure. DETAILED DESCRIPTION

[0053] In order to make the objects and advantages of the present application clearer, the present application will be further described below with reference to the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0054] The preferred embodiments of the present application will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application and are not intended to limit the protection scope of the present application.

[0055] It should be noted that in the description of the present application, the terms "upper", "lower", "left", "right", "inner", "outer" and the like indicate the direction or positional relationship terms based on the direction or positional relationship shown in the drawings, which are only for the convenience of description and do not indicate or imply that the device or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation of the present application.

[0056] In addition, it should also be noted that in the description of the present application, unless otherwise explicitly specified and limited, the terms "mounting", "connecting", "connection" should be understood broadly, for example, it can be fixed connection, or detachable connection, or integral connection; it can be mechanical connection, or electrical connection; it can be direct connection, or indirect connection through intermediate medium, or internal communication of two elements. Those skilled in the art can understand the specific meaning of the above terms in the present application according to the specific circumstances.

[0057] Please refer to Figs. 1 to 4 The present application provides an information collection and processing method for heat stroke pathological data integration based on deep learning, which comprises:

[0058] Obtaining user analysis data, determining the data state of the user analysis data according to the data richness and image transformation completion degree of the user analysis data;

[0059] determining the keyword selection mode according to the data state, the keyword selection mode being associated phrase selection generation or image conversion supplementary generation;

[0060] In the associated phrase selection generation, the number of selected keywords is determined according to the phrase influence coefficient, and the keywords are selected in the order from large to small according to the sensitive coefficient;

[0061] In the image conversion supplementary generation, the segmentation mode is determined according to the image disorder degree and the distribution trend degree of the analysis image in the user analysis data to generate a divided paragraph, and the keyword generation mode is determined according to the feature representative degree of the divided paragraph, that is, the keyword is generated according to the median point or the outlier point;

[0062] Under the initial search completion condition, optimization analysis is performed, the repetition keyword density is detected, and the repetition keyword analysis state is determined by using a reference deviation value, and the search optimization mode is determined according to the repetition keyword analysis state, that is, combined search or sequential search.

[0063] The application scenario of the present application is to search for analysis data through deep learning, and the intelligent evaluation of data state, the automatic generation of keywords and the optimization analysis of search results are performed through deep learning technology, so that the information collection result is more efficient, accurate and intelligent. The analysis data includes detected text and a plurality of analysis images, the detected text is a text related to heat stroke generated by the user himself, and the analysis image includes a plurality of two-dimensional coordinate images, including but not limited to a blood cell analysis image taking red blood cell volume as the horizontal coordinate and red blood cell number as the vertical coordinate, and a body temperature curve image taking time as the horizontal coordinate and body temperature as the vertical coordinate, which is easily understood by those skilled in the art and will not be described here.

[0064] The user analysis data is a single analysis data currently being analyzed, the keywords of the present application are keywords related to heat stroke, and the keywords include but are not limited to nuclear condensation, villus damage, intestinal microbiota and organ damage. The method described in the present application also needs to apply a pathology information library, which is used to store a plurality of to-be-selected data, the to-be-selected data being analysis data uploaded by the user himself in the historical process, and the above are common technical means in the art, which will not be described here.

[0065] In the present application, the initial search is performed after determining the keyword selection mode according to the data state, the initial search including: recording the keywords selected by the keyword selection mode as search keywords, and recording the to-be-selected data containing any search keyword as initial data; the initial search completion condition is that the initial data is determined to be completed;

[0066] The application corresponds to a plurality of historical records, any one of which records the data richness, image conversion completion degree, keyword quantity reference value, keyword distribution reference value, variation degree and image complexity reference value in the historical process of at least one data search, and each historical record corresponds to a qualified mark, which records whether the data search process meets the user's demand. The qualified mark can be recorded manually. It can be understood that the user can determine whether the data search process meets the demand according to the self-set index, which can be but is not limited to the error search index, which will not be described here. The error search index is the number of times of error search for selecting data.

[0067] Specifically, the keyword selection mode is determined according to the data state of the user analysis data, including:

[0068] When the data state is that the data richness is greater than the preset data richness, the keyword selection mode is associated word group selection generation;

[0069] When the data state is that the data richness is less than or equal to the preset data richness and the image conversion completion degree is greater than the preset image conversion completion degree, the keyword selection mode is image conversion supplement generation.

[0070] The data state includes a first data state, a second data state and a third data state. The first data state is that the data richness is greater than the preset data richness. The second data state is that the data richness is less than or equal to the preset data richness and the image conversion completion degree is greater than the preset image conversion completion degree. The third data state is that the data richness is less than or equal to the preset data richness and the image conversion completion degree is less than or equal to the preset image conversion completion degree.

[0071] It should be noted that in the third data state, the keyword selection mode is associated word group selection generation, and the initial search amount is increased. The increase of the initial search amount and the comprehensive evaluation value are in a negative correlation relationship.

[0072] The comprehensive evaluation value is the sum of the data richness and the image conversion completion degree. The initial search amount is the number of selected data in the initial search.

[0073] The values of the preset data richness and the preset image conversion completion degree can be determined by the user according to the actual application scenario. The greater the value of the preset data richness and the smaller the value of the preset image conversion completion degree, the greater the demand of the user for image conversion supplement generation of keywords. A value of the preset data richness and the preset image conversion completion degree is provided. The historical records of the user using image conversion supplement generation of keywords are detected. The average value of the data richness corresponding to the historical records that meet the user's demand is recorded as the preset data richness. The average value of the image conversion completion degree corresponding to the historical records that meet the user's demand is recorded as the preset image conversion completion degree.

[0074] Specifically, the data richness confirmation method comprises the following steps:

[0075] analyzing the user analysis data to obtain a text representation coefficient;

[0076] if the text representation coefficient is less than or equal to a preset text representation coefficient, determining the data richness according to a search frequency reference value and a keyword effective search reference value;

[0077] if the text representation coefficient is greater than the preset text representation coefficient, determining the data richness according to a keyword quantity reference value and a keyword distribution reference value.

[0078] wherein the text representation coefficient = the sum of the number of characters corresponding to each keyword in the user analysis data / the total amount of characters in the user analysis data;

[0079] The value of the preset text representation coefficient can be determined by the user according to the actual application scenario. The greater the value of the preset text representation coefficient, the greater the demand of the user for determining the data richness according to the search frequency reference value and the keyword effective search reference value. A value of the preset text representation coefficient is provided, and the preset text representation coefficient is 70%.

[0080] if the text representation coefficient is less than or equal to the preset text representation coefficient, the data richness = the search frequency reference value + the keyword effective search reference value;

[0081] if the text representation coefficient is greater than the preset text representation coefficient, the data richness = (the keyword quantity reference value / the preset keyword quantity reference value) + (the keyword distribution reference value / the preset keyword distribution reference value).

[0082] The search frequency reference value is the average value of the sub-search frequencies corresponding to each keyword in the user analysis data. The sub-search frequency corresponding to a single keyword is the number of times of initial search using the keyword in the historical records that can meet the user's demand. The keyword effective search reference value is the average value of the effective coefficients corresponding to each keyword in the user analysis data.

[0083] The confirmation method of the effective coefficient is as follows: for a single keyword, the keyword is recorded as a target keyword, the historical records that can meet the user's demand and use the target keyword for initial search are recorded as reference historical records, the to-be-selected data obtained by using the target keyword for initial search in the reference historical records is recorded as first reference pathological data, and the to-be-selected data obtained by searching in the reference historical records in an optimized manner and being the same as the first reference pathological data is recorded as second reference pathological data. The effective coefficient corresponding to the target keyword = the total amount of the second reference pathological data / the total amount of the first reference pathological data.

[0084] The keyword quantity reference value is the total number of keywords in the user analysis data, the keyword distribution reference value is the average of the reference distances corresponding to each keyword in the user analysis data, the reference distance corresponding to a single keyword is the average of the character interval reference values of other keywords in the user analysis data, and the character interval reference value is the total number of characters between any two keywords;

[0085] The values of the preset keyword quantity reference value and the preset keyword distribution reference value can be determined by the user according to historical records. It can be understood that the smaller the values of the preset keyword quantity reference value and the preset keyword distribution reference value, the greater the data richness determination result, and the greater the user's demand for selecting and generating keywords using associated word groups. A value of the preset keyword quantity reference value and the preset keyword distribution reference value is provided, and historical records for determining data richness according to the keyword quantity reference value and the keyword distribution reference value are detected. The average of the keyword quantity reference values corresponding to the historical records that can meet the user's demand is recorded as the preset keyword quantity reference value, and the average of the keyword distribution reference values corresponding to the historical records that can meet the user's demand is recorded as the preset keyword distribution reference value.

[0086] Specifically, the confirmation method of the image conversion completion degree comprises:

[0087] Determine the preset image associated text length according to the complexity reference value ratio of the image complexity reference value and the preset image complexity reference value.

[0088] Determine the image conversion completion degree according to the image associated text length difference value.

[0089] The complexity reference value ratio and the preset image associated text length have a positive correlation.

[0090] The complexity reference value ratio = image complexity reference value / preset image complexity reference value.

[0091] The image conversion completion degree and the image associated text length difference value have a positive correlation, and the image associated text length difference value = image associated text length-preset image associated text length.

[0092] The image complexity reference value is the average of the complexity thresholds corresponding to each analysis image in the user analysis data, and the complexity threshold corresponding to a single analysis image = image disorder degree corresponding to the analysis image + distribution degree corresponding to the analysis image.

[0093] For a single analysis image, the analysis image is recorded as a target image, the abscissa interval corresponding to the target image is equally divided by n, n equal division points can be obtained, the value of n is in positive correlation with the total length of the abscissa, the abscissa interval is the range between the abscissa value corresponding to the first point in the target image and the abscissa value corresponding to the last point, the total length of the abscissa = the abscissa corresponding to the last point in the target image - the abscissa corresponding to the first point in the target image, and each equal division point and the first point and the last point in the target image are recorded as feature points;

[0094] The image disorder degree is the standard deviation of the ordinate values corresponding to the feature points in the target image;

[0095] The distribution trend degree = 1 - [abnormal point number / (n + 2)], the abnormal point number is the number of feature points with a change degree greater than a preset change degree, for a single feature point, the feature point is recorded as a target feature point, and the feature points adjacent to the target feature point are recorded as adjacent points, the change degree corresponding to the target feature point is the larger value in the difference threshold values corresponding to each adjacent point, and the difference threshold value corresponding to a single adjacent point is the absolute value of the difference between the ordinate value corresponding to the target feature point and the ordinate value corresponding to the adjacent point;

[0096] In the present application, each analysis image corresponds to a figure number, the figure number can be a number 1, a number a and a number 1(a), which is easily understood by those skilled in the art, and will not be described in detail;

[0097] The image associated text length is the average value of the associated length reference values corresponding to each analysis image in the user analysis data, for a single analysis image, the position of the figure number corresponding to the analysis image in the detected text corresponding to the user analysis data is detected, and the sum of the number of keywords in each sentence in which the figure number is located is recorded as the associated length reference value corresponding to the analysis image, the confirmation method of the sentence in which the figure number is located is that, for a single figure number in each figure number corresponding to a single analysis image, the figure number is recorded as a target figure number, the characters between the first period before the reading order of the target figure number and the first period after the reading order of the target figure number are recorded as the sentence in which the target figure number is located, and the reading order is the order from front to back when reading the text;

[0098] The values of the preset variation degree, the preset image complexity reference value and the preset image associated text length can be determined by the user according to the actual application scene. The greater the value of the preset variation degree is, the greater the demand of the user for determining the feature point as an abnormal point is. The average value of the variation degree corresponding to the historical record that can meet the demand of the user is recorded as the preset variation degree by detecting the historical record of determining the feature point as an abnormal point. The greater the values of the preset image complexity reference value and the preset image associated text length are, the smaller the determined image conversion completion degree is, that is, the greater the demand of the user for selecting and generating the keyword by using the associated word group is. A preset image complexity reference value and a preset image associated text length are provided. The average value of the image complexity reference value corresponding to the historical record that can meet the demand of the user is recorded as the preset image complexity reference value by detecting the historical record of the user for generating the keyword by using the image conversion supplement. The average value of the image associated text length corresponding to the historical record that can meet the demand of the user is recorded as the preset image associated text length.

[0099] Specifically, the associated word group selection generation comprises:

[0100] For a single associated word group,

[0101] The number of selected keywords is determined according to the word group influence coefficient, and the keywords are selected in the order from large to small according to the sensitive coefficient;

[0102] The number of selected keywords in a single associated word group and the word group influence coefficient corresponding to the associated word group are in a positive correlation relationship.

[0103] The associated word group is determined according to the character interval reference value and the similarity threshold.

[0104] The associated word group is determined according to the character interval reference value and the similarity threshold. In the order from front to back according to the reading order, each keyword in the user analysis data is subjected to association analysis. When a single keyword is subjected to association analysis, the keyword is recorded as a target keyword, the keywords other than the target keyword and not recorded in the associated word group are recorded as reference keywords, the set of reference keywords with a character interval reference value less than a preset character interval reference value and a similarity threshold greater than a preset similarity threshold is recorded as an associated word group, and the association analysis is continued for the keywords not recorded in the associated word group until all the keywords are recorded in the associated word group.

[0105] The confirmation method of the similarity threshold is that, for any two keywords, which are recorded as analysis keywords, the number of qualified data in which the two analysis keywords exist simultaneously is recorded as the similarity threshold. The qualified data is each to-be-selected data selected by the search optimization method in the historical record that can meet the demand of the user.

[0106] The values of the preset character interval reference value and the preset similarity threshold value can be determined by the user according to an actual application scenario. The greater the user's demand for improving the search accuracy is, the smaller the value of the preset character interval reference value is, and the greater the value of the preset similarity threshold value is. A value of a preset character interval reference value is 20 characters, and a value of a preset similarity threshold value is 50% of the number of qualified data.

[0107] For a single associated phrase, the associated phrase is recorded as a target associated phrase, the phrase influence coefficient corresponding to the target associated phrase = the total amount of key words in the target associated phrase + the radiation coefficient corresponding to the target associated phrase, each key word in the target associated phrase is recorded as a first key word, a key word other than the first key word is recorded as a second key word, the same second key word as the first key word and the first key word are recorded as analysis words, and the number of characters between the first analysis key word and the last analysis key word in the reading order from front to back is recorded as the radiation coefficient corresponding to the target associated phrase.

[0108] The confirmation method of the sensitive coefficient is that, for a single key word, the key word is recorded as a first target key word, the data to be selected in the historical record that meets the user's demand and is selected by the search optimization method is recorded as first data, and other data to be selected other than the first data is recorded as second data. The sensitive coefficient corresponding to the first target key word = (the number of times of appearance of the first target key word in the first data / the total amount of the first data) - (the number of times of appearance of the first target key word in the second data / the total amount of the second data).

[0109] Specifically, for a single analysis image in the user analysis data, a segmentation method is determined according to the image disorder degree and the distribution trend degree of the analysis image.

[0110] If the image disorder degree is greater than a preset image disorder degree or the distribution trend degree is less than or equal to a preset distribution trend degree, the segmentation method is associated division.

[0111] If the image disorder degree is less than or equal to a preset image disorder degree and the distribution trend degree is greater than a preset distribution trend degree, the segmentation method is uniform division.

[0112] The values of the preset image disorder degree and the preset distribution trend degree can be determined by the user according to an actual application scenario. The greater the value of the preset image disorder degree is, the smaller the value of the preset distribution trend degree is, and the greater the user's demand for uniform division is. A value of a preset image disorder degree is the average value of the image disorder degrees of the historical records that meet the user's demand, and a value of a preset distribution trend degree is 80%.

[0113] In the association division, for a single analysis image, the association paragraph analysis is performed on each feature paragraph in the order of the horizontal coordinate values of the analysis image from small to large. When performing the association paragraph analysis on a single feature paragraph, the feature paragraph is recorded as a target feature paragraph. The slope difference between each feature paragraph located on the right side of the target feature paragraph and the target feature paragraph is detected in sequence. If the slope difference between the feature paragraph and the target feature paragraph is less than a preset slope difference, the feature paragraph is recorded as an associated feature paragraph of the target feature paragraph. If the slope difference between the feature paragraph and the target feature paragraph is greater than or equal to the preset slope difference, the association paragraph analysis on the target feature paragraph is stopped, the target feature paragraph and the associated feature paragraph corresponding thereto are recorded as a division paragraph, and the association paragraph analysis on the feature paragraph not recorded in the division paragraph is continued.

[0114] For two feature paragraphs, the slope difference is equal to the absolute value of the difference between the paragraph slope of one feature paragraph and the paragraph slope of the other feature paragraph. For a single feature paragraph, the vertical coordinate value corresponding to the leftmost feature point of the feature paragraph is recorded as a first value, the vertical coordinate value corresponding to the rightmost feature point of the feature paragraph is recorded as a second value, the paragraph slope of the feature paragraph is (second value-first value) / feature paragraph length, and the feature paragraph length is the length of the horizontal coordinate corresponding to the feature paragraph. The user can determine the value of the preset slope difference according to the actual application scenario. The smaller the value of the preset slope difference, the greater the user's demand for improving the association degree of the feature paragraphs in the division paragraph. A value of the preset slope difference is provided. The historical records of the association division are detected. The average value of the slope difference corresponding to the historical records that can meet the user's demand is recorded as the preset slope difference.

[0115] In the uniform division, the analysis image is uniformly divided into several division paragraphs with the same horizontal coordinate length. The number of division paragraphs and the complexity threshold value have a positive correlation.

[0116] Specifically, for a single division paragraph, the keyword generation mode is determined according to the feature representative degree of the division paragraph.

[0117] If the feature representative degree is greater than or equal to a preset feature representative degree, the keyword generation mode is to generate a keyword according to a median point.

[0118] If the feature representative degree is less than the preset feature representative degree, the keyword generation mode is to generate a keyword according to an outlier.

[0119] The feature representative degree is determined by taking a single divided paragraph in a single analysis image as a target divided paragraph, and the feature representative degree = deviation index + correlation index, taking other divided paragraphs in the analysis image as reference divided paragraphs, the deviation index = average value of the ordinate values of the feature points of the target divided paragraph - average value of the ordinate values of the feature points in the analysis image, and the correlation index is the average value of the sub-correlation coefficients corresponding to the target divided paragraph and each reference divided paragraph, and the calculation formula of the sub-correlation coefficient r of any two divided paragraphs is:

[0120] The number of feature points corresponding to the two divided paragraphs is detected, if the number of feature points is different, the smaller value of the number of feature points corresponding to the two divided paragraphs is taken as n, and the first n feature points in the two divided paragraphs are taken as reference points in the order of decreasing horizontal coordinates; if the number of feature points is the same, all feature points in the two divided paragraphs are taken as reference points;

[0121]

[0122] Wherein, n is the number of reference points corresponding to a single divided paragraph; x i and y i are the ordinate values of the i-th reference point in the monitoring time corresponding to the two divided paragraphs, is the average value of the ordinate values of the reference points of the divided paragraph corresponding to x i , is the average value of the ordinate values of the reference points of the divided paragraph corresponding to y i , i = 1, 2, 3, …, n;

[0123] The value of the preset feature representative degree can be determined by the user according to the actual application scene, the smaller the value of the preset feature representative degree, the greater the user's demand for generating keywords according to the median point, and a value of the preset feature representative degree is provided, the historical records of the user generating keywords according to the median point are detected, and the average value of the feature representative degrees corresponding to the historical records meeting the user's demand is taken as the preset feature representative degree;

[0124] When generating keywords according to the median point, the horizontal coordinate value corresponding to the horizontal coordinate name of the median point of the feature paragraph and the vertical coordinate value corresponding to the vertical coordinate name of the median point are taken as the generated keywords;

[0125] The median point corresponding to a single divided paragraph is a point with a reference value as the horizontal coordinate, and the reference value = (the horizontal coordinate value of the leftmost point of the divided paragraph + the horizontal coordinate value of the rightmost point of the divided paragraph) / 2;

[0126] According to the outlying point, the keyword is generated, and the transverse coordinate value corresponding to the transverse coordinate name of the outlying point corresponding to the feature paragraph and the longitudinal coordinate value corresponding to the longitudinal coordinate name of the outlying point are taken as the generated keyword.

[0127] The outlying point is the feature point with the largest change degree in a single divided paragraph.

[0128] Specifically, the optimization analysis includes:

[0129] According to the influence range of the repeated keyword, the use reference value corresponding to the repeated keyword is determined;

[0130] Detecting the repeated keyword density and determining the repeated keyword analysis state by using the reference deviation value;

[0131] According to the repeated keyword analysis state, the search optimization mode is determined.

[0132] Wherein, the repeated keyword is a search keyword appearing in each initial data, and the use reference value corresponding to the repeated keyword has a positive correlation with the influence range of the repeated keyword;

[0133] The repeated keyword analysis state includes a first repeated keyword analysis state and a second repeated keyword analysis state, the first repeated keyword analysis state is that the repeated keyword density is greater than the preset repeated keyword density or the use reference deviation value is less than or equal to the preset use reference deviation value, and the second repeated keyword analysis state is that the repeated keyword density is less than or equal to the preset repeated keyword density and the use reference deviation value is greater than the preset use reference deviation value.

[0134] Specifically, for a single repeated keyword, the confirmation mode of the influence range corresponding to the repeated keyword includes:

[0135] Extracting the text segment corresponding to the repeated keyword, and performing sub-analysis on each text segment;

[0136] In the sub-analysis, the initial analysis range of the repeated keyword is determined according to the segmentation coefficient of the text segment, the effective data amount in the initial analysis range is detected, and the initial analysis range of the repeated keyword is adjusted by increasing according to the difference between the effective data amount and the preset effective data amount;

[0137] The initial analysis range after the increasing adjustment is recorded as the influence range.

[0138] Wherein, for a single repeated keyword, the repeated keyword is recorded as a target repeated keyword, each initial data obtained by the initial search is recorded as initial data, the target repeated keyword corresponds to a text segment in each initial data, and the target repeated keyword in a single initial data corresponds to a text segment between a first target repeated keyword and a last target repeated keyword appearing in the initial data in a reading order from front to back in the initial data.

[0139] The segmentation coefficient = 1-[(the number of target repeated keywords in a single text segment+2) / (the number of keywords in a single text segment+2)];

[0140] The initial analysis range of the repeated keyword and the segmentation coefficient of the text segment are in a positive correlation relationship, and the initial analysis range of the repeated keyword is the number of characters selected for analysis in a single text segment in a reading order from front to back;

[0141] The effective data amount is the number of keywords with a collocation use frequency greater than a preset collocation use frequency in a single text segment corresponding to the target repeated keyword, and the confirmation method of the collocation use frequency is that, for a single keyword in a text segment, the keyword is recorded as a segment keyword, and the number of first data in which the segment keyword and the target repeated keyword coexist is recorded as the collocation use frequency corresponding to the segment keyword.

[0142] The initial analysis range of the repeated keyword is increased according to the difference between the effective data amount and a preset effective data amount, and the increase value of the initial analysis range of the repeated keyword and the difference between the effective data amount and the preset effective data amount are in a positive correlation relationship; the difference between the effective data amount and the preset effective data amount = the effective data amount-the preset effective data amount.

[0143] The values of the preset collocation use frequency and the preset effective data amount can be determined by a user according to an actual application scenario, the greater the demand of the user for improving the search efficiency, the greater the values of the preset collocation use frequency and the preset effective data amount, and a value of the preset collocation use frequency and the preset effective data amount is provided, the preset collocation use frequency is 50% of the total amount of first data, and the preset effective data amount is 60% of the total amount of keywords in a single text segment.

[0144] Specifically, the search optimization method is determined according to the repeated keyword analysis state, including:

[0145] If the repeated keyword analysis state is that the repeated keyword density is greater than a preset repeated keyword density or the use reference deviation value is less than or equal to a preset use reference deviation value, the search optimization method is a combined search.

[0146] If the repeated keyword analysis state is that the repeated keyword density is less than or equal to a preset repeated keyword density and the use reference deviation value is greater than a preset use reference deviation value, the search optimization mode is sequential search.

[0147] The repeated keyword analysis state includes a first repeated keyword analysis state and a second repeated keyword analysis state. The first repeated keyword analysis state is that the repeated keyword density is greater than a preset repeated keyword density or the use reference deviation value is less than or equal to a preset use reference deviation value. The second repeated keyword analysis state is that the repeated keyword density is less than or equal to a preset repeated keyword density and the use reference deviation value is greater than a preset use reference deviation value.

[0148] The repeated keyword density is determined by detecting the position of each repeated keyword in the user analysis data. The repeated keyword density is the area of the smallest rectangle that can contain each repeated keyword in the user analysis data divided by the area of the smallest rectangle that can contain each search keyword in the user analysis data. The use reference deviation value is the standard deviation of the number of searches corresponding to each initial data. The number of searches is the number of search keywords in a single initial data.

[0149] The preset repeated keyword density and the preset use reference deviation value can be determined by the user according to the actual application scenario. The smaller the preset repeated keyword density value is, the larger the preset use reference deviation value is, and the greater the user's demand for combined search is. A preset repeated keyword density and a preset use reference deviation value are provided. The average value of the repeated keyword density corresponding to the historical records that can meet the user's demand is recorded as the preset repeated keyword density. The average value of the use reference deviation value corresponding to the historical records that can meet the user's demand is recorded as the preset use reference deviation value.

[0150] When the search optimization mode is combined search, each search keyword is analyzed. When a single search keyword is analyzed, the search keyword is recorded as a target search keyword. Each search keyword other than the target search keyword is recorded as a reference search keyword. A combination of the target search keyword and the reference search keyword with a combination coefficient greater than a preset combination coefficient is recorded as a keyword combination. The search keyword analysis is continued until each search keyword is recorded in a keyword combination. The search is performed on each keyword combination in descending order of the combination influence coefficient. When a single keyword combination is searched, the data to be selected that contains each search keyword corresponding to the keyword combination is used as search data.

[0151] When the search optimization mode is sequential search, the search is sequentially performed on each search keyword in descending order of the emergence coefficient of the search keyword; when the search is performed on a single search keyword, the to-be-selected data in which the search keyword exists is taken as the search data;

[0152] The confirmation method of the collocation coefficient is that, for any two search keywords, the collocation coefficient = the number of initial data in which the two search keywords exist / the total amount of initial data; the value of the preset collocation coefficient can be determined by the user according to the actual application scenario; the smaller the value of the preset collocation coefficient is, the greater the user's demand for improving the search range is; a value of the preset collocation coefficient is provided, and the preset collocation coefficient is 40%;

[0153] The confirmation method of the combination influence coefficient is that, for a single keyword combination, the search keyword corresponding to the keyword combination is taken as a to-be-analyzed search keyword, and the total amount of the to-be-analyzed search keyword existing in the user analysis data is taken as the combination influence coefficient.

[0154] For a single search keyword, the emergence coefficient is the total amount of the search keyword existing in the user analysis data.

[0155] So far, the technical solutions of the present application have been described in combination with the preferred embodiments shown in the drawings, but those skilled in the art can easily understand that the protection scope of the present application is obviously not limited to these specific embodiments. Those skilled in the art can make equivalent changes or replacements to the related technical features without departing from the principles of the present application, and the technical solutions after the changes or replacements will fall within the protection scope of the present application.

[0156] The above only describes the preferred embodiments of the present application and is not used to limit the present application; those skilled in the art can make various changes and modifications to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method for information collection and processing of heatstroke pathological data integration based on deep learning, characterized in that, include: Acquire user analysis data and determine the data status of the user analysis data based on the data richness and image conversion completion rate; The keyword selection method is determined based on the data status, and the keyword selection method is either generated by selecting related word groups or generated by image conversion and supplementation. In the process of generating related keyword selection, the number of keywords to be selected is determined based on the keyword influence coefficient, and the keywords are selected in descending order of sensitivity coefficient; In the image conversion and supplementation process, the segmentation method is determined based on the image disorder and distribution trend of the analyzed images in the user analysis data to generate segmented paragraphs, and the keyword generation method is determined based on the feature representativeness of the segmented paragraphs, which is to generate keywords based on median points or outliers. Under the condition that the initial search is completed, optimization analysis is performed to detect the density of duplicate keywords and use reference deviation values ​​to determine the status of duplicate keyword analysis. Based on the status of duplicate keyword analysis, the search optimization method is determined to be either combined search or sequential search. The keyword selection method is determined based on the data status of user analysis data, including: When the data abundance is greater than the preset data abundance, the keyword selection method is generated by selecting related phrases; When the data status is that the data abundance is less than or equal to the preset data abundance and the image conversion completion is greater than the preset image conversion completion, the keyword selection method is image conversion supplement generation; For a single segment, the keyword generation method is determined based on the representativeness of the segment's features. If the feature representativeness is greater than or equal to the preset feature representativeness, the keyword generation method is to generate keywords based on the median point. If the feature representativeness is less than the preset feature representativeness, the keyword generation method is to generate keywords based on outliers; Optimization analysis includes: Determine the reference values ​​for the use of repeated keywords based on the scope of their influence. Detect the density of duplicate keywords and use reference deviation values ​​to determine the status of duplicate keyword analysis; Determine search optimization methods based on the status of duplicate keyword analysis; If the duplicate keyword analysis status shows that the duplicate keyword density is greater than the preset duplicate keyword density or the reference deviation value is less than or equal to the preset reference deviation value, then the search optimization method is combined search. If the duplicate keyword analysis status shows that the duplicate keyword density is less than or equal to the preset duplicate keyword density and the reference deviation value is greater than the preset reference deviation value, then the search optimization method is sequential search. Methods for confirming data sufficiency include: Analyze user analytics data to obtain text representation coefficients; If the text representation coefficient is less than or equal to the preset text representation coefficient, the data richness is determined based on the search frequency reference value and the effective search reference value of the keywords. If the text representation coefficient is greater than the preset text representation coefficient, the data sufficiency is determined based on the reference values ​​for the number of keywords and the keyword distribution. Methods for confirming the completion of image conversion include: The length of the text associated with the preset image is determined based on the ratio of the image's complex reference value to the preset image's complex reference value. The image conversion completion rate is determined based on the difference in length between the image and the associated text. The ratio of the complex reference value is positively correlated with the length of the preset image-associated text. The image conversion completion rate is positively correlated with the difference in image-related text length. The difference in image-related text length = image-related text length - preset image-related text length. Duplicate keywords are search keywords that appear in all initial data. The usage reference value corresponding to the duplicate keyword is positively correlated with the influence range of the duplicate keyword. The reference deviation is used as the standard deviation of the number of searches corresponding to each initial data point, where the number of searches is the number of search keywords present in a single initial data point. The search frequency reference value is the average of the sub-search frequencies corresponding to each keyword in the user analysis data; the sub-search frequency corresponding to a single keyword is the number of times that keyword was used for the first search in the historical records that can meet the user's needs; The effective search reference value for keywords is the average of the effectiveness coefficients corresponding to each keyword in the user analysis data; The reference value for the number of keywords is the total number of keywords in the user analysis data; The keyword distribution reference value is the average of the reference distances corresponding to each keyword in the user analysis data. The reference distance for a single keyword is the average of the character interval reference values ​​between that keyword and other keywords in the user analysis data. For any two keywords, the character interval reference value is the total number of characters between the two keywords. The image complexity reference value is the average of the complexity thresholds corresponding to each analysis image in the user analysis data. The complexity threshold corresponding to a single analysis image = the image disorder degree corresponding to that analysis image + the distribution trend degree corresponding to that analysis image.

2. The information collection and processing method for integrating pathological data of heatstroke based on deep learning according to claim 1, characterized in that, The generation of related phrase selection includes: For a single related phrase, The number of keywords to be selected is determined based on the phrase influence coefficient, and the keywords are selected in descending order of sensitivity coefficient; The number of keywords selected in a single related phrase is positively correlated with the influence coefficient of the corresponding phrase. The associated word groups are determined based on character spacing reference values ​​and similarity thresholds.

3. The information collection and processing method for integrating pathological data of heatstroke based on deep learning according to claim 2, characterized in that, For a single analysis image in the user analysis data, the segmentation method is determined based on the image disorder and distribution trend of the analysis image. If the image disorder degree is greater than the preset image disorder degree or the distribution trend degree is less than or equal to the preset distribution trend degree, the segmentation method is associative division. If the image disorder is less than or equal to the preset image disorder and the distribution trend is greater than the preset distribution trend, the segmentation method is uniform division.

4. The information collection and processing method for integrating heatstroke pathological data based on deep learning according to claim 1, characterized in that, Methods for confirming the scope of influence of a single repeating keyword include: Extract the text segments corresponding to repeated keywords, and perform sub-analysis on each text segment; In the sub-analysis, the initial analysis range of repeated keywords is determined based on the segmentation coefficient of the text segment, the effective data volume within the initial analysis range is detected, and the initial analysis range for extracting repeated keywords is increased based on the difference between the effective data volume and the preset effective data volume. The initial analysis range after the adjustment is completed is denoted as the influence range.

Citation Information

Patent Citations

  • Pathological literature searching and dialogue system based on large language model and RAG technology

    CN118643128A

  • Comprehensive information analysis method and system based on keyword extraction

    CN117993392A

  • High-order medical statistical data processing method and system based on intelligent medical engineering

    CN118571498A