Disease Content Push Method, Device, Equipment and Medium Based on Intelligent Association

Through the disease content push method based on intelligent association, the forward maximum matching word segmentation and inverted indexing technology are used to solve the problem that the existing search system cannot search related but does not contain keywords, and the accurate and fast disease content push is achieved to meet the personalized needs of users.

CN114860887BActive Publication Date: 2025-07-29KANG JIAN INFORMATION TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210589107.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-26
Publication Date
2025-07-29
Estimated Expiration
2042-05-26

AI Technical Summary

Technical Problem

Most existing search systems can only search for content containing keywords, but cannot search for content related to keywords but not keywords. When there are too few keywords to search, more content cannot be retrieved, and search content cannot be pushed to users according to the person.

Method used

The disease content push method based on intelligent association is adopted, and the forward maximum matching word segmentation algorithm is used to process word segmentation on the searched statements entered by the user. Through the inverted indexing technology in the mapping document, the disease content corresponding to the search terms is found and saved, and the results are pushed by sorting according to the preset weights.

Benefits of technology

It realizes the push of disease content that varies from person to person, and users can retrieve the results that match their own best, improves the search efficiency and accuracy, avoids the problem of a single search, and realizes efficient and fast real-time retrieval of massive data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114860887B_ABST
    Figure CN114860887B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and provides a disease content push method, device, equipment and medium based on intelligent association. The method includes: in response to a user's retrieval request, obtaining a to-be-retrieved statement input by the user; using a forward maximum matching word segmentation algorithm to perform word segmentation processing on the to-be-retrieved statement, obtaining a plurality of retrieval words included in the to-be-retrieved statement, and each retrieval word is composed of a plurality of consecutive different Chinese characters in the to-be-retrieved statement; searching and saving disease content corresponding to each retrieval word in a pre-stored mapping document; after sorting the disease content according to a preset weight, pushing the disease content to the user in ascending order of weight. The present invention matches the user's tags with the tags of the retrieved content, realizes personalized push content, and accurately pushes query results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method, device, equipment and medium for pushing disease content based on intelligent association. Background Art

[0002] At present, most products of many companies on the market basically contain a search function, such as product search in the mall module, disease search and drug retrieval in the prescription system, patient disease retrieval, doctor search for consultation, etc. Users usually input multiple keywords, and the retrieval system pushes the corresponding retrieved content to the user by matching the retrieved content with the keywords.

[0003] The inventor realizes that most of the above-mentioned retrieval systems can only search for content containing keywords, but cannot search for content related to the keywords but not containing the keywords. When the number of retrieval keywords is too small, more content cannot be retrieved, and the problem that the retrieved content cannot be pushed to users according to individual differences. Summary of the Invention

[0004] The purpose of the present invention is to provide a method, device, equipment and medium for pushing disease content based on intelligent association, so as to solve the technical problem that most of the existing retrieval systems can only search for content containing keywords, but cannot search for content related to the keywords but not containing the keywords. When the number of retrieval keywords is too small, more content cannot be retrieved, and the retrieved content cannot be pushed to users according to individual differences.

[0005] In the first aspect, a method for pushing disease content based on intelligent association is provided, including:

[0006] Responding to a retrieval request of a user, obtaining a to-be-retrieved statement input by the user;

[0007] Using the forward maximum matching word segmentation algorithm to perform word segmentation processing on the to-be-retrieved statement, obtaining a plurality of retrieval words included in the to-be-retrieved statement, and each retrieval word is composed of a plurality of consecutive different Chinese characters in the to-be-retrieved statement;

[0008] Searching and saving disease content corresponding to each retrieval word in a pre-stored mapping document, where each retrieval word corresponds to one or more disease contents, and the disease content represents information associated with the retrieval word;

[0009] After sorting the disease contents according to a preset weight, pushing the disease contents to the user in the order of decreasing weight.

[0010] In the second aspect, a device for pushing disease content based on intelligent association is provided, including:

[0011] A retrieval statement acquisition module, configured to respond to a retrieval request of a user and obtain a to-be-retrieved statement input by the user;

[0012] A search term acquisition module, which is used to perform word segmentation processing on the to-be-searched statement by using the forward maximum matching word segmentation algorithm, to obtain a plurality of search terms included in the to-be-searched statement, and each search term is composed of a plurality of consecutive different Chinese characters in the to-be-searched statement;

[0013] A disease content retrieval module, which is used to search for and save disease content corresponding to each search term in a pre-stored mapping document, and each search term corresponds to one or more disease contents, and the disease content represents information associated with the search term;

[0014] A disease content push module, which is used to sort each disease content according to a preset weight, and then push each disease content to the user in the order of decreasing weight.

[0015] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned intelligent question-answering processing method are implemented.

[0016] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned intelligent question-answering processing method are implemented.

[0017] The disease content push method, system, device and medium based on intelligent association of the present invention obtain the to-be-searched statement input by the user by responding to the user's search request. Use the forward maximum matching word segmentation algorithm to perform word segmentation processing on the to-be-searched statement to obtain a plurality of search terms included in the to-be-searched statement, and each search term is composed of a plurality of consecutive different Chinese characters in the to-be-searched statement. Then search for and save the disease content corresponding to each search term in a pre-stored mapping document, and each search term corresponds to one or more disease contents, and the disease content represents information associated with the search term. After sorting each disease content according to a preset weight, each disease content is pushed to the user in the order of decreasing weight, realizing personalized disease content push. Avoiding the problem of single search, realizing personalized push content for thousands of people, accurately pushing query results, and enabling users to retrieve the results most matching themselves in the first time. When retrieving, a small number of keywords input by the user can be retrieved, or more content can be retrieved. Information without keywords is also retrieved, so as to ensure that if the user only remembers a small number of keywords, satisfactory results can also be retrieved, realizing efficient, fast and real-time retrieval of massive data. Description of the Drawings

[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings:

[0019] Figure 1 It shows a schematic diagram of an application environment of the disease content push method based on intelligent association in an embodiment of the present invention;

[0020] Figure 2 It shows a schematic flowchart of the disease content push method based on intelligent association in an embodiment of the present invention;

[0021] Figure 3 It shows a schematic flowchart of step S20 in an embodiment of the present invention;

[0022] Figure 4 It shows a schematic flowchart of step S30 in an embodiment of the present invention;

[0023] Figure 5 It shows a schematic flowchart of step S32 in an embodiment of the present invention;

[0024] Figure 6 It shows a structural block diagram of the disease content push device based on intelligent association in an embodiment of the present invention;

[0025] Figure 7 It is a schematic structural diagram of a computer device in an embodiment of the present invention;

[0026] Figure 8 It is another schematic structural diagram of a computer device in an embodiment of the present invention. Detailed implementation manners

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0028] The disease content push method based on intelligent association provided by the embodiments of the present invention can be applied in, for example Figure 1In the application environment, the client communicates with the server through the network. The server can push corresponding disease content based on the retrieval request input by the client. By responding to the user's retrieval request, the to-be-retrieved statement input by the user is obtained. Using the forward maximum matching word segmentation algorithm, the to-be-retrieved statement is segmented to obtain multiple retrieval words included in the to-be-retrieved statement, and each retrieval word is composed of multiple consecutive different Chinese characters in the to-be-retrieved statement. Then, in the pre-stored mapping document, the disease content corresponding to each retrieval word is searched for and saved. Each retrieval word corresponds to one or more disease contents, and the disease content represents the information associated with the retrieval word. After sorting the disease contents according to a preset weight, the disease contents are pushed to the user in the order of decreasing weight, realizing personalized disease content push. In the present invention, for the current disease content push method for diseases, it is usually a single retrieval method, resulting in the user being able to only search for the content containing the keyword, and unable to search for the content related to the keyword but not containing the keyword. And the retrieval speed is slow. When the data reaches a certain volume, the retrieval speed drops, and the front-end loading is delayed, affecting the user experience. To solve the above problems, by matching the user tags with the tags of the retrieval content pre-stored in the mapping document maintained in the system, the inverted index method is used to retrieve the relevant disease content, and according to the different weights of the retrieval content, the retrieval content with a higher weight is preferentially pushed. Thus, the problem of single search is avoided, the push content of "thousands of people, thousands of faces" is realized, the query result is accurately pushed, and the user can retrieve the result most matching himself at the first time. When retrieving, a small number of keywords input by the user can be retrieved, or more content can be retrieved. The information not containing the keyword is also retrieved, so as to ensure that if the user only remembers a small number of keywords, a satisfactory result can also be retrieved, realizing efficient, fast and real-time retrieval of massive data. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.

[0029] Please refer to Figure 2 as shown Figure 2 FIG. is a schematic flowchart of a method for pushing disease content based on intelligent association provided by an embodiment of the present invention, including the following steps:

[0030] S10. In response to the user's retrieval request, obtain the to-be-retrieved statement input by the user.

[0031] In this embodiment, when a user wants to retrieve the content they want in a retrieval software, they can enter the statement to be retrieved in the search box of the retrieval software, such as "What is diabetes?", "How to prevent diabetes?", etc. It can be understood that the retrieval software in this embodiment can be any software that can meet the user's retrieval needs, including but not limited to various browsers, commodity search software, or Ping An Health APP, etc. The retrieval software can receive the statement to be retrieved input by the user in real time, and after analyzing and processing the statement to be retrieved, push the retrieved content to the user. Those skilled in the art can adaptively select the corresponding retrieval software according to actual retrieval needs.

[0032] In this embodiment, after step S1, it further includes: obtaining the user's retrieval account, and obtaining the user label of the user according to the retrieval account, where the user label is doctor, patient, or tourist. Since the permissions of doctors, patients, and tourists are different when logging in to the browsing system, the corresponding retrieval accounts are also different. The account information of the user can be obtained by means of data tracking. When the user logs in to the browsing system, the identity of the user can be distinguished by the account. Among them, users with the user label of tourist can directly view disease content in the browsing system without registration. After a user with the user label of doctor logs in to the browsing system, they can query various drugs for treating the same disease according to their permissions, so as to better prescribe suitable drugs for patients. After a user with the user label of patient logs in to the browsing system, the system will push the common treatment methods and corresponding departments for this disease to them, so as to facilitate the user to find the corresponding department in time and make an early diagnosis.

[0033] S20. Use the forward maximum matching word segmentation algorithm to perform word segmentation processing on the statement to be retrieved, and obtain multiple retrieval words included in the statement to be retrieved, where each retrieval word is composed of multiple consecutive different Chinese characters in the statement to be retrieved.

[0034] In step S20, the process of using the forward maximum matching word segmentation algorithm to perform word segmentation processing on the statement to be retrieved and obtaining multiple retrieval words included in the statement to be retrieved includes the following steps:

[0035] S21. Take the first Chinese character of the statement to be retrieved as the starting point;

[0036] S22. Retrieve the longest phrase that matches the statement to be retrieved from the pre-stored dictionary, and save it as the current retrieval word;

[0037] S23. Cut out the current retrieval word from the statement to be retrieved to obtain the segmented statement to be retrieved, and repeat steps S21 to S22 to obtain retrieval words until the statement to be retrieved is completely segmented.

[0038] In this embodiment, the forward maximum matching word segmentation algorithm is used to segment the search sentence, and the word segmentation can be performed by the elasticsearch-analysis-ik word segmenter, and the words in the dictionary pre-stored in the word segmenter can be updated at any time. Elasticsearch is an open source search server based on Lucene, which provides a distributed multi-user full-text search engine that can realize the real-time retrieval and analysis of massive data in this application. It uses existing database components or search components, and can easily integrate the retrieval method provided in this embodiment with the existing database system and web page publishing system to improve the stability of the system. Specifically, the first Chinese character of the search sentence can be used as the starting point, and several consecutive Chinese characters in the search sentence can be matched with the pre-stored phrases in the dictionary from left to right. If they match, the above-mentioned Chinese characters are used as the current search terms and are segmented out from the search sentence to obtain the segmented search sentence. For example, if the search phrase is "how to prevent diabetes," starting with the character "如" in the search phrase, and matching each character against multiple pre-stored phrases in the dictionary, multiple phrases beginning with "如" are retrieved, such as "how to prevent," "if you have diabetes," "how to treat," and "how to diagnose." The search phrase "何" in the search phrase is then matched against these multiple phrases one by one, yielding several phrases beginning with "如何," such as "how to prevent," "how to treat," and "how to diagnose." These phrases are then searched for words whose third character is "预." Repeating the matching process, the dictionary query "如何遵守" (how to prevent) matches the search term consisting of the first four characters in the search phrase. Therefore, "如何遵守" can be segmented from the search phrase, yielding the segmented search phrase "粘度." The segmented search phrase is then searched using the same search method, ultimately yielding the two search terms "how to prevent" and "粘度." Since the Chinese word segmentation function of the elasticsearch-analysis-ik word segmenter is relatively mature, by using the built-in positive maximum matching word segmentation algorithm in elasticsearch-analysis-ik, it is convenient to accurately segment the Chinese content of the search statement, with high word segmentation efficiency.

[0039] In this embodiment, before performing word segmentation on the to-be-retrieved statement, it further includes: cleaning the to-be-retrieved statement, and deleting punctuation marks, special symbols, and spaces in the to-be-retrieved statement. Exemplarily, if the to-be-retrieved statement is "How to treat diabetes?", first, necessary data cleaning is performed on this retrieval statement to remove the question mark and space therein, obtaining "How to treat diabetes". Among them, the punctuation marks in this embodiment can be various symbols that assist in recording languages such as commas, periods, colons, question marks, etc., and the special symbols can be tabulation characters, unit symbols, arrow symbols, etc. By cleaning the to-be-retrieved statement, various non-Chinese characters can be removed from the to-be-retrieved statement, thereby greatly improving the retrieval efficiency so as to quickly push disease content to users.

[0040] S30. Search for and save disease content corresponding to each retrieval term in a pre-stored mapping document. Each retrieval term corresponds to one or more disease contents, and the disease content represents information associated with the retrieval term.

[0041] In step S30, searching for and saving disease content corresponding to each retrieval term in the pre-stored mapping document includes the following process:

[0042] S31. Search the mapping document, match one of the retrieval terms with each first-level category pre-stored in the mapping document to obtain the first-level category corresponding to the current retrieval term. Among them, multiple first-level categories are pre-stored in the mapping document, each first-level category contains multiple second-level categories, the first-level category represents the type of disease, the second-level category represents the treatment drugs, preventive measures, and specific disease types corresponding to the current first-level category, each second-level category pre-stores a second-level category label and an address, the second-level category label is doctor, patient, or tourist, and the address represents the storage location of the disease content corresponding to the second-level category;

[0043] S32. Adopt the inverted index method to search for the disease content corresponding to the current retrieval term in the multiple second-level categories included in the first-level category;

[0044] S33. Select another retrieval term, and repeat steps S31 to S32 until all retrieval terms are searched, obtaining the disease content corresponding to each retrieval term.

[0045] Further, step S32 includes the following process:

[0046] S321. Match the user label of the user with the second-level category labels corresponding to each second-level category to obtain the second-level category whose second-level category label is consistent with the user label as the to-be-retrieved second-level category;

[0047] S322. According to the address pre-stored in the to-be-retrieved second-level category, search for the disease content corresponding to the current retrieval term.

[0048] In this embodiment, after obtaining multiple search terms by means of word segmentation, the inverted index method is used to search for the disease content corresponding to each search term in the mapping document. The inverted index is used to store the mapping of the storage location of a phrase in a document or a group of documents under full-text search. It is the most commonly used data structure in document retrieval systems. Through the inverted index, the document list containing this phrase can be quickly obtained according to the phrase, and the corresponding document information can be obtained. Specifically, by means of keyword matching, each search term is matched with each first-level category in the mapping document one by one to obtain the first-level category that matches the search term. The first-level categories are the names of various diseases, such as diabetes, heart disease, etc. Exemplarily, as shown in Table 1, Table 1 schematically lists the corresponding relationships between the first-level categories and the second-level categories in the dictionary. When the search term is diabetes, the first-level category that matches it is diabetes. When the second-level category is a specific disease type, it can be type 1 diabetes, type 2 diabetes, and gestational diabetes. When the second-level category is a therapeutic drug, it can be metformin, pioglitazone rosiglitazone, and voglibose. When the second-level category is a preventive measure, it can be weight control, persistent exercise, and regular blood glucose detection. Among them, type 1 diabetes (doc2) means that the relevant disease content of type 1 diabetes is stored in document 2. By matching the user tags with the second-level category tags, the second-level categories that meet the user tags are obtained, and the content in the address corresponding to the second-level category is used as the disease content of the search term.

[0049] Table 1

[0050]

[0051] As an example, when the user's tag is doctor and the search statement to be retrieved entered by the user is "diabetes", the relevant disease content of type 2 diabetes, the therapeutic drug being pioglitazone rosiglitazone, and the preventive measure being persistent exercise can be retrieved. For example, the introduction of type 2 diabetes stored in document 3, the dosage, usage, and side effects of pioglitazone rosiglitazone pre-stored in document 6, and the duration of exercise stored in document 9. When the user's tag is patient and the same search statement is entered, the relevant information of type 1 diabetes, the therapeutic drug being metformin, and the preventive measure being weight control can be retrieved. It can be understood that the second-level categories can also include the departments for treating diseases, the basic manifestations of diseases, etc. Those skilled in the art can adaptively change the second-level categories according to actual needs. By matching the second-level category tags with the user tags, the query results are accurately and quickly pushed, allowing the user to retrieve the results that are most suitable for themselves in the shortest time. Using the inverted index method, even information that does not contain the search term can be retrieved, thus realizing personalized push for thousands of people, making the pushed disease content more targeted.

[0052] S40. After sorting the disease content according to a preset weight, push the disease content to the user in the order of decreasing weight.

[0053] In this embodiment, after obtaining the disease content of the user, since each disease content is preset with a weight, the disease content can be sorted in descending order according to the corresponding weight, and the disease content with a higher weight is preferentially displayed. For example, when the user is a tourist and searches for diabetes, the obtained disease content and the corresponding weights are: medicines for treating diabetes, with a weight of 0.7, and preventive measures for diabetes, with a weight of 0.85. Then the disease content seen by the user will first present the preventive measures for diabetes, and then present the content such as medicines for treating diabetes. It realizes personalized push content and accurately pushes the query results, allowing the user to retrieve the most matching results for themselves in the first time. In addition, this retrieval method uses the ElasticeSearch technology, and the retrieval speed for a large amount of data is relatively fast. For a data volume of tens of millions, when searching for diabetes, the time taken by the traditional method to present the retrieval content is more than 5 seconds, while when using this method to search for diabetes, the time taken to present the retrieval content is less than 1 second.

[0054] In this embodiment, the process of obtaining the mapping document is as follows:

[0055] S301. Obtain the keywords of the sentence to be retrieved, and use the distributed crawler technology to crawl the data corresponding to the keywords.

[0056] S302. Parse and process the crawled data to construct the mapping document.

[0057] Exemplarily, when the retrieval target is "treatment of heart disease", its keywords are "heart disease" and "treatment". Using the distributed crawler technology, according to the above keywords, crawl on various network channels such as news websites, Weibo, and WeChat to obtain a large amount of data related to the keywords. Optionally, the distributed crawler can choose the Acrap framework. After obtaining the data corresponding to the retrieval target, perform tagging processing according to the content of the data. For example, if the data is a medicine for treating heart disease, such as Suxiao Jiuxin Pills, since it is an over-the-counter medicine and a relatively common medicine for treating heart disease, the label of this medicine can be marked as a patient. Through this parsing and processing method, the mapping document is constructed.

[0058] It can be understood that the disease content push method based on intelligent association in this embodiment can be applied to a variety of different fields, such as: speech recognition, medical diagnosis, testing of application programs, etc.

[0059] It can be seen that in the above solution, corresponding disease content is pushed based on the retrieval request input by the user. By responding to the user's retrieval request, the to-be-retrieved statement input by the user is obtained. The forward maximum matching word segmentation algorithm is used to perform word segmentation on the to-be-retrieved statement, and a plurality of retrieval words included in the to-be-retrieved statement are obtained. Each retrieval word is composed of a continuous plurality of different Chinese characters in the to-be-retrieved statement. Then, the disease content corresponding to each retrieval word is searched for and saved in the pre-stored mapping document. Each retrieval word corresponds to one or more disease contents, and the disease content represents information associated with the retrieval word. After sorting the disease contents according to a preset weight, the disease contents are pushed to the user in the order of decreasing weight, realizing personalized disease content push. In the present invention, for the current way of pushing disease content for diseases, it is usually a single retrieval method, which causes the user to only be able to search for content containing keywords, and cannot search for content related to the keywords but not containing the keywords. And the retrieval speed is slow. When the data reaches a certain volume, the retrieval speed drops, and the front-end loading is delayed, affecting the user experience. To solve the above problems, by matching the user label with the label of the retrieval content pre-stored in the mapping document maintained in the system, the inverted index method is used to retrieve the relevant disease content, and according to the different weights of the retrieval content, the retrieval content with a higher weight is preferentially pushed. Thus, the problem of single search is avoided, the push content for thousands of people with thousands of faces is realized, the query result is accurately pushed, and the user can retrieve the result most matching himself in the first time. When retrieving, a small number of keywords input by the user can be retrieved, or more content can be retrieved. Information not containing keywords is also retrieved, so as to ensure that if the user only remembers a small number of keywords, a satisfactory result can also be retrieved, realizing efficient, fast and real-time retrieval of massive data.

[0060] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0061] In one embodiment, a disease content push device based on intelligent association is provided. The disease content push device based on intelligent association corresponds one-to-one with the disease content push method based on intelligent association in the above embodiment. As Figure 6 shown, the disease content push device based on intelligent association includes a retrieval statement acquisition module 111, a retrieval word acquisition module 112, a disease content retrieval module 113, and a disease content push module 114. The detailed description of each functional module is as follows:

[0062] The retrieval statement acquisition module 111 is configured to obtain the to-be-retrieved statement input by the user in response to the user's retrieval request. The user has a user label, and the user label is a doctor, a patient, or a tourist.

[0063] A retrieval term acquisition module 112, which is used to perform word segmentation on the to-be-retrieved statement by using the forward maximum matching word segmentation algorithm, so as to obtain a plurality of retrieval terms included in the to-be-retrieved statement, and each retrieval term is composed of a plurality of consecutive different Chinese characters in the to-be-retrieved statement;

[0064] A disease content retrieval module 113, which is used to search for and save disease content corresponding to each retrieval term in a pre-stored mapping document, and each retrieval term corresponds to one or more disease contents, and the disease content represents information associated with the retrieval term;

[0065] A disease content push module 114, which is used to sort each disease content according to a preset weight, and then push each disease content to the user in the order of decreasing weight.

[0066] In one embodiment, the retrieval statement acquisition module 111 is further used for:

[0067] Obtain the user's retrieval account, and obtain the user label of the user according to the retrieval account, and the user label is a doctor, a patient or a tourist.

[0068] In one embodiment, the retrieval term acquisition module 112 is specifically used for:

[0069] Taking the first Chinese character of the to-be-retrieved statement as the starting point;

[0070] Retrieve the phrase with the most characters that matches the to-be-retrieved statement from a pre-stored dictionary as the current retrieval term;

[0071] Cut out the current retrieval term from the to-be-retrieved statement to obtain the cut to-be-retrieved statement, and repeat the above steps to obtain retrieval terms until the to-be-retrieved statement is completely cut.

[0072] In one embodiment, the retrieval term acquisition module 112 is further used for:

[0073] Clean the to-be-retrieved statement, and delete punctuation marks, special symbols and spaces in the to-be-retrieved statement.

[0074] In one embodiment, the disease content retrieval module 113 is specifically used for:

[0075] Search for the mapping document, match one of the search terms with each first-level category pre-stored in the mapping document, and obtain the first-level category corresponding to the current search term. Among them, multiple first-level categories are pre-stored in the mapping document, each first-level category contains multiple second-level categories, the first-level category represents the type of disease, the second-level category represents the treatment drugs, preventive measures and specific disease types corresponding to the current first-level category, and each second-level category pre-stores a second-level category label and an address. The second-level category label is doctor, patient or tourist, and the address represents the storage location of the disease content corresponding to the second-level category;

[0076] Adopt the inverted index method to search for the disease content corresponding to the current search term among the multiple second-level categories included in the first-level category;

[0077] Select another search term and repeat the above steps until all search terms are searched, and obtain the disease content corresponding to each search term.

[0078] In one embodiment, the disease content retrieval module 113 is further configured to:

[0079] Match the user label of the user with the second-level category labels corresponding to each second-level category, and obtain the second-level category whose second-level category label is consistent with the user label as the second-level category to be retrieved;

[0080] According to the address pre-stored in the second-level category to be retrieved, search for the disease content corresponding to the current search term.

[0081] In one embodiment, the disease content retrieval module 113 is further configured to:

[0082] Obtain the keywords of the statement to be retrieved, and adopt the distributed crawler technology to crawl the data corresponding to the keywords;

[0083] Parse and process the crawled data to construct the mapping document.

[0084] The present invention provides a disease content push device based on intelligent association. By responding to a user's retrieval request, the to-be-retrieved statement input by the user is obtained. Using the forward maximum matching word segmentation algorithm, the to-be-retrieved statement is segmented to obtain multiple retrieval words included in the to-be-retrieved statement, and each retrieval word is composed of multiple consecutive different Chinese characters in the to-be-retrieved statement. Then, the disease content corresponding to each retrieval word is searched for and saved in a pre-stored mapping document. Each retrieval word corresponds to one or more disease contents, and the disease content represents information associated with the retrieval word. After sorting the disease contents according to a preset weight, the disease contents are pushed to the user in ascending order of weight, realizing personalized disease content push. In the present invention, for the current disease content push method for diseases, it is usually a single retrieval method, which causes the user to only be able to search for content containing keywords, and cannot search for content related to the keywords but not containing the keywords. And the retrieval speed is slow. When the data reaches a certain volume, the retrieval speed decreases, and the front-end loading is delayed, affecting the user experience. To solve the above problems, by matching the user tags with the tags of the retrieval content pre-stored in the mapping document maintained in the system, the inverted index method is used to retrieve the relevant disease content. According to the different weights of the retrieval content, the retrieval content with a higher weight is preferentially pushed. Thus, the problem of single search is avoided, and the push content for each person is realized, and the query result is accurately pushed, allowing the user to retrieve the result most matching themselves in the first time. During retrieval, a small number of keywords input by the user can be retrieved, or more content can be retrieved. Information not containing keywords is also retrieved, so as to ensure that if the user only remembers a small number of keywords, a satisfactory result can also be retrieved, realizing efficient, fast and real-time retrieval of massive data.

[0085] For the specific limitations of the disease content push device based on intelligent association, reference can be made to the limitations of the disease content push method based on intelligent association in the above text, which will not be elaborated here. Each module in the above disease content push device based on intelligent association can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor in the computer device in hardware form or independent of it, or stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0086] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 7As shown in the figure. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a disease content push method based on intelligent association.

[0087] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 8 shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a disease content push method based on intelligent association.

[0088] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0089] In response to a user's retrieval request, obtain the retrieval statement input by the user;

[0090] Use the forward maximum matching word segmentation algorithm to perform word segmentation on the retrieval statement, and obtain multiple retrieval words included in the retrieval statement. Each retrieval word is composed of multiple consecutive different Chinese characters in the retrieval statement;

[0091] Search for and save the disease content corresponding to each retrieval word in a pre-stored mapping document. Each retrieval word corresponds to one or more disease contents, and the disease content represents information associated with the retrieval word;

[0092] After sorting the disease contents according to a preset weight, push the disease contents to the user in descending order of weight.

[0093] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0094] In response to a retrieval request of a user, obtain the statement to be retrieved input by the user;

[0095] Use the forward maximum matching word segmentation algorithm to perform word segmentation processing on the statement to be retrieved, and obtain a plurality of retrieval words included in the statement to be retrieved. Each retrieval word is composed of a plurality of consecutive different Chinese characters in the statement to be retrieved;

[0096] Search for and save the disease content corresponding to each retrieval word in a pre-stored mapping document. Each retrieval word corresponds to one or more disease contents, and the disease content represents information associated with the retrieval word;

[0097] After sorting the disease contents according to a preset weight, push the disease contents to the user in ascending order of the weight.

[0098] It should be noted that for the functions or steps that the above computer-readable storage medium or computer device can implement, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0099] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. The non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. The volatile memory can include random access memory (RAM) or an external cache. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0100] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0101] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A method for pushing disease content based on intelligent association, characterized in that, Including: In response to a user's retrieval request, obtain the statement to be retrieved input by the user; Use the forward maximum matching word segmentation algorithm to perform word segmentation on the statement to be retrieved, and obtain multiple retrieval words included in the statement to be retrieved. Each retrieval word is composed of multiple consecutive different Chinese characters in the statement to be retrieved; Search for and save the disease content corresponding to each retrieval word in a pre-stored mapping document. Each retrieval word corresponds to one or more disease contents, and the disease content represents information associated with the retrieval word. Among them, multiple first-level categories are pre-stored in the mapping document, each first-level category contains multiple second-level categories. The first-level category represents the type of disease, and the second-level category represents the treatment drugs, preventive measures, and specific disease types corresponding to the current first-level category. Each second-level category pre-stores a second-level category label and an address. The second-level category label is doctor, patient, or tourist, and the address represents the storage location of the disease content corresponding to the second-level category; After sorting the disease contents according to a preset weight, push the disease contents to the user in descending order of weight; After obtaining the statement to be retrieved input by the user, it further includes: obtaining the user's retrieval account, and obtaining the user label of the user according to the retrieval account. The user label is doctor, patient, or tourist; Match the user label with the label of the retrieved content pre-stored in the mapping document, and use the inverted index to retrieve the disease content.

2. The method for pushing disease content based on intelligent association according to claim 1, wherein The process of searching for and saving the disease content corresponding to each retrieval word in the pre-stored mapping document includes the following steps: S31. Search the mapping document, match one of the retrieval words with each first-level category pre-stored in the mapping document, and obtain the first-level category corresponding to the current retrieval word; S32. Adopt the inverted index method to search for the disease content corresponding to the current retrieval word in the multiple second-level categories included in the first-level category; S33. Select another retrieval word, and repeat steps S31 to S32 until all retrieval words are searched, and obtain the disease content corresponding to each retrieval word.

3. The method for pushing disease content based on intelligent association according to claim 2, wherein The process of adopting the inverted index method to search for the disease content corresponding to the current retrieval word in the multiple second-level categories included in the first-level category includes the following steps: Match the user label of the user with the second-level category label corresponding to each second-level category, and obtain the second-level category with the second-level category label consistent with the user label as the second-level category to be retrieved; According to the address pre-stored in the second-level category to be retrieved, search for the disease content corresponding to the current retrieval word.

4. The method for pushing disease content based on intelligent association according to claim 1, wherein, The process of using the forward maximum matching word segmentation algorithm to perform word segmentation on the statement to be retrieved and obtaining multiple retrieval words included in the statement to be retrieved includes the following steps: S21. Start from the first Chinese character of the statement to be retrieved; S22. Retrieve the phrase with the most characters that matches the statement to be retrieved from the pre-stored dictionary, and save it as the current retrieval word; S23. Cut out the current retrieval word from the statement to be retrieved, obtain the segmented statement to be retrieved, and repeat steps S21 to S22 to obtain retrieval words until the statement to be retrieved is segmented completely.

5. The method for pushing disease content based on intelligent association according to claim 1, wherein Before performing word segmentation on the to-be-retrieved statement, the method further includes: cleaning the to-be-retrieved statement, and deleting punctuation marks, special symbols, and spaces in the to-be-retrieved statement.

6. The disease content push method based on intelligent association according to claim 1, wherein The process of obtaining the mapping document is as follows: Obtain the keywords of the to-be-retrieved statement, and use the distributed crawler technology to crawl the data corresponding to the keywords; Perform parsing processing on the crawled data to construct the mapping document.

7. A disease content push device based on intelligent association, characterized in that, It includes: A retrieval statement acquisition module, configured to obtain the to-be-retrieved statement input by the user in response to the user's retrieval request; A retrieval term acquisition module, configured to perform word segmentation on the to-be-retrieved statement using the forward maximum matching word segmentation algorithm to obtain multiple retrieval terms included in the to-be-retrieved statement, and each retrieval term is composed of multiple consecutive different Chinese characters in the to-be-retrieved statement; A disease content retrieval module, configured to search for and save the disease content corresponding to each retrieval term in a pre-stored mapping document, where each retrieval term corresponds to one or more disease contents, and the disease content represents information associated with the retrieval term; wherein, multiple first-level categories are pre-stored in the mapping document, each first-level category includes multiple second-level categories, the first-level category represents the type of disease, the second-level category represents the treatment drugs, prevention measures, and specific disease types corresponding to the current first-level category, a second-level category label and an address are pre-stored for each second-level category, the second-level category label is doctor, patient, or tourist, and the address represents the storage location of the disease content corresponding to the second-level category; A disease content push module, configured to sort the disease contents according to a preset weight, and push the disease contents to the user in the order of decreasing weight; After obtaining the to-be-retrieved statement input by the user, the method further includes: obtaining the user's retrieval account, and obtaining the user label of the user according to the retrieval account, where the user label is doctor, patient, or tourist; Match the user label with the label of the retrieved content pre-stored in the mapping document, and use the inverted index to retrieve the disease content.

8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Data processing method and device based on search engine, terminal and medium

    CN111949697A