Method and system for monitoring medical device network risks based on large language model
Through the drug-mechanical network risk monitoring method based on the large language model, the evaluation data of drug-mechanical products is extracted and analyzed, effective users and key evaluation data are screened out, and product characteristic data is correlated to the problem that it is difficult to effectively screen and classify evaluation data in the existing technology, and efficient monitoring and early warning of the quality of drug-mechanical products is achieved.
Patent Information
- Application Number
- CN202410988411.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-07-23
AI Technical Summary
It is difficult for the existing technology to effectively screen and classify evaluation data of pharmaceutical and mechanical products, making it difficult for merchants to discover product quality in a timely manner, affecting consumer trust and merchants' competitive advantages.
A drug-machine network risk monitoring method based on a large language model is adopted to extract and analyze the characteristics and key information in the evaluation data, and to filter out effective users and key evaluation data, and conduct correlation analysis of product feature data, and hierarchical warning prompts.
It realizes efficient analysis and screening of drug and mechanical product evaluation data, timely identify and warn of risky sales behaviors, improves the accuracy and reliability of product quality monitoring, and provides merchants with an important reference for improving products.
Smart Images

Figure CN118941082B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pharmaceutical and mechanical network risk monitoring, and in particular to a pharmaceutical and mechanical network risk monitoring method and system based on a large language model. Background Art
[0002] With the development of the Internet, the pharmaceutical and medical Internet business and distribution industries have developed rapidly in recent years. Drugs, medical devices and cosmetics can be seen everywhere through Internet platforms. The booming Internet economy has also brought great regulatory difficulties and challenges to regulatory authorities. After selling goods, some merchants may face a large amount of evaluation data and find it difficult to effectively screen and classify the evaluation data. In addition, since users have different evaluations of product quality, if the data is not screened in time, it will be difficult for merchants to discover the quality deficiencies of the products. If it is not improved in time, it will lead to further growth of negative comments from consumers on pharmaceutical and medical products. Therefore, analyzing the evaluation data after the sale, timely discovering and warning of risky sales behaviors, and correcting them can help consumers buy products with greater confidence and enhance the competitive advantage of merchants.
[0003] However, the existing data has problems such as scattered sources and low data quality, which makes it impossible to timely and accurately grasp the regulatory base and discover risky sales behaviors. Therefore, it is very important to organize and analyze the sales data of merchants and the evaluation data of users on pharmaceutical and medical products, find useful data, timely identify and warn risky sales behaviors, and provide important references for merchants to further improve their products. Summary of the invention
[0004] The purpose of the present invention is to provide a method and system for monitoring medical and mechanical network risks based on a large language model to solve the problems raised in the above-mentioned background technology.
[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0006] The pharmaceutical and medical network risk monitoring method based on a large language model includes the following steps:
[0007] Step S100: Determine the monitored medical and mechanical products, collect the evaluation data of the medical and mechanical products after online sales, judge and extract the characteristic evaluation data therein according to the large language model, and use the users who publish the characteristic evaluation data as characteristic users;
[0008] Step S200: Analyze the evaluation data posted by the characteristic users according to the large language model, screen out valid users from the characteristic users, and use the characteristic evaluation data posted by the valid users as key evaluation data;
[0009] Step S300: Collect the promotional information of the medical device product, obtain a number of product feature data according to the large language model, and obtain the associated evaluation data corresponding to each product feature data;
[0010] Step S400: Analyze the associated evaluation data corresponding to the product feature data according to the large language model to obtain the edit distance between the associated evaluation data and the corresponding product feature data, and classify the product feature data into different levels according to the edit distance, and provide different early warning prompts.
[0011] Furthermore, step S100 includes:
[0012] Step S110: the evaluation data is text evaluation data; a large language model is pre-trained, and the large language model is used to extract keywords from the evaluation data;
[0013] Step S120: Obtain the evaluation data re corresponding to the medical device product. If the content of the evaluation data re is not empty, obtain the number of characters NOC corresponding to the evaluation data re. re And the number of keywords KWQ re If NOC is met re ≥NOC0 and KWQ re ≥KWQ0, NOC0 and KWQ0 are the word quantity threshold and keyword quantity threshold respectively, then the evaluation data re is used as the characteristic evaluation data, and the user who publishes the evaluation data re is used as the characteristic user FUR re .
[0014] After receiving the goods, some users may not comment or the comments they make are of no reference value. In addition, there are many users who do not comment. Therefore, it is necessary to filter out the empty comments first, and then filter out the comments with no reference value, leaving only the comments with reference value.
[0015] Further, step S200 includes:
[0016] Step S210: Taking the release time corresponding to the characteristic evaluation data re as the starting time, obtaining the COP of the time period starting from the starting time re Corresponding feature user FUR re All evaluation data of the time period COP re According to the chronological order, it is divided into T sub-time periods, and the feature weight corresponding to the t-th sub-time period is Where 1≤t≤T; the number of evaluations corresponding to the t-th sub-time period is taken as Q t , get the characteristic user FUR re Evaluation concentration factor And normalize it;
[0017] The evaluation concentration coefficient characterizes the density of recent user - released evaluations. The larger the concentration coefficient, the more recent evaluations by users. Then, such customers are most likely to have false evaluation data, and customers with false evaluation data cannot be effective customers.
[0018] Step S220: If CTF re >CTF 0 and Q T >Q0>2, randomly obtain Q random evaluation data within the T - th sub - time period, and through the large - language model, obtain all keywords corresponding to each random evaluation data. Here, CTF 0 is the concentration coefficient threshold, Q T is the number of evaluations corresponding to the T - th sub - time period, Q0 is the evaluation quantity threshold, and Q>2; respectively obtain the keyword sets KW q and KW q-1 corresponding to the adjacent q - th and (q + 1) - th random evaluation data, where 1≤q<q + 1≤Q, and denote the sets with the minimum and maximum number of keywords in the sets KW q and KW q-1 as KW min and KW max ;
[0019] Step S230: The large - language model is also used to calculate the semantic similarity between any two keywords and normalize the result of the semantic similarity; set the initial correlation coefficient to 0. Taking the first keyword corresponding to KW min as the reference word, according to the large - language model, if the semantic similarity between any keyword in KW max and the reference word is greater than the similarity threshold, the correlation coefficient is incremented by 1, and so on, to obtain the final correlation coefficient Taking the number of keywords corresponding to the set KW min as CN min , the evaluation similarity corresponding to the q - th and (q + 1) - th random evaluation data is:
[0020]
[0021] In this solution, the large - language model is a word - vector model. The large - language model refers to a class of models that can process natural language, and the word - vector model is one of them. The existing Word2Vec can calculate the semantic similarity between two words, and the specific technology can be obtained through cosine similarity and normalize the semantic similarity. The specific implementation process is not elaborated here;
[0022] Step S240: Furthermore, based on the evaluation similarities corresponding to all adjacent random evaluation data, obtain the characteristic user FURre Corresponding evaluation data similarity DAS re If DAS re >DAS 0 , DAS 0 To evaluate the data similarity threshold, the feature user FUR re As abnormal users, the remaining characteristic users except the abnormal users are regarded as valid users.
[0023] Furthermore, step S300 includes:
[0024] Step S310: The promotional information of the medical device product includes text information and picture information. The text data is extracted from the picture information, and the text information is combined for parsing to obtain each product feature data corresponding to the promotional content, and several keywords in each product feature data and several keywords in each key evaluation data are extracted;
[0025] Extracting text information from image information and video is an existing technology and can be implemented based on OCR (optical character recognition) technology, VideoSrt and other technologies.
[0026] Product feature data is the advantages of products displayed by merchants in the form of text, pictures or videos on the sales interface for promotion. Finding evaluation data related to product feature data and analyzing the evaluation data can timely detect whether there is any falsehood in the merchant's promotional content. If there is any falsehood, the corresponding product feature data should be warned;
[0027] Step S320: Based on the keyword set KW corresponding to the a-th product feature data a , and the keyword set KW corresponding to the b-th key evaluation data b ; Set the initial correlation coefficient to 0, and use the large language model to calculate KW a The first corresponding keyword is the reference word. If there is a KW b If the semantic similarity between any keyword in and the reference word is greater than the similarity threshold, the correlation coefficient is increased by 1, and so on, to obtain the final correlation coefficient. Will gather KW a The corresponding number of keywords is used as SN min , get the proportion of the keywords corresponding to the a-th product feature data in the b-th key evaluation data
[0028] If the ratio is greater than the ratio threshold, the bth key evaluation data is used as the associated evaluation data of the ath product feature data.
[0029] Furthermore, step S400 includes:
[0030] Step S410: The large language model is also used to calculate the correlation between any two keywords, and the correlation value is between [-1, 1]. Through the large language model, the correlation between the xth keyword in the ath product feature data and the yth keyword in the bth associated evaluation data is obtained. Degree of association According to the correlation degree between the xth keyword in the ath product feature data and each keyword in the bth associated evaluation data If exists Will satisfy The maximum degree of association is recorded as DEG max ; if exists Will satisfy The minimum value of the correlation degree is recorded as DEG min ;
[0031] The specific technique for obtaining the degree of correlation can be obtained through the Euclidean distance. The closer to -1, the opposite meanings of the two are, such as safety and danger; the closer to 0, the irrelevant meanings of the two are; the closer to 1, the similar meanings of the two are, such as safety and firmness;
[0032] Step S420: Set the deviation data corresponding to the xth keyword in the ath product feature data and the bth associated evaluation data to If there is DEG max ≤|DEG min |, then the deviation data If there is DEG max >|DEG min |, then the deviation data Then we get the edit distance between the bth associated evaluation data and the ath product feature data: X a is the number of keywords in the feature data of the ath product, and e is the natural index; and then according to each corresponding edit distance, the product feature data are divided into different levels.
[0033] According to each corresponding edit distance, the product feature data are divided into different levels, including: according to the edit distance of each associated evaluation data corresponding to the a-th product feature data, and both are normalized; if the number of edit distances in [EDS'1,1] is greater than the threshold value QU1, a serious warning is given to the a-th product feature data; otherwise, if the number of edit distances in [EDS'2,EDS'1) is greater than the threshold value QU2, a general warning is given to the a-th product feature data; otherwise, if the number of edit distances in [EDS'3,EDS'2) is greater than the threshold value QU3, a slight warning is given to the a-th product feature data.
[0034] The pharmaceutical and medical network risk monitoring system includes a characteristic user module, a key evaluation data module, a related evaluation data module, and an early warning prompt module;
[0035] Feature user module: used to determine the monitored pharmaceutical and medical device products, collect the evaluation data of the pharmaceutical and medical device products after online sales, judge and extract the feature evaluation data according to the large language model, and use the users who publish the feature evaluation data as feature users;
[0036] Key evaluation data module: used to analyze the evaluation data published by characteristic users according to the large language model, screen out valid users from the characteristic users, and use the characteristic evaluation data published by valid users as key evaluation data;
[0037] Related evaluation data module: used to collect the promotional information of the medical device product, obtain a number of product feature data according to the large language model, and obtain the related evaluation data corresponding to each product feature data;
[0038] Early warning prompt module: used to analyze the associated evaluation data corresponding to the product feature data according to the large language model, obtain the edit distance between the associated evaluation data and the product feature data, and classify the product feature data into different levels according to the edit distance, and provide different early warning prompts.
[0039] Further, the characteristic user module includes a large language model unit and a characteristic user unit;
[0040] Large language model unit: used for pre-training a large language model, wherein the large language model is used for extracting keywords from evaluation data;
[0041] Feature user unit: used to analyze the content of the evaluation data, judge and extract the feature evaluation data therein, and regard the user who publishes the feature evaluation data as the feature user.
[0042] Further, the associated evaluation data module includes a publicity information unit and an associated evaluation data unit;
[0043] Promotional information unit: used to collect promotional information of the medical device product and obtain several product feature data according to the large language model;
[0044] Related evaluation data unit: used to obtain the related evaluation data corresponding to each product feature data.
[0045] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: the present invention provides a method and system for monitoring the risk of medical device network based on a large language model, including: determining the monitored medical device products, collecting evaluation data of medical device products, and judging and extracting the characteristic evaluation data therein, and then obtaining characteristic users corresponding to the characteristic evaluation data; screening out effective users and obtaining key evaluation data; collecting the promotional information of medical device products to obtain a number of product characteristic data; dividing the product characteristic data into different levels and performing different early warning prompts. The present invention parses the promotional information of medical device products to obtain a number of product characteristic data, and then performs early warning prompts for each product characteristic data, timely identifies and warns of risky promotional data, improves the reliability of promotional information, and provides important reference and guidance for merchants to further improve their products. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0047] Figure 1 It is a flow chart of the method for monitoring the medical and mechanical network risk based on the large language model of the present invention;
[0048] Figure 2 It is a structural diagram of the pharmaceutical and mechanical network risk monitoring system based on a large language model of the present invention. DETAILED DESCRIPTION
[0049] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0050] See also Figure 1-Figure 2 The present invention provides a method and system for monitoring the risk of medical and medical network based on a large language model. The corresponding flowchart of the method for monitoring the risk of medical and medical network based on a large language model is as follows: Figure 1 As shown. The following steps are included:
[0051] Step S100: Determine the monitored pharmaceutical and medical device products, collect the evaluation data of the pharmaceutical and medical device products after online sales, judge and extract the characteristic evaluation data therein according to the large language model, and use the users who publish the characteristic evaluation data as characteristic users.
[0052] Step S110: the evaluation data is text evaluation data; a large language model is pre-trained, and the large language model is used to extract keywords from the evaluation data.
[0053] The evaluation data in this solution is the text evaluation data of the user. The existing technology of extracting keywords from the evaluation data as a large language model will not be described in detail here.
[0054] Step S120: Obtain the evaluation data re corresponding to the medical device product. If the content of the evaluation data re is not empty, obtain the number of characters NOC corresponding to the evaluation data re. re And the number of keywords KWQ re If NOC is met re ≥NOC0 and KWQ re ≥KWQ0, NOC0 and KWQ0 are the word quantity threshold and keyword quantity threshold respectively, then the evaluation data re is used as the characteristic evaluation data, and the user who publishes the evaluation data re is used as the characteristic user FUR re .
[0055] After receiving the product, some users may not comment or the comments they make are of no reference value, and there are many users who do not comment. Therefore, it is necessary to filter out the empty comments first, and then filter out the comments with no reference value, leaving only the comments with reference value. In this embodiment, the word quantity threshold NOC0 and the keyword quantity threshold KWQ0 are 15 and 2 respectively.
[0056] Step S200: Analyze the evaluation data posted by the characteristic users according to the large language model, screen out valid users from the characteristic users, and use the characteristic evaluation data posted by the valid users as key evaluation data;
[0057] Step S210: Taking the release time corresponding to the characteristic evaluation data re as the starting time, obtaining the COP of the time period starting from the starting time re Corresponding feature user FUR re All evaluation data of the time period COP re According to the chronological order, it is divided into T sub-time periods, and the feature weight corresponding to the t-th sub-time period is Where 1≤t≤T; the number of evaluations corresponding to the t-th sub-time period is taken as Q t , get the characteristic user FUR re Evaluation concentration factor And normalize it;
[0058] The evaluation concentration coefficient characterizes the density of recent user - released evaluations. The larger the concentration coefficient, the more recent evaluations from users. Then, such customers are most likely to have false evaluation data. Customers with false evaluation data cannot be valid customers. The purpose of setting the feature weight for each sub - time period is as follows: Since the larger the \(t\), the closer the corresponding time period is to the starting moment, and recent evaluations can better reflect the current true attitude and behavior of customers, thus more accurately assessing the possibility of whether customers have false evaluation data, the weight should be relatively large; in addition, as time goes by, old data may gradually lose its accurate reflection of the current situation, and the weight should be relatively small. In this embodiment, the time period \(COP\) re is 3 months, \(T = 3\), then each month is divided into a sub - time period.
[0059] Step S220: If it satisfies \(CTF\) re \(>CTF\) 0 and \(Q\) T \(>Q_0>2\), randomly obtain \(Q\) random evaluation data within the \(T\) - th sub - time period, and through the large - language model, obtain all keywords corresponding to each random evaluation data. Among them, \(CTF\) 0 is the concentration - coefficient threshold, \(Q\) T is the number of evaluations corresponding to the \(T\) - th sub - time period, \(Q_0\) is the evaluation - number threshold, \(Q>2\); respectively obtain the keyword sets \(KW\) q and \(KW\) q-1 corresponding to the adjacent \(q\) - th and \((q + 1)\) - th random evaluation data, \(1\leq q<q + 1\leq Q\), and record the sets with the minimum and maximum number of keywords in the sets \(KW\) q and \(KW\) q-1 as \(KW\) min and \(KW\) max .
[0060] In this embodiment, the concentration - coefficient threshold \(CTF\) 0 value is 0.5. Exceeding 0.5 indicates relatively high concentration and a greater possibility of being a non - valid customer. Then, it is necessary to further analyze the evaluation data of this part of customers. The evaluation - number threshold \(Q_0\) is 10.
[0061] Step S230: The large - language model is also used to calculate the semantic similarity between any two keywords and normalize the results of the semantic similarity; set the initial correlation coefficient to 0, use the first keyword corresponding to \(KW\) min as the reference word. According to the large - language model, if there is any keyword in \(KW\) max whose semantic similarity with the reference word is greater than the similarity threshold, the correlation coefficient is incremented by 1, and so on, to obtain the final correlation coefficient Take the number of keywords corresponding to the set \(KW\) min as \(CN\)min , the evaluation similarity corresponding to the qth and q+1th random evaluation data is obtained as follows:
[0062]
[0063] In this solution, the large language model is a word vector model. The large language model refers to a type of model that can process natural language, and the word vector model is one of them. The existing Word2Vec can calculate the semantic similarity between two words. The specific technology can be obtained through cosine similarity, and the semantic similarity is normalized. The specific implementation process is not repeated here. Since the semantic similarity between two words ranges from [0,1], the larger the value, the more similar it is. In this embodiment, the similarity threshold is 0.6.
[0064] Step S240: Then, according to the evaluation similarities corresponding to all adjacent random evaluation data, the characteristic user FUR is obtained. re Corresponding evaluation data similarity DAS re If DAS re >DAS 0 , DAS 0 To evaluate the data similarity threshold, the feature user FUR re As abnormal users, the remaining characteristic users except the abnormal users are regarded as valid users.
[0065] Evaluation data similarity DAS re In order to find the average value of all evaluation similarities, in this embodiment, the evaluation data similarity threshold DAS 0 =0.7.
[0066] Step S300: Collect the promotional information of the medical device product, obtain a number of product feature data according to the large language model, and obtain the associated evaluation data corresponding to each product feature data;
[0067] Step S310: The promotional information of the pharmaceutical and medical device product includes text information and picture information. The text data is extracted from the picture information and analyzed in combination with the text information to obtain each product feature data corresponding to the promotional content, and several keywords in each product feature data and several keywords in each key evaluation data are extracted.
[0068] Extracting text information from image information and video is an existing technology and can be implemented based on OCR (optical character recognition) technology, VideoSrt and other technologies.
[0069] Product feature data is used by merchants to display the advantages of products in the form of text, pictures or videos on the sales interface for promotion. By finding evaluation data related to the product feature data and analyzing the evaluation data, it is possible to promptly discover whether there is any falsehood in the merchant's promotional content. If there is any falsehood, an early warning should be issued for the corresponding product feature data.
[0070] Step S320: Based on the keyword set KW corresponding to the a-th product feature data a , and the keyword set KW corresponding to the b-th key evaluation data b ; Set the initial correlation coefficient to 0, and use the large language model to calculate KW a The first corresponding keyword is the reference word. If there is a KW b If the semantic similarity between any keyword in and the reference word is greater than the similarity threshold, the correlation coefficient is increased by 1, and so on, to obtain the final correlation coefficient. Will gather KW a The corresponding number of keywords is used as SN min , get the proportion of the keywords corresponding to the a-th product feature data in the b-th key evaluation data If the ratio is greater than the ratio threshold, the bth key evaluation data is used as the associated evaluation data of the ath product feature data.
[0071] In this embodiment, the similarity threshold is 0.6 and the ratio threshold is 0.75.
[0072] Step S400: Analyze the associated evaluation data corresponding to the product feature data according to the large language model to obtain the edit distance between the associated evaluation data and the corresponding product feature data, and classify the product feature data into different levels according to the edit distance, and provide different early warning prompts.
[0073] Step S410: The large language model is also used to calculate the correlation between any two keywords, and the correlation value is between [-1, 1]. Through the large language model, the correlation between the xth keyword in the ath product feature data and the yth keyword in the bth associated evaluation data is obtained. Degree of association According to the correlation degree between the xth keyword in the ath product feature data and each keyword in the bth associated evaluation data If exists Will satisfy The maximum degree of association is recorded as DEG max ; if exists Will satisfy The minimum value of the correlation degree is recorded as DEG min .
[0074] Among them, the specific technology for obtaining the degree of correlation can be obtained through Euclidean distance. The closer to -1, the two have opposite meanings, such as safety and danger; the closer to 0, the two have no correlation; the closer to 1, the two have similar meanings, such as safety and firmness.
[0075] It should be noted that, in this embodiment, the semantic similarity and the degree of association are different: the semantic similarity represents the degree of relationship between two words, such as safety and danger, they are often used to describe the same situation or topic, and the semantic similarity is large, close to 1; the degree of association represents the degree of meaning of the two words, and safety and danger are opposite in meaning, one is safe and the other is unsafe, so the degree of association is small, close to -1.
[0076] Step S420: Set the deviation data corresponding to the xth keyword in the ath product feature data and the bth associated evaluation data to If there is DEG max ≤|DEG min |, then the deviation data If there is DEG max >|DEG min |, then the deviation data Then we get the edit distance between the bth associated evaluation data and the ath product feature data: X a is the number of keywords in the feature data of the ath product, and e is the natural index; and then according to each corresponding edit distance, the product feature data are divided into different levels.
[0077] The edit distance represents the distance between the keywords of the associated evaluation data and the product feature data. It can also be understood as the distance required to convert the associated evaluation data into product feature data. If the distance is small, it means that the two are relatively close. According to each corresponding edit distance, the product feature data is divided into different levels, including:
[0078] According to the edit distance of each associated evaluation data corresponding to the a-th product feature data, they are normalized; if the number of edit distances in [EDS'1,1] is greater than the threshold value QU1, a serious warning is given to the a-th product feature data; otherwise, if the number of edit distances in [EDS'2,EDS'1) is greater than the threshold value QU2, a general warning is given to the a-th product feature data; otherwise, if the number of edit distances in [EDS'3,EDS'2) is greater than the threshold value QU3, a slight warning is given to the a-th product feature data. In this embodiment, EDS'1 = 0.8, QU1 = 10; EDS'2 = 0.6, QU2 = 9; EDS'3 = 0.4, QU1 = 8.
[0079] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0080] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A drug-mechanical network risk monitoring method based on a large language model, characterized in that: The steps include: Step S100: Determine the monitored medical and mechanical products, collect the evaluation data of the medical and mechanical products after online sales, judge and extract the characteristic evaluation data therein according to the large language model, and use the users who publish the characteristic evaluation data as characteristic users; Step S200: Analyze the evaluation data posted by the characteristic users according to the large language model, screen out valid users from the characteristic users, and use the characteristic evaluation data posted by the valid users as key evaluation data; Step S300: Collect the promotional information of the medical device product, obtain a number of product feature data according to the large language model, and obtain the associated evaluation data corresponding to each product feature data; Step S400: Analyze the associated evaluation data corresponding to the product feature data according to the large language model to obtain the edit distance between the associated evaluation data and the corresponding product feature data, and classify the product feature data into different levels according to the edit distance, and perform different early warning prompts; Step S210: Taking the release time corresponding to the characteristic evaluation data re as the starting time, obtaining the COP of the time period starting from the starting time re Corresponding feature user FUR re All evaluation data of the time period COP re According to the chronological order, it is divided into T sub-time periods, and the feature weight corresponding to the t-th sub-time period is Where 1≤t≤T; the number of evaluations corresponding to the t-th sub-time period is taken as Q t , get the characteristic user FUR re Evaluation concentration factor And normalize it; Step S220: If CTF is satisfied re > CTF 0 and Q T > Q0 > 2, randomly obtain Q random evaluation data within the T-th sub-time period, and use the large language model to obtain all keywords corresponding to each random evaluation data, where CTF 0 is the concentration coefficient threshold, Q T is the number of evaluations corresponding to the T-th sub-time period, Q0 is the evaluation number threshold, Q > 2; respectively obtain the keyword sets KW q and KW q-1 corresponding to the adjacent q-th and (q + 1)-th random evaluation data, 1 ≤ q < q + 1 ≤ Q, and denote the sets with the smallest and largest number of keywords in the sets KW q and KW q-1 as KW min and KW max ; Step S230: The large language model is also used to calculate the semantic similarity between any two keywords and normalize the result of the semantic similarity; set the initial correlation coefficient to 0, and use KW min The first keyword is the reference word. According to the large language model, if there is a KW max If the semantic similarity between any keyword in and the reference word is greater than the similarity threshold, the correlation coefficient is increased by 1, and so on, to obtain the final correlation coefficient. Will gather KW min The corresponding number of keywords is used as CN min , the evaluation similarity corresponding to the qth and q+1th random evaluation data is obtained as follows: Step S240: Then, according to the evaluation similarities corresponding to all adjacent random evaluation data, the characteristic user FUR is obtained. re Corresponding evaluation data similarity DAS re If DAS re >DAS 0 , DAS 0 To evaluate the data similarity threshold, the feature user FUR re As abnormal users, other characteristic users except the abnormal users are regarded as valid users; Step S410: The large language model is also used to calculate the correlation between any two keywords, and the correlation value is between [-1, 1]. Through the large language model, the correlation between the xth keyword in the ath product feature data and the yth keyword in the bth associated evaluation data is obtained. Degree of association According to the correlation degree between the xth keyword in the ath product feature data and each keyword in the bth associated evaluation data If exists Will satisfy The maximum degree of association is recorded as DEG max ; if exists Will satisfy The minimum value of the correlation degree is recorded as DEG min ; Step S420: Set the deviation data corresponding to the xth keyword in the ath product feature data and the bth associated evaluation data to If there is DEG max ≤|DEG min |, then the deviation data If there is DEG max >|DEG min |, then the deviation data Then we get the edit distance between the bth associated evaluation data and the ath product feature data: X a is the number of keywords in the feature data of the ath product, and e is the natural index; and then according to each corresponding edit distance, the product feature data are divided into different levels.
2. The method for monitoring medical device network risks based on a large language model according to claim 1 is characterized in that: Step S100 includes: Step S110: the evaluation data is text evaluation data; a large language model is pre-trained, and the large language model is used to extract keywords from the evaluation data; Step S120: Obtain the evaluation data re corresponding to the medical device product. If the content of the evaluation data re is not empty, obtain the number of characters NOC corresponding to the evaluation data re. re And the number of keywords KWQ re If NOC is met re ≥NOC0 and KWQ re ≥KWQ0, NOC0 and KWQ0 are the word quantity threshold and keyword quantity threshold respectively, then the evaluation data re is used as the characteristic evaluation data, and the user who publishes the evaluation data re is used as the characteristic user FUR re .
3. The method for monitoring medical device network risks based on a large language model according to claim 1 is characterized in that: Step S300 includes: Step S310: The promotional information of the medical device product includes text information and picture information. The text data is extracted from the picture information, and the text information is combined for parsing to obtain each product feature data corresponding to the promotional content, and several keywords in each product feature data and several keywords in each key evaluation data are extracted; Step S320: Based on the keyword set KW corresponding to the a-th product feature data a , and the keyword set KW corresponding to the b-th key evaluation data b ; Set the initial correlation coefficient to 0, and use the large language model to calculate KW a The first corresponding keyword is the reference word. If there is a KW b If the semantic similarity between any keyword in and the reference word is greater than the similarity threshold, the correlation coefficient is increased by 1, and so on, to obtain the final correlation coefficient. Will gather KW a The corresponding number of keywords is used as SN min , get the proportion of the keywords corresponding to the a-th product feature data in the b-th key evaluation data If the ratio is greater than the ratio threshold, the bth key evaluation data is used as the associated evaluation data of the ath product feature data.
4. A pharmaceutical mechanization network risk monitoring system, used to implement the pharmaceutical mechanization network risk monitoring method based on a large language model as described in any one of claims 1 to 3, characterized in that: The system includes a characteristic user module, a key evaluation data module, a related evaluation data module and an early warning prompt module; Feature user module: used to determine the monitored pharmaceutical and medical device products, collect the evaluation data of the pharmaceutical and medical device products after online sales, judge and extract the feature evaluation data according to the large language model, and use the users who publish the feature evaluation data as feature users; Key evaluation data module: used to analyze the evaluation data published by characteristic users according to the large language model, screen out valid users from the characteristic users, and use the characteristic evaluation data published by valid users as key evaluation data; Related evaluation data module: used to collect the promotional information of the medical device product, obtain a number of product feature data according to the large language model, and obtain the related evaluation data corresponding to each product feature data; Early warning prompt module: used to analyze the associated evaluation data corresponding to the product feature data according to the large language model, obtain the edit distance between the associated evaluation data and the product feature data, and classify the product feature data into different levels according to the edit distance, and provide different early warning prompts.
5. The pharmaceutical mechanization network risk monitoring system according to claim 4 is characterized in that: The characteristic user module includes a large language model unit and a characteristic user unit; Large language model unit: used for pre-training a large language model, wherein the large language model is used for extracting keywords from evaluation data; Feature user unit: used to analyze the content of the evaluation data, judge and extract the feature evaluation data therein, and regard the user who publishes the feature evaluation data as the feature user.
6. The pharmaceutical mechanization network risk monitoring system according to claim 4 is characterized in that: The associated evaluation data module includes a publicity information unit and an associated evaluation data unit; Promotional information unit: used to collect promotional information of the medical device product and obtain several product feature data according to the large language model; Related evaluation data unit: used to obtain the related evaluation data corresponding to each product feature data.
Citation Information
Patent Citations
Risk assessing method for e-commerce product quality
CN107977798A
Epidemic disease monitoring and early warning method and device and electronic equipment
CN112562863A