Methods, devices, equipment, and storage media for displaying text
Patent Information
- Application Number
- CN202310450179.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-24
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2043-04-24
AI Technical Summary
[0003]本公开提供了一种用于文本的展示方法、装置、设备以及存储介质,以至少解决相关技术中搜索文本的展示的效果较差的技术问题
Smart Images

Figure CN116467434B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of display, specifically to the field of displaying search text, and more particularly to a method for displaying text. Background Technology
[0002] With the continuous development of artificial intelligence technology, AI models can be used to classify and display search text. However, because the effect of using AI models to classify search text is relatively limited, the final display effect of the search text is poor. Summary of the Invention
[0003] This disclosure provides a method, apparatus, device, and storage medium for displaying text, to at least solve the technical problem of poor display effect of searched text in related technologies.
[0004] According to one aspect of this disclosure, a method for displaying text is provided, comprising: in response to receiving search text in a preset domain, determining an initial classification result of the search text; extracting at least one keyword belonging to the preset domain from the search text and obtaining the page views of at least one keyword; normalizing the at least one keyword in the search text to obtain target text, wherein the target text contains target keywords that are semantically identical to at least one keyword; and displaying the target text according to the initial classification result and page views.
[0005] According to another aspect of this disclosure, a text display apparatus is provided, comprising: a determining module, configured to determine an initial classification result of the search text in response to receiving search text in a preset domain; an extraction module, configured to extract at least one keyword belonging to the preset domain from the search text and obtain the page views of the at least one keyword; a processing module, configured to normalize the at least one keyword in the search text to obtain target text, wherein the target text contains target keywords that are semantically identical to the at least one keyword; and a display module, configured to display the target text according to the initial classification result and page views.
[0006] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the text display method proposed in this disclosure.
[0007] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to execute the text display method proposed in this disclosure.
[0008] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the text display method proposed in this disclosure.
[0009] In this disclosure, in response to receiving search text from a preset domain, an initial classification result of the search text is determined; at least one keyword belonging to the preset domain is extracted from the search text, and the pageviews of at least one keyword are obtained; the at least one keyword in the search text is normalized to obtain target text, wherein the target text contains target keywords that are semantically identical to at least one keyword; and the target text is displayed according to the initial classification result and pageviews. It is readily apparent that by normalizing multiple keywords in the search text, semantically identical keywords are normalized to obtain corresponding target text that can be uniformly expressed. This, in turn, allows for the display of the target text and pageviews, achieving a clear and visual representation of search text information from the preset domain, thus solving the technical problem of poor search text display in related technologies.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0012] Figure 1 This is a flowchart of a text display method according to an embodiment of the present disclosure;
[0013] Figure 2 This is a schematic diagram illustrating the changes in browsing popularity based on user search indexes according to embodiments of this disclosure;
[0014] Figure 3 This is a schematic diagram showing the page displayed based on the user search index according to an embodiment of this disclosure;
[0015] Figure 4 A flowchart illustrating the text display based on user search indexes according to embodiments of this disclosure;
[0016] Figure 5 This is a schematic diagram illustrating the changes in the popularity of the doctor-mentioned index according to embodiments of this disclosure;
[0017] Figure 6 This is a flowchart illustrating the text display of index implementation based on embodiments of this disclosure;
[0018] Figure 7This is a structural block diagram of entity mining, entity construction, and entity recognition according to embodiments of this disclosure;
[0019] Figure 8 This is a structural block diagram of a text display device according to an embodiment of the present disclosure;
[0020] Figure 9 This is a hardware structure block diagram of a computer terminal (or mobile device) for a text display method according to an embodiment of the present disclosure. Detailed Implementation
[0021] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0022] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0023] The method embodiments provided in this disclosure can be performed in a mobile terminal, computer terminal, or similar electronic device. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the disclosure described and / or claimed herein.
[0024] According to embodiments of this disclosure, a method for displaying text is provided. It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0025] Figure 1 This is a flowchart of a text display method according to an embodiment of this disclosure. Figure 1 As shown, the method includes the following steps:
[0026] Step S10: In response to receiving search text from a preset domain, determine the initial classification result of the search text.
[0027] Specifically, the aforementioned preset field can be a pre-determined field. Optionally, the preset field can be the medical field, the catering field, the education field, etc., which are only illustrative examples and are not intended to limit the preset field.
[0028] The search text mentioned above can be text that needs to be processed within a preset domain.
[0029] Optionally, taking the pharmaceutical field mentioned above as an example, the search text can be either text describing the "user search index of a certain drug" or text describing the "user mention index of a certain drug". It should be noted that the search text is not uniquely limited and can be determined according to the actual situation of the preset field.
[0030] The initial classification results described above can be used to represent the results obtained by classifying the search text.
[0031] One optional implementation method is to classify the parts of speech in the search text to obtain initial classification results. Optionally, the system can iterate through popular topics discussed by users on a specific application or website corresponding to a preset domain, and then classify each popular topic by part of speech. This can be done through model classification, manual annotation, or a collaborative approach of model classification and manual verification. No single method for part-of-speech classification is set here; it can be analyzed according to specific circumstances. That is, when the amount of popular topics discussed by users is large, model classification is needed to obtain the initial classification results of the search text, thereby improving classification efficiency. When the amount of popular topics discussed by users is small, classification based on manual annotation is needed to obtain the initial classification results of the search text, thereby improving classification accuracy.
[0032] Alternatively, the search can be conducted by classifying the functional characteristics of each search keyword in the search text. Optionally, the search text may contain search keywords such as "drug A," "drug B," and "drug C." By classifying the indications and functions of each search keyword, a classification result can be determined with indication 'a' as one category, or with indication 'b' as another category, etc. Here, the functional characteristic classification result is not uniquely set; it needs to be classified according to the optional indications and functions.
[0033] In another optional implementation, by classifying the search text by part of speech or by the functional features of search keywords, multiple classification results of the search text can be determined. It should be noted that the initial classification result can be any one of the multiple classification results. In order to display text in a preset domain based on the classification results, for example, multiple classification results in the preset domain can be displayed simultaneously, and then the user views of multiple classification results within the same preset time period can be statistically analyzed to obtain the view popularity of each classification result. By arranging the view popularity in a preset order, a preset number of view popularity values in the sequence can be determined, and then the classification results of the search text corresponding to the preset number of view popularity values can be determined and displayed. This achieves the technical effect of dynamically displaying the classification results of the search text based on the user's view volume.
[0034] Step S11: Extract at least one keyword belonging to the preset domain from the search text, and obtain the pageviews of the at least one keyword.
[0035] Specifically, at least one of the aforementioned keywords can be used to represent keywords extracted from the search text.
[0036] In one optional embodiment, a preset set of statement fields is determined based on a preset domain; it is determined whether the statement fields in the search text satisfy the preset set of statement fields; and statement fields that satisfy the preset set of statement fields are determined as keywords.
[0037] In another optional embodiment, the sentence information in the search text is structurally analyzed, and the sentence fields that meet the preset structure fields are determined as keywords.
[0038] In a third optional embodiment, the frequency of occurrence of statement fields in the search text is collected to determine the frequency of statement fields; it is determined whether the frequency of statement fields meets the preset frequency of occurrence; and statement fields that meet the preset frequency of occurrence are determined as keywords.
[0039] Optionally, keyword extraction from search text can be based on a preset domain. For example, in the case of a preset domain like medicine, fields related to drug names can be identified as keywords. In addition, fields related to indications and functions can also be identified as keywords. It should be noted that the above is only an illustrative description and does not impose a unique requirement on the keywords.
[0040] In addition, during the extraction of keywords from search text, on the one hand, structural analysis can be performed on the sentence information in the search text to identify keywords such as "drug A granules," "drug A capsules," "drug B," or "drug a efficacy." On the other hand, by traversing the words in the search text and statistically analyzing the frequency of word occurrences, words that meet the preset frequency can be identified and extracted as the aforementioned keywords.
[0041] For example, when the search text contains information such as "Drug A granules can achieve drug effect a", "Drug A capsules can achieve drug effect a", "Drug B can achieve drug effect a", etc., at least one of the above keywords includes, but is not limited to, "Drug A granules", "Drug A capsules", "Drug B" or "drug effect a".
[0042] The aforementioned pageviews can be used to represent pageviews for at least one of the aforementioned keywords. It should be noted that the aforementioned pageviews can be pageviews by users for the keywords, or pageviews by users after the doctor mentions the aforementioned keywords.
[0043] In one optional embodiment, in the process of obtaining the views of at least one keyword, it is necessary to dynamically monitor the user's browsing data through a backend server in a preset domain and provide real-time feedback on the user's browsing volume.
[0044] Step S12: Normalize at least one keyword in the search text to obtain target text, wherein the target text contains target keywords that have the same semantics as the at least one keyword.
[0045] Specifically, the target text mentioned above can be used to represent the text obtained after processing the search text. Unlike the search text mentioned above, which is a comprehensive description of a preset domain, the target text is a more focused summary obtained by normalizing semantically similar fields in the comprehensive description text, but without changing the main information of the search text.
[0046] Optionally, directly displaying the preset domain based on fields in the search text might result in unclear display. Since search text may contain multiple descriptions of the same key field, displaying each description individually could lead to unclear descriptions and wasted display space. Therefore, to avoid this problem, it's necessary to normalize the multiple descriptions of the same key field, summarizing them into a single, unified description. This normalization process achieves a clearer display of the search text while also saving display space.
[0047] In one optional embodiment, by normalizing at least one keyword in the search text, a keyword corresponding to the normalized keyword is obtained. For example, when the search text contains information such as "drug A granules can achieve drug effect a", "drug A capsules can achieve drug effect a", "drug B can achieve drug effect a", etc., by normalizing the keywords "drug A granules" and "drug A capsules", the normalized keyword "drug A" can be obtained. Then, the target text obtained after processing the search text will contain "drug A can achieve drug effect a".
[0048] The aforementioned target keywords can be used to represent keywords obtained by normalizing at least one keyword, namely the keyword "Drug A".
[0049] In addition, although "Drug A Granules" and "Drug A Capsules" exist in different forms, since they both belong to "Drug A", the corresponding "Drug A" can be obtained by normalizing the two keywords. At the same time, the user page views corresponding to "Drug A Granules" and "Drug A Capsules" can also be normalized. The user page views a1 of "Drug A Granules" and a2 of "Drug A Capsules" can be added together to obtain the user page views of "Drug A".
[0050] Step S13: Display the target text according to the initial classification result and the number of views.
[0051] Specifically, after determining the initial classification results and pageviews, the target text needs to be displayed. It is important to note that the initial classification results and pageviews should be consistent.
[0052] In one optional embodiment, by classifying the parts of speech in the search text, an initial classification result based on part-of-speech classification can be obtained. At the same time, the page views including part-of-speech classifications such as "negative", "neutral", and "positive" can be obtained. By normalizing the "negative" keywords, "negative" target text can be obtained. By normalizing the "neutral" keywords, "neutral" target text can be obtained. By normalizing the "positive" keywords, "positive" target text can be obtained. Then, the page views of the "negative" target text and its superimposed processing, the page views of the "neutral" target text and its superimposed processing, and the page views of the "positive" target text and its superimposed processing are visualized.
[0053] In another optional embodiment, by classifying the functional features of each search keyword in the search text, an initial classification result based on functional feature classification can be obtained. Simultaneously, the pageviews for functional feature classifications such as "fever," "headache," and "cough" are acquired. By normalizing the keyword "fever" and its pageviews, the target text for "fever," "headache," and "cough" can be obtained. The target texts for "fever," "headache," and "cough" and their normalized pageviews are then visualized. Generally, the classification methods for search text include, but are not limited to, the two methods described above; the classification results are not uniquely limited here.
[0054] According to steps S10 to S13 of this disclosure, in response to receiving search text from a preset domain, an initial classification result of the search text is determined; at least one keyword belonging to the preset domain is extracted from the search text, and the pageviews of at least one keyword are obtained; the at least one keyword in the search text is normalized to obtain target text, wherein the target text contains target keywords that have the same semantic meaning as at least one keyword; the target text is displayed according to the initial classification result and pageviews. It is readily apparent that by normalizing multiple keywords in the search text, the semantically identical keywords are normalized to obtain corresponding target text that can be uniformly expressed. This, in turn, allows for the display of the target text and pageviews, achieving a clear and visual display of search text information from the preset domain, thus solving the technical problem of poor display effects of search text in related technologies.
[0055] The method described in this embodiment will be further described below.
[0056] Optionally, the target text can be displayed according to the initial classification results and pageviews, including: mapping the pageviews using a preset mapping logic to obtain mapped pageviews; and displaying the target text according to the initial classification results and mapped pageviews.
[0057] Specifically, the aforementioned preset mapping logic can be used to represent a pre-defined logic for mapping actual pageviews.
[0058] The mapped pageviews mentioned above can be used to represent pageviews obtained by mapping actual pageviews.
[0059] In one optional embodiment, a mapping formula can be embedded in the preset mapping logic. That is, by inputting the actual pageviews into the mapping formula for calculation, the mapped pageviews can be obtained. For example, when a user's search index is finalized, and the pageviews a1 of the search query keyword are mapped, the preset mapping logic can be: a1*5*(0.975+0.01*b), where b(0,1), where a1 is the actual pageviews of the keyword, and b is a preset error value between 0 and 1. Based on the above preset mapping logic, given the actual pageviews a1, the corresponding mapped pageviews can be obtained by calculating a1 using the preset mapping logic.
[0060] In another optional embodiment, when the user mention index is finalized and the pageviews a2 of the keywords mentioned in the content are mapped, a preset mapping logic can be used: a2*5*(0.975+0.01*b), where b(0,1), where a2 is the actual pageviews of the keyword, and b is a preset error value between 0 and 1. Based on the above preset mapping logic, given the actual pageviews a2, the corresponding mapped pageviews can be obtained by calculating a2 using the preset mapping logic.
[0061] In summary, by establishing a pre-defined mapping logic, the actual pageview data can be mapped, thus achieving the technical effect of masking the real data and preventing data leakage.
[0062] Optionally, the target text is displayed according to the initial classification results and the mapped pageviews, including: filtering fields in the target text that do not meet the preset conditions to obtain filtered text; aggregating the filtered text using an aggregation model to obtain aggregated text; and displaying the aggregated text according to the initial classification results and the mapped pageviews.
[0063] Specifically, the aforementioned preset conditions can be used to represent pre-defined conditions that conform to the text display, including but not limited to blacklists, the requirement that the text contains entities and their corresponding semantics, word count limits, and standardization of punctuation marks.
[0064] The filtered text mentioned above can be used to represent the text obtained by filtering the fields of the target text.
[0065] Optionally, based on preset conditions, fields such as blacklists, entities without semantics, semantics without entities, exceeding the word limit, or non-standard punctuation can be filtered in the target text to obtain filtered text that meets the preset conditions.
[0066] The aggregation model described above can be used to represent a model for aggregating similar fields in filtered text.
[0067] The aggregated text mentioned above can be used to represent the text obtained after aggregating similar fields in the filtered text.
[0068] In one optional embodiment, during the process of aggregating the filtered text using an aggregation model, the semantics of each field in the filtered text can be retrieved to determine the semantic information of each field. Then, similar semantic information can be aggregated to obtain aggregated text, thereby avoiding the repetition of fields corresponding to similar semantic information and achieving the technical effect of reducing the occurrence of semantically repetitive search keywords.
[0069] Optionally, the aggregated text can be displayed according to the initial classification results and the mapped pageviews, including: sorting the aggregated text based on the mapped pageviews to obtain a sorting result; and displaying the aggregated text according to the initial classification results based on the sorting result.
[0070] Specifically, the sorting results described above can be used to represent the results obtained by sorting aggregated text.
[0071] In one optional embodiment, by arranging the mapped pageviews corresponding to each text in the aggregated text, an arranged sequence of aggregated text can be obtained, thereby enabling the text corresponding to the mapped pageviews in the aggregated text sequence to be highlighted or prioritized.
[0072] Optionally, when the mapped pageviews of each text in the aggregated text are sorted in descending order, the text with the highest-ranking mapped pageviews in the aggregated text sequence can be identified, thereby highlighting or prioritizing that text. Conversely, when the mapped pageviews of each text in the aggregated text are sorted in ascending order, the text with the lowest-ranking mapped pageviews in the aggregated text sequence can be identified, thereby highlighting or prioritizing that text.
[0073] Figure 2 This is a schematic diagram illustrating the changes in browsing popularity based on user search indexes according to embodiments of this disclosure. For example... Figure 2As shown, the horizontal axis reflects the display date, the vertical axis reflects browsing popularity, curve A reflects the popularity change curve of user search index, and curve B reflects the popularity change curve of doctor mention index. Specifically, at time point t, the popularity value of doctor mention index B1 is higher than the popularity value of user search index A1.
[0074] Figure 3 This is a schematic diagram illustrating the page display based on the user search index according to an embodiment of this disclosure. For example... Figure 3 As shown, on the user search index landing page, negative topics account for 4%, neutral topics account for 99%, and positive topics account for 1%.
[0075] Figure 4 This is a flowchart illustrating the text display based on user search indexes according to embodiments of this disclosure, such as... Figure 4 As shown:
[0076] S41, retrieve the search text for the preset domain;
[0077] Search text includes, but is not limited to, user search query fields, query date fields, drug mapping fields, page view mapping fields, and polarity tag fields.
[0078] S42, Filter the search text based on preset conditions to obtain the filtered text;
[0079] The preset conditions include, but are not limited to, blacklists, the requirement that the text contains entities and their corresponding semantics, word count limits, and standardization of punctuation marks.
[0080] S43, aggregate the filtered text to obtain aggregated text;
[0081] S44, Sort the aggregated text based on the mapped pageviews to obtain the sorting results;
[0082] S45, Based on the sorting results, display the aggregated text according to the initial classification results.
[0083] Based on the sorting results, a preset number of search query fields with different polarities for different drugs are displayed.
[0084] Figure 5 This is a schematic diagram illustrating the changes in the popularity of the doctor-mentioned index according to an embodiment of this disclosure. For example... Figure 5 As shown, A, B, C, D, and E represent different drugs; 'a' represents the real-time popularity of A, 'b' represents the real-time popularity of B, 'c' represents the real-time popularity of C, 'd' represents the real-time popularity of D, and 'e' represents the real-time popularity of E. A descending arrow indicates a decreasing real-time popularity, an ascending arrow indicates a increasing real-time popularity, and a horizontal bar indicates that the real-time popularity remains in dynamic equilibrium. Figure 5 As can be seen, the center of the ellipse represents drug A, the circles on the first ellipse closest to the center represent the indications and functions of drug A, and the squares represent other drugs with the same indications and functions.
[0085] Figure 6 This is a flowchart illustrating the text display of index landing according to the embodiments of this disclosure, such as... Figure 6 As shown:
[0086] S61, retrieve the search text for the preset domain;
[0087] Search text includes, but is not limited to, content mention fields, mention date fields, drug fields, and disease fields.
[0088] S62, filter the search text using a blacklist based on preset conditions to obtain the blacklist filtered text;
[0089] The preset conditions include, but are not limited to, words related to food and other non-pharmaceutical products, as well as disease-related terms.
[0090] S63, obtain the whitelist filtered text by performing whitelist filtering on the blacklist filtered text;
[0091] S64: Determine if the whitelist filter text exists in the whitelist. If it exists, output the number of drug views and the number of disease views. If it does not exist, jump to S63.
[0092] By identifying texts with high pageview volumes within the aggregated text, the technology enables the highlighting or priority display of those texts.
[0093] Optionally, the target text can be displayed based on the initial classification results and pageviews, including: in response to a pageview count greater than or equal to a preset pageview count, displaying the target text according to the initial classification results and pageview count.
[0094] If the number of views is less than the preset number of views, at least one keyword will be displayed according to the initial category results and the number of views.
[0095] Specifically, the aforementioned preset pageview count can be used to represent a pre-set number of text views.
[0096] In one optional embodiment, if the number of views for each keyword is greater than or equal to the preset number of views, it indicates that the base number of views for each keyword is large. If the text corresponding to each keyword is displayed sequentially, the display space occupied will be large. In order to save display space, each keyword can be normalized. While reducing the display space of the text, it also ensures that the content of the text is not affected, thus achieving better focus on the text information.
[0097] In another optional embodiment, if the number of views for each keyword is less than the preset number of views, it indicates that the base number of views for each keyword is small, and there is no need to normalize the keywords. In order to clearly display the text information corresponding to each keyword, it is necessary to display the text information corresponding to each keyword according to the number of views, thus achieving clear display of the text information corresponding to each keyword even when the number of views is small.
[0098] Optionally, determining the initial classification result of the search text includes: retrieving the scene classification model corresponding to the preset domain; classifying the search text using the scene classification model to obtain the initial classification result, wherein the scene classification model is obtained by model distillation of the preset model based on the first training data.
[0099] Specifically, the scene classification model described above can be used to represent a model for classifying search text.
[0100] The first training data mentioned above can be used to represent text data with a large number of bases in a preset domain.
[0101] The aforementioned preset model can be used to represent a pre-defined model for classification, including but not limited to the Wenxin model.
[0102] In one optional embodiment, since the Wenxin Big Model originates from and serves the industry, and is an industry-level knowledge growth big model, it is necessary to match the industry it serves with the corresponding preset domain. That is, the Wenxin Big Model is model distilled using a large amount of first training data to obtain a scene classification model that matches the preset domain. This facilitates the classification of search text based on the scene classification model, thus achieving accurate classification of search text.
[0103] Optionally, the method further includes: adjusting the preset model using the second training data to obtain an adjusted model, wherein the amount of data in the second training data is less than the amount of data in the first training data; and performing model distillation on the adjusted model based on the first training data to obtain a scene classification model.
[0104] Specifically, the second training data mentioned above can be used to represent text data with a small number of bases in a preset domain.
[0105] The aforementioned adjustment model can be used to represent the model obtained by fine-tuning the large text model using text data with a small base number.
[0106] The scene classification model described above can be used to represent a classification model obtained by distilling a model finely tuned using text data with a large base number. The purpose of distillation is to transform a large model into a smaller one, thereby improving the model's prediction efficiency.
[0107] In one alternative embodiment, the large text model is first adjusted by labeling a small amount of scene data to obtain an adjusted model adapted to a preset domain. However, since the adjusted model is a large model, in order to improve the prediction efficiency of the model, a large amount of text data is needed to distill the adjusted model to obtain a smaller model after distillation, which is the scene classification model mentioned above.
[0108] By adjusting the model and performing model distillation, a scene classification model is obtained, which achieves the technical effect of improving the classification prediction efficiency of the scene classification model.
[0109] Figure 7 This is a structural block diagram of entity mining, entity construction, and entity recognition according to embodiments of this disclosure. For example... Figure 7 As shown, the diagram includes three parts: entity mining (71), entity construction (72), and entity recognition (73). In entity mining (71), a recognition model (711) is used to identify massive amounts of data (712) to enhance entities (713). In entity construction (72), data is cleaned (722) by introducing data source providers (721), and manually verified (724) after data crawling (723) to build an entity database (725). In entity recognition (73), sample data is mined (733) based on entity matching (731) and rule matching (732), and the mined sample data is used for model training (734). During model training, model evaluation (735) and model optimization (736) are performed on the entity model.
[0110] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0111] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this disclosure.
[0112] This disclosure also provides a text display device for implementing the above embodiments and preferred embodiments, which will not be repeated hereafter. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0113] Figure 8 This is a structural block diagram of a text display device according to an embodiment of the present disclosure, such as... Figure 8 As shown, a text display device 800 includes: a determining module 80, configured to determine an initial classification result of the search text in response to receiving search text in a preset domain; an extraction module 81, configured to extract at least one keyword belonging to the preset domain from the search text and obtain the page views of at least one keyword; a processing module 82, configured to normalize at least one keyword in the search text to obtain target text, wherein the target text contains target keywords that have the same semantics as at least one keyword; and a display module 83, configured to display the target text according to the initial classification result and page views.
[0114] Optionally, the display module includes: a mapping unit for mapping pageviews using preset mapping logic to obtain mapped pageviews; and a first display unit for displaying the target text according to the initial classification results and the mapped pageviews.
[0115] Optionally, the display module includes: a filtering unit for filtering fields in the target text that do not meet preset conditions to obtain filtered text; an aggregation unit for aggregating the filtered text using an aggregation model to obtain aggregated text; and a second display unit for displaying the aggregated text according to the initial classification results and mapped pageviews.
[0116] Optionally, the display module includes: a sorting unit for sorting the aggregated text based on the mapped pageviews to obtain the sorting results; and a third display unit for displaying the aggregated text according to the initial classification results based on the sorting results.
[0117] Optionally, the display module includes: a fourth display unit, used to display the target text according to the initial classification results and the number of views in response to the number of views being greater than or equal to a preset number of views.
[0118] Optionally, the display module includes: a fifth display unit, used to display at least one keyword according to the initial classification results and the number of views in response to the view count being less than the preset view count.
[0119] Optionally, the determining module includes: a retrieval unit for retrieving a scene classification model corresponding to a preset domain; and a classification unit for classifying the search text using the scene classification model to obtain an initial classification result, wherein the scene classification model is obtained by model distillation of the preset model based on the first training data.
[0120] Optionally, the device further includes: an adjustment module for adjusting a preset model using second training data to obtain an adjusted model, wherein the amount of data in the second training data is less than that in the first training data; and a distillation module for performing model distillation on the adjusted model based on the first training data using an initial classification model to obtain a scene classification model.
[0121] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.
[0122] The method embodiments provided in this disclosure can be executed in a mobile terminal, computer terminal, or similar electronic device.
[0123] Figure 9 This is a hardware structure block diagram of a computer terminal (or mobile device) according to an embodiment of the present disclosure of a method for displaying text. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0124] like Figure 9 As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded into random access memory (RAM) 903 from storage unit 908. RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 909 is also connected to bus 904.
[0125] Multiple components in device 900 are connected to I / O interface 909, including: input unit 909, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0126] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as methods for screening users. For example, in some embodiments, the text display method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the text display method described above can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to perform methods for screening users by any other suitable means (e.g., by means of firmware).
[0127] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0128] According to embodiments of this disclosure, this disclosure also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps in any of the above method embodiments.
[0129] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0130] Optionally, in this disclosure, the processor described above can be configured to perform the following steps via a computer program:
[0131] S1, in response to receiving search text from a preset domain, determines the initial classification result of the search text;
[0132] S2, extract at least one keyword belonging to the preset domain from the search text, and obtain the page views of at least one keyword;
[0133] S3, normalize at least one keyword in the search text to obtain the target text, wherein the target text contains target keywords that have the same semantics as at least one keyword;
[0134] S4 displays the target text according to the initial classification results and pageviews.
[0135] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.
[0136] According to embodiments of the present disclosure, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are configured to perform the steps in any of the above method embodiments at runtime.
[0137] Optionally, in this embodiment, the non-volatile storage medium described above can be configured to store a computer program for performing the following steps:
[0138] S1, in response to receiving search text from a preset domain, determines the initial classification result of the search text;
[0139] S2, extract at least one keyword belonging to the preset domain from the search text, and obtain the page views of at least one keyword;
[0140] S3, normalize at least one keyword in the search text to obtain the target text, wherein the target text contains target keywords that have the same semantics as at least one keyword;
[0141] S4 displays the target text according to the initial classification results and pageviews.
[0142] Optionally, in this embodiment, the aforementioned non-transitory computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. More specific examples of readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0143] According to embodiments of this disclosure, a computer program product is also provided. Program code for implementing the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0144] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0145] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0146] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0147] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0148] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0149] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for displaying text, comprising: In response to receiving search text from a preset domain, determine the initial classification result of the search text; Extract at least one keyword belonging to the preset domain from the search text, and obtain the page views of the at least one keyword; Normalize at least one keyword in the search text to obtain target text, wherein the target text contains target keywords that have the same semantics as the at least one keyword; The target text is displayed according to the initial classification results and the number of views. The process of displaying the target text according to the initial classification result and the number of views includes: determining a preset mapping logic based on the number of views and a preset error value; mapping the number of views using the preset mapping logic to obtain a mapped number of views; filtering fields in the target text that do not meet preset conditions to obtain filtered text, wherein the preset conditions include: the field contains an entity and the semantics corresponding to the entity, the number of characters in the field is limited, and the punctuation of the field is standardized; aggregating fields with similar semantic information in the filtered text using an aggregation model to obtain aggregated text; sorting the aggregated text based on the mapped number of views to obtain a sorting result; and displaying the aggregated text according to the initial classification result based on the sorting result. Displaying the target text based on the initial classification result and the number of views further includes: in response to the number of views being greater than or equal to a preset number of views, displaying the target text according to the initial classification result and the number of views, wherein the initial classification result is obtained by classifying the search text based on part-of-speech or functional features; The method further includes: in response to the pageview count being less than a preset pageview count, displaying the at least one keyword according to the initial classification result and the pageview count.
2. The text display method according to claim 1, determining the initial classification result of the search text, includes: Retrieve the scene classification model corresponding to the preset domain; The search text is classified using the scene classification model to obtain the initial classification result, wherein the scene classification model is obtained by model distillation of a preset model based on the first training data.
3. The text display method according to claim 2, further comprising: The preset model is adjusted using the second training data to obtain an adjusted model, wherein the amount of data in the second training data is less than the amount of data in the first training data. Based on the first training data, the adjusted model is distilled to obtain the scene classification model.
4. A text display device, comprising: A determination module is used to determine the initial classification result of the search text in response to receiving search text in a preset domain; An extraction module is used to extract at least one keyword belonging to the preset domain from the search text and obtain the page views of the at least one keyword; The processing module is used to normalize at least one keyword in the search text to obtain target text, wherein the target text contains target keywords that have the same semantics as the at least one keyword; The display module is used to display the target text according to the initial classification results and the number of views; The display module includes: A mapping unit is used to determine a preset mapping logic based on the number of views and a preset error value; and to map the number of views using the preset mapping logic to obtain the mapped number of views. The first display unit is used to display the target text according to the initial classification result and the mapped pageviews; The display module also includes: A filtering unit is used to filter fields in the target text that do not meet preset conditions to obtain filtered text. The preset conditions include: the field contains an entity and the semantics corresponding to the entity, the number of characters in the field is limited, and the punctuation of the field is standardized. An aggregation unit is used to aggregate fields with similar semantic information in the filtered text using an aggregation model to obtain aggregated text; A sorting unit is used to sort the aggregated text based on the mapped pageviews to obtain a sorting result; The third display unit is used to display the aggregated text according to the initial classification result based on the sorting result; The display module also includes: The fourth display unit is used to display the target text according to the initial classification result and the number of views in response to the view count being greater than or equal to a preset view count. The initial classification result is obtained by classifying the search text based on part-of-speech or functional features. The display module also includes: The fifth display unit is used to display the at least one keyword according to the initial classification result and the number of views in response to the view count being less than the preset view count.
5. The apparatus according to claim 4, wherein the determining module comprises: The retrieval unit is used to retrieve the scene classification model corresponding to the preset domain; A classification unit is used to classify the search text using the scene classification model to obtain the initial classification result, wherein the scene classification model is obtained by model distillation of a preset model based on the first training data.
6. The apparatus according to claim 5, further comprising: An adjustment module is used to adjust the preset model using the second training data to obtain an adjusted model, wherein the amount of data in the second training data is less than that in the first training data. The distillation module is used to perform model distillation on the adjusted model based on the first training data using the initial classification model to obtain the scene classification model.
7. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-3.
8. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-3.
9. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-3.
Citation Information
Patent Citations
Public opinion information classification and assessment system
CN106257458A
Language model training method and device thereof, equipment and storage medium
CN113515948A