Semantic recognition method, device, electronic device and readable medium for text materials
By combining word segmentation matching and topic recognition models, the focus is optimized, solving the problems of manpower and time costs in the generation of shopping guide materials, achieving efficient and accurate semantic recognition, and reducing the impact of biased topics.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING WODONG TIANJUN INFORMATION TECH CO LTD
- Filing Date
- 2022-05-06
- Publication Date
- 2026-05-26
Smart Images

Figure CN114896982B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of text recognition technology, and more specifically, to a method, apparatus, electronic device, and readable medium for semantic recognition of text materials. Background Technology
[0002] Currently, as a type of text material used to guide customer consumption, shopping guide materials not only need to be vivid and interesting, but also need to meet the requirements of efficiency and diversification in order to attract consumers' attention and reading interest. However, writing shopping guide materials requires not only a lot of manpower but also a lot of time.
[0003] In related technologies, in order to generate more flexible shopping guide materials targeting different concerns and to cater to different consumer groups, the automatic writing of personalized shopping guide materials has emerged. The first step is based on a large amount of training data of intelligent models. Secondly, the process of collecting this training data requires extracting the concerns of shopping guide materials for any category of products. Finally, the shopping guide materials are classified according to the extracted concerns.
[0004] However, the process of classifying and training sales guide materials relies on extensive labeling, which not only leads to high manpower and time costs and is unable to cope with massive amounts of training materials, but also fails to comprehensively extract the key points of interest from the sales guide materials, and may even introduce biased information.
[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this disclosure is to provide a method, apparatus, electronic device, and readable medium for semantic recognition of text materials, which at least to some extent overcomes the problem of poor accuracy in semantic recognition of text materials due to limitations and defects in related technologies.
[0007] According to a first aspect of the present disclosure, a semantic recognition method for text material is provided, comprising: performing word segmentation on the text material to be processed to obtain word segments; performing matching processing on the word segments according to a preset correspondence between attention points and keywords; if the matching fails, inputting the text material into a trained topic recognition model, wherein the topic recognition model outputs the probability of the text material corresponding to each of the attention points; and determining the semantic topic of the text material based on the probability of the text material corresponding to each of the attention points.
[0008] In one exemplary embodiment of this disclosure, the semantic recognition method for text material further includes: if a match is successful, determining the semantic topic of the text material based on the focus corresponding to the word segmentation.
[0009] In one exemplary embodiment of this disclosure, before performing word segmentation on the text material to be processed, the method further includes: determining a topic recognition model to be trained; inputting samples of the text material and a preset number of attention points into the topic recognition model to be trained for training, and recording the consistency score for each training session; and determining the topic recognition model corresponding to the largest consistency score as the topic recognition model after initial training.
[0010] In one exemplary embodiment of this disclosure, before performing word segmentation on the text material to be processed, the method further includes: after completing the initial training of the topic recognition model, determining the number of samples of text material corresponding to the focus; and merging or splitting the focus according to the number of samples of the text material.
[0011] In one exemplary embodiment of this disclosure, merging or splitting the points of interest based on the number of samples of the text material includes: determining the size relationship between the number of samples of the text material and a preset number of samples; determining points of interest where the number of samples of the text material is less than the preset number of samples as first-type points of interest; and merging multiple first-type points of interest.
[0012] In one exemplary embodiment of this disclosure, merging or splitting the points of interest based on the number of samples of the text material further includes: determining the size relationship between the number of samples of the text material and a preset number of samples; determining points of interest whose number of samples of the text material is greater than or equal to the preset number of samples as second type of points of interest; merging the first type of points of interest into the second type of points of interest; segmenting the second type of points of interest into words; and splitting the second type of points of interest based on the word segmentation results of the second type of points of interest.
[0013] In one exemplary embodiment of this disclosure, before performing word segmentation on the text material to be processed, the method further includes: after completing the merging or splitting of the points of interest, updating the probability of the text material sample corresponding to the points of interest; and after completing the probability update of all the points of interest, determining that the topic recognition model training is complete.
[0014] In one exemplary embodiment of this disclosure, before performing word segmentation on the text material to be processed, the method further includes: after completing the training of the topic recognition model, performing clustering processing on the samples of the text material corresponding to the point of interest; and extracting keywords from the samples of the clustered text material based on word frequency.
[0015] In an exemplary embodiment of this disclosure, determining the semantic topic of the text material based on the probability of the text material corresponding to each of the points of interest includes: determining the probability of the text material corresponding to each of the points of interest; determining the point of interest with the highest probability as a first type of point of interest, and determining a first topic of the text material based on the first type of point of interest; determining the points of interest in the topic recognition model other than the first type of point of interest as a second type of point of interest; calculating the probability difference between the probability of the first type of point of interest and the probability of the second type of point of interest; determining whether the probability difference is less than or equal to a preset probability difference; if the probability difference is determined to be less than or equal to the preset probability difference, then determining a second topic of the text material based on the second type of point of interest, and determining the semantics of the text material based on the first topic and the second topic; if the probability difference is determined to be greater than the preset probability difference, then determining the semantics of the text material based on the first topic.
[0016] According to a second aspect of the present disclosure, a semantic recognition device for text materials is provided, comprising: a word segmentation module configured to perform word segmentation processing on the text material to be processed to obtain word segments; a matching module configured to perform matching processing on the word segments according to a preset correspondence between attention points and keywords; a recognition module configured to input the text material into a trained topic recognition model if the matching fails, wherein the topic recognition model outputs the probability that the text material corresponds to each of the attention points; and a determination module configured to determine the semantic topic of the text material based on the probability that the text material corresponds to each of the attention points.
[0017] In one exemplary embodiment of this disclosure, the determining module is further configured to: if a match is successful, determine the semantic theme of the text material based on the focus corresponding to the word segmentation.
[0018] In one exemplary embodiment of this disclosure, the determining module is further configured to: determine a topic recognition model to be trained; input the samples of the text material and the preset number of attention points into the topic recognition model to be trained for training, and record the consistency score of each training session; and determine the topic recognition model corresponding to the largest consistency score as the topic recognition model after initial training.
[0019] In an exemplary embodiment of this disclosure, the determining module is further configured to: after completing the initial training of the topic recognition model, determine the number of samples of text material corresponding to the point of interest; and merge or split the point of interest according to the number of samples of the text material.
[0020] In an exemplary embodiment of this disclosure, the determining module is further configured to: determine the size relationship between the number of samples of the text material and the preset number of samples; determine the points of interest where the number of samples of the text material is less than the preset number of samples as first type of points of interest; and merge multiple first type of points of interest.
[0021] In an exemplary embodiment of this disclosure, the determining module is further configured to: determine the size relationship between the number of samples of the text material and the preset number of samples; determine the points of interest whose number of samples of the text material is greater than or equal to the preset number of samples as the second type of points of interest; merge the first type of points of interest into the second type of points of interest; perform word segmentation on the second type of points of interest; and split the text material according to the word segmentation result of the second type of points of interest.
[0022] In an exemplary embodiment of this disclosure, the determining module is further configured to: after merging or splitting the points of interest, update the probability of the text material sample corresponding to the points of interest; and after updating the probability of all the points of interest, determine that the topic recognition model training is complete.
[0023] In an exemplary embodiment of this disclosure, the determining module is further configured to: after completing the training of the topic recognition model, perform clustering processing on the samples of text materials corresponding to the point of interest; and extract keywords from the samples of the clustered text materials based on word frequency.
[0024] In an exemplary embodiment of this disclosure, the determining module is further configured to: determine the probability of the text material corresponding to each of the points of interest; determine the point of interest with the highest probability as a first type of point of interest, and determine a first theme of the text material based on the first type of point of interest; determine the points of interest in the theme recognition model other than the first type of point of interest as a second type of point of interest; calculate the probability difference between the probability of the first type of point of interest and the probability of the second type of point of interest; determine whether the probability difference is less than or equal to a preset probability difference; if the probability difference is determined to be less than or equal to the preset probability difference, then determine a second theme of the text material based on the second type of point of interest, and determine the semantics of the text material based on the first theme and the second theme; if the probability difference is determined to be greater than the preset probability difference, then determine the semantics of the text material based on the first theme.
[0025] According to a third aspect of this disclosure, an electronic device is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to perform the method as described in any of the preceding methods based on instructions stored in the memory.
[0026] According to a fourth aspect of this disclosure, a computer-readable storage medium is provided having a program stored thereon that, when executed by a processor, implements a semantic recognition method for text material as described in any of the preceding claims.
[0027] In this embodiment, the text material is first segmented and matched. If the match is successful, the semantic topic of the text material can be directly determined based on the matched points of interest, which improves the efficiency and reliability of semantic processing. If the match fails, the probability of the text material corresponding to each point of interest can be further determined based on the topic recognition model. The semantic type of the text material is determined by combining the probability of each point of interest. This not only effectively improves the accuracy of identifying the semantic topic of the text material, but also reduces the probability of introducing biased topics into the text material for text materials with non-standard grammar through matching and topic recognition model processing.
[0028] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0029] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0030] Figure 1 A schematic diagram of an exemplary system architecture for a semantic recognition scheme of text materials to which embodiments of the present invention can be applied is shown;
[0031] Figure 2 This is a flowchart of a semantic recognition method for text materials according to an exemplary embodiment of this disclosure;
[0032] Figure 3 This is a flowchart of another semantic recognition method for text materials in an exemplary embodiment of this disclosure;
[0033] Figure 4 This is a flowchart of another semantic recognition method for text materials in an exemplary embodiment of this disclosure;
[0034] Figure 5 This is a flowchart of another semantic recognition method for text materials in an exemplary embodiment of this disclosure;
[0035] Figure 6 This is a flowchart of another semantic recognition method for text materials in an exemplary embodiment of this disclosure;
[0036] Figure 7This is a flowchart of another semantic recognition method for text materials in an exemplary embodiment of this disclosure;
[0037] Figure 8 This is a flowchart of another semantic recognition method for text materials in an exemplary embodiment of this disclosure;
[0038] Figure 9 This is a flowchart of another semantic recognition method for text materials in an exemplary embodiment of this disclosure;
[0039] Figure 10 This is a flowchart of a semantic recognition scheme for text materials in an exemplary embodiment of this disclosure;
[0040] Figure 11 This is a schematic diagram illustrating the attention optimization process of a semantic recognition scheme for text materials in an exemplary embodiment of this disclosure;
[0041] Figure 12 This is a schematic diagram illustrating the correspondence between attention points and keywords in a semantic recognition scheme for text materials according to an exemplary embodiment of this disclosure;
[0042] Figure 13 This is a schematic diagram of keyword matching in a semantic recognition scheme for text material according to an exemplary embodiment of this disclosure;
[0043] Figure 14 This is a schematic diagram illustrating the processing of a topic recognition model in a semantic recognition scheme for text materials according to an exemplary embodiment of this disclosure;
[0044] Figure 15 This is a block diagram of a semantic recognition device for text material in an exemplary embodiment of this disclosure;
[0045] Figure 16 This is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0046] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., can be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0047] Furthermore, the accompanying drawings are merely illustrative of this disclosure, and the same reference numerals in the drawings denote the same or similar parts, thus repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0048] Figure 1 A schematic diagram of an exemplary system architecture for a semantic recognition scheme of text materials to which embodiments of the present invention can be applied is shown.
[0049] like Figure 1 As shown, system architecture 100 may include one or more of terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.
[0050] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, there can be any number of terminal devices, networks, and servers. For example, server 105 could be a server cluster composed of multiple servers.
[0051] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Terminal devices 101, 102, and 103 can be various electronic devices with displays, including but not limited to smartphones, tablets, laptops, and desktop computers, etc.
[0052] In some embodiments, the semantic recognition method for text materials provided in this invention is generally executed by terminal 105, and correspondingly, the semantic recognition device for text materials is generally located in terminal device 103 (or terminal device 101 or 102). In other embodiments, some servers may have functions similar to those of terminal devices to execute this method. Therefore, the semantic recognition method for text materials provided in this invention is not limited to execution on a terminal device.
[0053] The exemplary embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0054] Figure 2 This is a flowchart of a semantic recognition method for text materials in an exemplary embodiment of this disclosure.
[0055] refer to Figure 2 Semantic recognition methods for text materials can include:
[0056] Step S202: Perform word segmentation on the text material to be processed to obtain word segments.
[0057] Step S204: Match the word segmentation according to the preset correspondence between the focus and the keywords.
[0058] Step S206: If the matching fails, the text material is input into the trained topic recognition model, and the topic recognition model outputs the probability that the text material corresponds to each of the points of interest.
[0059] Step S208: Determine the semantic theme of the text material based on the probability that the text material corresponds to each of the points of interest.
[0060] In this embodiment, the text material is first segmented and matched. If the match is successful, the semantic topic of the text material can be directly determined based on the matched points of interest, which improves the efficiency and reliability of semantic processing. If the match fails, the probability of the text material corresponding to each point of interest can be further determined based on the topic recognition model. The semantic type of the text material is determined by combining the probability of each point of interest. This not only effectively improves the accuracy of identifying the semantic topic of the text material, but also reduces the probability of introducing biased topics into the text material for text materials with non-standard grammar through matching and topic recognition model processing.
[0061] In the above embodiments, the focus refers to the characteristics of a product described in different dimensions. A product contains multiple focus points, just as a mobile phone contains multiple dimensions such as screen size, color, camera function, and chip.
[0062] The following section provides a detailed explanation of each step in the semantic recognition method for text materials.
[0063] In one exemplary embodiment of this disclosure, the method further includes: if a match is successful, determining the semantic theme of the text material based on the focus corresponding to the word segmentation.
[0064] In one exemplary embodiment of this disclosure, such as Figure 3 As shown, before performing word segmentation on the text material to be processed, the following steps are also included:
[0065] Step S302: Determine the topic recognition model to be trained.
[0066] Step S304: Input the text material samples and the preset number of attention points into the topic recognition model to be trained for training, and record the consistency score for each training session.
[0067] Step S306: The topic recognition model corresponding to the largest consistency score is determined as the topic recognition model after initial training.
[0068] In the above embodiments, the topic recognition model to be trained is an unsupervised text classification. In practical applications, the sample labels cannot be known in advance in many cases. Various application problems can be solved based on unlabeled training samples.
[0069] In one exemplary embodiment of this disclosure, such as Figure 4 As shown, before performing word segmentation on the text material to be processed, the following steps are also included:
[0070] Step S402: After completing the initial training of the topic recognition model, determine the number of samples of text material corresponding to the point of interest.
[0071] Step S404: Merge or split the points of interest based on the number of samples of the text material.
[0072] In the above embodiments, merging or splitting the points of interest by the number of samples of the text material can improve the accuracy and reliability of the points of interest, thereby improving the accuracy and reliability of matching text material according to the points of interest.
[0073] In one exemplary embodiment of this disclosure, such as Figure 5 As shown, merging or splitting the points of interest based on the number of samples of the text material includes:
[0074] Step S502: Determine the size relationship between the number of samples of the text material and the preset number of samples.
[0075] Step S504: Determine the points of interest where the number of samples of the text material is less than the preset number of samples as the first type of points of interest.
[0076] Step S506: Merge multiple first-type concerns.
[0077] In one exemplary embodiment of this disclosure, as shown in 6, merging or splitting the points of interest based on the number of samples of the text material further includes:
[0078] Step S602: Determine the size relationship between the number of samples of the text material and the preset number of samples.
[0079] Step S604: Determine the points of interest where the number of samples of the text material is greater than or equal to the preset number of samples as the second type of points of interest.
[0080] Step S606: Merge the first type of concern into the second type of concern.
[0081] Step S608: Segment the second type of interest into words.
[0082] Step S610: Segment the words based on the word segmentation results of the second type of interest.
[0083] In the above embodiments, by adjusting the size relationship between the number of text material samples and the preset number of samples, it is possible not only to merge fewer concerns to reduce the number of concerns to be maintained and the operational pressure, but also to split larger concerns, that is, to refine the division of concerns, making the description of concerns clearer, more accurate and more diverse, thereby improving the accuracy of matching concerns with text materials and further reducing the probability of introducing biased topics.
[0084] In one exemplary embodiment of this disclosure, such as Figure 7 As shown, before performing word segmentation on the text material to be processed, the following steps are also included:
[0085] Step S702: After merging or splitting the points of interest, update the probability of the text material sample corresponding to the points of interest.
[0086] Step S704: After completing the probability update of all the points of interest, determine that the topic recognition model training is complete.
[0087] In the above embodiment, if the probabilities of a text material belonging to focus 1 and 2 are p1 and p2 respectively, then the probability of belonging to focus 1 in the newly optimized topic recognition model is p1+p2.
[0088] In the above embodiment, if the probability of a material belonging to attention point 4 is p4, the probabilities of the optimized attention points 2 and 4 are (1 / 2)*p4 and (1 / 2)*p4 respectively. Then, based on this probability, the clustering result of the text material can be obtained again.
[0089] In one exemplary embodiment of this disclosure, such as Figure 8 As shown, before performing word segmentation on the text material to be processed, the following steps are also included:
[0090] Step S802: After completing the training of the topic recognition model, cluster the samples of text materials corresponding to the points of interest.
[0091] Step S804: Extract keywords from the clustered text samples based on word frequency.
[0092] In one exemplary embodiment of this disclosure, such as Figure 9 As shown, determining the semantic theme of the text material based on the probability that the text material corresponds to each of the aforementioned points of interest includes:
[0093] Step S902: Determine the probability that the text material corresponds to each of the points of interest.
[0094] Step S904: Determine the most probable point of interest as the first type of point of interest, and determine the first theme of the text material based on the first type of point of interest.
[0095] Step S906: Identify the points of interest in the topic recognition model other than the first type of points of interest as the second type of points of interest.
[0096] Step S908: Calculate the probability difference between the probability of the first type of concern and the probability of the second type of concern.
[0097] Step S910: Determine whether the probability difference is less than or equal to the preset probability difference. If yes, proceed to step S912; otherwise, proceed to step S914.
[0098] Step S912: If the probability difference is determined to be less than or equal to the preset probability difference, then the second topic of the text material is determined according to the second type of concern, and the semantics of the text material is determined according to the first topic and the second topic.
[0099] In the above embodiments, by determining the second topic of the text material based on the second type of concern when the probability difference is less than or equal to the preset probability difference, and determining the semantics of the text material based on the first topic and the second topic, the comprehensiveness, accuracy and reliability of the semantic topic of the text material are improved.
[0100] Step S914: If it is determined that all probability differences are greater than the preset probability difference, then the semantics of the text material is determined according to the first topic.
[0101] In the above embodiments, the preset probability difference can be 5%, 10%, and 15%, etc., but is not limited to this.
[0102] In one exemplary embodiment of this disclosure, such as Figure 10 As shown, the semantic recognition scheme for text materials includes two stages: the attention point extraction process 1002 and the shopping guide material classification process 1004. The input of the shopping guide material classification process 1004 depends on the output after attention point matching.
[0103] In one exemplary embodiment of this disclosure, during the attention point extraction process 1002, all shopping guide materials are first modeled using a text topic recognition model to extract attention points, and the topic recognition model is saved. Then, the extracted initial attention points and the topic recognition model are optimized and selected to obtain the final attention points. After obtaining the attention points, keywords related to each attention point are extracted and saved.
[0104] In one exemplary embodiment of this disclosure, in the shopping guide material classification process 1004, the materials to be classified are first classified according to keyword matching. The keywords in this step are derived from the keywords saved in step 1. For materials that do not successfully match the keywords, they are then sent to the topic recognition model for soft classification.
[0105] In one exemplary embodiment of this disclosure, two points of interest are first extracted from the text material: “body appearance” and “charging and battery life”. Then, all materials are classified to determine whether each material belongs to “body appearance” or “charging and battery life”.
[0106] In one exemplary embodiment of this disclosure, the topic recognition model is trained as follows: the input to the training module of the topic recognition model is all text materials and a given number of attention points, and the output is the trained model and the extracted attention points.
[0107] In one exemplary embodiment of this disclosure, the topic identification model is an LDA (Latent Dirichlet Allocation, three-layer Bayesian probability) model. Since the optimal number of attention points is unknown beforehand, the number of categories is initially set to 2-50, and the performance of the LDA model with each category number is tested. The optimal number of categories is selected based on the model's coherence score, and the LDA model with that number of categories is saved. The saved LDA model can then cluster existing materials. The LDA model automatically groups materials with similar content into one category and determines the initial attention points based on some high-frequency words in the materials within each category.
[0108] For example, when training a topic recognition model on text materials from a mobile phone, it was found that the coherent score was the highest when the number of attention points was set to 13. So the model corresponding to the number of attention points was saved. This model divided the materials into 13 categories, namely "performance-related, appearance-related, shooting-related, beauty-related, charging-related, battery capacity-related, etc."
[0109] In one exemplary embodiment of this disclosure, the topic recognition model optimization process includes: the input of the topic recognition model optimization module is the initial topic recognition model and its initial focus points from the previous step, and the output is the optimized topic recognition model and the optimized focus points.
[0110] In one exemplary embodiment of this disclosure, the focus points are optimized, i.e., merged and split, based on the number of materials in each focus point of the currently saved topic recognition model.
[0111] In one exemplary embodiment of this disclosure, concerns involving a small amount of material are merged into similar concern categories, while concerns involving an excessive amount of material are split into two concerns. The optimized concerns are then used as the final concerns. In this case, the topic recognition model is also effectively optimized. For example, concerns related to "charging" and concerns related to "battery capacity" can be merged.
[0112] In one exemplary embodiment of this disclosure, such as Figure 11 As shown, the initial focus points 1102 of the topic recognition model include “1”, “2”, “3”, “4” and “5”. After merging and splitting, focus points 1 and 2 are merged, and focus point 4 is split. The initial focus points 1104 of the topic recognition model include “A”, “B”, “3”, “C” and “5”.
[0113] The diagram below illustrates how to optimize the topic identification model LDA and how to calculate the probability value of each piece of text corresponding to each type of interest after optimization. This probability helps the topic identification model cluster the materials.
[0114] In one exemplary embodiment of this disclosure, if the probabilities of a text element belonging to attention point 1 and 2 are p1 and p2 respectively, then the probability of belonging to attention point 1 in the newly optimized model is p1+p2. If the probability of a text element belonging to attention point 4 is p4, then the probabilities of attention points 2 and 4 in the optimized model are (1 / 2)*p4 and (1 / 2)*p4 respectively. The clustering results of the text elements can then be obtained again based on these probabilities.
[0115] Keyword Extraction: The keyword extraction module takes an optimized topic recognition model and focus points as input, and outputs several keywords corresponding to each focus point. Based on the optimized topic recognition model and focus points, a clustering result of the existing materials can be obtained, meaning that each focus point category contains materials belonging to that focus point.
[0116] In one exemplary embodiment of this disclosure, such as Figure 11 As shown, the topic recognition model includes five concerns: "A", "B", "3", "C", and "5". For each concern category, high-frequency words are extracted from the materials it contains. For example, for the concern "charging and battery life", high-frequency words such as "fast charging", "charging", "battery", "long-lasting", "charge", "phone", and "of" might be extracted. Then, some meaningless or irrelevant words are removed (such as "of" and "phone" in the example). This yields each concern category, and the final keywords are denoted as "a", "b", "d", "c", and "e", respectively.
[0117] Categorization of shopping guide materials: Transfer the topic recognition model (including focus points) and keywords saved in the previous step to this step to categorize any newly emerging text materials.
[0118] Keyword Matching Classification: The input for Keyword Matching Classification is a new text document, and the output is the category of interest to which that text document belongs. First, the input text document is segmented into words, and then keywords are matched against each interest point in turn. The interest points corresponding to the matched keywords are taken as the interest category of the text document.
[0119] In one exemplary embodiment of this disclosure, such as Figure 13 As shown, the semantic recognition scheme for text materials includes the following steps:
[0120] In step S1302, the input text is "This phone has an ultra-long standby time and can be fully charged in 2 hours".
[0121] Step S1304: Segment the text material, and the segmented words include "model", "mobile phone", "ultra", "long", "standby", "charging", "hours", "fully charged" and "resurrected", etc.
[0122] Step S1306: Construct focus 2 as "body appearance", and the keywords include body, appearance, shell, and material.
[0123] Step S1308: Construct focus 1 as "charging and battery life", and the keywords include battery, point, battery life, durability, and charging.
[0124] Among them, "standby" successfully matches with "durability", which is one of the keywords included in the focus corresponding to "charging and battery life" in focus 1, and "charging" successfully matches with "charging", which is one of the keywords included in the focus corresponding to "charging and battery life" in focus 1. Therefore, the theme semantics of this text material is identified as "charging and battery life". If there is no successful match, it enters the next link; if there is a successful match, the classification is completed.
[0125] In an exemplary embodiment of the present disclosure, as Figure 14 shown, the classification of the theme recognition model includes the following steps:
[0126] Step S1402: The input copy of the classification of the theme recognition model is the shopping guide material that was not successfully matched in the previous step, and the output is the focus category to which this text material belongs. Among them, the input copy is: "This mobile phone is small and exquisite, easy to carry".
[0127] Step S1404: Input the text material into the optimized LDA theme recognition model in the first stage, and this model will output the probability that this material belongs to each focus category.
[0128] In an exemplary embodiment of the present disclosure, the focus with the maximum probability can be taken as the final focus of this text material.
[0129] Step S1404: The LDA theme recognition model will output the probability that this text material is assigned to each focus. The probability corresponding to focus 1 (charging and battery life) is 0.15, and the probability corresponding to focus 2 (body shell) is 0.85. The "body appearance" with the maximum probability of 0.85 can be selected as the final focus category of this text material.
[0130] In an exemplary embodiment of the present disclosure, if the difference in the probabilities of the top 2 focuses is less than a preset probability difference (threshold), it means that this text material contains multiple focuses, and both focus categories can be assigned to this text material.
[0131] Corresponding to the above method embodiments, this disclosure also provides a semantic recognition device for text materials, which can be used to execute the above method embodiments.
[0132] Figure 15 This is a block diagram of a semantic recognition device for text material in an exemplary embodiment of this disclosure.
[0133] refer to Figure 15 The semantic recognition device 1500 for text materials may include:
[0134] The word segmentation module 1502 is set to perform word segmentation on the text material to be processed in order to obtain the word segments.
[0135] The matching module 1504 is configured to perform matching processing on the word segmentation according to the preset correspondence between the focus and the keywords.
[0136] The recognition module 1506 is configured to input the text material into a trained topic recognition model if the matching fails, and the topic recognition model outputs the probability that the text material corresponds to each of the points of interest.
[0137] The determination module 1508 is configured to determine the semantic theme of the text material based on the probability that the text material corresponds to each of the points of interest.
[0138] In one exemplary embodiment of this disclosure, the determining module 1506 is further configured to: if a match is successful, determine the semantic topic of the text material based on the focus corresponding to the word segmentation.
[0139] In an exemplary embodiment of this disclosure, the determining module 1506 is further configured to: determine the topic recognition model to be trained; input the samples of the text material and the preset number of attention points into the topic recognition model to be trained for training, and record the consistency score of each training session; and determine the topic recognition model corresponding to the largest consistency score as the topic recognition model after initial training.
[0140] In an exemplary embodiment of this disclosure, the determining module 1506 is further configured to: after completing the initial training of the topic recognition model, determine the number of samples of text material corresponding to the point of interest; and merge or split the point of interest according to the number of samples of the text material.
[0141] In an exemplary embodiment of this disclosure, the determining module 1506 is further configured to: determine the size relationship between the number of samples of the text material and the preset number of samples; determine the points of interest where the number of samples of the text material is less than the preset number of samples as first type of points of interest; and merge multiple first type of points of interest.
[0142] In an exemplary embodiment of this disclosure, the determining module 1506 is further configured to: determine the size relationship between the number of samples of the text material and the preset number of samples; determine the points of interest whose number of samples of the text material is greater than or equal to the preset number of samples as the second type of points of interest; merge the first type of points of interest into the second type of points of interest; perform word segmentation on the second type of points of interest; and perform splitting based on the word segmentation result of the second type of points of interest.
[0143] In an exemplary embodiment of this disclosure, the determining module 1506 is further configured to: after completing the merging or splitting of the points of interest, update the probability of the text material sample corresponding to the points of interest; and after completing the probability update of all the points of interest, determine that the topic recognition model training is complete.
[0144] In an exemplary embodiment of this disclosure, the determining module 1506 is further configured to: after completing the training of the topic recognition model, perform clustering processing on the samples of text materials corresponding to the points of interest; and extract keywords from the samples of the clustered text materials based on word frequency.
[0145] In an exemplary embodiment of this disclosure, the determining module 1506 is further configured to: determine the probability of the text material corresponding to each of the points of interest; determine the point of interest with the highest probability as a first type of point of interest, and determine a first theme of the text material based on the first type of point of interest; determine the points of interest in the theme recognition model other than the first type of point of interest as a second type of point of interest; calculate the probability difference between the probability of the first type of point of interest and the probability of the second type of point of interest; determine whether the probability difference is less than or equal to a preset probability difference; if the probability difference is determined to be less than or equal to the preset probability difference, then determine a second theme of the text material based on the second type of point of interest, and determine the semantics of the text material based on the first theme and the second theme; if the probability difference is determined to be greater than the preset probability difference, then determine the semantics of the text material based on the first theme.
[0146] Since the functions of the apparatus 1500 have been described in detail in their respective method embodiments, they will not be repeated here.
[0147] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0148] In an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described method is also provided.
[0149] Those skilled in the art will understand that various aspects of the present invention can be implemented as systems, methods, or program products. Therefore, various aspects of the present invention can be specifically implemented in the following forms: entirely hardware implementations, entirely software implementations (including firmware, microcode, etc.), or implementations combining hardware and software aspects, collectively referred to herein as “circuits,” “modules,” or “systems.”
[0150] The following reference Figure 16 To describe an electronic device 1600 according to this embodiment of the present invention. Figure 16 The electronic device 1600 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of the present invention.
[0151] like Figure 16 As shown, the electronic device 1600 is manifested in the form of a general-purpose computing device. The components of the electronic device 1600 may include, but are not limited to: at least one processing unit 1610, at least one storage unit 1620, and a bus 1630 connecting different system components (including storage unit 1620 and processing unit 1610).
[0152] The storage unit stores program code that can be executed by the processing unit 1610, causing the processing unit 1610 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of the present invention. For example, the processing unit 1610 can perform the method shown in the embodiments of this disclosure.
[0153] Storage unit 1620 may include readable media in the form of volatile storage units, such as random access memory (RAM) 16201 and / or cache memory 16202, and may further include read-only memory (ROM) 16203.
[0154] Storage unit 1620 may also include a program / utility 16204 having a set (at least one) program module 16205, such program module 16205 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.
[0155] Bus 1630 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0156] Electronic device 1600 can also communicate with one or more external devices 1640 (e.g., keyboard, pointing device, Bluetooth device, etc.), one or more devices that enable a user to interact with electronic device 1600, and / or any device that enables electronic device 1600 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 1650. Furthermore, electronic device 1600 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 1660. As shown, network adapter 1660 communicates with other modules of electronic device 1600 via bus 1630. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 1600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0157] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0158] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible embodiments, various aspects of the invention may also be implemented as a program product comprising program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps of the various exemplary embodiments of the invention described in the "Exemplary Methods" section of this specification.
[0159] The program product for implementing the above-described method according to embodiments of the present invention may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0160] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0161] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0162] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0163] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0164] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0165] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and concept of this disclosure are indicated by the claims.
Claims
1. A semantic recognition method for text materials, characterized in that, include: Determine the topic recognition model to be trained; The text material samples and the preset number of attention points are input into the topic recognition model to be trained for training, and the consistency score of each training session is recorded. The topic recognition model corresponding to the highest consistency score is determined as the topic recognition model after initial training. After completing the initial training of the topic recognition model, determine the number of text material samples corresponding to the point of interest; The focus points are merged or split based on the number of samples of the text material; After merging or splitting the points of interest, update the probability of the text material sample corresponding to the points of interest; After completing the probability update of all the aforementioned points of interest, the training of the topic recognition model is determined to be complete. The text material to be processed is segmented into words to obtain the segmented words; The word segmentation is matched according to the preset correspondence between the focus and the keywords; If the matching fails, the text material is input into the trained topic recognition model, and the topic recognition model outputs the probability that the text material corresponds to each of the points of interest; The semantic theme of the text material is determined based on the probability that the text material corresponds to each of the aforementioned points of interest.
2. The semantic recognition method for text materials as described in claim 1, characterized in that, Also includes: If a match is successful, the semantic theme of the text material is determined based on the focus corresponding to the word segmentation.
3. The semantic recognition method for text materials as described in claim 1, characterized in that, Merging or splitting the points of interest based on the number of samples of the text material includes: Determine the relationship between the number of samples of the text material and the preset number of samples; The points of interest where the number of samples of the text material is less than the preset number of samples are identified as the first type of points of interest. Multiple concerns of the first type are merged.
4. The semantic recognition method for text materials as described in claim 3, characterized in that, Merging or splitting the points of interest based on the number of samples of the text material also includes: Determine the relationship between the number of samples of the text material and the preset number of samples; The points of interest whose number of samples of the text material is greater than or equal to the preset number of samples are identified as the second type of points of interest; The first category of concerns is merged into the second category of concerns; Segment the second type of concern into words; The words are segmented based on the word segmentation results of the second type of concern.
5. The semantic recognition method for text materials as described in any one of claims 1-4, characterized in that, Before performing word segmentation on the text material to be processed, the following is also included: After training the topic recognition model, the samples of text materials corresponding to the points of interest are clustered. Keyword extraction is performed on samples of clustered text materials based on word frequency.
6. The semantic recognition method for text materials as described in claim 1, characterized in that, Determining the semantic theme of the text material based on the probability that the text material corresponds to each of the aforementioned points of interest includes: Determine the probability that the text material corresponds to each of the aforementioned points of interest; The point of interest with the highest probability is identified as the first type of point of interest, and the first theme of the text material is determined based on the first type of point of interest. The points of interest in the topic recognition model other than the first type of points of interest are identified as the second type of points of interest. Calculate the probability difference between the probability of the first type of concern and the probability of the second type of concern; Determine whether the probability difference is less than or equal to a preset probability difference; If the probability difference is determined to be less than or equal to the preset probability difference, then the second theme of the text material is determined based on the second type of concern, and the semantics of the text material is determined based on the first theme and the second theme. If it is determined that all the probability differences are greater than the preset probability difference, then the semantics of the text material are determined based on the first theme.
7. A semantic recognition device for text materials, characterized in that, include: The module is set to determine the topic recognition model to be trained. The text material samples and the preset number of attention points are input into the topic recognition model to be trained for training, and the consistency score of each training session is recorded. The topic recognition model corresponding to the highest consistency score is determined as the topic recognition model after initial training. After completing the initial training of the topic recognition model, determine the number of text material samples corresponding to the point of interest; The focus points are merged or split based on the number of samples of the text material; After merging or splitting the points of interest, update the probability of the text material sample corresponding to the points of interest; After completing the probability update of all the aforementioned points of interest, the training of the topic recognition model is determined to be complete. The word segmentation module is set up to perform word segmentation on the text material to be processed, in order to obtain the word segments; The matching module is configured to perform matching processing on the word segments according to the preset correspondence between the focus and the keywords; The recognition module is configured to input the text material into a trained topic recognition model if a match fails, and the topic recognition model outputs the probability that the text material corresponds to each of the points of interest. The determining module is further configured to determine the semantic theme of the text material based on the probability that the text material corresponds to each of the points of interest.
8. The semantic recognition device for text materials as described in claim 7, characterized in that, The determining module is further configured as follows: If a match is successful, the semantic theme of the text material is determined based on the focus corresponding to the word segmentation.
9. The semantic recognition device for text materials as described in claim 7, characterized in that, The determining module is further configured as follows: Determine the relationship between the number of samples of the text material and the preset number of samples; The points of interest where the number of samples of the text material is less than the preset number of samples are identified as the first type of points of interest. Multiple concerns of the first type are merged.
10. The semantic recognition device for text materials as described in claim 9, characterized in that, The determining module is further configured as follows: Determine the relationship between the number of samples of the text material and the preset number of samples; The points of interest whose number of samples of the text material is greater than or equal to the preset number of samples are identified as the second type of points of interest; The first category of concerns is merged into the second category of concerns; Segment the second type of concern into words; The words are segmented based on the word segmentation results of the second type of concern.
11. The semantic recognition device for text materials as described in any one of claims 7-10, characterized in that, The determining module is further configured as follows: After training the topic recognition model, the samples of text materials corresponding to the points of interest are clustered. Keyword extraction is performed on samples of clustered text materials based on word frequency.
12. The semantic recognition device for text materials as described in any one of claims 7-10, characterized in that, The determining module is further configured as follows: Determine the probability that the text material corresponds to each of the aforementioned points of interest; The point of interest with the highest probability is identified as the first type of point of interest, and the first theme of the text material is determined based on the first type of point of interest. The points of interest in the topic recognition model other than the first type of points of interest are identified as the second type of points of interest. Calculate the probability difference between the probability of the first type of concern and the probability of the second type of concern; Determine whether the probability difference is less than or equal to a preset probability difference; If the probability difference is determined to be less than or equal to the preset probability difference, then the second theme of the text material is determined based on the second type of concern, and the semantics of the text material is determined based on the first theme and the second theme. If it is determined that all the probability differences are greater than the preset probability difference, then the semantics of the text material are determined based on the first theme.
13. An electronic device, characterized in that, include: Memory; as well as A processor coupled to the memory, the processor being configured to execute a semantic recognition method for text material as described in any one of claims 1-6 based on instructions stored in the memory.
14. A computer-readable storage medium having a program stored thereon that, when executed by a processor, implements a semantic recognition method for text material as described in any one of claims 1-6.