Method, device, electronic device and storage medium for determining tag information
The semantic matching technology obtains the inside and outside of the text of the pending text, and combines the characteristics of candidate keywords, the problem of insufficient label information in the existing technology is solved, and the comprehensive expansion of user label information is achieved, which is suitable for the determination of label information in artificial intelligence and big data environments.
Patent Information
- Application Number
- CN202110126378.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-29
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2041-01-29
AI Technical Summary
The existing methods of determining the tags that users are interested in are insufficient information rich, resulting in insufficient comprehensive tag information.
By obtaining the semantic matching of keywords in text and keywords outside the text of the to be processed, and combining the keyword characteristics of candidate keywords, the user's tag information is determined, which expands the richness and comprehensiveness of tag information.
It improves the richness and comprehensiveness of user tag information, ensures that tag information is more complete, and is suitable for natural language processing and machine learning in the field of artificial intelligence, especially in cloud computing and big data environments.
Smart Images

Figure CN114817697B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular to a method, device, electronic device, and storage medium for determining tag information. Background Art
[0002] Currently, when mining corresponding tags of interest for users, it is necessary to process the text corresponding to the user's information of interest and extract keywords from the text to determine the user's tags of interest based on the extracted keywords. Although this method can determine the user's tags of interest to a certain extent, the determined tags of interest are not rich in information to a certain extent. In other words, the information richness of the existing method of determining user tags of interest still needs to be improved. Summary of the Invention
[0003] The embodiments of the present application provide a method, device, electronic device, and storage medium for determining tag information, which expand the user's tag information and improve the richness and comprehensiveness of the user's tag information.
[0004] In one aspect, an embodiment of the present application provides a method for determining tag information, the method comprising:
[0005] Obtaining the text to be processed corresponding to the information of interest of the object, and extracting keywords from the text to be processed;
[0006] Extracting text features of the above text to be processed;
[0007] Obtaining keyword features of each candidate keyword, wherein the candidate keyword is a keyword in a keyword thesaurus;
[0008] Based on the semantic matching between the text features of the text to be processed and the features of each of the keyword, determining the non-textual keywords corresponding to the text to be processed from the candidate keywords;
[0009] The tag information of the object is determined based on text keywords, wherein the text keywords include the in-text keywords and the out-text keywords.
[0010] In one aspect, an embodiment of the present application provides a device for determining tag information, the device comprising:
[0011] A keyword processing module is used to obtain the text to be processed corresponding to the information of interest of the object and extract keywords from the text to be processed;
[0012] A text feature processing module, used to extract text features of the text to be processed;
[0013] The keyword processing module is configured to obtain keyword features of each candidate keyword and determine, from each candidate keyword, non-textual keywords corresponding to the text to be processed based on the semantic matching between the text features of the text to be processed and each keyword feature, wherein the candidate keywords are keywords in the keyword lexicon;
[0014] The tag information determination module is used to determine the tag information of the object based on text keywords, wherein the text keywords include the keywords within the text and the keywords outside the text.
[0015] In an optional embodiment, the keyword processing module is further configured to:
[0016] Extracting keyword features of each of the candidate keywords, and storing the keyword features of each of the candidate keywords in a keyword thesaurus;
[0017] When obtaining the keyword features of each candidate word, the pre-stored keyword features of each candidate keyword are obtained from the keyword database.
[0018] In an optional embodiment, the text features of the to-be-processed text include first word features of each word contained in the to-be-processed text, and the keyword features include second word features of each word contained in the candidate keywords; the keyword processing module is configured to:
[0019] For any of the above keyword features, the semantic matching degree between the text features of the above text to be processed and the above keyword features is determined in the following manner:
[0020] Determine the similarity between each of the first word features and each of the second word features, respectively, wherein the pairwise features include a first word feature and a second word feature;
[0021] Based on the similarity between each of the two features, the semantic matching degree between the text features of the to-be-processed text and the keyword features is determined.
[0022] In an optional embodiment, the keyword processing module is configured to:
[0023] For any first word feature, determine the weights corresponding to the second word features based on the first similarities corresponding to the first word feature, and perform weighted summation on the second word features based on the weights corresponding to the second word features to obtain an updated first word feature;
[0024] For any second word feature, determine the weight corresponding to each of the first word features based on each second similarity corresponding to the second word feature, and perform weighted summation of each of the first word features based on the weight corresponding to each of the first word features to obtain an updated second word feature;
[0025] Based on the updated first word features and the updated second word features, the semantic matching degree between the text features of the to-be-processed text and the keyword features is determined.
[0026] In an optional embodiment, the keyword processing module is configured to:
[0027] For any of the first word features, based on the correlation between the first word feature and the updated first word feature, obtain a local word feature corresponding to the first word feature; based on the first word feature, the updated first word feature, and the local word feature corresponding to the first word feature, obtain a global word feature corresponding to the first word feature;
[0028] For any of the second word features, based on the correlation between the second word feature and the updated second word feature, obtain the local word feature corresponding to the second word feature; based on the second word feature, the updated second word feature and the local word feature corresponding to the second word feature, obtain the global word feature corresponding to the second word feature;
[0029] Based on the global word features corresponding to each first word feature and the global word features corresponding to each second word feature, the semantic matching degree between the text features of the to-be-processed text and the keyword features is determined.
[0030] In an optional embodiment, the keyword processing module is configured to:
[0031] Determining a first difference feature between the first word feature and the updated first word feature;
[0032] Determining a first similarity feature between the first word feature and the updated first word feature;
[0033] The local word features corresponding to the first word feature include a first difference feature and a first similarity feature;
[0034] determining a second difference feature between the second word feature and the updated second word feature;
[0035] determining a second similarity feature between the second word feature and the updated second word feature;
[0036] Among them, the local word features corresponding to the above-mentioned second word features include a second difference feature and a second similarity feature.
[0037] In an optional embodiment, the text feature processing module is configured to:
[0038] Based on the above text to be processed, extracting text features of the above text to be processed through a feature extraction model;
[0039] The above-mentioned determination of the non-textual keywords corresponding to the above-mentioned text to be processed from the above-mentioned candidate keywords based on the semantic matching between the text features of the above-mentioned text to be processed and the features of each of the above-mentioned keywords includes:
[0040] Based on the text features of the text to be processed and the features of each of the above keywords, the non-text keywords corresponding to the text to be processed are determined from the candidate keywords through a semantic matching model.
[0041] In one aspect, an embodiment of the present application provides an information recommendation method, the method comprising:
[0042] Obtaining the information to be recommended and the text of the information to be recommended;
[0043] Obtaining label information of each candidate recommendation object, wherein the label information of the candidate recommendation object is determined by using any of the optional embodiments of the label information determination method described above;
[0044] Determine a target recommendation object from the candidate recommendation objects based on the matching degree between the text and the tag information;
[0045] Recommend the above-mentioned information to be recommended to the above-mentioned target recommendation object.
[0046] In one aspect, an embodiment of the present application provides an information recommendation device, comprising:
[0047] The information processing module to be recommended is used to obtain the information to be recommended and the text of the information to be recommended;
[0048] A tag information acquisition module, configured to acquire tag information of each candidate recommendation object, wherein the tag information is determined using the tag information determination method provided in any optional embodiment of the present application;
[0049] A target recommendation object determination module is used to determine a target recommendation object from the candidate recommendation objects based on the matching degree between the text and the tag information;
[0050] The information recommendation module is used to recommend the above-mentioned information to be recommended to the above-mentioned target recommendation object.
[0051] In an optional embodiment, the target recommendation object determination module is further configured to perform at least one of the following:
[0052] For any of the candidate recommendation objects, obtaining text features of the text and object label features of the label information of the candidate recommendation object, and determining a matching degree between the text and the label information based on the text features of the text and the object label features of the label information;
[0053] Extract text keywords of the above text, and determine the matching degree between the above text and the above label information based on the text keywords of the above text and the above label information, wherein the above text keywords include keywords within the text and keywords outside the text, and the above text keywords are determined by the method of determining the label information provided in any optional embodiment of the present application.
[0054] On the one hand, an embodiment of the present application provides an electronic device, which includes a processor and a memory, which are connected to each other; the memory is used to store a computer program; the processor is configured to execute the method provided by any possible implementation of the above-mentioned tag information determination method and / or the method provided by any possible implementation of the information recommendation method when calling the above-mentioned computer program.
[0055] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method provided by any possible implementation of the method for determining label information and / or the method provided by any possible implementation of the information recommendation method.
[0056] In one aspect, embodiments of the present application provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method provided in any possible implementation of the tag information determination method and / or the method provided in any possible implementation of the information recommendation method.
[0057] The beneficial effects of the embodiments of the present application are:
[0058] In an embodiment of the present application, when processing the to-be-processed text corresponding to the information of interest of any object, in addition to obtaining the in-text keywords of the to-be-processed text, the text features of the to-be-processed text can also be semantically matched with the keyword features of each candidate keyword. Through the semantic matching degree, the out-text keywords corresponding to the to-be-processed text are obtained, and the label information of the object is determined based on the in-text keywords and the out-text keywords. In this way, in addition to considering the in-text key information of the to-be-processed text of the object's information of interest, the out-text keywords corresponding to the to-be-processed text are also considered, making the keyword information corresponding to the to-be-processed text more complete and comprehensive. That is to say, when determining the user's label information, in addition to determining the user's label information based on the in-text keywords, the user's label information can also be determined based on the in-text keywords and the out-text keywords, which expands the user's label information and improves the richness and comprehensiveness of the user's label information. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0060] Figure 1 This is a schematic diagram of the structure of an optional tag information determination system provided in an embodiment of the present application;
[0061] Figure 2 This is a flow chart of an optional method for determining tag information provided in an embodiment of the present application;
[0062] Figure 3a This is a schematic diagram of an optional principle for determining keywords outside of a text provided in an embodiment of the present application;
[0063] Figure 3b This is a schematic diagram of an optional principle of processing text to be processed / candidate keywords through a Bert model provided in an embodiment of the present application;
[0064] Figure 3c This is a schematic diagram of an optional method for establishing a keyword database provided in an embodiment of the present application;
[0065] Figure 4 This is a schematic diagram of an optional semantic matching principle provided in an embodiment of the present application;
[0066] Figure 5 This is a flow chart of an optional information recommendation method provided in an embodiment of the present application;
[0067] Figure 6This is a schematic diagram of the structure of an optional device for determining tag information provided in an embodiment of the present application;
[0068] Figure 7 This is a schematic diagram of the structure of an optional information recommendation device provided in an embodiment of the present application;
[0069] Figure 8 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0070] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0071] The method for determining label information provided in the embodiments of the present application can be applied to fields such as natural language processing and machine learning in the field of artificial intelligence, and can also be applied to various fields of cloud technology, such as cloud computing and cloud services in cloud technology, and can also be applied to related data computing and processing fields in the field of big data.
[0072] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0073] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0074] Natural language processing (NLP) is a key area of research in computer science and artificial intelligence. It studies the theories and methods that enable effective communication between humans and computers using natural language. Natural language processing integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language we use in everyday life—and is closely linked to the study of linguistics.
[0075] Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.
[0076] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0077] Cloud technology refers to a hosting technology that unifies hardware, software, network and other resources within a wide area network or local area network to achieve data computing, storage, processing and sharing. The tag information determination method provided in the embodiment of the present application can be implemented based on cloud computing in cloud technology.
[0078] Cloud computing refers to obtaining required resources on demand and in an easily scalable manner through the Internet. It is the product of the integration of traditional computer and network technologies such as grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balancing.
[0079] Artificial intelligence cloud services, also commonly referred to as AIaaS (AI as a Service), are a mainstream AI platform service model. Specifically, AIaaS platforms break down several common AI services and provide independent or packaged services in the cloud, such as processing resource conversion requests.
[0080] Big data refers to data sets that cannot be captured, managed, and processed within a certain timeframe using conventional software tools. These are massive, high-growth, and diverse information assets that require new processing models to enhance decision-making power, insight discovery, and process optimization. With the advent of the cloud era, big data has attracted increasing attention. Special technologies are required to effectively implement the tag information determination method provided in this embodiment based on big data. Technologies suitable for big data include massively parallel processing databases, data mining, distributed file systems, distributed databases, and the aforementioned cloud computing.
[0081] As an example, Figure 1 FIG. 1 shows a schematic diagram of a system for determining tag information applicable to an embodiment of the present application. It can be understood that the method for determining tag information provided in the embodiment of the present application can be applied to, but not limited to, Figure 1 In the application scenario shown.
[0082] In this example, Figure 1 As shown, the text processing system in this example may include but is not limited to an application server 101, a network 102, and a user terminal 103 on which a client program of the application is installed. The user terminal 103 can communicate with the server 101 via the network 102, and the server 101 can determine the tag information of the user (such as an object). The server 101 includes a database 1011 and a processing engine 1012. The above-mentioned user terminal 103 includes a human-computer interaction screen 1031 (user interface of the application), a processor 1032 and a memory 1033. The human-computer interaction screen 1031 is used for a user (such as an object) to browse information of interest through the human-computer interaction screen. The processor 1032 is used to process the relevant operations of the user. The memory 1033 is used to store the information of interest.
[0083] like Figure 1 As shown, the specific implementation process of the method for determining tag information in this application may include steps S1-S4:
[0084] Step S1: For any object (ie user), the user can interact with the information of interest through the human-computer interaction screen 1031 of the user terminal, such as browsing, clicking, copying, and collecting the information of interest, etc., wherein the memory 1033 is used to store the information of interest.
[0085] Step S2, for any object, the processing engine 1012 in the server 101 obtains the text to be processed corresponding to the object's information of interest, and extracts the keywords in the text of the above-mentioned text to be processed, as well as extracts the text features of the above-mentioned text to be processed; wherein, the database 1011 in the server 101 can be used to store the object's information of interest, the text to be processed, the keywords in the text, and the text features.
[0086] In step S3, the processing engine 1012 in the server 101 obtains the keyword features of each candidate keyword, wherein the above-mentioned candidate keywords are keywords in the keyword thesaurus, and based on the semantic matching between the above-mentioned text features and each of the above-mentioned keyword features, determines the non-text keywords corresponding to the above-mentioned text to be processed from each of the above-mentioned candidate keywords, wherein the database 1011 in the server 101 can also be used to store non-text keywords.
[0087] The keyword features of each candidate keyword may be pre-stored in the local database 1011 . When in use, the processing engine 1012 may directly obtain the keyword features from the local database 1011 .
[0088] Step S4: The processing engine 1012 in the server 101 determines the tag information of the object based on the text keywords, wherein the text keywords include the in-text keywords and the out-text keywords.
[0089] It is understood that the above is only an example and is not limited to this embodiment.
[0090] The server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server or server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The aforementioned network can include, but is not limited to, wired networks and wireless networks. The wired network includes a local area network, a metropolitan area network, and a wide area network, and the wireless network includes Bluetooth, Wi-Fi, and other networks that enable wireless communication. The user terminal can be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a laptop computer, a digital broadcast receiver, a MID (Mobile Internet Device), a PDA (Personal Digital Assistant), a desktop computer, an in-vehicle terminal (such as an in-vehicle navigation terminal), a smart speaker, a smartwatch, etc. The user terminal and the server can be connected directly or indirectly via wired or wireless communication, but are not limited to such. The specific method can also be determined based on the actual application scenario requirements and is not limited here.
[0091] See also Figure 2 , Figure 2 This is a flow chart of a method for determining tag information provided by an embodiment of the present application. The method can be executed by any electronic device, such as a server, or by interaction between a user terminal and a server. Optionally, it can be executed by a server, such as Figure 2 As shown, the method for determining tag information provided in the embodiment of the present application includes the following steps:
[0092] Step S201 : obtaining a text to be processed corresponding to the information of interest of the object, and extracting keywords from the text to be processed.
[0093] Step S202: extracting text features of the text to be processed.
[0094] Step S203: Obtain keyword features of each candidate keyword, wherein the candidate keyword is a keyword in the keyword database.
[0095] Step S204 : Based on the semantic matching degree between the text features of the text to be processed and the features of each keyword, determine the non-textual keywords corresponding to the text to be processed from the candidate keywords.
[0096] Step S205 : determining the tag information of the object based on the text keywords, wherein the text keywords include the keywords within the text and the keywords outside the text.
[0097] Optionally, the above-mentioned information of interest can be understood as historical information browsed by the user. For example, the information of interest can be advertising information, APP description information, news article information, game promotion information, etc. The form of the information of interest can include at least one of video, voice, picture, text, etc., without any limitation here.
[0098] By processing the information of interest, the corresponding text to be processed can be obtained. For example, by processing advertising information, app description information, news article information, and game promotion information, corresponding data such as ad titles, app description text information, and news article text information can be obtained, and these data can be used as the text to be processed. Then, through keyword extraction technology, words or phrases that are thematically relevant to the text to be processed and have commercial significance are extracted from the text to be processed to obtain the in-text keywords corresponding to the text to be processed. The in-text keywords are original keywords obtained based on the original text information of the text to be processed.
[0099] Specifically, multi-industry named entity recognition can be performed on texts from multiple scenarios, such as advertisements clicked by users, articles read by users, titles of products purchased by users, and descriptions of apps installed or downloaded by users, to obtain entity labels, which can be used as keywords in the text of the object.
[0100] Then, through feature extraction, the text features of the text to be processed are obtained. The keyword features of each candidate keyword are also obtained, wherein these candidate keywords are keywords in a keyword vocabulary. Each candidate keyword can be a keyword in a pre-established keyword vocabulary. The keyword vocabulary can include professional vocabulary in multiple industries, such as the gaming industry, the financial industry, the advertising industry, the film and television industry, etc., which are not limited here.
[0101] A semantic match is performed between the text features and the features of each candidate keyword to obtain a semantic match degree. Based on the semantic match degree, at least one non-textual keyword is determined from the candidate keywords to have a high semantic match degree with the text to be processed. The non-textual keyword is a keyword related to the original keyword in the original text information of the text to be processed, in addition to the original keyword.
[0102] Then, based on the keywords in the text and the keywords outside the text, the label information of the object is determined (this label information can also be called a keyword label). For example, the keywords in the text and the keywords outside the text can be directly determined as the label information of the object, or feature extraction can be performed on the keywords in the text and the keywords outside the text respectively, and the feature information of the extracted keywords in the text and the keywords outside the text can be used as the label information of the object. This is not limited here.
[0103] In one example, in the gaming industry, for object 1, assuming that the object 1 label obtained by the in-text keywords corresponding to the information that object 1 is interested in is "real person", "arcade", and "fishing", the above method can be used to determine the out-of-text keywords corresponding to the in-text keywords. The labels corresponding to the out-of-text keywords are "arcade fishing" and "fishing master". Then, "real person", "arcade", "fishing", "arcade fishing" and "fishing master" can all be determined as the user's label information. At this time, the user's label information is expanded, and the richness of the user's label information is improved.
[0104] As an example, when a user interacts with information of interest (such as advertising information, etc.) (such as browsing the advertising information), the text to be processed can be extracted from the information of interest, and then the keywords in the text can be extracted from the text to be processed, and the non-text keywords corresponding to the advertising information can be used as an extension of the keyword extraction results in the advertising information text, and the non-text keywords can be used as the user's tag information to supplement the user's tag information, thereby achieving the purpose of supplementing the mining of user tag information under the advertising business, and using the user tag information for targeted application in the targeted delivery of the advertising business, or applied to the rough and fine ranking models of the advertising system, or to assist in the audience mining of the advertising industry, etc.
[0105] Specifically, the text to be processed is first obtained from the ad title. Then, the non-textual keywords of the text to be processed are obtained using the method described above. The in-text keywords and non-textual keywords are merged. The user's tag information is determined based on the merged text keywords, and the user's tag information is associated with the user. When advertising information is delivered, users can be recalled based on their tag information. For example, if a user is associated with a certain region, when an advertisement is targeted to that region, users in that region who are related to the advertisement can be recalled.
[0106] It is understood that the above is only an example and this embodiment does not impose any limitation thereto.
[0107] Through this embodiment, when processing the to-be-processed text corresponding to the information of interest of any object, in addition to obtaining the in-text keywords of the to-be-processed text, the text features of the to-be-processed text can also be semantically matched with the keyword features of each candidate keyword. Through the semantic matching degree, the out-text keywords corresponding to the to-be-processed text are obtained, and the label information of the object is determined based on the in-text keywords and the out-text keywords. In this way, in addition to considering the in-text key information of the to-be-processed text of the object's information of interest, the out-text keywords corresponding to the to-be-processed text are also considered, making the keyword information corresponding to the to-be-processed text more complete and comprehensive. That is to say, when determining the user's label information, in addition to determining the user's label information based on the in-text keywords, the user's label information can also be determined based on the in-text keywords and the out-text keywords, which expands the user's label information and improves the richness and comprehensiveness of the user's label information.
[0108] In the related art, when extracting keywords from input text, only the keywords contained in the input text itself are considered, which has the problem of insufficient information.
[0109] To solve this problem, the non-textual keywords of the text to be processed can be determined in the following manner to expand the information of the text to be processed. In an optional embodiment, the above-mentioned extraction of text features of the text to be processed includes:
[0110] Based on the above text to be processed, extracting text features of the above text to be processed through a feature extraction model;
[0111] The above-mentioned determination of the non-textual keywords corresponding to the above-mentioned text to be processed from the above-mentioned candidate keywords based on the semantic matching between the text features of the above-mentioned text to be processed and the features of each of the above-mentioned keywords includes:
[0112] Based on the text features of the text to be processed and the features of each of the above keywords, the non-text keywords corresponding to the text to be processed are determined from the candidate keywords through a semantic matching model.
[0113] Optionally, a feature extraction model may be used to extract features from the text to be processed to obtain text features of the text to be processed. The feature extraction model may be a bidirectional encoder representation from transformers (Bert model for short).
[0114] A semantic matching model can be used to perform semantic matching between the text features of the text to be processed and the features of each keyword, thereby determining the non-textual keywords corresponding to the text to be processed from the candidate keywords. The semantic matching model can be an Enhanced LSTM for Natural Language Inference (ESIM) natural language understanding model.
[0115] See also Figure 3a , Figure 3a This is a schematic diagram of an optional principle for determining keywords outside of a text provided by an embodiment of the present application, such as Figure 3a As shown, the text to be processed is segmented to obtain the corresponding segmented words (i.e. Figure 3a The various circles shown in the left part of the figure) are used to input the various word segments corresponding to the text to be processed into the representation layer, that is, input into the Bert model shown in the figure, and the text to be processed is predicted by the Bert model (that is, the feature extraction is performed on the text to be processed) to obtain the text features corresponding to the text to be processed.
[0116] For any candidate keyword among the candidate keywords, there are two ways to obtain the keyword features corresponding to the candidate keyword, as follows:
[0117] Method 1: If Figure 3a As shown, the pre-stored keyword features of the candidate keyword are directly obtained from the keyword database.
[0118] Method 2: If Figure 3a As shown in the dotted part, the candidate keyword is segmented to obtain the corresponding segmented words (i.e. Figure 3a The circles shown in the right part of the figure) are used to input each word segment corresponding to the candidate keyword into the representation layer, that is, input into the Bert model shown in the figure, and the candidate keyword is predicted by the Bert model (that is, the feature of the candidate keyword is extracted) to obtain the keyword feature corresponding to the candidate keyword.
[0119] According to the above method 1 or method 2, the keyword features of all candidate keywords can be obtained, and the text features and the keyword features of all candidate keywords are semantically matched through the ESIM model shown in the figure, that is, the text features and any keyword features of all candidate keywords are semantically matched to obtain the semantic matching degree between the text features and all keyword features. Based on the semantic matching degree, the scores of the above text features and the keyword features of all candidate keywords are obtained. Based on the scores, at least one keyword feature of part or all is selected from the keyword features of all candidate keywords, and the candidate keywords corresponding to the at least one keyword feature selected as part or all are used as the non-text keywords of the text features.
[0120] When calculating semantic matching, cosine similarity can be used to calculate the semantic matching between text features and any keyword features to obtain the semantic matching between the text features and all keyword features. Cosine similarity, also known as cosine similarity, evaluates the semantic matching between two keyword features by calculating the cosine value of the angle between them.
[0121] Alternatively, the text features can be subjected to semantic matching calculations with all keyword features through nearest neighbor search (NN), k-nearest neighbor search (K-NN), or approximate nearest neighbor search (ANN), and at least one keyword feature of part or all can be selected from all keyword features, and the candidate keywords corresponding to the at least one keyword feature of part or all selected can be used as out-of-text keywords.
[0122] It should be noted that this application Figure 3a In the example shown, keywords outside the text are obtained, which can expand the keyword information of the text to be processed. Compared with the method of extracting keywords by only considering the information of the text to be processed itself, the amount of information of the keywords contained in the text to be processed is enriched, and the comprehensiveness of the keywords of the text to be processed is improved.
[0123] In actual business, full-lexicon matching technology (i.e., the aforementioned technology that semantically matches text features with candidate keywords in a keyword lexicon) can be used to compensate for the insufficient number of keywords extracted from a text. Furthermore, for crowd-sourced tag mining in various industries, a vocabulary suitable for each industry can be used in conjunction with this technology to perform keyword mining for a specific industry within a text (also known as tag information mining).
[0124] In the present application, the Figure 3aThe semantic matching model of Bert-ESIM in this example is used for keyword extraction. Compared with the use of the recurrent neural network model (RNN) for keyword extraction in the above example, the effect on the test data is shown in Table 1 after experimental verification. It can be seen from Table 1 that the use of the Bert-ESIM model in this example for keyword extraction will significantly improve the accuracy (Precision), recall rate (Recall) and F-Measure (i.e., F value) of keyword extraction.
[0125] Precision is also known as "accuracy," "correctness," or "precision rate." Recall is also known as "recall rate." F-Measure is the weighted average of precision (P) and recall (R).
[0126] Table 1
[0127]
[0128] Furthermore, using keyword caching technology significantly improves prediction efficiency. Table 2 shows a comparison of the time required to predict keywords for a document using a 2.58 million-word online vocabulary before and after using caching technology.
[0129] Table 2
[0130] Time required to predict 2.58 million candidate keywords Before caching 120 minutes After caching 1 minute
[0131] Furthermore, keywords within the text and keywords outside the text can be used to determine the label information of the object. When you want to recommend information to be recommended, you can determine the target recommendation object from each candidate recommendation object based on the matching degree between the text of the information to be recommended and the label information of each candidate recommendation object, and recommend the information to the target recommendation object.
[0132] For example, the information to be recommended may be advertising information that an advertiser wants to recommend. The text to be processed is extracted from the advertising information, for example, the title of the advertisement is extracted from the advertising information as the text to be processed, and so on. Then, keywords within the text to be processed (such as "leisure" and "fishing master") are extracted to obtain label information for each candidate recommendation object. For example, if object 1 has labels such as "real person," "arcade," "fishing," "arcade fishing," and "fishing master," since the keywords within the text to be processed match the label of object 1 (i.e., "fishing master"), the information to be recommended can be recommended to object 1.
[0133] In the above process, as described above, before object 1's tag information is expanded, object 1's tag information is "real person," "arcade," and "fishing." After expansion, object 1's tag information is "real person," "arcade," "fishing," "arcade fishing," and "fishing master." When recommending the recommended information, if object 1's tags are not expanded, since object 1's tag information does not include "leisure" and "fishing master," the recommended information may not be recommended to object 1. However, after expanding object 1's tag information, object 1's tag information includes "fishing master," ensuring that the recommended information is recommended to object 1, thereby expanding the promotion range of the recommended information.
[0134] In one example, the label information of the object may be updated according to the keyword information corresponding to the text of the information to be recommended. For example, “leisure” in the information to be recommended may be added to the label information of object 1 .
[0135] It is understood that the above is only an example and this embodiment does not impose any limitation thereto.
[0136] The following example illustrates how to use the Bert model to obtain the text features corresponding to the text to be processed, and how to use the Bert model to obtain the keyword features of the candidate keywords. Figure 3b , Figure 3b This is a schematic diagram of an optional principle of processing the text to be processed / candidate keywords through the Bert model provided in an embodiment of the present application.
[0137] like Figure 3b As shown in the figure, take the text to be processed as an example, the text to be processed is input into the Bert model, where [CLS], Tok1, Tok 2...Tok N represent the input of the text to be processed. The Bert model can clearly represent the text to be processed in a token sequence. [CLS] is a special symbol of the Bert model, usually inserted before the text. [CLS] 、E1、E2……E N Respectively represent the input embeddings of [CLS], Tok 1, Tok 2, ... Tok N. The circles shown in the figure are the core of the Bert model, that is, the bidirectional encoder represents the Transformer. Only two layers are shown in the figure. In practical applications, it can be set as needed and is not limited here. The bidirectional encoder represents the Transformer to process the text to be processed, and the output representation corresponding to the text to be processed can be output, that is, C, H1, H2 ... H shown in the figure. N , where C is the output result corresponding to [CLS], and C can be used as the output result corresponding to the text to be processed, that is, the text feature.
[0138] For the implementation method of obtaining the keyword features corresponding to the candidate keywords through the Bert model, please refer to the above process and will not be repeated here.
[0139] In an optional embodiment, the above method further includes:
[0140] Extracting keyword features of each of the candidate keywords, and storing the keyword features of each of the candidate keywords in a keyword thesaurus;
[0141] The keyword features of each candidate word are obtained as follows:
[0142] The pre-stored keyword features of each of the candidate keywords are obtained from the keyword database.
[0143] Optionally, a keyword thesaurus can be pre-built and the keyword features of each candidate keyword can be pre-stored in the keyword thesaurus so that when the keyword feature is used, the keyword feature can be directly obtained from the keyword thesaurus, avoiding the process of keyword extraction for each candidate keyword, which can greatly reduce the amount of calculation and improve calculation efficiency.
[0144] Optionally, you also need to build a keyword database in advance. Figure 3c This is a schematic diagram of an optional method for establishing a keyword thesaurus provided in an embodiment of the present application. In actual business, the training data corresponding to the business type can be determined according to the business type. For example, the training data can be constructed through scenarios such as advertising business, APP description, e-commerce title, information articles, etc., and the candidate keywords appearing in the texts of these businesses are used to train the Bert model and the ESIM model. For training, a supervised training method can be used, and positive and negative examples can be obtained through manual annotation. After the Bert model is trained, the keyword features of each candidate keyword can be extracted separately through the Bert model, and the extracted keyword features can be pre-stored in the keyword thesaurus. When in use, the keyword features can be directly obtained to reduce the amount of calculation. Specifically, the keyword thesaurus can be an online thesaurus of 2.58 million and a commercial thesaurus of 20 million, etc., which are not limited here.
[0145] Through this embodiment, the keyword features can be pre-stored in the keyword database, thereby avoiding the process of keyword extraction for each candidate keyword, which can greatly reduce the amount of calculation and improve the calculation efficiency.
[0146] The following details the process of semantic matching between text features and keyword features. Figure 4 , Figure 4 This is a schematic diagram of an optional semantic matching principle provided by the embodiment of the present application. Figure 4The neural network BiLSTM (Bi-directional Long Short-Term Memory, bidirectional LSTM) or the tree-long short-term memory network (Tree-LSTM) shown in the figure is combined with the forward LSTM and the backward LSTM to perform semantic matching. Figure 4 When Premise represents text features, Hypothesis represents keyword features. Figure 4 In the expression, Premise represents keyword features, and Hypothesis represents text features.
[0147] in, Figure 4 The input encoding shown is used to convert the text features and keyword features obtained from the BERT model into text features and keyword features suitable for the BiLSTM model. The relevant processing is performed based on the converted text features and keyword features. The specific process is as follows:
[0148] In an optional embodiment, the text features of the to-be-processed text include first word features of each word contained in the to-be-processed text, and the keyword features include second word features of each word contained in the candidate keywords.
[0149] For any of the above keyword features, the semantic matching degree between the text features of the above text to be processed and the above keyword features is determined in the following manner:
[0150] Determine the similarity between each of the first word features and each of the second word features, respectively, wherein the pairwise features include a first word feature and a second word feature;
[0151] Based on the similarity between each of the two features, the semantic matching degree between the text features of the to-be-processed text and the keyword features is determined.
[0152] Optionally, in this example, a specific example is used to illustrate the above process. For example, for the text to be processed T = [t1, t2, ..., t n ] (such as [real person, arcade, fishing]) and the word sequence P of candidate keywords outside the text to be processed (such as "arcade fishing") = [p1, p2, ..., p m ] (for example, [arcade, fishing]), use the Bert model to model the two and obtain the representation sequences X and K. The specific formula is as follows:
[0153] X=x1,x2,…,x n ,X=f Bert([t1,t2,…,t n ]) (1)
[0154] K=k1,k2,…,k m ,K=f Bert ([p1,p2,…,p m ]) (2)
[0155] Among them, X is the text feature, n is the number of words contained in the text to be processed, and x i is the i-th word in the text to be processed, 1≤i≤n, K is the keyword feature, m is the number of words contained in the candidate keyword, k j is the jth word in the candidate keywords, 1≤j≤m.
[0156] Furthermore, for text features and keyword features, the ESIM model is further used for semantic matching. Based on the first word features of the text features and the second word features of the keyword features, the similarity between each first word feature and each second word feature is determined, wherein the pairwise features include a first word feature and a second word feature; based on the similarity between each pairwise feature, the semantic matching degree between the text feature X and the keyword feature K is determined. The similarity between the pairwise features is
[0157] in, Figure 4 The local inference modeling shown is used to determine the global text features and global keyword features corresponding to the text features and keyword features, namely the following global text features and global keyword features The “-” and “⊙” shown in the figure correspond one-to-one with the “-” and “⊙” in the following formulas (5) and (6). The specific process is as follows:
[0158] In an optional embodiment, for any of the above keyword features, determining the semantic matching degree between the text features of the to-be-processed text and the above keyword features based on the similarity between each pair of the above features includes:
[0159] For any first word feature, determine the weights corresponding to the second word features based on the first similarities corresponding to the first word feature, and perform weighted summation on the second word features based on the weights corresponding to the second word features to obtain an updated first word feature;
[0160] For any second word feature, determine the weight corresponding to each of the first word features based on each second similarity corresponding to the second word feature, and perform weighted summation of each of the first word features based on the weight corresponding to each of the first word features to obtain an updated second word feature;
[0161] Based on the updated first word features and the updated second word features, the semantic matching degree between the text features of the to-be-processed text and the keyword features is determined.
[0162] Optionally, the similarity between the two features is obtained in the above manner. Further, based on the similarity between the two features, the text feature X and the keyword feature K of the to-be-processed text can be weighted to obtain the updated text feature of the to-be-processed text corresponding to the text feature X and the keyword feature K. and keyword features in, and Each word in is determined according to the following formula:
[0163]
[0164]
[0165] in, is the first word feature of the i-th word after update, is the second word feature of the j-th word after update. By each composition, By each composition.
[0166] Through the updated text features and keyword features Determine the semantic matching degree between the above text features and the above keyword features.
[0167] In an optional embodiment, determining the semantic matching degree between the text features of the to-be-processed text and the keyword features based on the updated first word features and the updated second word features includes:
[0168] For any of the first word features, based on the correlation between the first word feature and the updated first word feature, obtain a local word feature corresponding to the first word feature; based on the first word feature, the updated first word feature, and the local word feature corresponding to the first word feature, obtain a global word feature corresponding to the first word feature;
[0169] For any of the second word features, based on the correlation between the second word feature and the updated second word feature, obtain the local word feature corresponding to the second word feature; based on the second word feature, the updated second word feature and the local word feature corresponding to the second word feature, obtain the global word feature corresponding to the second word feature;
[0170] Based on the global word features corresponding to each first word feature and the global word features corresponding to each second word feature, the semantic matching degree between the text features of the to-be-processed text and the keyword features is determined.
[0171] In an optional embodiment, for any of the first word features, obtaining the local word feature corresponding to the first word feature based on the correlation between the first word feature and the updated first word feature includes:
[0172] Determining a first difference feature between the first word feature and the updated first word feature;
[0173] Determining a first similarity feature between the first word feature and the updated first word feature;
[0174] The local word features corresponding to the first word feature include a first difference feature and a first similarity feature;
[0175] For any of the above second word features, based on the above second word feature, the updated second word feature and the local word feature corresponding to the above second word feature, including:
[0176] determining a second difference feature between the second word feature and the updated second word feature;
[0177] determining a second similarity feature between the second word feature and the updated second word feature;
[0178] Among them, the local word features corresponding to the above-mentioned second word features include a second difference feature and a second similarity feature.
[0179] Optionally, after obtaining the updated text features and keyword features of the to-be-processed text in the above manner, in order to further enhance the local information, the following operations are performed on the updated text features and keyword features of the to-be-processed text respectively:
[0180]
[0181]
[0182] Among them, - represents bitwise subtraction, ⊙ represents bitwise multiplication, and is the local word feature corresponding to the first word feature, and is the local word feature corresponding to the second word feature. is the global word feature corresponding to the first word feature. is the global word feature corresponding to the second word feature. is the first difference feature corresponding to the first word feature, is the first similarity feature corresponding to the first word feature. is the second difference feature corresponding to the second word feature, is the second similar feature corresponding to the second word feature. Global text features corresponding to the constituent text features By each Global keyword features corresponding to the constituent keyword features
[0183] Based on global text features and global keyword features Determine the semantic matching degree between the above text features and the above keyword features.
[0184] in, Figure 4 The inference composition shown is used to determine the final text feature representation and keyword feature representation corresponding to the text features and keyword features of the text to be processed, that is, the following final text feature representation y doc and keyword feature representation y keyword The specific process is as follows:
[0185] Furthermore, global text features can be and global keyword features Use the Long Short-Term Memory (LSTM) network for further processing and use average / max pooling to get the final text feature representation y doc and keyword feature representation y keyword .
[0186] Further, if Figure 4 The prediction process shown is based on the text feature representation y doc and keyword feature representation y keyword Determine the similarity score between text features and keyword features. The specific formula is as follows:
[0187] y=W[y doc ;y keyword ]+b (7)
[0188] label=argmax l∈y sofmax(y) (8)
[0189] Among them, y doc Through Perform average pooling or maximum pooling to obtain, y keyword Through The result is obtained by performing average pooling or maximum pooling. y is used to represent the similarity score between text features and keyword features. W and b are network parameters, representing the weight matrix and bias, respectively. Label is used to represent the result after y is normalized. Label is a value greater than or equal to 0 and less than or equal to 1. argmax is a function. For example, the function Y = f(x), x0 = argmax(f(x)) means that the parameter x0 satisfies f(x0) as the maximum value of f(x). In other words, argmax(f(x)) is the variable x that makes f(x) reach its maximum value. sofmax(y) is a normalization function.
[0190] In one example, in semantic matching mode, text features and keyword features can be modeled and represented separately. When performing keyword prediction on the text to be processed, the keyword features of each candidate keyword in the keyword vocabulary can be obtained through the model, and then semantic matching can be performed with each text to be processed, which can improve the overall prediction efficiency. Specifically, for each candidate keyword, the keyword features obtained by the Bert model of each candidate keyword are pre-saved in the local keyword vocabulary, and then used directly by loading the keyword features during actual prediction. In actual application, for online vocabulary of 2.58 million words and commercial vocabulary of 20 million words, the semantic matching method is used, and the candidate keywords are pre-cached through this keyword caching technology to accelerate the prediction efficiency. And when actually matching, the vector matching efficiency can be further improved by using cosine similarity or ANN nearest neighbor retrieval. Taking cosine similarity as an example, in the prediction stage, the text feature X output by the Bert model and the keyword feature K read into the local cache are directly averaged and pooled to obtain the corresponding text feature y. doc and keyword feature y keyword , and then judge whether the candidate keyword is a keyword by the size of the cosine similarity between the two features.
[0191] See also Figure 5 , Figure 5 This is a flow chart of an optional information recommendation method provided by an embodiment of the present application. The method can be executed by any electronic device, such as a server or a user terminal, or can be completed by the interaction between the user terminal and the server. Optionally, it can be executed by the user terminal, such as Figure 5 As shown, the information recommendation method provided in the embodiment of the present application includes the following steps:
[0192] Step S501: Acquire information to be recommended and the text of the information to be recommended.
[0193] Step S502: Obtain label information of each candidate recommendation object.
[0194] Step S503: determining a target recommended object from the candidate recommended objects based on the matching degree between the text and the tag information.
[0195] Step S504: recommend the information to be recommended to the target recommendation object.
[0196] Optionally, the recommended information can be understood as information that is intended to be promoted, and the recommended information can be determined based on the actual application scenario. For example, the recommended information can be advertising information, app description information, news article information, game promotion information, etc. The form of the information of interest can include at least one of video, audio, picture, text, etc., without any limitation here.
[0197] The label information of each candidate recommendation object (such as the above-mentioned object) can be obtained by referring to the above description and will not be repeated here.
[0198] According to the matching degree between the text of the information to be recommended and each candidate recommendation information, a target recommendation object may be determined from each candidate recommendation object, and then the information to be recommended may be recommended to the target recommendation object.
[0199] In an optional embodiment, at least one of the following is further included:
[0200] For any of the candidate recommendation objects, obtaining text features of the text and object label features of the label information of the candidate recommendation object, and determining a matching degree between the text and the label information based on the text features of the text and the object label features of the label information;
[0201] Extract text keywords of the above text, and determine the matching degree between the above text and the above tag information based on the text keywords of the above text and the above tag information, wherein the above text keywords include keywords within the text and keywords outside the text, and the above text keywords are determined by any possible embodiment of the tag information determination method.
[0202] Optionally, in one example, feature extraction can be used to extract text features of the text of the information to be recommended and object label features of the label information of the candidate recommendation objects. Then, based on the matching degree between the text features of the text and the object label features, the candidate recommendation objects corresponding to the text features of the text whose matching degree exceeds a certain threshold (such as 90%) are selected as target recommendation objects, and then the information to be recommended is recommended to the target recommendation object.
[0203] In one example, the text keywords of the text can also be obtained to determine the matching degree between the text keywords of the text and the label information, and the candidate recommendation objects corresponding to the text keywords whose matching degree exceeds a certain threshold (such as 90%) are selected as target recommendation objects, and then the information to be recommended is recommended to the target recommendation objects.
[0204] For the process of recommending the information to be recommended, please refer to the above description.
[0205] See also Figure 6 , Figure 6 Schematic diagram of a device for determining tag information provided in an embodiment of the present application. The device 1 for determining tag information provided in an embodiment of the present application includes:
[0206] The keyword processing module 11 is used to obtain the text to be processed corresponding to the information of interest of the object and extract keywords from the text to be processed;
[0207] A text feature processing module 12 is used to extract text features of the text to be processed;
[0208] The keyword processing module 11 is configured to obtain keyword features of each candidate keyword and determine, from each candidate keyword, non-textual keywords corresponding to the text to be processed based on the semantic matching between the text features of the text to be processed and the keyword features, wherein the candidate keywords are keywords in the keyword lexicon;
[0209] The tag information determining module 13 is configured to determine the tag information of the object based on text keywords, wherein the text keywords include the in-text keywords and the out-text keywords.
[0210] In an optional embodiment, the keyword processing module is further configured to:
[0211] Extracting keyword features of each of the candidate keywords, and storing the keyword features of each of the candidate keywords in a keyword thesaurus;
[0212] When obtaining the keyword features of each candidate word, the pre-stored keyword features of each candidate keyword are obtained from the keyword database.
[0213] In an optional embodiment, the text features of the to-be-processed text include first word features of each word contained in the to-be-processed text, and the keyword features include second word features of each word contained in the candidate keywords; the keyword processing module is configured to:
[0214] For any of the above keyword features, the semantic matching degree between the text features of the above text to be processed and the above keyword features is determined in the following manner:
[0215] Determine the similarity between each of the first word features and each of the second word features, respectively, wherein the pairwise features include a first word feature and a second word feature;
[0216] Based on the similarity between each of the two features, the semantic matching degree between the text features of the to-be-processed text and the keyword features is determined.
[0217] In an optional embodiment, the keyword processing module is configured to:
[0218] For any first word feature, determine the weights corresponding to the second word features based on the first similarities corresponding to the first word feature, and perform weighted summation on the second word features based on the weights corresponding to the second word features to obtain an updated first word feature;
[0219] For any second word feature, determine the weight corresponding to each of the first word features based on each second similarity corresponding to the second word feature, and perform weighted summation of each of the first word features based on the weight corresponding to each of the first word features to obtain an updated second word feature;
[0220] Based on the updated first word features and the updated second word features, the semantic matching degree between the text features of the to-be-processed text and the keyword features is determined.
[0221] In an optional embodiment, the keyword processing module is configured to:
[0222] For any of the first word features, based on the correlation between the first word feature and the updated first word feature, obtain a local word feature corresponding to the first word feature; based on the first word feature, the updated first word feature, and the local word feature corresponding to the first word feature, obtain a global word feature corresponding to the first word feature;
[0223] For any of the second word features, based on the correlation between the second word feature and the updated second word feature, obtain the local word feature corresponding to the second word feature; based on the second word feature, the updated second word feature and the local word feature corresponding to the second word feature, obtain the global word feature corresponding to the second word feature;
[0224] Based on the global word features corresponding to each first word feature and the global word features corresponding to each second word feature, the semantic matching degree between the text features of the to-be-processed text and the keyword features is determined.
[0225] In an optional embodiment, the keyword processing module is configured to:
[0226] Determining a first difference feature between the first word feature and the updated first word feature;
[0227] Determining a first similarity feature between the first word feature and the updated first word feature;
[0228] The local word features corresponding to the first word feature include a first difference feature and a first similarity feature;
[0229] determining a second difference feature between the second word feature and the updated second word feature;
[0230] determining a second similarity feature between the second word feature and the updated second word feature;
[0231] Among them, the local word features corresponding to the above-mentioned second word features include a second difference feature and a second similarity feature.
[0232] In an optional embodiment, the text feature processing module is configured to:
[0233] Based on the above text to be processed, extracting text features of the above text to be processed through a feature extraction model;
[0234] The above-mentioned determination of the non-textual keywords corresponding to the above-mentioned text to be processed from the above-mentioned candidate keywords based on the semantic matching between the text features of the above-mentioned text to be processed and the features of each of the above-mentioned keywords includes:
[0235] Based on the text features of the text to be processed and the features of each of the above keywords, the non-text keywords corresponding to the text to be processed are determined from the candidate keywords through a semantic matching model.
[0236] In an embodiment of the present application, when processing the to-be-processed text corresponding to the information of interest of any object, in addition to obtaining the in-text keywords of the to-be-processed text, the text features of the to-be-processed text can also be semantically matched with the keyword features of each candidate keyword. Through the semantic matching degree, the out-text keywords corresponding to the to-be-processed text are obtained, and the label information of the object is determined based on the in-text keywords and the out-text keywords. In this way, in addition to considering the in-text key information of the to-be-processed text of the object's information of interest, the out-text keywords corresponding to the to-be-processed text are also considered, making the keyword information corresponding to the to-be-processed text more complete and comprehensive. That is to say, when determining the user's label information, in addition to determining the user's label information based on the in-text keywords, the user's label information can also be determined based on the in-text keywords and the out-text keywords, which expands the user's label information and improves the richness and comprehensiveness of the user's label information.
[0237] See also Figure 7 , Figure 7 Schematic diagram of the structure of an information recommendation device provided in an embodiment of the present application. The information recommendation device 2 provided in an embodiment of the present application includes:
[0238] The information processing module 21 is used to obtain the information to be recommended and the text of the information to be recommended;
[0239] A tag information acquisition module 22 is configured to acquire tag information of each candidate recommendation object, wherein the tag information is determined using the tag information determination method provided in any optional embodiment of the present application;
[0240] A target recommendation object determination module 23 is configured to determine a target recommendation object from the candidate recommendation objects based on the matching degree between the text and the tag information;
[0241] The information recommendation module 23 is used to recommend the information to be recommended to the target recommendation object.
[0242] In an optional embodiment, the target recommendation object determination module is further configured to perform at least one of the following:
[0243] For any of the candidate recommendation objects, obtaining text features of the text and object label features of the label information of the candidate recommendation object, and determining a matching degree between the text and the label information based on the text features of the text and the object label features of the label information;
[0244] Extract text keywords of the above text, and determine the matching degree between the above text and the above label information based on the text keywords of the above text and the above label information, wherein the above text keywords include keywords within the text and keywords outside the text, and the above text keywords are determined by the method of determining the label information provided in any optional embodiment of the present application.
[0245] In a specific implementation, the above-mentioned tag information determination device 1 can execute the above-mentioned various function modules built in it. Figure 2 For the implementation methods provided in each step, please refer to the implementation methods provided in the above steps for details, which will not be repeated here.
[0246] The information recommendation device 2 can execute the above-mentioned Figure 5 For the implementation methods provided in each step, please refer to the implementation methods provided in the above steps for details, which will not be repeated here.
[0247] The above mainly introduces that the execution subject is hardware to implement the label information determination method and / or information recommendation method in this application, but the execution subject of the label information determination method and / or information recommendation method in this application is not limited to hardware. The execution subject of the label information determination method and / or information recommendation method in this application can also be software. The above-mentioned label information determination device and / or information recommendation device can be a computer program (including program code) running in a computer device. For example, the label information determination device and / or information recommendation device is an application software; the device can be used to execute the corresponding steps in the method provided in the embodiment of the present application.
[0248] In some embodiments, the tag information determination device and / or information recommendation device provided by the embodiments of the present invention can be implemented in a combination of software and hardware. As an example, the tag information determination device and / or information recommendation device provided by the embodiments of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the tag information determination method and / or information recommendation method provided by the embodiments of the present invention. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0249] In other embodiments, the tag information determination device and / or information recommendation device provided by the embodiments of the present invention may be implemented in software. Figure 6 The device 1 for determining tag information shown can be software in the form of a program or plug-in, and includes a series of modules, including a keyword processing module 11, a text feature processing module 12, and a tag information determination module 13, for implementing the tag information determination method provided by the embodiment of the present invention. And / or, Figure 7 The information recommendation device 2 shown can be software in the form of a program and plug-in, and includes a series of modules, including a to-be-recommended information processing module 21, a label information acquisition module 22, a target recommendation object determination module 23 and an information recommendation module 24 for implementing the information recommendation method provided in an embodiment of the present invention.
[0250] See also Figure 8 , Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 8 As shown, the electronic device 1000 in this embodiment may include: a processor 1001, a network interface 1004 and a memory 1005. In addition, the above-mentioned electronic device 1000 may also include: a user interface 1003, and at least one communication bus 1002. The communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), a keyboard (Keyboard), and the user interface 1003 may optionally include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1004 may be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk memory. The memory 1005 may optionally be at least one storage device located away from the aforementioned processor 1001. As Figure 8 As shown, the memory 1005 as a computer-readable storage medium may include an operating system, a network communication module, a user interface module, and a device control application.
[0251] exist Figure 8 In the electronic device 1000 shown, the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 can be used to call the computer program stored in the memory 1005.
[0252] It should be understood that in some feasible embodiments, the processor 1001 may be a central processing unit (CPU), or may be another general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc. The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store device type information.
[0253] In a specific implementation, the electronic device 1000 can execute the above-mentioned functions through its built-in functional modules. Figure 2 、 Figure 5 For the implementation methods provided in each step, please refer to the implementation methods provided in the above steps for details, which will not be repeated here.
[0254] The present invention also provides a computer-readable storage medium that stores a computer program and is executed by a processor to implement Figure 2 、 Figure 5 For the methods provided in each step, please refer to the implementation methods provided in the above steps for details, which will not be repeated here.
[0255] The above-mentioned computer-readable storage medium can be the internal storage unit of the task processing device provided by any of the aforementioned embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. The above-mentioned computer-readable storage medium can also include a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc. Further, the computer-readable storage medium can also include both the internal storage unit of the electronic device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or is to be output.
[0256] The embodiment of the present application provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above-mentioned Figure 2 、 Figure 5 The method provided by any possible embodiment.
[0257] The terms "first," "second," and the like in the claims, specification, and drawings of this application are used to distinguish between different objects, not to describe a specific order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or device. Reference herein to an "embodiment" means that a particular feature, structure, or characteristic described in conjunction with an embodiment may be included in at least one embodiment of the present application. The presence of such a phrase in various locations in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive with other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments. The term "and / or," as used in this specification and the appended claims, refers to any and all possible combinations of one or more of the associated listed items, including, but not limited to, those listed.
[0258] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the above description generally describes the components and steps of each example according to their functions. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0259] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. A method for determining tag information, characterized in that: include: Obtaining a text to be processed corresponding to the information of interest of the object, and extracting keywords within the text to be processed; Extracting text features of the text to be processed; the text features of the text to be processed include first word features of each word contained in the text to be processed; Obtaining keyword features of each candidate keyword, wherein the candidate keyword is a keyword in a keyword thesaurus, and the keyword features of each candidate keyword include second word features of each word included in the candidate keyword; For each of the keyword features, updating each of the first word features according to the similarity between each of the first word features and each of the second word features of the keyword features to obtain updated first word features, wherein the pairwise features include a first word feature and a second word feature; for each of the first word features, obtaining a local word feature corresponding to the first word feature based on the correlation between the first word feature and the updated first word feature, wherein the local word feature corresponding to the first word feature includes a first difference feature and a first similarity feature, and concatenating the first word feature, the updated first word feature, and the local word feature corresponding to the first word feature to obtain a global word feature corresponding to the first word feature; determining a semantic matching degree between the text feature of the to-be-processed text and the keyword feature according to the global word feature corresponding to each of the first word features and the keyword feature; Determining, from the candidate keywords, the non-textual keywords corresponding to the text to be processed based on the semantic matching between the text features of the text to be processed and the features of each keyword; The tag information of the object is determined based on text keywords, wherein the text keywords include the in-text keywords and the out-text keywords.
2. The method according to claim 1, characterized in that The method further comprises: Extracting keyword features of each candidate keyword and storing the keyword features of each candidate keyword in a keyword thesaurus; The step of obtaining keyword features of each candidate word includes: The pre-stored keyword features of each candidate keyword are obtained from the keyword database.
3. The method according to claim 1, characterized in that For each of the keyword features, updating each of the first word features according to the similarity between each of the first word features and each of the second word features of the keyword features to obtain each updated first word feature includes: For each first word feature, determining a weight corresponding to each second word feature based on a similarity between the first word feature and each second word feature, and performing a weighted summation of each second word feature based on the weight corresponding to each second word feature to obtain an updated first word feature; The method further comprises: For each second word feature, determine the weight corresponding to each first word feature based on the similarity between the second word feature and each first word feature, and perform weighted summation of each first word feature based on the weight corresponding to each first word feature to obtain an updated second word feature; based on the correlation between the second word feature and the updated second word feature, obtain the local word feature corresponding to the second word feature, and concatenate the second word feature, the updated second word feature, and the local word feature corresponding to the second word feature to obtain the global word feature corresponding to the second word feature; wherein the local word feature corresponding to the second word feature includes a second difference feature and a second similarity feature; For each of the keyword features, determining the semantic matching degree between the text features of the to-be-processed text and the keyword features based on the global word features corresponding to each of the first word features and the keyword features, includes: Based on the global word features corresponding to the first word features and the global word features corresponding to the second word features of the keyword features, the semantic matching degree between the text features of the to-be-processed text and the keyword features is determined.
4. The method according to claim 3, characterized in that For each of the keyword features, determining the semantic matching degree between the text features of the to-be-processed text and the keyword features based on the global word features corresponding to the first word features and the global word features corresponding to the second word features of the keyword features includes: determining a global text feature based on the global word features corresponding to each of the first word features; determining a global keyword feature based on the global word feature corresponding to each of the second word features; Based on the global text features and the global keyword features, a semantic matching degree between the text features of the to-be-processed text and the keyword features is determined.
5. The method according to claim 1, wherein: For each first word feature, the local word feature corresponding to the first word feature is determined in the following manner: Subtracting the first word feature from the updated first word feature bit by bit to obtain a first difference feature; The first word feature and the updated first word feature are bitwise multiplied to obtain a first similarity feature.
6. The method according to any one of claims 1 to 5, characterized in that The extracting text features of the to-be-processed text includes: Based on the text to be processed, extracting text features of the text to be processed through a feature extraction model; The determining of the non-textual keywords corresponding to the text to be processed from the candidate keywords based on the semantic matching between the text features of the text to be processed and the features of each keyword includes: Based on the text features of the text to be processed and the features of each keyword, a non-textual keyword corresponding to the text to be processed is determined from each candidate keyword through a semantic matching model.
7. An information recommendation method, characterized in that: include: Obtaining information to be recommended and the text of the information to be recommended; Obtaining label information of each candidate recommendation object, wherein the label information of the candidate recommendation object is determined by the method of any one of claims 1 to 6; Determining a target recommendation object from the candidate recommendation objects based on a degree of matching between the text and the tag information; Recommend the information to be recommended to the target recommendation object.
8. The method according to claim 7, characterized in that Also include at least one of the following: For any candidate recommendation object, obtaining text features of the text and object label features of the label information of the candidate recommendation object, and determining a matching degree between the text and the label information based on the text features of the text and the object label features of the label information; Extract text keywords of the text, and determine the degree of match between the text and the label information based on the text keywords of the text and the label information, wherein the text keywords include keywords within the text and keywords outside the text, and the text keywords are determined by the method described in any one of claims 1 to 6.
9. A device for determining tag information, characterized in that: The device comprises: A keyword processing module is used to obtain the text to be processed corresponding to the information of interest of the object and extract keywords in the text to be processed; A text feature processing module, configured to extract text features of the text to be processed; the text features of the text to be processed include first word features of each word contained in the text to be processed; The keyword processing module is further used to obtain keyword features of each candidate keyword, and for each keyword feature, update each first word feature according to the similarity between each first word feature and each second word feature of the keyword feature to obtain each updated first word feature, and the two-two features include a first word feature and a second word feature; for each first word feature, based on the correlation between the first word feature and the updated first word feature, obtain the local word feature corresponding to the first word feature, and the local word feature corresponding to the first word feature includes a first difference feature and a first similarity feature, and the first word feature, the updated first word feature and the local word feature are combined. splicing the first word feature of the text to be processed and the local word feature corresponding to the first word feature to obtain the global word feature corresponding to the first word feature; determining the semantic matching degree between the text feature of the text to be processed and the keyword feature according to the global word feature corresponding to each of the first word features and the keyword feature; and based on the semantic matching degree between the text feature of the text to be processed and each of the keyword features, determining the out-of-text keyword corresponding to the text to be processed from each of the candidate keywords, wherein the candidate keyword is a keyword in a keyword thesaurus, and the keyword feature of each of the candidate keywords includes the second word feature of each word contained in the candidate keyword; The tag information determination module is configured to determine the tag information of the object based on text keywords, wherein the text keywords include the in-text keywords and the out-text keywords.
10. The device according to claim 9, characterized in that The keyword processing module is further used to: Extracting keyword features of each candidate keyword and storing the keyword features of each candidate keyword in a keyword thesaurus; When obtaining the keyword features of each candidate word, the pre-stored keyword features of each candidate keyword are obtained from the keyword database.
11. The device according to claim 9, characterized in that The keyword processing module is further used to: For each first word feature, determining a weight corresponding to each second word feature based on a similarity between the first word feature and each second word feature, and performing a weighted summation on each second word feature based on the weight corresponding to each second word feature to obtain an updated first word feature; For each second word feature, determining a weight corresponding to each first word feature based on a similarity between the second word feature and each first word feature, and performing a weighted summation of each first word feature based on the weight corresponding to each first word feature to obtain an updated second word feature; Based on the correlation between the second word feature and the updated second word feature, a local word feature corresponding to the second word feature is obtained, and the second word feature, the updated second word feature, and the local word feature corresponding to the second word feature are concatenated to obtain a global word feature corresponding to the second word feature; wherein the local word feature corresponding to the second word feature includes a second difference feature and a second similarity feature; Based on the global word features corresponding to the first word features and the global word features corresponding to the second word features of the keyword features, the semantic matching degree between the text features of the to-be-processed text and the keyword features is determined.
12. The device according to claim 11, characterized in that The keyword processing module is further used to: determining a global text feature based on the global word features corresponding to each of the first word features; determining a global keyword feature based on the global word feature corresponding to each of the second word features; Based on the global text features and the global keyword features, a semantic matching degree between the text features of the to-be-processed text and the keyword features is determined.
13. The device according to claim 9, characterized in that When determining the local word feature corresponding to the first word feature, the keyword processing module is configured to: Subtracting the first word feature from the updated first word feature bit by bit to obtain a first difference feature; The first word feature and the updated first word feature are bitwise multiplied to obtain a first similarity feature.
14. The device according to any one of claims 9 to 13, characterized in that The text feature processing module is used to: Based on the text to be processed, extracting text features of the text to be processed through a feature extraction model; The determining of the non-textual keywords corresponding to the text to be processed from the candidate keywords based on the semantic matching between the text features of the text to be processed and the features of each keyword includes: Based on the text features of the text to be processed and the features of each keyword, a non-textual keyword corresponding to the text to be processed is determined from each candidate keyword through a semantic matching model.
15. An information recommendation device, characterized in that: include: The information processing module to be recommended is used to obtain the information to be recommended and the text corresponding to the information to be recommended; a label information acquisition module, configured to acquire label information of each candidate recommendation object, wherein the label information is determined using the method of any one of claims 1 to 6; a target recommendation object determination module, configured to determine a target recommendation object from the candidate recommendation objects based on a degree of matching between the text and the tag information; The information recommendation module is used to recommend the information to be recommended to the target recommendation object.
16. The device according to claim 15, characterized in that The text feature processing module is further configured to perform at least one of the following: For any candidate recommendation object, obtaining text features of the text and object label features of the label information of the candidate recommendation object, and determining a matching degree between the text and the label information based on the text features of the text and the object label features of the label information; Extract text keywords of the text, and determine the degree of match between the text and the label information based on the text keywords of the text and the label information, wherein the text keywords include keywords within the text and keywords outside the text, and the text keywords are determined by the method described in any one of claims 1 to 6.
17. An electronic device, characterized in that: comprising a processor and a memory, wherein the processor and the memory are connected to each other; The memory is used to store computer programs; The processor is configured to execute the method as claimed in any one of claims 1 to 6 or any one of claims 7 to 8 when calling the computer program.
18. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method described in any one of claims 1 to 6 or any one of claims 7 to 8.
Citation Information
Patent Citations
Online text label real-time adding method and device and related equipment
CN110795911A
Document recommendation method and device based on semantic tag
US20200210468A1