Text processing method and device and storage medium
By obtaining and analyzing search records, the relationship between upper and lower word is automatically determined, which solves the problems of high cost and low accuracy of word list construction in the prior art, and realizes efficient and accurate word pair recognition.
Patent Information
- Application Number
- CN202510088631.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art has high cost and poor recognition accuracy when constructing a vocabulary list of upper and lower terms.
By obtaining multiple search records, word pairs with upper and lower relationships are determined based on the query text set. The method includes obtaining search records, dividing a query text set, and determining the upper and lower relationship of candidate word pairs based on the number of clicks in the click text.
It realizes unsupervised mining of word pairs with upper and lower relationships using search record logs, reducing the cost of manual recognition and improving the recognition accuracy and coverage of word pairs.
Smart Images

Figure CN119962524A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a text processing method, device and storage medium. Background Art
[0002] Hyponyms and hyponyms refer to pairs of words that have a conceptual relationship between the upper and lower parts. This relationship is essentially a relationship between the general and the particular. For example, "convolutional neural network" is a particular to "deep learning", and "deep learning" is a particular to "machine learning". As a basic feature of linguistics, hyponyms and hyponyms have a wide range of applications in the field of natural language processing. They can not only be used for query understanding in search engines, but also for understanding the basic features of algorithms to improve algorithm capabilities. At present, the industry usually constructs a vocabulary for the relationship between hyponyms and hyponyms manually, or uses manual annotation and training of neural network models to identify and construct them. This is costly and has poor accuracy in identifying hyponyms and hyponyms. Summary of the invention
[0003] In view of this, the present disclosure proposes a text processing method, device and storage medium.
[0004] According to one aspect of the present disclosure, a text processing method is provided. The method comprises:
[0005] Obtain multiple search records, each of which includes a query text and a click text, wherein the query text represents the text entered during the search, and the click text represents the text corresponding to the content clicked during the search;
[0006] Dividing the query text into at least one query text set based on the plurality of search records; wherein the query text in each query text set corresponds to the same click text;
[0007] Determine word pairs with a hyponym relationship based on a query text set.
[0008] In a possible implementation, the search record also includes the number of clicks on the click text, and determining the word pairs with a hyponym relationship based on the query text set includes:
[0009] Determine a plurality of candidate word pairs in the query text set;
[0010] According to the number of clicks on the click text corresponding to the query text related to the candidate word pairs, word pairs with a hyponymous relationship among the plurality of candidate word pairs are determined.
[0011] In a possible implementation, determining a word pair having a hyponymous relationship among multiple candidate word pairs according to the number of clicks on the click text corresponding to the query text related to the candidate word pair includes:
[0012] According to the number of clicks of the click text corresponding to the query text related to any candidate word pair, determine the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to each word in the candidate word pair;
[0013] Based on the number of clicks corresponding to the multiple candidate word pairs and the number of clicks corresponding to each word in the candidate word pairs, word pairs with a hyponymous and hyponymous relationship among the multiple candidate word pairs are determined.
[0014] In a possible implementation, the number of clicks corresponding to any candidate word pair and the number of clicks corresponding to each word in the candidate word pair are determined according to the number of clicks of the click text corresponding to the query text related to any candidate word pair, including:
[0015] If the words in the candidate word pair appear together in at least one query text, the sum of the click counts of the click texts corresponding to at least one query text is taken as the click count corresponding to the candidate word pair;
[0016] If each word in the candidate word pair appears in different query texts, the sum of the click counts of the click texts corresponding to the different query texts is taken as the click count corresponding to the candidate word pair;
[0017] The number of clicks corresponding to any word in the candidate word pair is determined according to the sum of the number of clicks of the click text corresponding to at least one query text to which the word belongs.
[0018] In a possible implementation, multiple candidate word pairs are determined in the query text set, including:
[0019] Using a sliding window of a preset size to process the query text in each query text set, a plurality of candidate word pairs are obtained;
[0020] Merge the same candidate word pairs in multiple query text sets.
[0021] In a possible implementation, based on the number of clicks corresponding to the multiple candidate word pairs and the number of clicks corresponding to each word in the candidate word pairs, determining the word pairs with a hyponymous relationship among the multiple candidate word pairs includes:
[0022] Based on the number of clicks corresponding to each word in the candidate word pair, determine the hypernym and hyponym in the candidate word pair, wherein the number of clicks corresponding to the hyponym in the candidate word pair is greater than the number of clicks corresponding to the hypernym in the candidate word pair;
[0023] Based on the number of clicks corresponding to the multiple candidate word pairs and the number of clicks corresponding to the hypernyms and hyponyms in the candidate word pairs, word pairs having a hypernymy relationship among the multiple candidate word pairs are determined.
[0024] In a possible implementation, based on the number of clicks corresponding to the multiple candidate word pairs and the number of clicks corresponding to the hypernyms and hyponyms in the candidate word pairs, determining the word pairs with a hypernymy relationship among the multiple candidate word pairs includes:
[0025] Based on the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym and hyponym in the candidate word pair, determine whether the candidate word pair satisfies the co-occurrence relationship;
[0026] In response to the candidate word pair satisfying the co-occurrence relationship, determining that the candidate word pair is a word pair having a hyponymy relationship;
[0027] Among them, the co-occurrence relationship means that when the hypernym appears, the hyponym will also appear, but when the hyponym appears, the hypernym may not appear.
[0028] In a possible implementation, based on the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym and hyponym in the candidate word pair, determining whether the candidate word pair satisfies the co-occurrence relationship includes:
[0029] Determine a first co-occurrence feature value according to the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym in the candidate word pair, and determine that the candidate word pair satisfies a first co-occurrence relationship in response to the first co-occurrence feature value being greater than a first preset threshold;
[0030] Determine a second co-occurrence feature value according to the number of clicks corresponding to the candidate word pair, and the number of clicks corresponding to the hypernym and the number of clicks corresponding to the hyponym in the candidate word pair, and determine that the candidate word pair satisfies a second co-occurrence relationship in response to the second co-occurrence feature value being greater than a second preset threshold;
[0031] In the case where the candidate word pair satisfies both the first co-occurrence relationship and the second co-occurrence relationship, it is determined that the candidate word pair satisfies the co-occurrence relationship.
[0032] In a possible implementation, the first co-occurrence feature value is the quotient of the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym in the candidate word pair;
[0033] The second co-occurrence feature value is the absolute value of the difference between the quotient of the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hyponym in the candidate word pair and the first co-occurrence feature value.
[0034] In a possible implementation, the first co-occurrence feature value is the logarithm of the quotient of the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym in the candidate word pair;
[0035] The second co-occurrence feature value is the absolute value of the difference between the quotient of the logarithm of the number of clicks corresponding to the candidate word pair and the logarithm of the number of clicks corresponding to the hyponym in the candidate word pair and the first co-occurrence feature value.
[0036] In a possible implementation, multiple candidate word pairs are determined in the query text set, including:
[0037] Determine a plurality of initial candidate word pairs in the query text set;
[0038] A plurality of initial candidate word pairs are screened to obtain a plurality of candidate word pairs.
[0039] In a possible implementation, a plurality of initial candidate word pairs are screened to obtain a plurality of candidate word pairs, including:
[0040] Among the multiple initial candidate word pairs, initial candidate word pairs whose hyponyms or synonyms of the hyponyms do not appear in the corresponding click text are removed to obtain multiple candidate word pairs.
[0041] In a possible implementation, a plurality of initial candidate word pairs are screened to obtain a plurality of candidate word pairs, including:
[0042] Using a large language model to determine whether the click text corresponding to the initial candidate word pair describes the content of the hyponym in the initial candidate word pair;
[0043] When the click text corresponding to the initial candidate word pair does not describe the content about the hyponym, the initial candidate word pair is removed to obtain multiple candidate word pairs.
[0044] In a possible implementation, a plurality of initial candidate word pairs are screened to obtain a plurality of candidate word pairs, including:
[0045] Among the multiple initial candidate word pairs, the initial candidate word pairs containing words of a preset type are removed to obtain multiple candidate word pairs. The words of the preset type include: one or more of auxiliary words, prepositions, quantifiers, and conjunctions.
[0046] According to another aspect of the present disclosure, a text processing device is provided. The device includes:
[0047] An acquisition module is used to acquire multiple search records, each of which includes a query text and a click text, wherein the query text represents the text input during the search, and the click text represents the text corresponding to the content clicked during the search;
[0048] A division module, used to divide the query text into at least one query text set based on the multiple search records; wherein the query text in each query text set corresponds to the same click text;
[0049] The determination module is used to determine word pairs with a hyponym relationship based on a query text set.
[0050] In a possible implementation, the search record further includes the number of clicks on the clicked text, and the determination module is used to:
[0051] Determine a plurality of candidate word pairs in the query text set;
[0052] According to the number of clicks on the click text corresponding to the query text related to the candidate word pairs, word pairs with a hyponymous relationship among the plurality of candidate word pairs are determined.
[0053] In a possible implementation, determining a word pair having a hyponymous relationship among multiple candidate word pairs according to the number of clicks on the click text corresponding to the query text related to the candidate word pair includes:
[0054] According to the number of clicks of the click text corresponding to the query text related to any candidate word pair, determine the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to each word in the candidate word pair;
[0055] Based on the number of clicks corresponding to the multiple candidate word pairs and the number of clicks corresponding to each word in the candidate word pairs, word pairs with a hyponymous and hyponymous relationship among the multiple candidate word pairs are determined.
[0056] In a possible implementation, the number of clicks corresponding to any candidate word pair and the number of clicks corresponding to each word in the candidate word pair are determined according to the number of clicks of the click text corresponding to the query text related to any candidate word pair, including:
[0057] If the words in the candidate word pair appear together in at least one query text, the sum of the click counts of the click texts corresponding to at least one query text is taken as the click count corresponding to the candidate word pair;
[0058] If each word in the candidate word pair appears in different query texts, the sum of the click counts of the click texts corresponding to the different query texts is taken as the click count corresponding to the candidate word pair;
[0059] The number of clicks corresponding to any word in the candidate word pair is determined according to the sum of the number of clicks of the click text corresponding to at least one query text to which the word belongs.
[0060] In a possible implementation, a module is determined to:
[0061] Using a sliding window of a preset size to process the query text in each query text set, a plurality of candidate word pairs are obtained;
[0062] Merge the same candidate word pairs in multiple query text sets.
[0063] In a possible implementation, based on the number of clicks corresponding to the multiple candidate word pairs and the number of clicks corresponding to each word in the candidate word pairs, determining the word pairs with a hyponymous relationship among the multiple candidate word pairs includes:
[0064] Based on the number of clicks corresponding to each word in the candidate word pair, determine the hypernym and hyponym in the candidate word pair, wherein the number of clicks corresponding to the hyponym in the candidate word pair is greater than the number of clicks corresponding to the hypernym in the candidate word pair;
[0065] Based on the number of clicks corresponding to the multiple candidate word pairs and the number of clicks corresponding to the hypernyms and hyponyms in the candidate word pairs, word pairs having a hypernymy relationship among the multiple candidate word pairs are determined.
[0066] In a possible implementation, based on the number of clicks corresponding to the multiple candidate word pairs and the number of clicks corresponding to the hypernyms and hyponyms in the candidate word pairs, determining the word pairs with a hypernymy relationship among the multiple candidate word pairs includes:
[0067] Based on the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym and hyponym in the candidate word pair, determine whether the candidate word pair satisfies the co-occurrence relationship;
[0068] In response to the candidate word pair satisfying the co-occurrence relationship, determining that the candidate word pair is a word pair having a hyponymy relationship;
[0069] Among them, the co-occurrence relationship means that when the hypernym appears, the hyponym will also appear, but when the hyponym appears, the hypernym may not appear.
[0070] In a possible implementation, based on the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym and hyponym in the candidate word pair, determining whether the candidate word pair satisfies the co-occurrence relationship includes:
[0071] Determine a first co-occurrence feature value according to the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym in the candidate word pair, and determine that the candidate word pair satisfies a first co-occurrence relationship in response to the first co-occurrence feature value being greater than a first preset threshold;
[0072] Determine a second co-occurrence feature value according to the number of clicks corresponding to the candidate word pair, and the number of clicks corresponding to the hypernym and the number of clicks corresponding to the hyponym in the candidate word pair, and determine that the candidate word pair satisfies a second co-occurrence relationship in response to the second co-occurrence feature value being greater than a second preset threshold;
[0073] In the case where the candidate word pair satisfies both the first co-occurrence relationship and the second co-occurrence relationship, it is determined that the candidate word pair satisfies the co-occurrence relationship.
[0074] In a possible implementation, the first co-occurrence feature value is the quotient of the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym in the candidate word pair;
[0075] The second co-occurrence feature value is the absolute value of the difference between the quotient of the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hyponym in the candidate word pair and the first co-occurrence feature value.
[0076] In a possible implementation, the first co-occurrence feature value is the logarithm of the quotient of the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym in the candidate word pair;
[0077] The second co-occurrence feature value is the absolute value of the difference between the quotient of the logarithm of the number of clicks corresponding to the candidate word pair and the logarithm of the number of clicks corresponding to the hyponym in the candidate word pair and the first co-occurrence feature value.
[0078] In a possible implementation, a module is determined to:
[0079] Determine a plurality of initial candidate word pairs in the query text set;
[0080] A plurality of initial candidate word pairs are screened to obtain a plurality of candidate word pairs.
[0081] In a possible implementation, a plurality of initial candidate word pairs are screened to obtain a plurality of candidate word pairs, including:
[0082] Among the multiple initial candidate word pairs, initial candidate word pairs whose hyponyms or synonyms of the hyponyms do not appear in the corresponding click text are removed to obtain multiple candidate word pairs.
[0083] In a possible implementation, a plurality of initial candidate word pairs are screened to obtain a plurality of candidate word pairs, including:
[0084] Using a large language model to determine whether the click text corresponding to the initial candidate word pair describes the content of the hyponym in the initial candidate word pair;
[0085] When the click text corresponding to the initial candidate word pair does not describe the content about the hyponym, the initial candidate word pair is removed to obtain multiple candidate word pairs.
[0086] In a possible implementation, a plurality of initial candidate word pairs are screened to obtain a plurality of candidate word pairs, including:
[0087] Among the multiple initial candidate word pairs, the initial candidate word pairs containing words of a preset type are removed to obtain multiple candidate word pairs. The words of the preset type include: one or more of auxiliary words, prepositions, quantifiers, and conjunctions.
[0088] According to another aspect of the present disclosure, a text processing device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.
[0089] According to another aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.
[0090] According to another aspect of the present disclosure, a computer program product is provided, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.
[0091] According to an embodiment of the present application, by obtaining multiple search records to determine word pairs with hyponymous and hyponymous relationships therein, it is possible to ensure a wide coverage of hyponymous and hyponymous words, wherein the query text is divided into at least one query text set based on the multiple search records, the query text in each query text set corresponds to the same click text, and the word pairs with hyponymous and hyponymous relationships are determined based on the query text set, wherein the query text represents the text input during the search, and the click text represents the text corresponding to the content clicked and browsed during the search. It is possible to utilize the search record log to determine the word pairs with hyponymous and hyponymous relationships based on the content input and browsed during the search, thereby reducing the cost of manual recognition and improving the recognition accuracy of the hyponymous and hyponymous word pairs while ensuring that the word pairs have a large coverage.
[0092] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.
[0094] Figure 1 A schematic diagram showing an application scenario according to an embodiment of the present application.
[0095] Figure 2 A flowchart of a text processing method according to an embodiment of the present application is shown.
[0096] Figure 3 A flowchart of a text processing method according to an embodiment of the present application is shown.
[0097] Figure 4 A structural diagram of a text processing device according to an embodiment of the present application is shown.
[0098] Figure 5 It is a block diagram of a device 1900 for text processing according to an exemplary embodiment. DETAILED DESCRIPTION
[0099] Various exemplary embodiments, features and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0100] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0101] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the subject matter of the present disclosure.
[0102] Hyponyms and hyponyms refer to pairs of words that have a conceptual relationship between the upper and lower words. This relationship is essentially a relationship between the general and the specific. For example, "convolutional neural network" is specific to "deep learning", and "deep learning" is specific to "machine learning". As a basic feature of linguistics, hyponyms and hyponyms have a wide range of applications in the field of natural language processing. They can not only be used for query understanding in search engines, but also for understanding the basic features of algorithms to improve algorithm capabilities. At present, the industry usually constructs a hyponym relationship vocabulary manually, or uses manual annotation and training of neural network models to identify and construct it. This is costly and has poor accuracy in identifying hyponyms and hyponyms.
[0103] In view of this, the present application provides a text processing method, device and storage medium. The method of the embodiment of the present application can determine the word pairs with a hierarchical relationship by obtaining multiple search records, which can ensure that the coverage of the hierarchical and hyponymous words is large, wherein the query text is divided into at least one query text set based on the multiple search records, and the query text in each query text set corresponds to the same click text, and the word pairs with a hierarchical relationship are determined based on the query text set, wherein the query text represents the text input during the search, and the click text represents the text corresponding to the content clicked and browsed during the search, and the search record log can be used to determine the word pairs with a hierarchical relationship based on the content input and browsed during the search, thereby reducing the cost of manual recognition, and improving the recognition accuracy of the word pairs with a hierarchical relationship while ensuring that the word pairs have a large coverage.
[0104] Figure 1A schematic diagram of an application scenario according to an embodiment of the present application is shown. The text processing system of the embodiment of the present application can be used in scenarios where word pairs with a hyponym relationship are mined, such as Figure 1 As shown, the text processing system of the embodiment of the present application can collect search record logs, which may include multiple search records, and identify word pairs with a hyponym relationship in the search record text. Thus, it is possible to realize unsupervised mining of word pairs with a hyponym relationship based on the search record log.
[0105] A plurality of word pairs with a hyponym relationship can be identified by a text processing system. In an application scenario, a hyponym relationship vocabulary can be constructed based on these word pairs with a hyponym relationship, and the hyponym relationship vocabulary can be used for subsequent natural language processing tasks. The word pairs with a hyponym relationship identified in the embodiment of the present application can also be used in other scenarios where hyponyms and hyponyms are required, and the present application does not limit this.
[0106] The text processing system of the embodiment of the present application can be deployed in a terminal device or a server. The terminal device can be any one or more of a mobile phone, a foldable electronic device, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), and a vehicle-mounted device. The embodiment of the present disclosure does not impose any special restrictions on the specific type of the terminal device, and the terminal device can have a wired or wireless communication function.
[0107] The server can be located locally or in the cloud, and can be a physical device or a virtual device, such as a virtual machine, a container, etc., and has a wireless communication function, wherein the wireless communication function can be set in the chip (system) or other parts or components of the server. The wireless communication function can be realized, for example, through mobile communication technologies such as 2G / 3G / 4G / 5G, as well as Wi-Fi, Bluetooth, frequency modulation (FM), digital radio, satellite communication, etc. Communication can also be carried out through wired connection to achieve interaction with other devices.
[0108] Figure 2 FIG. 1 is a flowchart of a text processing method according to an embodiment of the present application. The method can be used in the above-mentioned text processing system, such as Figure 2 As shown, the method may include:
[0109] Step S201, obtaining multiple search records.
[0110] Search records can be provided by search engines, indicating the interaction between users and search engines or search results (such as search behavior and browsing behavior). Multiple search records can be stored in a search record log, which can be a systematic file or database for recording and managing historical search / browsing behaviors. Multiple search records can come from one or more users.
[0111] Each search record may include a query text and a click text. The search record may also include the number of clicks on the click text.
[0112] The query text may represent the text entered during the search, for example, the content entered by the user in the search box of the search engine; the click text may represent the text corresponding to the content clicked during the search, and the content clicked may be the title and / or text content of the link clicked during the search. The query text and the click text may be pre-processed, for example, by word segmentation, removal of special characters (punctuation marks), etc., and each may consist of one or more words.
[0113] The number of clicks may represent the cumulative number of clicks by one or more users on the same click text (e.g., a link or multiple links with the same click text (e.g., the same text)) in the search results page when searching based on a query text within a preset time period.
[0114] Each search record can be represented by a triple, which is: query text-click text-click count. An example of a search record is:
[0115] Query text: [Guangdong Shenzhen Senior High School] - Click text: [Introduction to Shenzhen Senior High School] - Number of clicks: 50.
[0116] The spaces in the query text and the click text may indicate the result of word segmentation. In the example, the query text and the click text each include two words.
[0117] Step S202: divide the query text into at least one query text set based on multiple search records.
[0118] The query texts in each query text set correspond to the same click texts, so that a key value aggregated with click texts can be formed based on the texts clicked by users, and different click texts correspond to different query text sets. For the same click text, the candidate word pairs obtained based on the query text set can be used as the key, and the number of clicks on the click text corresponding to the query text related to the candidate word pair can be used as the value.
[0119] In one example, six search records corresponding to the same click text are as follows:
[0120] Query text 1: [Guangdong Shenzhen Senior High School] - Click text: [Introduction to Shenzhen Senior High School] - Click count: 50;
[0121] Query text 2: [Shenzhen Senior High School in Shenzhen] - Click text: [Introduction to Shenzhen Senior High School] - Click count: 30;
[0122] Query text 3: [Encyclopedia of famous high schools] - Click text: [Introduction to Shenzhen Senior High School] - Number of clicks: 10;
[0123] Query text 4: [Shenzhen Senior High School, Shenzhen, Guangdong] - Click text: [Introduction to Shenzhen Senior High School] - Click count: 30;
[0124] Query text 5: [Shenzhen Senior High School Encyclopedia] - Click text: [Introduction to Shenzhen Senior High School] - Number of clicks: 10.
[0125] Query text 6: [Shenzhen Senior High School] - Click text: [Introduction to Shenzhen Senior High School] - Number of clicks: 60.
[0126] Based on the above, we can construct the query text set corresponding to the click text [Introduction to Shenzhen Senior High School]:
[0127] ([Shenzhen Senior High School in Guangdong], [Shenzhen Senior High School in Shenzhen], [Encyclopedia of high school departments of famous high schools], [Shenzhen Senior High School in Shenzhen, Guangdong], [Encyclopedia of Shenzhen Senior High School], [Shenzhen Senior High School]).
[0128] The query text set may include different query texts, but corresponds to the same click text.
[0129] Step S203: determining word pairs having a hyponymous relationship based on the query text set.
[0130] Each word pair having a hyponym relationship may include a hypernym and an associated hyponym.
[0131] According to an embodiment of the present application, by obtaining multiple search records to determine word pairs with hyponymous and hyponymous relationships therein, it is possible to ensure a wide coverage of hyponymous and hyponymous words, wherein the query text is divided into at least one query text set based on the multiple search records, the query text in each query text set corresponds to the same click text, and the word pairs with hyponymous and hyponymous relationships are determined based on the query text set, wherein the query text represents the text input during the search, and the click text represents the text corresponding to the content clicked and browsed during the search. It is possible to utilize the search record log to determine the word pairs with hyponymous and hyponymous relationships based on the content input and browsed during the search, thereby reducing the cost of manual recognition and improving the recognition accuracy of the hyponymous and hyponymous word pairs while ensuring that the word pairs have a large coverage.
[0132] Figure 3 FIG. 1 is a flowchart of a text processing method according to an embodiment of the present application. Figure 3 As shown, in step S203, it may include:
[0133] Step S301, determining a plurality of candidate word pairs in a query text set.
[0134] Each candidate word pair may include a hypernym and a hyponym.
[0135] Since the distance between words with a hyponymy relationship in the text is generally not too far apart, in the present application, a sliding window can be selected to process the query text in the query text set to generate candidate word pairs. In step S301, the following can be performed:
[0136] The query text in each query text set is processed using a sliding window of a preset size to obtain multiple candidate word pairs; the same candidate word pairs in multiple query text sets are merged.
[0137] The size of the sliding window can be set as needed, indicating the number of words to be picked up each time. For example, when the size of the sliding window is 3, three adjacent words can be selected from a query text in the query text set as the word picking result each time; when the number of words in the query text is less than the number of sliding windows (for example, when there is only one word in the query text, the query text is not picked up, such as the above query text 6), all the words in the query text can be selected as the word picking result. For a certain word picking result (for example, 3 (or more or less) words), when the word picking result includes only one word, the word picking result can be deleted, when the word picking result includes two words, the two words can be combined to obtain a candidate word pair, and when the word picking result includes 3 or more words, the words can be combined in pairs to obtain one or more candidate word pairs, for example, for 3 words, three candidate word pairs can be obtained by combining them in pairs. When the number of words in a query text is greater than the number of sliding windows, the sliding window can be used to extract words from the query text, move a predetermined number of steps to the right, and then extract words again. The predetermined number of steps for each movement can also be set as needed, for example, 1, which means that the sliding window slides one step to the right in the query text each time to extract words. For example, for the query text [Shenzhen Senior High School in Shenzhen, Guangdong], when the size of the sliding window is 3 and the number of steps for each movement is 1, the result of the first word extraction can be [Shenzhen, Guangdong], and the result of the second word extraction can be [Shenzhen Senior High School in Shenzhen]. For example, for the second word extraction result, 3 candidate word pairs can be obtained: [Shenzhen], [Shenzhen Senior High School in Shenzhen], [Shenzhen Shenzhen Senior High School in Shenzhen]. The same candidate word pairs in multiple query text sets can be merged into one candidate word pair.
[0138] Step S302: determining word pairs having a hyponymous relationship among a plurality of candidate word pairs according to the number of clicks on the click text corresponding to the query text related to the candidate word pairs.
[0139] Therefore, by determining multiple candidate word pairs in the query text set and using the number of clicks on the click text corresponding to the query text related to the candidate word pairs to determine the word pairs with hyponymy in the candidate word pairs, it is possible to use logs to unsupervisedly mine words with hyponymy.
[0140] Among them, the hypernyms and hyponyms in each candidate word pair can be determined first according to the number of clicks of the click text corresponding to the query text related to the candidate word pair. It should be understood that the hyponyms and hypernyms determined here can be regarded as a hypothetical initial result, and then the initial result can be verified to determine the final result, that is, whether the hypernym and hyponym relationship in the candidate word pair is established. If it is established, the corresponding candidate word pair can be regarded as a word pair with a hypernym and hyponym relationship.
[0141] In this process, the number of clicks on the click text corresponding to the query text related to the candidate word pair can be used to count the co-occurrence of each word in the candidate word pair in the query text, thereby determining the word pairs with a hyponymous relationship in the candidate word pair. The present application does not limit the method of using the number of clicks on the click text corresponding to the query text related to the candidate word pair to count the co-occurrence of each word in the candidate word pair in the query text. In a possible implementation, in step S302, it is possible to:
[0142] According to the number of clicks of the click text corresponding to the query text related to any candidate word pair, the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to each word in the candidate word pair are determined; based on the number of clicks corresponding to multiple candidate word pairs and the number of clicks corresponding to each word in the candidate word pairs, the word pairs with a hierarchical relationship among the multiple candidate word pairs are determined.
[0143] Therefore, by determining the number of clicks corresponding to the candidate word pairs and the number of clicks corresponding to each word in the candidate word pairs, the co-occurrence relationship of each word can be reflected, so that word pairs with a hyponymous relationship can be accurately screened out from multiple candidate word pairs.
[0144] The number of clicks corresponding to a candidate word pair can be determined based on the number of clicks of a click text corresponding to a query text related to the candidate word pair. The number of clicks corresponding to each word in a candidate word pair can be determined based on the number of clicks of a click text corresponding to a related query text. The present application does not limit the method for determining the number of clicks corresponding to a candidate word pair and the number of clicks corresponding to each word in a candidate word pair. In one possible implementation, in the process of determining the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to each word in the candidate word pair based on the number of clicks of a click text corresponding to a query text related to any candidate word pair, it is possible to:
[0145] If the words in the candidate word pair appear together in at least one query text, the sum of the number of clicks of the click texts corresponding to at least one query text shall be taken as the number of clicks corresponding to the candidate word pair; if the words in the candidate word pair appear in different query texts, the sum of the number of clicks of the click texts corresponding to the different query texts shall be taken as the number of clicks corresponding to the candidate word pair; the number of clicks corresponding to any word in the candidate word pair shall be determined based on the sum of the number of clicks of the click texts corresponding to at least one query text to which the word belongs.
[0146] In this way, it is possible to more accurately screen out word pairs with a hyponymous and hyponymous relationship from multiple candidate word pairs in the subsequent step.
[0147] The different query texts may be query texts in the same query text set, or may be query texts in different query text sets.
[0148] For example, if all the words in the candidate word pair appear in query text A and query text B, that is, all the words in the candidate word pair appear in query text A and query text B, then the candidate word pair can be considered to be related to query text A and query text B, and the sum of the number of clicks of the click texts corresponding to query text A and query text B can be used as the number of clicks corresponding to the candidate word pair. If word 1 and word 2 in the candidate word pair do not appear in any query text together, but in query text C and query text D respectively, it can also be considered that the candidate word pair is related to query text C and query text D, and the sum of the number of clicks of the click texts corresponding to query text C and query text D can be used as the number of clicks corresponding to the candidate word pair.
[0149] For another example, if word 1 in the candidate word pair belongs to query text E and query text F respectively, the candidate word pair can be regarded as being related to query text E and query text F. The sum of the number of clicks of the click texts corresponding to query text E and query text F can be used as the number of clicks corresponding to word 1. For the same word in different candidate word pairs, their corresponding numbers of clicks can be added together as the number of clicks corresponding to the word.
[0150] Thus, based on the example given in step S202 above, the number of clicks corresponding to each candidate word pair is obtained as follows:
[0151] Candidate word pair 1: [Guangdong Shenzhen Senior High School] - clicks: 80;
[0152] Candidate word pair 2: [Shenzhen Shenzhen Senior High School] - click count: 60;
[0153] Candidate word pair 3: [well-known high school high school] - clicks: 10;
[0154] Candidate word pair 4: [Shenzhen Senior High School] - click count: 30;
[0155] Candidate word pair 5: [Shenzhen Senior High School Encyclopedia] - Number of clicks: 10.
[0156] Among them, only some candidate word pairs are shown as examples. Based on the examples of the above 5 candidate word pairs, the examples of the number of clicks corresponding to each word are:
[0157] Guangdong: 80; Shenzhen: 60; Shenzhen Senior High School: 180; De: 30; Encyclopedia: 20; Famous middle schools: 10; High school: 10.
[0158] For example, the number of clicks on candidate word pair 1: [Guangdong Shenzhen Senior High School] is the sum of the number of clicks corresponding to query text 1 and query text 4, the number of clicks on the word "Shenzhen" is the sum of the number of clicks on query text 2 and query text 4, and so on.
[0159] Next, in the embodiment of the present application, the hypernyms and hyponyms in each candidate word pair can be determined based on the number of clicks corresponding to multiple candidate word pairs and the number of clicks corresponding to each word in the candidate word pair, and further, the word pairs with a hypernymy relationship can be determined. In the embodiment of the present application, in the process of determining the hypernyms and hyponyms in each candidate word pair based on the number of clicks of the click text corresponding to the query text related to the candidate word pair, we can assume that the hypernyms and hyponyms in each candidate word pair must satisfy the following relationship: when the hypernym appears, the hyponym must appear, and when the hyponym appears, the hypernym does not necessarily appear.
[0160] In a possible implementation, in the process of determining the word pairs having a hyponymous relationship among the multiple candidate word pairs based on the number of clicks corresponding to the multiple candidate word pairs and the number of clicks corresponding to each word in the candidate word pairs, the following method may be used:
[0161] Based on the number of clicks corresponding to each word in the candidate word pair, the hypernyms and hyponyms in the candidate word pair are determined; based on the number of clicks corresponding to multiple candidate word pairs, and the number of clicks corresponding to the hypernyms and hyponyms in the candidate word pairs, the word pairs with a hypernym relationship in the multiple candidate word pairs are determined.
[0162] Thus, the relationship between the hypernym and hyponym in the preliminary candidate word pairs can be obtained, and based on this, the word pairs with hyponymy relationship among multiple candidate word pairs can be determined unsupervised to obtain more accurate results.
[0163] Among them, the number of clicks corresponding to the hyponym in the candidate word pair is greater than the number of clicks corresponding to the hypernym in the candidate word pair. For example, for the hypernym a and hyponym b in a certain candidate word pair, when the hypothesis relationship that when the hypernym appears, the hyponym must appear, and when the hyponym appears, the hypernym does not necessarily appear is satisfied, there is p(b|a)>p(a|b), where p(b|a)=count(ab) / count(a), p(a|b)=count(ab) / count(b), count(ab) represents the number of clicks corresponding to the candidate word pair including words a and b, count(a) represents the number of clicks corresponding to word a, and count(b) represents the number of clicks corresponding to word b. According to the above formula, it can be obtained that when count(b)>count(a), p(b|a)>p(a|b). Thus, the word with a larger corresponding number of clicks in the candidate word pair can be used as the hyponym, and the other word can be used as the hypernym. When the number of clicks corresponding to the two words in the candidate word pair is the same, either word (such as the word with more characters) can be selected as the hyponym and the other word as the hypernym, or the candidate word pair can be selected to be deleted.
[0164] For the above five examples of candidate word pairs, the examples of hypernyms and hyponyms in each candidate word pair are as follows:
[0165] Candidate word pair 1: Guangdong (hypernym) - Shenzhen Senior High School (hyponym)
[0166] Candidate word pair 2: Shenzhen (hypernym) - Shenzhen Senior High School (hyponym)
[0167] Candidate word pair 3: well-known middle school (hyponym) - senior high school department (hypernym)
[0168] Candidate word pair 4: of (hypernym) - Shenzhen Senior High School (hyponym)
[0169] Candidate word pair 5: Shenzhen Senior High School (hyponym) - encyclopedia (hypernym)
[0170] Thus, it is possible to further verify whether the relationship between the hypernym and hyponym in each candidate word pair holds to obtain word pairs with hyponymy relationship.
[0171] In the embodiments of the present application, in order to further improve the recognition accuracy and recognition efficiency of hyponymy relationship word pairs, the above candidate word pairs can also be screened. In step S301, it is possible to:
[0172] A plurality of initial candidate word pairs are determined in the query text set; and the plurality of initial candidate word pairs are screened to obtain a plurality of candidate word pairs.
[0173] Among them, the multiple initial candidate word pairs can be the multiple candidate word pairs obtained by using the sliding window method mentioned above.
[0174] In the process of screening multiple initial candidate word pairs to obtain multiple candidate word pairs, it is possible to: remove the initial candidate word pairs whose hyponyms or synonyms of the hyponyms do not appear in the corresponding click text from the multiple initial candidate word pairs to obtain multiple candidate word pairs.
[0175] In this way, irrelevant results can be quickly filtered out and the recognition efficiency of hyponymous and hyponymous word pairs can be improved.
[0176] The synonyms of the hyponyms may be determined based on existing word lists or other methods, and the present application does not limit the method for determining the synonyms of the hyponyms.
[0177] For example, it can be determined whether the hyponyms or synonyms of the hyponyms in the initial candidate word pair appear directly in the corresponding click text. For example, if the click text corresponding to the initial candidate word pair: famous middle school (hyponym) - senior high school (hypernym) does not contain "famous middle school" or a synonym of "famous middle school" (such as "famous middle school"), it can be considered that the user clicked the click text accidentally, which is not representative and can be directly filtered.
[0178] In the process of screening multiple initial candidate word pairs to obtain multiple candidate word pairs, we can: use the large language model to determine whether the click text corresponding to the initial candidate word pair describes the content about the hyponym in the initial candidate word pair; if the click text corresponding to the initial candidate word pair does not describe the content about the hyponym, remove the initial candidate word pair to obtain multiple candidate word pairs.
[0179] In this way, the capabilities of a large language model can be used to help filter out irrelevant results and improve the recognition accuracy and efficiency of hyponymous and hyponymous word pairs.
[0180] Among them, it is possible to input a prompt word in the large language model or use the natural language processing capabilities of the large language model in other ways to determine whether the click text corresponding to the initial candidate word pair describes the content about the hyponym in the initial candidate word pair. For example, you can input in the large language model: "Does [click text] describe the content about [hyponym]?" to determine whether the [click text] corresponding to the initial candidate word pair describes the content about the [hyponym] in the initial candidate word pair. For example, "Shen Gao" is the abbreviation of "Shenzhen Senior High School". When the hyponym is "Shenzhen Senior High School" and "Shen Gao" exists in the corresponding click text, the large language model can be used to determine the content related to "Shenzhen Senior High School" described in the click text.
[0181] In the process of screening multiple initial candidate word pairs to obtain multiple candidate word pairs, it is possible to: remove the initial candidate word pairs containing words of a preset type from the multiple initial candidate word pairs to obtain multiple candidate word pairs.
[0182] In this way, irrelevant results can be quickly filtered out and the recognition efficiency of hyponymous and hyponymous word pairs can be improved.
[0183] The preset words may include one or more of auxiliary words, prepositions, quantifiers, conjunctions, etc. For example, words such as one, two, come, etc., etc., left and right, up and down, and, and so on. These words do not have a concept of upper and lower positions, and can be considered as unimportant words, so they can be directly filtered. Existing dictionaries such as auxiliary words, prepositions, quantifiers, conjunctions, etc. can be used to determine whether these types of words appear in the candidate word pairs.
[0184] Referring back to the above process of further verifying whether the hypernym and hyponym relationship in each candidate word pair is established to obtain word pairs with a hypernym and hyponym relationship, in the process of determining the word pairs with a hypernym and hyponym relationship in the multiple candidate word pairs based on the number of clicks corresponding to the multiple candidate word pairs and the number of clicks corresponding to the hypernym and hyponym in the candidate word pairs, the following can be done:
[0185] Based on the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym and hyponym in the candidate word pair, determining whether the candidate word pair satisfies the co-occurrence relationship; in response to the candidate word pair satisfying the co-occurrence relationship, determining that the candidate word pair is a word pair with a hypernym relationship;
[0186] Among them, the co-occurrence relationship can mean that when the hypernym appears, the hyponym also appears, and when the hyponym appears, the hypernym may not appear. When the candidate word pair does not satisfy the co-occurrence relationship, it can be considered that the candidate word pair is a word pair without a hypernym relationship, that is, the relationship between the hypernym and the hyponym in the candidate word pair does not hold.
[0187] The present application does not limit the method for determining whether a candidate word pair satisfies a co-occurrence relationship. For example, the Bayesian formula or other methods can be used to determine whether the candidate word pair satisfies a co-occurrence relationship. In the process of determining whether the candidate word pair satisfies a co-occurrence relationship based on the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym and hyponym in the candidate word pair, the following can be done:
[0188] Determine a first co-occurrence feature value according to the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym in the candidate word pair, and determine that the candidate word pair satisfies a first co-occurrence relationship in response to the first co-occurrence feature value being greater than a first preset threshold;
[0189] Determine a second co-occurrence feature value according to the number of clicks corresponding to the candidate word pair, and the number of clicks corresponding to the hypernym and the number of clicks corresponding to the hyponym in the candidate word pair, and determine that the candidate word pair satisfies a second co-occurrence relationship in response to the second co-occurrence feature value being greater than a second preset threshold;
[0190] In the case where the candidate word pair satisfies both the first co-occurrence relationship and the second co-occurrence relationship, it is determined that the candidate word pair satisfies the co-occurrence relationship.
[0191] In this way, the recognition accuracy of hyponymous and hyponymous word pairs can be further improved.
[0192] Among them, the first co-occurrence relationship can indicate that when the hypernym appears, the hyponym will also appear; the second co-occurrence relationship can indicate that when the hyponym appears, the hypernym may not appear. The first preset threshold and the second preset threshold can be the same or different thresholds pre-set as needed. In the embodiment of the present application, there is no limitation on the method for determining the first co-occurrence feature value and the second co-occurrence feature value. For example, the first co-occurrence feature value and the second co-occurrence feature value can be determined using statistical methods such as the Bayesian formula.
[0193] Among them, the first co-occurrence feature value can be the quotient of the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym in the candidate word pair; the second co-occurrence feature value can be the quotient of the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hyponym in the candidate word pair, and the absolute value of the difference between the first co-occurrence feature value.
[0194] For example, according to the Bayesian formula, A can be defined as a hypernym and B as a hyponym. In this case, the first co-occurrence feature value can be expressed as P(B|A), where P(B|A)=Count(AB) / Count(A), Count(AB) can represent the number of clicks corresponding to the candidate word pair, and Count(A) can represent the number of clicks corresponding to the hypernym in the candidate word pair. The second co-occurrence feature value can be expressed as |P(A|B)-P(B|A)|, where P(A|B)=Count(AB) / Count(B), and Count(B) can represent the number of clicks corresponding to the hyponym in the candidate word pair.
[0195] Based on the examples of candidate word pairs given above, for example, we can obtain the first co-occurrence feature value of candidate word pair 1: Guangdong (hypernym)-Shenzhen Senior High School (hypernym) is 1, and the second co-occurrence feature value is 0.6; the first co-occurrence feature value of candidate word pair 2: Shenzhen (hypernym)-Shenzhen Senior High School (hypernym) is 1, and the second co-occurrence feature value is 0.7; the first co-occurrence feature value of candidate word pair 5: Shenzhen Senior High School (hypernym)-Encyclopedia (hypernym) is 0.5, and the second co-occurrence feature value is 0.9.
[0196] Since the number of clicks in actual search engines is usually very large, when calculating the first co-occurrence feature value and the second co-occurrence feature value, the number of clicks can also be log-operated to reduce the impact of excessive clicks, to prevent the number of clicks corresponding to a certain click text from affecting the recognition result when it is too large. At this time, the first co-occurrence feature value can be the logarithm of the quotient of the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym in the candidate word pair; the second co-occurrence feature value can be the quotient of the logarithm of the number of clicks corresponding to the candidate word pair and the logarithm of the number of clicks corresponding to the hyponym in the candidate word pair, and the absolute value of the difference with the first co-occurrence feature value.
[0197] For example, the above P(B|A) and P(A|B) can be expressed as:
[0198] P(B|A)=log(Count(AB)) / log(Count(A)), P(A|B)=log(Count(AB)) / log(Count(B)). Thus, we can get the first co-occurrence feature value as log(Count(AB)) / log(Count(A)), and the second co-occurrence feature value as |log(Count(AB)) / log(Count(B))-log(Count(AB)) / log(Count(A))|.
[0199] Among them, the first preset threshold and the second preset threshold can also be set based on a small number of recognition results, and a large number of subsequent candidate word pairs can be judged based on the set first preset threshold and the second preset threshold. Taking the above three calculation results as an example, since the correct recognition results of hyponyms and hyponyms should be Guangdong (hypernym)-Shenzhen Senior High School (hypernym) and Shenzhen (hypernym)-Shenzhen Senior High School (hypernym), since they both meet the first co-occurrence feature value close to 1 and the second co-occurrence feature value remains at a large level (>0.5), the first preset threshold can be set to 0.9 and the second preset threshold can be set to 0.5, so that the two correct results can be retained and the incorrect candidate Shenzhen Senior High School (hypernym)-Encyclopedia (hypernym) can be removed.
[0200] Figure 4FIG. 2 shows a structural diagram of a text processing device according to an embodiment of the present application. Figure 4 As shown, the device comprises:
[0201] The acquisition module 401 is used to acquire multiple search records, each of which includes a query text and a click text, wherein the query text represents the text input during the search, and the click text represents the text corresponding to the content clicked during the search;
[0202] A division module 402 is used to divide the query text into at least one query text set based on the multiple search records; wherein the query text in each query text set corresponds to the same click text;
[0203] The determination module 403 is used to determine word pairs having a hyponymy relationship based on the query text set.
[0204] In a possible implementation, the search record further includes the number of clicks on the clicked text, and the determination module 403 is used to:
[0205] Determine a plurality of candidate word pairs in the query text set;
[0206] According to the number of clicks on the click text corresponding to the query text related to the candidate word pairs, word pairs with a hyponymous relationship among the plurality of candidate word pairs are determined.
[0207] In a possible implementation, determining a word pair having a hyponymous relationship among multiple candidate word pairs according to the number of clicks on the click text corresponding to the query text related to the candidate word pair includes:
[0208] According to the number of clicks of the click text corresponding to the query text related to any candidate word pair, determine the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to each word in the candidate word pair;
[0209] Based on the number of clicks corresponding to the multiple candidate word pairs and the number of clicks corresponding to each word in the candidate word pairs, word pairs with a hyponymous relationship among the multiple candidate word pairs are determined.
[0210] In a possible implementation, the number of clicks corresponding to any candidate word pair and the number of clicks corresponding to each word in the candidate word pair are determined according to the number of clicks of the click text corresponding to the query text related to any candidate word pair, including:
[0211] If all words in the candidate word pair appear together in at least one query text, the sum of the click counts of the click texts corresponding to at least one query text is taken as the click count corresponding to the candidate word pair;
[0212] If each word in the candidate word pair appears in different query texts, the sum of the click counts of the click texts corresponding to the different query texts is taken as the click count corresponding to the candidate word pair;
[0213] The number of clicks corresponding to any word in the candidate word pair is determined according to the sum of the number of clicks of the click text corresponding to at least one query text to which the word belongs.
[0214] In a possible implementation, the determination module 403 is configured to:
[0215] Using a sliding window of a preset size to process the query text in each query text set, a plurality of candidate word pairs are obtained;
[0216] Merge the same candidate word pairs in multiple query text sets.
[0217] In a possible implementation, based on the number of clicks corresponding to the multiple candidate word pairs and the number of clicks corresponding to each word in the candidate word pairs, determining the word pairs with a hyponymous relationship among the multiple candidate word pairs includes:
[0218] Based on the number of clicks corresponding to each word in the candidate word pair, determine the hypernym and hyponym in the candidate word pair, wherein the number of clicks corresponding to the hyponym in the candidate word pair is greater than the number of clicks corresponding to the hypernym in the candidate word pair;
[0219] Based on the number of clicks corresponding to the multiple candidate word pairs and the number of clicks corresponding to the hypernyms and hyponyms in the candidate word pairs, word pairs having a hypernymy relationship among the multiple candidate word pairs are determined.
[0220] In a possible implementation, based on the number of clicks corresponding to the multiple candidate word pairs and the number of clicks corresponding to the hypernyms and hyponyms in the candidate word pairs, determining the word pairs with a hypernymy relationship among the multiple candidate word pairs includes:
[0221] Based on the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym and hyponym in the candidate word pair, determine whether the candidate word pair satisfies the co-occurrence relationship;
[0222] In response to the candidate word pair satisfying the co-occurrence relationship, determining that the candidate word pair is a word pair having a hyponymy relationship;
[0223] Among them, the co-occurrence relationship means that when the hypernym appears, the hyponym will also appear, but when the hyponym appears, the hypernym may not appear.
[0224] In a possible implementation, based on the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym and the hyponym in the candidate word pair, determining whether the candidate word pair satisfies the co-occurrence relationship includes:
[0225] Determine a first co-occurrence feature value according to the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym in the candidate word pair, and determine that the candidate word pair satisfies a first co-occurrence relationship in response to the first co-occurrence feature value being greater than a first preset threshold;
[0226] Determine a second co-occurrence feature value according to the number of clicks corresponding to the candidate word pair, and the number of clicks corresponding to the hypernym and the number of clicks corresponding to the hyponym in the candidate word pair, and determine that the candidate word pair satisfies a second co-occurrence relationship in response to the second co-occurrence feature value being greater than a second preset threshold;
[0227] In the case where the candidate word pair satisfies both the first co-occurrence relationship and the second co-occurrence relationship, it is determined that the candidate word pair satisfies the co-occurrence relationship.
[0228] In a possible implementation, the first co-occurrence feature value is the quotient of the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym in the candidate word pair;
[0229] The second co-occurrence feature value is the absolute value of the difference between the quotient of the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hyponym in the candidate word pair and the first co-occurrence feature value.
[0230] In a possible implementation, the first co-occurrence feature value is the logarithm of the quotient of the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym in the candidate word pair;
[0231] The second co-occurrence feature value is the absolute value of the difference between the quotient of the logarithm of the number of clicks corresponding to the candidate word pair and the logarithm of the number of clicks corresponding to the hyponym in the candidate word pair and the first co-occurrence feature value.
[0232] In a possible implementation, the determination module 403 is configured to:
[0233] Determine a plurality of initial candidate word pairs in the query text set;
[0234] A plurality of initial candidate word pairs are screened to obtain a plurality of candidate word pairs.
[0235] In a possible implementation, a plurality of initial candidate word pairs are screened to obtain a plurality of candidate word pairs, including:
[0236] Among the multiple initial candidate word pairs, initial candidate word pairs whose hyponyms or synonyms of the hyponyms do not appear in the corresponding click text are removed to obtain multiple candidate word pairs.
[0237] In a possible implementation, a plurality of initial candidate word pairs are screened to obtain a plurality of candidate word pairs, including:
[0238] Using a large language model to determine whether the click text corresponding to the initial candidate word pair describes the content of the hyponym in the initial candidate word pair;
[0239] When the click text corresponding to the initial candidate word pair does not describe the content about the hyponym, the initial candidate word pair is removed to obtain multiple candidate word pairs.
[0240] In a possible implementation, a plurality of initial candidate word pairs are screened to obtain a plurality of candidate word pairs, including:
[0241] Among the multiple initial candidate word pairs, the initial candidate word pairs containing words of a preset type are removed to obtain multiple candidate word pairs. The words of the preset type include: one or more of auxiliary words, prepositions, quantifiers, and conjunctions.
[0242] According to an embodiment of the present application, by obtaining multiple search records to determine word pairs with hyponymous and hyponymous relationships therein, it is possible to ensure a wide coverage of hyponymous and hyponymous words, wherein the query text is divided into at least one query text set based on the multiple search records, the query text in each query text set corresponds to the same click text, and the word pairs with hyponymous and hyponymous relationships are determined based on the query text set, wherein the query text represents the text input during the search, and the click text represents the text corresponding to the content clicked and browsed during the search. It is possible to utilize the search record log to determine the word pairs with hyponymous and hyponymous relationships based on the content input and browsed during the search, thereby reducing the cost of manual recognition and improving the recognition accuracy of the hyponymous and hyponymous word pairs while ensuring that the word pairs have a large coverage.
[0243] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0244] The embodiment of the present disclosure also provides a computer-readable storage medium on which computer program instructions are stored, and the computer program instructions implement the above method when executed by a processor. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0245] The embodiment of the present disclosure also proposes a text processing device, comprising: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.
[0246] The embodiments of the present disclosure also provide a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.
[0247] Figure 51 is a block diagram of a device 1900 for text processing according to an exemplary embodiment. For example, the device 1900 can be provided as a server or a terminal device. Figure 5 , the apparatus 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions, such as an application, that can be executed by the processing component 1922. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0248] The device 1900 may also include a power supply component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to a network, and an input / output interface 1958 (I / O interface). The device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server 2000. TM , MacOS X TM , Unix TM ,Linux TM , FreeBSD TM or similar.
[0249] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, which can be executed by the processing component 1922 of the device 1900 to perform the above method.
[0250] The present disclosure may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0251] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples of computer-readable storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not to be interpreted as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.
[0252] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.
[0253] The computer program instructions for performing the operation of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed completely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be customized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0254] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.
[0255] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0256] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0257] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of the module, program segment or instruction includes one or more executable instructions for realizing the specified logical function. In some alternative implementations, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of special hardware and computer instructions.
[0258] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A text processing method, characterized in that: The method comprises: Acquire multiple search records, each search record includes a query text and a click text, wherein the query text represents the text input during the search, and the click text represents the text corresponding to the content clicked and browsed during the search; Dividing the query text into at least one query text set based on multiple search records; wherein the query text in each query text set corresponds to the same click text; Based on the query text set, word pairs having a hyponymy relationship are determined.
2. The method according to claim 1, characterized in that The search record also includes the number of clicks on the click text. The determining of word pairs having a hyponymous relationship based on the query text set includes: Determine a plurality of candidate word pairs in the query text set; According to the number of clicks on the click text corresponding to the query text related to the candidate word pair, a word pair having a hyponymous relationship among the plurality of candidate word pairs is determined.
3. The method according to claim 2, characterized in that The step of determining a word pair having a hyponymous relationship among a plurality of candidate word pairs according to the number of clicks of the click text corresponding to the query text related to the candidate word pair comprises: According to the number of clicks of the click text corresponding to the query text related to any candidate word pair, determine the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to each word in the candidate word pair; Based on the number of clicks corresponding to the multiple candidate word pairs and the number of clicks corresponding to each word in the candidate word pairs, word pairs with a hyponymous and hyponymous relationship among the multiple candidate word pairs are determined.
4. The method according to claim 3, characterized in that The method of determining the number of clicks corresponding to any candidate word pair and the number of clicks corresponding to each word in the candidate word pair according to the number of clicks of the click text corresponding to the query text related to any candidate word pair includes: If the words in the candidate word pair appear together in at least one query text, the sum of the click counts of the click texts corresponding to the at least one query text is taken as the click count corresponding to the candidate word pair; If each word in the candidate word pair appears in different query texts, the sum of the click counts of the click texts corresponding to the different query texts is used as the click count corresponding to the candidate word pair; The number of clicks corresponding to any word in the candidate word pair is determined according to the sum of the number of clicks of the click text corresponding to at least one query text to which the word belongs.
5. The method according to claim 2, characterized in that: The step of determining a plurality of candidate word pairs in the query text set includes: Using a sliding window of a preset size to process the query text in each query text set, a plurality of candidate word pairs are obtained; Merge the same candidate word pairs in multiple query text sets.
6. The method according to claim 3, characterized in that The determining of word pairs having a hyponymous relationship among the plurality of candidate word pairs based on the number of clicks corresponding to the plurality of candidate word pairs and the number of clicks corresponding to each word in the candidate word pairs includes: Based on the number of clicks corresponding to each word in the candidate word pair, determine the hypernym and hyponym in the candidate word pair, wherein the number of clicks corresponding to the hyponym in the candidate word pair is greater than the number of clicks corresponding to the hypernym in the candidate word pair; Based on the number of clicks corresponding to the plurality of candidate word pairs and the number of clicks respectively corresponding to the hypernym and the hyponym in the candidate word pairs, word pairs having a hypernym relationship among the plurality of candidate word pairs are determined.
7. The method according to claim 6, characterized in that The determining of word pairs having a hyponymy relationship among the plurality of candidate word pairs based on the number of clicks corresponding to the plurality of candidate word pairs and the number of clicks respectively corresponding to the hypernym and the hyponym in the candidate word pairs comprises: Based on the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym and the hyponym in the candidate word pair, respectively, determining whether the candidate word pair satisfies a co-occurrence relationship; In response to the candidate word pair satisfying the co-occurrence relationship, determining that the candidate word pair is a word pair having a hyponymy relationship; The co-occurrence relationship indicates that when a hypernym appears, a hyponym will also appear, but when a hyponym appears, a hypernym may not appear.
8. The method according to claim 7, characterized in that The step of judging whether the candidate word pair satisfies a co-occurrence relationship based on the number of clicks corresponding to the candidate word pair and the number of clicks respectively corresponding to the hypernym and the hyponym in the candidate word pair includes: Determine a first co-occurrence feature value according to the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym in the candidate word pair, and determine that the candidate word pair satisfies a first co-occurrence relationship in response to the first co-occurrence feature value being greater than a first preset threshold; Determine a second co-occurrence feature value according to the number of clicks corresponding to the candidate word pair, and the number of clicks corresponding to the hypernym and the number of clicks corresponding to the hyponym in the candidate word pair, and determine that the candidate word pair satisfies a second co-occurrence relationship in response to the second co-occurrence feature value being greater than a second preset threshold; In the case where the candidate word pair satisfies both the first co-occurrence relationship and the second co-occurrence relationship, it is determined that the candidate word pair satisfies the co-occurrence relationship.
9. The method according to claim 8, characterized in that The first co-occurrence feature value is the quotient of the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym in the candidate word pair; The second co-occurrence feature value is the absolute value of the difference between the quotient of the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hyponym in the candidate word pair and the first co-occurrence feature value.
10. The method according to claim 8, characterized in that The first co-occurrence feature value is the logarithm of the quotient of the number of clicks corresponding to the candidate word pair and the number of clicks corresponding to the hypernym in the candidate word pair; The second co-occurrence feature value is the absolute value of the difference between the quotient of the logarithm of the number of clicks corresponding to the candidate word pair and the logarithm of the number of clicks corresponding to the hyponym in the candidate word pair and the first co-occurrence feature value.
11. The method according to claim 2, characterized in that The step of determining a plurality of candidate word pairs in the query text set includes: Determine a plurality of initial candidate word pairs in the query text set; The multiple initial candidate word pairs are screened to obtain the multiple candidate word pairs.
12. The method according to claim 11, characterized in that The screening process of the plurality of initial candidate word pairs to obtain the plurality of candidate word pairs includes: Among the multiple initial candidate word pairs, initial candidate word pairs whose hyponyms or synonyms of the hyponyms do not appear in the corresponding click text are removed to obtain the multiple candidate word pairs.
13. The method according to claim 11, characterized in that The screening process of the plurality of initial candidate word pairs to obtain the plurality of candidate word pairs includes: Using a large language model to determine whether the click text corresponding to the initial candidate word pair describes the content of the hyponym in the initial candidate word pair; When the click text corresponding to the initial candidate word pair does not describe the content of the hyponym, the initial candidate word pair is removed to obtain the multiple candidate word pairs.
14. The method according to claim 11, characterized in that The screening process of the plurality of initial candidate word pairs to obtain the plurality of candidate word pairs includes: Among the multiple initial candidate word pairs, the initial candidate word pairs containing words of a preset type are removed to obtain the multiple candidate word pairs, where the words of the preset type include one or more of auxiliary words, prepositions, quantifiers, and conjunctions.
15. A text processing device, characterized in that: The device comprises: An acquisition module, used to acquire multiple search records, each search record includes a query text and a click text, wherein the query text represents the text input during the search, and the click text represents the text corresponding to the content clicked and browsed during the search; A division module, used to divide the query text into at least one query text set based on multiple search records; wherein the query text in each query text set corresponds to the same click text; The determination module is used to determine word pairs with a hyponymous relationship based on the query text set.
16. A text processing device, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the method described in any one of claims 1 to 14 when executing the instructions stored in the memory.
17. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 14 is implemented.