Text evolution information extraction method and device, electronic equipment and storage medium

CN117494695BActive Publication Date: 2026-09-11MASHANG CONSUMER FINANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310559802.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-17
Publication Date
2026-09-11
Estimated Expiration
2043-05-17

AI Technical Summary

Technical Problem

[0003]然而,目前对文字的研究技术,均面向目标文字本身涉及的发音、语义、用法等方面进行,无对文字演变过程的研究技术

Benefits of technology

[0017]第五方面,本公开提供了一种计算机程序,该计算机程序存储在计算机可读存储介质中,所述计算机程序在被处理器执行时实现上述的文字演变信息提取方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117494695B_ABST
    Figure CN117494695B_ABST
Patent Text Reader

Abstract

The present disclosure provides a character evolution information extraction method and device, electronic equipment and storage medium. The method comprises: determining a target character and a starting era and a terminal era of an evolution period corresponding to the target character; obtaining p candidate characters from a target component vector of the target character, a target radical vector of the target character and a preset component-radical vector library, wherein the p candidate characters are all characters of the starting era; determining an origin character from the p candidate characters according to a first text set and a second text set; extracting character evolution information from the origin character to the target character from a cross-era text set corresponding to the origin character, wherein the character evolution information comprises: an associated character set that has an influence on the evolution of the origin character to the target character, and an influence degree parameter corresponding to each associated character in the associated character set. It can be seen that the present embodiment provides a character evolution process research technology with high feasibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence (AI) technology, and in particular to a method, apparatus, electronic device, and storage medium for extracting text evolution information. Background Technology

[0002] With the development of science and technology, there are now many different research techniques for text, and the techniques from various research perspectives have become relatively mature, such as character recognition technology, text-to-speech technology, and semantic extraction technology.

[0003] However, current research techniques on writing systems focus on the pronunciation, semantics, and usage of the target script itself, lacking techniques for studying the evolution of writing systems. Therefore, there is an urgent need in this field to develop a technology for learning and researching the evolution of writing systems. Summary of the Invention

[0004] This disclosure provides a method, apparatus, electronic device, and storage medium for extracting text evolution information.

[0005] Firstly, this disclosure provides a method for extracting text evolution information, the method comprising:

[0006] Determine the target character and the start and end eras of its corresponding evolution period;

[0007] Based on the target radical vector of the target character, the target component vector of the target character, and a preset radical vector library, obtain p candidate characters, where each of the p candidate characters is a character from the starting era, and p is an integer greater than or equal to 1.

[0008] The origin character is determined from the p candidate characters according to the first text set and the second text set. The origin character refers to the starting character of the target character in the starting era. The first text set includes at least one text corresponding to the ending era obtained by filtering according to the target character. The second text set includes at least two texts corresponding to the starting era obtained by filtering according to the p candidate characters.

[0009] The text evolution information from the origin character to the target character is extracted from the cross-era text set corresponding to the origin character. The text evolution information includes: a set of related characters that affect the evolution of the origin character to the target character, and an influence degree parameter corresponding to each related character in the set of related characters. The cross-era text set includes the text of the origin character contained between the starting era and the ending era.

[0010] Secondly, this disclosure provides a device for extracting text evolution information, the device comprising:

[0011] The determination module is used to determine the target character and the start and end eras of the corresponding evolution period of the target character;

[0012] The acquisition module is used to acquire p candidate characters based on the target radical vector of the target character, the target radical vector of the target character, and a preset radical vector library. The p candidate characters are all characters from the starting era, and p is an integer greater than or equal to 1.

[0013] The determining module is further configured to determine the origin character from the p candidate characters according to the first text set and the second text set. The origin character refers to the starting character of the target character in the starting era. The first text set includes at least one text corresponding to the ending era obtained by filtering according to the target character. The second text set includes at least two texts corresponding to the starting era obtained by filtering according to the p candidate characters.

[0014] An extraction module is used to extract the text evolution information from the origin character to the target character from the cross-era text set corresponding to the origin character. The text evolution information includes: a set of related characters that affect the evolution of the origin character to the target character, and an influence degree parameter corresponding to each related character in the set of related characters. The cross-era text set includes the text of the origin character contained between the starting era and the ending era.

[0015] Thirdly, this disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the above-described text evolution information extraction method.

[0016] Fourthly, this disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described method for extracting text evolution information.

[0017] Fifthly, this disclosure provides a computer program stored in a computer-readable storage medium, which, when executed by a processor, implements the above-described method for extracting text evolution information.

[0018] In the embodiments provided in this disclosure, after determining the target character and the start and end eras of its evolution period, p candidate characters can be obtained based on the target radical vector, the target component vector, and a preset radical vector library. Each of the p candidate characters represents a character from the start era, and p is an integer greater than or equal to 1. In other words, after obtaining the target character and the start and end eras of its evolution period, the technical solution of this disclosure can select multiple characters from the start era as candidate characters based on their glyphs, thereby facilitating the determination of the origin character of the start era corresponding to the target character from these candidate characters. Subsequently, an origin character is determined from the p candidate characters based on a first text set and a second text set. The origin character refers to the starting character of the target character in the starting era. The first text set includes at least one text corresponding to the ending era selected based on the target character. The second text set includes at least two texts corresponding to the starting era selected based on the p candidate characters. Then, text evolution information from the origin character to the target character is extracted from the cross-era text set corresponding to the origin character. This text evolution information includes: a set of related characters that influence the evolution of the origin character to the target character, and an influence degree parameter for each related character in the set. The cross-era text set includes texts containing the origin character from the starting era to the ending era. That is, after determining multiple candidate characters, the technical solution of this disclosure obtains text evolution information from the origin character to the target character from a semantic dimension. Specifically, the origin character is determined based on the degree of semantic change of the text after character substitution between the text sets of the starting era and the text sets of the ending era. Furthermore, by using the semantic evolution of the cross-era texts of the origin character, the text evolution information from the origin character to the target character can be extracted. As can be seen, in this embodiment of the present disclosure, after determining the evolution era and the target character after evolution, it can first screen candidate characters of the starting era by glyph, and then select the origin character corresponding to the target character from the candidate characters by the semantics of each candidate character. Then, it can further extract the text evolution information from the origin character to the target character by the semantics of the text environment in which the origin character is located, thereby providing a highly feasible technology for studying the text evolution process.

[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0020] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which:

[0021] Figure 1 A flowchart illustrating a method for extracting text evolution information provided in this embodiment of the disclosure;

[0022] Figure 2 A flowchart illustrating a method for generating a radical vector library provided in this embodiment of the disclosure;

[0023] Figure 3 A flowchart illustrating an exemplary method for extracting text evolution information provided in this disclosure embodiment;

[0024] Figure 4 A block diagram of a text evolution information extraction device provided in this embodiment of the present disclosure;

[0025] Figure 5 This is a block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation

[0026] To enable those skilled in the art to better understand the technical solutions of this disclosure, exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0027] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.

[0028] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.

[0029] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Words such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.

[0030] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.

[0031] This disclosure relates to the field of natural language processing (NLP) technology, and mainly to the technology for extracting information on the evolution of text over time.

[0032] Conventional text processing techniques based on NLP (Natural Language Processing) mostly focus on aspects of the target text itself, such as pronunciation, semantics, and usage. For example, Bidirectional Encoder Representation from Transformers (BERT) is used to predict text based on its semantics; and Optical Character Recognition (OCR) is used to translate characters into text. Currently, there are no research techniques that examine the evolution of text over time.

[0033] This disclosure provides a method for extracting text evolution information. After determining the target character and its corresponding evolution period, the method selects characters from the beginning of multiple evolution periods as candidate characters based on the radicals and components of the target character from the perspective of character shape. Then, the method selects the origin character corresponding to the target character from the semantics of each candidate character. Finally, the method extracts the text evolution information from the origin character to the target character through the semantics of the text environment in which the origin character is located. This provides a highly feasible technique for studying the text evolution process.

[0034] The text evolution information extraction method according to embodiments of this disclosure can be executed by an electronic device, which may be an in-vehicle device, user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, wearable device, etc. The method can be implemented by a processor calling computer-readable program instructions stored in memory. The electronic device can be used as a terminal device or a server.

[0035] Figure 1 A flowchart illustrating a method for extracting text evolution information provided in this embodiment of the disclosure. (Refer to...) Figure 1 The method includes:

[0036] In step S11, the target character and the start and end eras of its corresponding evolution period are determined.

[0037] The electronic device in this embodiment may include a human-computer interaction interface, through which the electronic device can receive a target character input by a user, as well as the start and end eras of the evolution period of the corresponding target character.

[0038] For example, the human-computer interaction interface may be a metaverse interaction interface for virtual reality.

[0039] The evolution period described in this embodiment can be a time period from the beginning era to the end era. The beginning era and the end era can be any historical era or dynasty, for example, the beginning era is the Qin Dynasty and the end era is the Ming Dynasty.

[0040] In some implementations, the number of target words for the termination era can be at least one, and this disclosure does not limit this.

[0041] In step S12, p candidate characters are obtained based on the target radical vector of the target character, the target component vector of the target character, and a preset radical vector library.

[0042] Wherein, the p candidate characters are all characters from the starting era, and p is an integer greater than or equal to 1.

[0043] It should be noted that the characters involved in the embodiments of the present disclosure are Chinese characters. Chinese characters are usually formed by combining radicals and components, wherein radicals and components can be used to indicate pronunciation and meaning respectively in a Chinese character, and radicals and components can be combined according to the positional relationships of top, bottom, left and right. For example, the Chinese character "江(river)" is formed by combining the radical "氵" on the left and the component "工" on the right. The radical "氵" can be used to indicate the meaning of the Chinese character "江", and the component "工" can be used to indicate the pronunciation of the Chinese character "江"; for another example, the Chinese character "驾(drive)" is formed by combining the radical "加" on the top and the component "马" on the bottom. The radical "加" can be used to indicate the pronunciation of the Chinese character "驾", and the component "马" can be used to indicate the meaning of the Chinese character "驾". It should be understood that some Chinese characters are not obtained by combining two parts, such as the Chinese characters "皿(dish)", "首(head)" and the like. For such Chinese characters, the Chinese character itself can be treated as a radical or a component.

[0044] In view of this, the embodiments of the present disclosure can decompose a Chinese character into the most basic radical part and component part, and then determine the character through the vector of the radical part and the vector of the component part. Optionally, before step S12, the embodiments of the present disclosure pre-set a radical vector library, and the preset radical vector library includes radical vectors and component vectors of a plurality of characters. Exemplarily, the radical vector library can be obtained by processing an existing character library by using a radical vector generation model.

[0045] In some implementation manners, if the application scenario of the embodiments of the present disclosure is a general scenario, the electronic device may invoke a preset radical vector generation model to process characters in a first preset character library, and generate the radical vector library. The first preset character library can be existing general character corpora such as dictionaries and ancient classics.

[0046] In other implementation manners, if the application scenario of the embodiments of the present disclosure includes a specified application scenario, for example, the specified application scenario is an Internet application scenario, the electronic device may first generate a general radical vector library, then optimize the general radical vector library according to specified semantics of some characters in the specified application scenario, and use the optimized radical vector library as the preset radical vector library.

[0047] For example, if the second preset character library contains a specified semantic character, the preset radical vector generation model is invoked to process the characters in the second preset character library to generate an initial radical vector library; the phonetic and semantic feature distributions of the radicals of the specified semantic character are obtained, as well as the phonetic and semantic feature distributions of the radicals of the specified semantic character; the phonetic and semantic feature vectors of the radicals of the specified semantic character are obtained by feature extraction from text sample pairs, and added to the initial radical vector library to obtain the radical vector library; the text sample pairs include positive text samples and negative text samples, the positive text samples include the specified semantic character, and the negative text samples include characters different from the specified semantic character. The phonetic feature distribution includes the feature distribution of pronunciation; the semantic feature distribution includes the feature distribution of semantics. The features of pronunciation include the features of the tongue in space, the articulation point, etc. during pronunciation. The features of semantics include features representing color, time, living or non-living objects, etc.

[0048] Examples of generating radical vector libraries for specific application scenarios are detailed below. Figure 2 The illustrated embodiments are described herein and will not be repeated here.

[0049] Further, after step S11, the electronic device can obtain the target radical vector corresponding to the radical of the target character and the target radical vector corresponding to the radical of the target character from the preset radical vector library. Then, based on vector similarity, it obtains x radical vectors from the preset radical vector library according to the target radical vector and y radical vectors according to the target radical vector, where x and y are both integers greater than or equal to 1. Then, according to the combination rules, the x radical vectors and the y radical vectors are combined, and p1 character vectors are obtained based on the combination result, where p1 is greater than or equal to p. Based on the preset character vector library, p character vectors are selected from the p1 character vectors, and the characters corresponding to the p character vectors are the P candidate characters.

[0050] For example, vector similarity can be implemented as the cosine distance between vectors. The larger the cosine distance between two vectors, the smaller the similarity between them; conversely, the smaller the cosine distance, the larger the similarity. Therefore, after determining the target radical vector and target component vector of the target character, the electronic device can calculate the cosine distance between the target radical vector and each radical vector in a preset radical vector library, and then sequentially determine x radical vectors according to the order of cosine distance from smallest to largest. Similarly, it can calculate the cosine distance between the target component vector and each component vector in the preset radical vector library, and then sequentially determine y component vectors according to the order of cosine distance from smallest to largest.

[0051] The combination rules for radical vectors and component vectors can be pre-set according to the positional relationship of the radicals and components of the characters, such as radicals to the left of the component, radicals above the component, radicals to the right of the component, radicals below the component, etc. After obtaining x radical vectors and y component vectors, the electronic device can use permutation and combination to combine the x radical vectors and y component vectors according to the combination rules to obtain p1 character vectors.

[0052] It should be noted that any one of the p1 character vectors may correspond to a character or a symbol rather than a character. Even if the corresponding character vector corresponds to a character, it may not be the character from the initial era. Based on this, in this embodiment of the disclosure, a character vector library can be pre-deployed, which includes multiple character vectors from various eras. After obtaining the p1 character vectors, the electronic device can filter out p character vectors from the p1 character vectors according to the character vector library. These p character vectors are the character vectors of the p candidate characters. For example, for each character vector in the p1 character vectors, the electronic device can detect whether the character vector is recorded in the character vectors of the initial era in the character vector library. If so, the character vector is retained; if not, the character vector is deleted from the p1 character vectors, so that p character vectors are ultimately retained.

[0053] By adopting this implementation method, several radicals similar to the radicals of the target character are obtained, and several radicals similar to the components of the target character are obtained. Then, they are combined to obtain p candidate characters. This can ensure that the shape of the obtained p candidate characters has a high similarity to the shape of the target character. That is, the embodiments of this disclosure first select multiple characters from the starting era as candidate characters by shape, so as to determine the origin character of the starting era corresponding to the target character from these candidate characters, which is beneficial to improving the accuracy of the origin character corresponding to the target character.

[0054] In step S13, the origin character is determined from p candidate characters based on the first text set and the second text set.

[0055] Wherein, the originating character refers to the starting character of the target character in the starting era. The first text set includes at least one text corresponding to the ending era obtained by filtering based on the target character, and the second text set includes at least two texts corresponding to the starting era obtained by filtering based on the p candidate characters.

[0056] It should be noted that the embodiments of this disclosure include a preset text corpus, which includes text data from multiple eras. In view of this, before determining the originating character from the p candidate characters based on the first text set and the second text set, the electronic device can select a first preset number of texts to form the first text set based on the word frequency of the target character in the corpus of the terminating era of the preset text corpus, and select a second preset number of texts to form the second text set based on the word frequency of the P candidate characters in the corpus of the starting era of the preset text corpus.

[0057] It should be understood that the first preset quantity and the second preset quantity can be the same or different. In some implementations, the second preset quantity can be greater than the first preset quantity. For example, the first preset quantity is 8 and the second preset quantity is 20.

[0058] In some implementations, after determining the target character in step S11, the electronic device can calculate the frequency of the target character in each text corpus of the termination era, i.e., the ratio of the total number of times the target character appears in the text corpus to the total number of characters in the text corpus. Then, a first preset number of texts can be determined sequentially according to the word frequency from high to low, and the determined first preset number of texts is used as the first text set. After obtaining p candidate characters in step S12, for each text corpus corresponding to the starting era, the electronic device can calculate the word frequency of each of the p candidate characters in the text corpus, i.e., the ratio of the total number of times each candidate character appears in the text corpus to the total number of characters in the text corpus. Then, the sum of the word frequencies of the p candidate characters can be used as the total word frequency of the p candidate characters in the text corpus. Then, a second preset number of texts can be determined sequentially according to the total word frequency from high to low, and the determined second preset number of texts is used as the second text set.

[0059] Furthermore, corresponding to the p candidate characters, the first text set, and the second text set, the electronic device generates replacement text sets respectively, resulting in p replacement text sets. Then, the electronic device can determine the acceptability of each replacement text set in the p replacement text sets. In this embodiment, the acceptability of each replacement text set includes the acceptability of each replacement text in the corresponding replacement text set, whereby the acceptability of each replacement text is used to characterize the degree of semantic change of the corresponding replacement text compared to its original semantics. Based on this, the electronic device can determine the candidate character corresponding to the replacement text set with the highest acceptability in the p replacement text sets as the origin character.

[0060] For example, taking the first replacement text set in the p replacement text sets as an example, the first replacement text set includes a first sub-text set and a second sub-text set. The first sub-text set is obtained by replacing the target word in the first text set with a first candidate word, and the second sub-text set is obtained by replacing the first candidate word in the second text set with the target word. The first candidate word is any one of the p candidate words.

[0061] For example, taking the second replacement text set in a set of p replacement texts as an example, determining the acceptability of each replacement text set in the set of p replacement texts includes: determining the acceptability of each replacement text in the second replacement text set respectively, and determining the sum of the acceptability of each replacement text as the acceptability of the second replacement text set.

[0062] For example, the acceptability of replacement text can be expressed as a perplexity (ppl) value. Correspondingly, the ppl value of the second set of replacement texts is the sum of the ppl values ​​of each replacement text in the second set. A larger ppl value indicates greater semantic "perplexity" of the corresponding text, indicating lower semantic acceptability; conversely, a smaller ppl value indicates less semantic "perplexity" of the corresponding text, indicating higher semantic acceptability. Therefore, in this example, the set of replacement texts with the highest acceptability is the set with the smallest ppl value.

[0063] It should be noted that in this embodiment, the electronic device can call a character model to process the radical vector and the component vector. Since the target character and candidate characters are from different eras, at least one of the radicals of the target character and the candidate characters, or the components of the target character and the candidate characters, will be different. Therefore, before calling the character model, the electronic device can adjust some parameters of the character model according to the characters of the corresponding era, so that the character model is suitable for processing characters of that era. For example, before obtaining the target radical vector and the target component vector of the target character, the electronic device can adjust some parameters of the character model to obtain a first character model suitable for processing characters from the end era, and then call the first character model to obtain the target radical vector and the target component vector. Before obtaining p candidate characters, the electronic device can adjust some parameters of the character model to obtain a second character model suitable for processing characters from the beginning era, and then call the second character model to process and obtain p candidate characters.

[0064] As can be seen, after determining p candidate characters based on their shapes, this implementation method can further filter out the origin characters corresponding to the target characters from the perspective of the usage of the characters in the text and their semantics, thereby obtaining origin characters that are similar to the target characters in both shape and semantics.

[0065] In step S14, the text evolution information from the origin character to the target character is extracted from the cross-era text set corresponding to the origin character.

[0066] The text evolution information includes: a set of related characters that influence the evolution of the origin character into the target character, and an influence degree parameter for each related character in the set. The cross-era text set includes texts from the starting era to the ending era that contain the origin character.

[0067] In some implementations, after determining the originating character, the electronic device can obtain the cross-era text set from the preset text corpus, and then extract the semantic features of each cross-era text in the cross-era text set. For example, the semantic features of each cross-era text can characterize the context in which the originating character is used. Then, the electronic device determines the related characters from the corresponding cross-era texts based on the semantic features of each cross-era text. All related characters corresponding to the cross-era text set constitute the related character set. Furthermore, the electronic device calculates the influence degree parameter of each related character in the related character set based on the semantic features of the target character. The related character set and the influence degree parameter corresponding to each related character in the related character set constitute the text evolution information.

[0068] For example, after acquiring the cross-era text set, the electronic device can perform word segmentation on each cross-era text, then determine the character vector of each segmented character, and then calculate the correlation parameter between each character vector and the character vector of the originating character through an attention mechanism. Furthermore, the electronic device can identify characters corresponding to character vectors with correlation parameters greater than a certain threshold as associated characters, and use the correlation parameters corresponding to these associated characters as influence parameters.

[0069] It should be understood that the correlation parameter can characterize the degree of influence of a character on the development of its origin character into its target character. For example, the correlation parameter can be implemented as a probability value. The larger the probability value, the greater the influence of the character on the development of its origin character into its target character; the smaller the probability value, the smaller the influence of the character on the development of its origin character into its target character.

[0070] Furthermore, after obtaining the information on the evolution of the characters, the electronic device can also visualize and display the information on the evolution of the characters based on the metaverse.

[0071] It should be understood that the implementation methods described in steps S11 to S14 above can be used in scenarios where the terminator is a single character or multiple characters. If the terminator is a single character, that character is the target character; if the terminator is multiple characters, each character can be used as a target character, and multiple threads can be used to extract the text evolution information of the multiple target characters respectively.

[0072] In summary, the embodiments of this disclosure, after determining the evolution period and the target character after evolution, can first screen candidate characters of the initial era by character shape, then select the origin character corresponding to the target character from the candidate characters by semantics, and then further extract the evolution information from the origin character to the target character by semantics of the text environment in which the origin character is located, thereby providing a highly feasible technology for studying the evolution process of characters.

[0073] The following describes the text evolution information extraction method provided in the embodiments of this disclosure with reference to examples.

[0074] Taking the metaverse display scene provided by an electronic device as an example, and combining it with the interaction between the user, the technology of the embodiments of this disclosure will be described.

[0075] Based on the above description of the implementation method, the implementation process of this embodiment can be divided into two implementation stages, which include the generation of the radical vector library and the extraction of character evolution information.

[0076] Generation of radical vector library

[0077] This example describes a method for generating a radical vector library by taking a specific application scenario as an example.

[0078] See Figure 2 , Figure 2 a method for generating a radical vector library is schematically illustrated, which is used to generate a radical vector library for use in a specific application scenario. The method may comprise the following steps:

[0079] Step S21, calling a preset radical vector generation model to process characters in a second preset character library, and generating an initial radical vector library.

[0080] Wherein, the second preset character library refers to a library containing specified semantic characters. Specified semantic characters refer to characters that have general meanings in general scenarios but express specific semantics when used in a specific application scenario. For example, in a general scenario, the general meaning of "luodi" (landing) is falling onto the ground, while when used in the Internet field, it means that a certain technology or product is about to be released. The initial radical vector library is a general radical vector library, that is, component vectors and radical vectors obtained based on the general semantics of each character. Both the component vectors and radical vectors contained in the initial radical vector library are two-dimensional vectors.

[0081] In this example, semantic basic sememes and semantic specified sememe sequences can be respectively labeled for the component vectors and radical vectors of specified semantic characters in the initial radical vector library. Semantic basic sememes may be referred to as explicit sememes, and semantic specified sememes may be referred to as implicit sememes.

[0082] Step S22, obtaining the phonetic sememe distribution and semantic sememe distribution of the component of the specified semantic character, and the phonetic sememe distribution and semantic sememe distribution of the radical of the specified semantic character.

[0083] After obtaining the semantic basic sememes and semantic specified sememe sequences of the component vector of the specified semantic character, as well as the semantic basic sememes and semantic specified sememe sequences of the radical vector, the phonetic sememe distribution and semantic sememe distribution of the component of the specified semantic character, and the phonetic sememe distribution and semantic sememe distribution of the radical of the specified semantic character can be obtained based on each sememe.

[0084] For example, the phonetic space of each phonetic sememe distribution of components and radicals can be divided according to initials, finals, tones, consonants and vowels. The phonetic space refers to a spatial feature that reflects the pronunciation position, the change trend and position of the tongue when each phoneme is pronounced.

[0085] For example, semantic feature distribution can be deployed across several dimensions, including a category space and an intensional space. Based on this, all semantic features in the semantic base features and semantic index feature sequences can be classified, for example, into syntactic features, category features, and intensional features. Then, a category space and an intensional space can be constructed based on each type of feature. For example, each type of feature can be mapped to a high-dimensional vector space containing all feature vectors of that type. For example, a category space and an intensional space can be constructed based on the acquired feature sequence. For example, the category feature space can include: biological (e.g., human) and non-biological (e.g., wood). The intensional feature space can include: a time dimension (e.g., very long), a spatial dimension (e.g., square), a color dimension (e.g., white), etc.

[0086] Step S23: Construct text sample pairs.

[0087] The text sample pairs are used for text comparison tasks. The purpose of text comparison tasks is to uncover implicit semantic features of words through word-to-word and text-to-text comparisons. Based on this, text sample pairs include positive and negative text samples. Positive text samples contain text with a specified semantic meaning, while negative text samples contain text with a different semantic meaning, or the specified semantic meaning in the negative text sample has a common meaning in the negative text sample. That is, the positive text sample is a valid sentence, and the negative text sample is an invalid sentence; they differ by only one word. For example, a positive text sample might be: "Product launch conference"; a negative text sample might be: "Two iron balls fall simultaneously."

[0088] Furthermore, the semantic vectors of the radicals and the semantic vectors of the components can be obtained from the phonetic and semantic distributions of the radicals and the phonetic and semantic distributions of the components obtained in step S22, in order to obtain the semantic vectors of the radicals and the semantic vectors of the components of each character in the positive and negative text samples.

[0089] Step S24: Perform feature extraction on the text sample pair to obtain the phonetic semantic vector and semantic semantic vector of the radical of the specified semantic character.

[0090] In this example, the text samples obtained in step S23 can be input into the machine learning model to train it. Since the semantic feature vectors of each character and the semantic feature sequences that distinguish between characters are used in the training of the machine learning model, the model can extract richer and more accurate semantic features that reflect the corresponding characters. Based on this, during the training of the machine learning model using text samples, the phonetic and semantic feature vectors of the radicals of a specified semantic character, as well as the phonetic and semantic feature vectors of the components of a specified semantic character, can be obtained.

[0091] In step S25, adding the phonetic sememe vector and semantic sememe vector of the radical of the specified semantic character to the initial radical vector library to obtain a radical vector library.

[0092] Wherein, both the component vectors and radical vectors contained in the radical vector library are two-dimensional vectors.

[0093] Adding the obtained phonetic sememe vector and semantic sememe vector of the component of the specified semantic character, as well as the phonetic sememe vector and semantic sememe vector of the radical of the specified semantic character, to the initial radical vector library, so as to optimize the component vector and radical vector of the specified character in the initial radical vector library, and obtain a radical vector library with stronger representational capability.

[0094] It can be seen that with this implementation method, by performing comparison when characters in text sample pairs are different or the same character has different semantics, the implicit sememes of characters with specified semantics can be further mined, so as to mine the sememe vectors of characters in specified application scenarios, so that the obtained radical vector library has stronger representational capability.

[0095] In addition, in some other implementation manners, after step S24, step S26 may further be included. In step S26, splicing the phonetic sememe vector and semantic sememe vector of the component of each character in a certain order to obtain a one-dimensional component vector of the character, and splicing the phonetic sememe vector and semantic sememe vector of the radical of each character in a certain order to obtain a one-dimensional radical vector of the character, then combining the one-dimensional component vector and the one-dimensional radical vector of the character to obtain a character vector of the character.

[0096] Wherein, for sememe vectors of components or radicals, splicing can be performed in the order of splicing phonetic sememe vectors first and then semantic sememe vectors. For example, in the process of splicing semantic sememe vectors, sememe vectors in the category space can be spliced first, and then sememe vectors in the connotation space can be spliced.

[0097] In addition, if a character has no component or no radical, the sememe vector of the non-existent part of the character is set to 0. For example, the character "反" has no component, then the component sememe vector of the character "反" can be set to 0.

[0098] In step S27, for each character vector, generating a mapping relationship between the character vector and the component vector and radical vector of the character in the radical vector library.

[0099] Wherein, the mapping relationship between the character vector and the component vector and radical vector of the character in the radical vector library can be implemented as a conversion equation between the character vector and the component vector and radical vector.

[0100] Extraction of character evolution information

[0101] After obtaining the radical vector library, the electronic device can respond to the user's input information and perform character evolution information extraction based on the obtained radical vector library.

[0102] See Figure 3 , Figure 3 A flowchart illustrating an exemplary method for extracting text evolution information provided in this disclosure. The method includes:

[0103] Step S30: Receive user input of Song Dynasty, Modern, and character a.

[0104] In this example, the Song Dynasty is the starting era, and the ending era, such as the modern era, is the ending era. Character 'a' is a modern character, which is the target character of the ending era.

[0105] Step S31: Determine the first character model, which is used to process modern characters.

[0106] Step S32: Call the first text model to select the top-n1 texts from the modern texts in the preset text corpus according to the word frequency of character 'a', and use the top-n1 texts as the first text set.

[0107] For example, n1 can be 8, then the first text set contains 8 texts.

[0108] Step S33: Call the first character model to obtain the target radical vector corresponding to the radical of character a and the target radical vector corresponding to the radical of character a from the radical vector library.

[0109] Step S34: Obtain x radical vectors that are similar to the target radical vector from the radical vector library in descending order of similarity, and obtain y radical vectors that are similar to the target radical vector from the radical vector library in descending order of similarity.

[0110] Step S35: Combine x radical vectors and x component vectors according to the character combination rules to obtain p1 character vectors.

[0111] Step S36: According to the mapping relationship between the radical vector library and the character vectors, select p character vectors from p1 character vectors that match the Song Dynasty character vectors in the character vector library. These p character vectors correspond to p Song Dynasty characters.

[0112] Among them, the p characters from the Song Dynasty are candidate characters for character a in the Song Dynasty. The following uses characters b1 to bp to represent the p candidate characters.

[0113] Step S37: Determine the second character model, which is used to process characters from the Song Dynasty.

[0114] Step S38: Call the second text model to select the top-n2 texts from the Song Dynasty texts in the preset text corpus according to the word frequency of characters b1 to bp, and use the top-n2 texts as the second text set.

[0115] For example, n2 could be 15, so the second text set contains 15 texts.

[0116] Step S39: For the corresponding character bi, replace character a in the first text set with character bi and replace character bi in the second text set with character a to obtain the replacement text set corresponding to character bi.

[0117] Among them, bi belongs to b1 to bp.

[0118] Step S310: Calculate the ppl value of the replacement text set corresponding to each character in b1 to bp.

[0119] Taking the replacement text set corresponding to the character bi as an example, the replacement text set corresponding to bi includes 23 replacement texts (the first text set includes 8 texts, and the second text set includes 15 texts). The ppl value of each of the 23 replacement texts is calculated to obtain 23 ppl values. The sum of the 23 ppl values ​​is the ppl value of the replacement text set corresponding to bi.

[0120] Step S311: The character bx corresponding to the replacement text set with the smallest ppl value is taken as the origin character of character a in the Song Dynasty.

[0121] Among them, bx belongs to b1 to bp.

[0122] Step S312: Obtain a cross-era text set containing the character bx from the Song Dynasty to the present from the text corpus.

[0123] Step S313: Extract the words that influence word a in each text of the cross-era text set as related words of word a.

[0124] Step S314: Determine the parameter of the degree of influence of the associated character on the evolution of character bx into character a.

[0125] Furthermore, in the metaverse, related words and their corresponding influence parameters can be displayed in a virtual reality manner.

[0126] As can be seen, by employing the character evolution information extraction method provided in this embodiment, after determining the target character and the start and end eras of its evolution period, p candidate characters can be obtained based on the target radical vector, the target component vector, and a preset radical vector library. Each of the p candidate characters represents a character from the start era, where p is an integer greater than or equal to 1. In other words, after obtaining the target character and the start and end eras of its evolution period, the technical solution of this disclosure can select multiple characters from the start era as candidate characters based on their glyphs, thereby facilitating the determination of the origin character of the target character's corresponding start era from these candidate characters. Subsequently, an origin character is determined from the p candidate characters based on a first text set and a second text set. The origin character refers to the starting character of the target character in the starting era. The first text set includes at least one text corresponding to the ending era selected based on the target character. The second text set includes at least two texts corresponding to the starting era selected based on the p candidate characters. Then, text evolution information from the origin character to the target character is extracted from the cross-era text set corresponding to the origin character. This text evolution information includes: a set of related characters that influence the evolution of the origin character to the target character, and an influence degree parameter for each related character in the set. The cross-era text set includes texts containing the origin character from the starting era to the ending era. That is, after determining multiple candidate characters, the technical solution of this disclosure obtains the evolution information from the origin character to the target character from a semantic dimension. Specifically, the origin character is determined based on the degree of semantic change of the text after character substitution between the text sets of the starting era and the text sets of the ending era. Furthermore, by using the semantic evolution of the cross-era texts of the origin character, the evolution information from the origin character to the target character can be extracted. As can be seen, in this embodiment of the present disclosure, after determining the evolution era and the target character after evolution, it is possible to first screen candidate characters of the starting era by glyph, and then select the origin character corresponding to the target character from the candidate characters by the semantics of each candidate character. Then, the evolution information from the origin character to the target character is further extracted by the semantics of the text environment in which the origin character is located, thereby providing a highly feasible technology for studying the evolution process of characters.

[0127] It is understood that the various method embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further. Those skilled in the art will understand that in the above methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.

[0128] In addition, this disclosure also provides a device for extracting text evolution information, an electronic device, and a computer-readable storage medium, all of which can be used to implement any of the text evolution information extraction methods provided in this disclosure. The corresponding technical solutions and descriptions are described in the corresponding records in the method section and will not be repeated here.

[0129] Figure 4 This is a block diagram of a text evolution information extraction device provided in an embodiment of the present disclosure.

[0130] Reference Figure 4 This disclosure provides a text evolution information extraction device, which includes a determining module 41, an acquiring module 42, and an extraction module 43. Each module, when running, can implement some or all of the functions described in the above method implementation.

[0131] For example, the determining module 41 is used to determine the target character and the start and end eras of the corresponding evolution period of the target character; the acquiring module 42 is used to acquire p candidate characters based on the target radical vector of the target character, the target component vector of the target character, and a preset radical vector library, wherein the p candidate characters are all characters of the start era, and p is an integer greater than or equal to 1; the determining module 41 is also used to determine the origin character from the p candidate characters based on a first text set and a second text set, wherein the origin character refers to the starting character of the target character in the start era, and the first text set includes characters based on... The target character is selected to obtain at least one text corresponding to the termination era, and the second text set includes at least two texts corresponding to the starting era selected based on the p candidate characters; the extraction module 43 is used to extract the text evolution information from the origin character to the target character from the cross-era text set corresponding to the origin character, the text evolution information includes: a set of related characters that affect the evolution of the origin character to the target character, and an influence degree parameter corresponding to each related character in the set of related characters, and the cross-era text set includes the texts containing the origin character between the starting era and the termination era.

[0132] For details on the specific implementation method, please refer to the above. Figures 1 to 3 The implementation method shown is not described in detail here.

[0133] It is understandable that the above division of modules is only a logical functional division. In actual implementation, each of the above modules can be integrated into the hardware implementation. For example, the function of the determination module 41 in the above implementation can be integrated into the I / O interface implementation, and the functions of the acquisition module 42 and the extraction module 43 can be integrated into the processor implementation.

[0134] Reference Figure 5 , Figure 5This disclosure provides an electronic device comprising: at least one processor 501; at least one memory 502; and one or more I / O interfaces 503 connected between the processor 501 and the memory 502; wherein the memory 502 stores one or more computer programs executable by the at least one processor 501, the one or more computer programs being executed by the at least one processor 501 to enable the at least one processor 501 to perform the above-described text evolution information extraction method.

[0135] This disclosure also provides a computer-readable storage medium, which may be a volatile or non-volatile computer-readable storage medium, on which a computer program is stored, wherein the computer program, when executed by the processor 501:

[0136] The process involves: determining the target character and the starting and ending eras of its evolution period; obtaining p candidate characters based on the target radical vector, the target component vector, and a pre-defined radical vector library, where each of the p candidate characters represents a character from the starting era, and p is an integer greater than or equal to 1; determining the origin character from the p candidate characters using a first text set and a second text set, where the origin character is the starting character of the target character in the starting era, the first text set includes at least one text corresponding to the ending era selected based on the target character, and the second text set includes at least two texts corresponding to the starting era selected based on the p candidate characters; extracting the character evolution information from the origin character to the target character from the cross-era text set corresponding to the origin character, where the character evolution information includes: a set of related characters that influence the evolution of the origin character to the target character, and an influence degree parameter for each related character in the set of related characters, and the cross-era text set includes texts containing the origin character between the starting era and the ending era.

[0137] In some embodiments, the processor 501 is further configured to: obtain, from the preset radical vector library, a target radical vector corresponding to the radical of the target character and a target radical vector corresponding to the radical of the target character; obtain x radical vectors from the preset radical vector library based on the target radical vectors and y radical vectors based on the target radical vectors, where x and y are both integers greater than or equal to 1; combine the x radical vectors and the y radical vectors according to a combination rule, and obtain p1 character vectors based on the combination result, where p1 is greater than or equal to p; and select p character vectors from the p1 character vectors according to the preset character vector library, where the characters corresponding to the p character vectors are the P candidate characters.

[0138] In some embodiments, the processor 501 is further configured to: select a first preset number of texts to form a first text set based on the word frequency of the target character in the corpus of the termination era of the preset text corpus; and select a second preset number of texts to form a second text set based on the word frequency of the P candidate characters in the corpus of the beginning era of the preset text corpus.

[0139] In some embodiments, the processor 501 is further configured to: generate replacement text sets corresponding to the p candidate characters, the first text set, and the second text set, respectively, to obtain p replacement text sets, wherein the first replacement text set in the p replacement text sets includes: a first sub-text set and a second sub-text set, wherein the first sub-text set is obtained by replacing the target character in the first text set with a first candidate character, and the second sub-text set is obtained by replacing the first candidate character in the second text set with the target character, wherein the first candidate character is any one of the p candidate characters; determine the acceptability of each replacement text set in the p replacement text sets, wherein the acceptability of each replacement text set includes the acceptability of each replacement text in the corresponding replacement text set, wherein the acceptability of each replacement text is used to characterize the degree of semantic change of the corresponding replacement text compared to the semantic change before the corresponding replacement text was replaced; and determine the candidate character corresponding to the replacement text set with the highest acceptability in the p replacement text sets as the origin character.

[0140] In some embodiments, the processor 501 is further configured to: determine the acceptability of each replacement text in the second replacement text set, wherein the second replacement text set is any one of the p replacement text sets; and determine the sum of the acceptability of each replacement text as the acceptability of the second replacement text set.

[0141] In some embodiments, the processor 501 is further configured to: obtain the cross-era text set from the preset text corpus; extract the semantic features of each cross-era text in the cross-era text set; determine the related characters from the corresponding cross-era texts based on the semantic features of each cross-era text, wherein all related characters corresponding to the cross-era text set constitute the related character set; calculate the influence degree parameter of each related character in the related character set based on the semantic features of the target character, wherein the related character set and the influence degree parameter corresponding to each related character in the related character set constitute the text evolution information.

[0142] In some embodiments, the processor 501 is further configured to: generate the radical vector library using any of the following processing methods, the processing methods including: calling a preset radical vector generation model to process the characters in a first preset character library to generate the radical vector library; or, if a second preset character library contains a specified semantic character, calling the preset radical vector generation model to process the characters in the second preset character library to generate an initial radical vector library; obtaining the phonetic semantic feature distribution and semantic semantic feature distribution of the radical of the specified semantic character, and the phonetic semantic feature distribution and semantic semantic feature distribution of the radical of the specified semantic character; obtaining the phonetic semantic feature vector and semantic semantic feature vector of the radical of the specified semantic character, and the phonetic semantic feature vector and semantic semantic feature vector of the radical of the specified semantic character by performing feature extraction on text sample pairs, and adding them to the initial radical vector library to obtain the radical vector library; the text sample pair includes positive text samples and negative text samples, the positive text samples include the specified semantic character, and the negative text samples include characters different from the specified semantic character.

[0143] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code is run in the processor of an electronic device, the processor in the electronic device executes the above-described text evolution information extraction method.

[0144] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).

[0145] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable program instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0146] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0147] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0148] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0149] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0150] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0151] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0152] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0153] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.

Claims

1. A method of extracting information on the evolution of characters, characterized by, The method includes: Determine the target character and the start and end eras of its corresponding evolution period; Based on the target radical vector of the target character, the target component vector of the target character, and a preset radical vector library, obtain p candidate characters, where each of the p candidate characters is a character from the starting era, and p is an integer greater than or equal to 1. The origin character is determined from the p candidate characters according to the first text set and the second text set. The origin character refers to the starting character of the target character in the starting era. The first text set includes at least one text corresponding to the ending era obtained by filtering according to the target character. The second text set includes at least two texts corresponding to the starting era obtained by filtering according to the p candidate characters. The text evolution information from the origin character to the target character is extracted from the cross-era text set corresponding to the origin character. The text evolution information includes: a set of related characters that affect the evolution of the origin character to the target character, and an influence degree parameter corresponding to each related character in the set of related characters. The cross-era text set includes the text of the origin character contained between the starting era and the ending era. The step of determining the origin character from the p candidate characters based on the first text set and the second text set includes: For each of the p candidate characters, the first text set, and the second text set, a replacement text set is generated to obtain p replacement text sets; Determine the acceptability of each of the p replacement text sets; The candidate character corresponding to the most acceptable replacement text set in the p replacement text sets is determined as the origin character.

2. The character evolution information extraction method according to claim 1, characterized by, The step of obtaining p candidate characters based on the target radical vector of the target character, the target component vector of the target character, and a preset radical vector library includes: From the preset radical vector library, obtain the target radical vector corresponding to the radical of the target character and the target radical vector corresponding to the radical of the target character; Based on vector similarity, x radical vectors are obtained from the preset radical vector library according to the target radical vector, and y radical vectors are obtained according to the target radical vector, where x and y are both integers greater than or equal to 1. According to the combination rules, the x radical vectors and the y component vectors are combined, and p1 character vectors are obtained based on the combination results, where p1 is greater than or equal to p; Based on a preset character vector library, p character vectors are selected from the p1 character vectors, and the characters corresponding to the p character vectors are the p candidate characters.

3. The character evolution information extraction method according to claim 1, characterized by, Before determining the origin character from the p candidate characters based on the first text set and the second text set, the method further includes: Based on the word frequency of the target character in the corpus of the terminology of the terminology of the terminology of the preset text corpus, a first preset number of texts are selected to form the first text set; Based on the word frequencies of the p candidate characters in the corpus of the initial era of the preset text corpus, a second preset number of texts are selected to form the second text set.

4. The character evolution information extraction method according to claim 1, characterized by, The first replacement text set in the p replacement text sets includes: a first sub-text set and a second sub-text set. The first sub-text set is obtained by replacing the target character in the first text set with a first candidate character. The second sub-text set is obtained by replacing the first candidate character in the second text set with the target character. The first candidate character is any one of the p candidate characters. The acceptability of each replacement text set includes the acceptability of each replacement text in the corresponding replacement text set. The acceptability of each replacement text is used to characterize the degree of semantic change of the corresponding replacement text compared to the semantic change of the original text.

5. The method for extracting text evolution information according to claim 4, characterized in that, Determining the acceptability of each of the p replacement text sets includes: For the second replacement text set, the acceptability of each replacement text in the second replacement text set is determined, wherein the second replacement text set is any one of the p replacement text sets; The sum of the acceptability of each replacement text is determined as the acceptability of the second replacement text set.

6. The character evolution information extraction method according to claim 1, characterized by, The extraction of text evolution information from the origin character to the target character from the cross-era text set corresponding to the origin character includes: The cross-generational text set is obtained from a pre-set text corpus; Extract the semantic features of each cross-era text in the cross-era text set; Based on the semantic features of each cross-era text, related words are determined from the corresponding cross-era texts, and all related words corresponding to the cross-era text set constitute the related word set; The influence degree parameter of each associated character in the associated character set is calculated based on the semantic features of the target character. The associated character set and the influence degree parameter corresponding to each associated character in the associated character set constitute the character evolution information.

7. The method according to any one of claims 1 to 6, wherein Also includes: The radical vector library is generated using any of the following processing methods, wherein the processing methods include: The preset radical vector generation model is invoked to process the characters in the first preset character library to generate the radical vector library. or, If the second preset character library contains specified semantic characters, the preset radical vector generation model is called to process the characters in the second preset character library and generate an initial radical vector library. Obtain the phonetic and semantic feature distributions of the radicals of the specified semantic character, as well as the phonetic and semantic feature distributions of the components of the specified semantic character. By extracting features from text sample pairs, the phonetic and semantic feature vectors of the radicals of the specified semantic character and the phonetic and semantic feature vectors of the components of the specified semantic character are obtained, and added to the initial radical vector library to obtain the radical vector library; the text sample pair includes positive text samples and negative text samples, the positive text samples include the specified semantic character, and the negative text samples include characters different from the specified semantic character.

8. A character evolution information extraction apparatus characterized by comprising: The device includes: The determination module is used to determine the target character and the start and end eras of the corresponding evolution period of the target character; The acquisition module is used to acquire p candidate characters based on the target radical vector of the target character, the target radical vector of the target character, and a preset radical vector library. The p candidate characters are all characters from the starting era, and p is an integer greater than or equal to 1. The determining module is further configured to determine the origin character from the p candidate characters according to the first text set and the second text set. The origin character refers to the starting character of the target character in the starting era. The first text set includes at least one text corresponding to the ending era obtained by filtering according to the target character. The second text set includes at least two texts corresponding to the starting era obtained by filtering according to the p candidate characters. An extraction module is used to extract the text evolution information from the origin character to the target character from the cross-era text set corresponding to the origin character. The text evolution information includes: a set of related characters that affect the evolution of the origin character to the target character, and an influence degree parameter corresponding to each related character in the set of related characters. The cross-era text set includes the text of the origin character contained between the starting era and the ending era. The determining module is also used for: For each of the p candidate characters, the first text set, and the second text set, a replacement text set is generated, resulting in p replacement text sets; Determine the acceptability of each of the p replacement text sets; The candidate character corresponding to the most acceptable replacement text set in the p replacement text sets is determined as the origin character.

9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the text evolution information extraction method as described in any one of claims 1-7.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by the processor, it implements the text evolution information extraction method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Text semantic analysis method, text semantic analysis terminal and storage medium

    CN107704453A

  • Character vector calculation method and device based on radicals

    CN113255318A