A text processing method, device, and storage medium

By constructing text glyph charts and generating text feature vectors, the problems of low efficiency and low accuracy of text similarity recognition in the prior art are solved, and more efficient text error correction and recognition are achieved.

CN113569851BActive Publication Date: 2025-07-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110078268.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-20
Publication Date
2025-07-22
Estimated Expiration
2041-01-20

AI Technical Summary

Technical Problem

In the prior art, in word processing, especially in text error correction, the method of finding close characters by obfuscating dictionary is inefficient and accurate, and lacks judgment standards.

Method used

By obtaining the split character shapes of reference characters, constructing a text glyph chart, generating text feature vectors, determining text similarity, and improving the efficiency and accuracy of text similarity recognition.

Benefits of technology

It improves the efficiency and accuracy of text similarity recognition, can correct typos more effectively, and improves the convenience and accuracy of text processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113569851B_ABST
    Figure CN113569851B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses a text processing method, device, and computer-readable storage medium. The method includes: obtaining at least two reference texts, splitting the at least two reference texts to obtain split glyphs of the at least two reference texts, constructing a text glyph graph based on the at least two reference texts and the split glyphs of the at least two reference texts, the text glyph graph including the association relationship between each reference text and its corresponding split glyph, determining each reference text and the split text as a text to be recognized, generating a text feature vector for each text to be recognized based on the text glyph graph, and determining the text similarity between each text to be recognized according to the text feature vector of each text to be recognized. By establishing the association relationship between each reference text through splitting glyphs and obtaining the text feature vector of each reference text based on this association relationship, the efficiency and accuracy of text similarity recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a text processing method, device and computer-readable storage medium. Background Art

[0002] With the continuous development of computer technology, a vast amount of information has emerged in the network. Some of this information is text-carrying information, and when processing this text-carrying information, text processing is involved; for example, semantic recognition, text correction, character recognition, etc. Taking text correction as an example, the process of processing text includes the process of determining the similar-shaped characters of the target character; for example, determining the similar-shaped characters of a misspelled character, and determining the correct character from the similar-shaped characters of the misspelled character to correct the misspelled character. Currently, mainly by using a confusion dictionary to find the similar-shaped characters of the target character, the efficiency is low. Due to the lack of a judgment standard, it is difficult to judge the similarity degree of the target character and each similar-shaped character through the confusion dictionary, and the accuracy is low. Summary of the Invention

[0003] Embodiments of the present invention provide a text processing method, device and storage medium, which can improve the efficiency and accuracy of text similarity recognition.

[0004] On the one hand, embodiments of the present application provide a text processing method, which includes:

[0005] Obtain at least two reference characters, and split the at least two reference characters to obtain the split glyphs of the at least two reference characters; the split glyphs of the at least two reference characters include split characters;

[0006] Based on the at least two reference characters and the split glyphs of the at least two reference characters, construct a text glyph graph; the text glyph graph includes the association relationship between each reference character and the split glyph to which it belongs;

[0007] Determine each reference character and the split character as a character to be recognized, and generate a text feature vector for each character to be recognized based on the text glyph graph;

[0008] According to the text feature vectors of each character to be recognized, determine the text similarity between each character to be recognized.

[0009] On the one hand, the present application provides a text processing device, which includes:

[0010] An acquisition unit, configured to obtain at least two reference characters, and split the at least two reference characters to obtain the split glyphs of the at least two reference characters; the split glyphs of the at least two reference characters include split characters;

[0011] A processing unit for constructing a character glyph map based on the at least two reference characters and the split glyphs of the at least two reference characters; the character glyph map includes the association relationship between each reference character and its corresponding split glyph; and for determining each reference character and the split characters as characters to be recognized, generating a character feature vector for each character to be recognized based on the character glyph map; and for determining the character similarity between each character to be recognized according to the character feature vector of each character to be recognized.

[0012] On the one hand, the present application provides an intelligent device, including a processor, a memory, and a communication interface, which are interconnected. Wherein, the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the above-mentioned text processing method.

[0013] On the one hand, the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned text processing method is implemented.

[0014] On the one hand, the present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned text processing method.

[0015] In the embodiments of the present application, at least two reference characters are obtained, the at least two reference characters are split to obtain the split glyphs of the at least two reference characters, a character glyph map is constructed based on the at least two reference characters and the split glyphs of the at least two reference characters, the character glyph map includes the association relationship between each reference character and its corresponding split glyph, each reference character and the split characters are determined as characters to be recognized, a character feature vector for each character to be recognized is generated based on the character glyph map, and the character similarity between each character to be recognized is determined according to the character feature vector of each character to be recognized. It can be seen that the character glyph map establishes the association relationship between each reference character by splitting the glyphs, and obtains the character feature vector of each reference character based on the association relationship between each reference character, and determines the similarity of each reference character through the character feature vector of each reference character, improving the efficiency and accuracy of character similarity recognition. Description of the Drawings

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0017] Figure 1 It is a scene architecture diagram for text processing provided by an embodiment of this application;

[0018] Figure 2 It is a schematic flow chart of a text processing method provided by an embodiment of this application;

[0019] Figure 3 It is a schematic diagram of a text glyph map provided by an embodiment of this application;

[0020] Figure 4 It is a schematic flow chart of a text processing method provided by an embodiment of this application;

[0021] Figure 5a It is a flow chart for determining an associated text sequence provided by an embodiment of this application;

[0022] Figure 5b It is a schematic diagram of the process of processing text to be recognized through a text glyph model provided by an embodiment of this application;

[0023] Figure 5c It is a schematic diagram of page correction with an incorrect text title provided by an embodiment of this application;

[0024] Figure 6 It is a schematic structural diagram of a text processing device provided by an embodiment of this application;

[0025] Figure 7 It is a schematic structural diagram of an intelligent device provided by an embodiment of this application. Detailed implementation manners

[0026] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0027] Embodiments of the present application relate to Artificial Intelligence (AI) and Machine Learning (ML). Among them, AI uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0028] AI technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, processing technologies for large application programs, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. Embodiments of the present application mainly relate to natural language processing technology.

[0029] Natural Language Processing (NLP). NLP is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language used by people in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs. Embodiments of the present application mainly relate to text processing technology in natural language processing technology.

[0030] ML is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. ML is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. ML and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. Embodiments of the present application mainly relate to training an initial word vector model using a glyph map to obtain a target word vector model.

[0031] Please refer toFigure 1 , Figure 1 This is a scenario architecture diagram for text processing provided by an embodiment of the present application. As Figure 1 shown, the scenario architecture diagram includes a terminal device 101 and a server 102. Among them, the terminal device 101 is the device used by the user. The terminal device 101 may include, but is not limited to: smart phones (such as Android phones, iOS phones, etc.), tablet computers, portable personal computers, Mobile Internet Devices (MID), and other devices; the terminal device is often configured with a display device, and the display device may also be a monitor, a display screen, a touch screen, etc. The touch screen may also be a touch control screen, a touch control panel, etc. The embodiments of the present invention do not make any limitations.

[0032] The server 102 refers to a background device that can provide text processing services for the terminal device 101. The server 102 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected through wired communication or wireless communication methods, and this application does not make any restrictions here.

[0033] It should be noted that Figure 1 the number of terminal devices and servers in the text processing scenario shown is only for example. For example, the number of terminal devices and servers can be multiple, and this application does not limit the number of terminal devices and servers. In one embodiment, the text processing system may only include the terminal device 101 equipped with a text processing device, or only include the server 102 equipped with a text processing device.

[0034] Figure 1In the text processing scenario shown, text may include but is not limited to: Chinese characters, English characters, Japanese characters, Korean characters, etc. The main flow of text processing of the present application is described below using Chinese characters as an example: (1) Obtain at least two reference characters, split the at least two reference characters, and obtain split glyphs of the at least two reference characters, where the split glyphs include split characters; for example, split the reference character "零" to obtain split glyphs "雨" and "令", and since "雨" and "令" are both characters, the split glyphs of "零" include the split characters "雨" and "令"; for another example, split the reference character "邻" to obtain split glyphs "令" and "阝", and since "令" is a character and "阝" is not a character, the split glyphs of "邻" include the split character "令". (2) Based on at least two reference characters and at least two split glyphs of the reference characters, a character glyph graph is constructed, wherein the character glyph graph includes the association relationship between each reference character and its corresponding split glyph; for example, since the reference character "零" and the reference character "邻" both include the split glyph (character) "令", the reference character "零" and the reference character "邻" are both associated with "令" in the character glyph graph (i.e., the reference character "零" and the reference character "邻" are connected by the split glyph (character) "令"). (3) Each reference character and split character is determined as a character to be recognized, and a character feature vector of each character to be recognized is generated based on the character glyph graph (i.e., according to the association relationship between each reference character in the character glyph graph) (e.g., the target character is represented by N characters associated with the target character). (4) Determine the text similarity between each character to be recognized based on the text feature vector of each character to be recognized; for example, target character 1 is represented by N first associated characters associated with target character 1, and target character 2 is represented by M second associated characters associated with target character 2. Determine the similarity between the target character and target character 2 based on the number of identical characters in the N first associated characters and the M second associated characters.

[0035] In the embodiment of the present application, at least two reference characters are obtained, the at least two reference characters are split, and the split glyphs of the at least two reference characters are obtained. Based on the at least two reference characters and the split glyphs of the at least two reference characters, a character glyph graph is constructed, and the character glyph graph includes the association relationship between each reference character and the split glyph to which it belongs. Each reference character and the split character are determined as characters to be recognized, and a character feature vector of each character to be recognized is generated based on the character glyph graph. According to the character feature vector of each character to be recognized, the character similarity between each character to be recognized is determined. It can be seen that the character glyph graph establishes the association relationship between each reference character by splitting the glyphs, and obtains the character feature vector of each reference character based on the association relationship between each reference character. The similarity of each reference character is determined by the character feature vector of each reference character, thereby improving the efficiency and accuracy of character similarity recognition.

[0036] The text processing solution provided in the embodiment of the present application is described in detail below with reference to the accompanying drawings, taking Chinese characters as an example.

[0037] See also Figure 2 , Figure 2 A flowchart of a text processing method provided in an embodiment of the present application is shown below. The text processing solution can be executed by a smart device, which can be Figure 1 The terminal device 101 or the server 102 in the embodiment; the scheme includes steps S201-S204, wherein:

[0038] S201. The smart device obtains at least two reference characters, splits the at least two reference characters, and obtains split glyphs of the at least two reference characters.

[0039] The splitting methods for splitting the reference characters include splitting according to the structure of the characters (such as upper and lower structures, left and right structures, internal and external structures, etc.) and splitting according to radicals. The split characters are the radicals and characters obtained after splitting the characters; for example, the character "邻" is split, and the split characters obtained include the character "令" and the radical "阝". It should be noted that after splitting each reference character, at least one split character can be obtained; for example, after splitting the reference character "林", the split character "木" can be obtained. If the reference character cannot be further split, the character is used as a split character; for example, if the reference character "一" cannot be further split, "一" is used as a split character.

[0040] In one embodiment, the smart device splits the reference text X times according to the structure of the text to obtain the split characters of the reference text; wherein X is a positive integer. For example, X=1, the reference text "霖" is split according to the structure of the text to obtain the split characters "雨" and "林", and since X=1, "雨" and "林" are not further split; for another example, X=2, the reference text "霖" is split according to the structure of the text to obtain the split characters "雨" and "林", and "林" is split for the second time to obtain the split character "木". In one embodiment, the smart device splits the reference text according to the structure of the text until the reference text is split into inseparable characters (cannot be split into other characters); for example, the reference text "藿" is split according to the structure of the text to obtain the split characters "艹" and "霍", and the split characters "霍" are further split according to the structure of the text to obtain the split characters "雨" and "隹", and since "雨" and "隹" are inseparable characters, "雨" and "隹" are not further split.

[0041] In another embodiment, the reference character is split according to the radical of the character to obtain the split character shapes of the reference character. For example, the radical of the reference character "藿" is "艹", and the reference character is split according to the radical to obtain the split character shapes "艹" and "霍".

[0042] S202: The smart device constructs a character glyph diagram based on at least two reference characters and at least two split glyphs of the reference characters.

[0043] The text glyph graph includes the association relationship between each reference text and the split glyphs of each reference text. In one embodiment, the association relationship between each reference text and the corresponding split glyphs (i.e., the split glyphs of each reference text) refers to a connection relationship. Specifically, the smart device constructs the text glyph graph by establishing a connection between each reference text and the corresponding split glyphs. Figure 3 A schematic diagram of a text glyph diagram provided in an embodiment of the present application. Figure 3 As shown, the reference character "霍", the reference character "需", the reference character "霖" and the reference character "零" all include the split character "雨", and the reference character "霍", the reference character "需", the reference character "霖" and the reference character "零" are respectively connected with the split character "雨"; similarly, the smart device establishes connections between other reference characters and the split character shapes, and can obtain Figure 3 The text glyph diagram shown.

[0044] It can be understood that the reference character "霍" and the reference character "霖" are connected to each other through the split character "雨", and the association between the two reference characters is inversely proportional to the number of split characters and / or reference characters included in the shortest path between the two characters; wherein the shortest path refers to the path that includes the least number of split characters and / or reference characters (e.g. Figure 3 As shown in , the shortest path between “藿” and “需” is “藿”→“霍”→“雨”→“需”). For example, Figure 3 As shown, since the number of split glyphs and / or reference characters included in the shortest path between "藿" and "扇" is 6, and the number of split glyphs and / or reference characters included in the shortest path between "藿" and "需" is 4, the correlation between "藿" and "需" is greater than the correlation between "藿" and "扇".

[0045] S203: The smart device determines each reference character and split character as a character to be recognized, and generates a character feature vector for each character to be recognized based on the character glyph graph.

[0046] The characters to be recognized include reference characters and characters that cannot be separated in the split glyphs. Generating a character feature vector for each character to be recognized based on the character glyph graph means generating a character feature vector for each character to be recognized based on the association relationship between the reference characters in the character glyph graph.Figure 3 As shown, each connection node in the text glyph map corresponds to a reference text, or a split glyph of a reference text. An associated text sequence of each reference text is obtained according to the text glyph map. The associated text sequence of reference text i includes M connection nodes that are sequentially connected to reference text i, where reference text i is any reference text in the text glyph map, and M is a positive integer less than or equal to the total number of connection nodes in the text glyph map.

[0047] The intelligent device trains the initial model based on the associated text sequence of each reference text obtained from the text glyph map to obtain a text glyph model. In one implementation, the intelligent device optimally trains the parameters in the initial model through the associated text sequence of the text for training to obtain a text glyph model. After obtaining the text glyph model, the text to be recognized is used as the input of the text glyph model, and the text feature vector of each text to be recognized output by the text glyph model is obtained.

[0048] S204. The intelligent device determines the text similarity between each text to be recognized according to the text feature vector of each text to be recognized.

[0049] In one implementation, the intelligent device calculates the vector distance between the text feature vectors of each text to be recognized, and determines the similarity between each text to be recognized according to the vector distance between the text feature vectors of each text to be recognized. Specifically, the intelligent device performs weighted processing or derivative processing on the vector distance between the text feature vectors of each text to be recognized to obtain the similarity between each text to be recognized.

[0050] In the embodiments of the present application, the intelligent device obtains at least two reference texts, splits the at least two reference texts to obtain split glyphs of the at least two reference texts, constructs a text glyph map based on the at least two reference texts and the split glyphs of the at least two reference texts. The text glyph map includes the association relationship between each reference text and its affiliated split glyph. Each reference text and split text are determined as texts to be recognized. A text feature vector of each text to be recognized is generated based on the text glyph map. The text similarity between each text to be recognized is determined according to the text feature vector of each text to be recognized. It can be seen that the text glyph map establishes the association relationship between each reference text by splitting the glyph, and obtains the text feature vector of each reference text based on the association relationship between each reference text. The similarity between each reference text is determined through the text feature vector of each reference text, improving the efficiency and accuracy of text similarity recognition.

[0051] Please refer to Figure 4 , Figure 4The flowchart of another text processing method provided by an embodiment of this application. This text processing solution can be executed by an intelligent device, which can specifically be the Figure 1 terminal device 101 or server 102 in; This solution includes steps S401 - S408, where:

[0052] S401. The intelligent device obtains at least two reference characters, splits the at least two reference characters, and obtains the split glyphs of the at least two reference characters.

[0053] For the specific implementation of step S401, reference can be made to the Figure 2 implementation of step S201 in, which will not be elaborated here.

[0054] S402. The intelligent device establishes an edge connection between the reference character i and the split glyph of the reference character i.

[0055] The reference character i is any one of the at least two reference characters, and the split glyph to which the reference character i belongs refers to the split glyph obtained by splitting the reference character i. In one implementation, the intelligent device connects each reference character among the at least two reference characters to the split glyph to which it belongs (i.e., establishes an edge connection) to obtain a character glyph graph.

[0056] It can be understood that if there is the same target split glyph 1 in the split glyphs to which reference character 1 belongs and the split glyphs to which reference character 2 belongs, then in the character glyph graph, reference character 1 and reference character 2 are connected through this target split glyph 1; similarly, if there is the same target split glyph 1 in the split glyphs to which reference character 1 belongs and the split glyphs to which reference character 2 belongs, and there is the same target split glyph 2 in the split glyphs to which reference character 2 belongs and the split glyphs to which reference character 3 belongs, then in the character glyph graph, reference character 1 and reference character 3 are connected through target split glyph 1, reference character 2, and target split glyph 2. Please refer to Figure 3 , because there is the same target split glyph 'zhui' in the split glyphs of the reference characters 'Jiao' and 'Huo', so the reference characters 'Jiao' and 'Huo' are connected through the split glyph 'zhui'; similarly, there is the same target split glyph 'Yu' in the split glyphs of the reference characters 'Huo' and 'Xu', so the reference characters 'Huo' and 'Xu' are connected through the split glyph 'Yu', and the reference characters 'Jiao' and 'Xu' are connected through the split glyph 'zhui', reference character 'Huo', and split glyph 'Yu'.

[0057] S403. The intelligent device determines the edge weight of the edge connection between the reference character i and the split glyph to which it belongs according to the character attribute information of the reference character i and the glyph attribute information of the split glyph of the reference character i.

[0058] The text attribute information of the reference text includes the number of strokes and pronunciation. If the split glyphs are text, the glyph attribute information of the split glyphs includes the number of strokes and pronunciation; if the split glyphs are non-text (such as "丬", "阝", etc.), the glyph attribute information of the split glyphs includes the number of strokes.

[0059] In one embodiment, the edge between the reference character i and the split glyph to which it belongs (i.e., the split glyph of the reference Chinese character i) is a directed edge, i.e., the edge between the reference character i and the split glyph to which it belongs includes the first edge between the reference character i and the split glyph to which it belongs (i.e., the edge from the reference character i to the split glyph to which it belongs is the first edge), and the second edge between the split glyph to which the reference character i belongs and the reference character i (i.e., the edge from the split glyph to which the reference character i belongs to the reference character i is the second edge). The smart device obtains the first edge weight of the reference character i for the split glyph to which it belongs, and obtains the second edge weight of the split glyph to which the reference character i belongs for the reference character i based on the text attribute information of the reference character i and the glyph attribute information of the split glyph to which it belongs; and determines the first edge weight and the second edge weight as the edge weight of the edge between the reference character i and the split glyph to which it belongs.

[0060] Specifically, the smart device determines each reference character and each split glyph of the reference character as a connection node in the character glyph graph. Then, the first stroke information of the reference character i (the first stroke information includes the number of strokes of the reference character i) and the first node number of the adjacent connection nodes of the reference character i in the character glyph graph are obtained, and the first stroke information and the first node number are determined as the character attribute information of the reference character i; similarly, the second stroke information of the split glyph of the reference character i (the first stroke information includes the number of strokes of the split glyph of the reference character i) and the second node number of the adjacent connection nodes of the split glyph of the reference character i in the character glyph graph are obtained, and the second stroke information and the second node number are determined as the glyph attribute information; according to the first stroke information of the reference character i and the glyph attribute information of the split glyph to which the reference character i belongs, the first connection weight is determined, and according to the second stroke information of the split glyph to which the reference character i belongs and the character attribute information of the reference character i, the second connection weight is determined. In one embodiment, the edge weight w from the connection node i (i.e., the connection node where the reference character i is located) to the connection node j (i.e., the connection node where the split character to which the reference character i belongs is located) is ij The calculation formula is:

[0061]

[0062] Among them, s i is the number of strokes of the reference character i corresponding to the connection node i, s j is the number of strokes of the split glyph corresponding to the connection node j, dj is the number of adjacent connected nodes of connected node j, and a is a dynamic parameter (such as a=0.5).

[0063] Please refer to Figure 3 , the first stroke information of the reference character "邻" includes the number of strokes of "邻" being 7, and the first node number of the adjacent connection node of the reference character "邻" being 2 (i.e., s i =7,d i =2); the number of strokes of the second stroke information package "令" of the split character "令" is 5, and the number of second nodes of the adjacent connection nodes of the split character "令" is 3 (ie, s j =5,d j =3). Based on the above edge weight formula, assuming a = 0.5, the edge weight of the directed edge from the reference character "邻" to the split character "令" can be calculated as: (5 / (2+0.5)×1 / 3) 2 =4 / 9; Similarly, the weight of the directed edge of the reference character "邻" can be calculated as: (5 / (2+0.5)×1 / 2) 2 =1.

[0064] In another embodiment, the edge between the reference character i and the corresponding split glyph (i.e., the split glyph of the reference character i) is an undirected edge. The intelligent device calculates the edge weight of the edge between the reference character i and the corresponding split glyph based on the text attribute information of the reference character i and the glyph attribute information of the corresponding split glyph, and based on the text attribute information of the reference character i and the glyph attribute information of the split glyph. In one embodiment, the edge weight of the edge between the reference character i and the corresponding split glyph is obtained by averaging the weights of the directed edges between the reference character i and the corresponding split glyph; for example, the edge weight of the edge between the reference character "邻" and the corresponding split glyph "令" is: (1+4 / 9) / 2=13 / 18. In another embodiment, the calculation formula for the edge weight w of the edge between the reference character i and the corresponding split glyph is:

[0065]

[0066] Among them, s i is the number of strokes of the reference character i corresponding to the connection node i, s j is the number of strokes of the split glyph corresponding to the connection node j, d i is the number of adjacent connected nodes of connected node i, d j is the number of adjacent connected nodes of connected node j, and a is a dynamic parameter (such as a=0.5).

[0067] In yet another embodiment, the character attribute information of the reference character and the glyph attribute information of the decomposed glyph of the reference character both include pronunciation. The intelligent device determines the connection weight of the connection between the reference character and the decomposed glyph of the reference character according to the similarity degree of the pronunciation of the reference character (such as the pitch of the pronunciation, the front nasal sound, the back nasal sound, etc.) (for example, the similarity degree between the reference character with exactly the same pronunciation and the decomposed glyph it belongs to > the similarity degree between the reference characters with the same pinyin but different pitches > the similarity degree between the reference characters with different nasal sounds and the decomposed glyph they belong to > the similarity degree between the reference characters with different flat tongue sounds and retroflex sounds and the decomposed glyph they belong to); wherein, the similarity degree of the pronunciation is directly proportional to the magnitude of the connection weight. Further, the number of strokes and pronunciation of the reference character, and the number of strokes and pronunciation of the decomposed glyph of the reference character can be combined; the connection weight of the connection between the reference character i and the decomposed glyph of the reference character is comprehensively determined from multiple dimensions (strokes and pronunciation).

[0068] Optionally, the intelligent device sets the connection weight of each connection in the character glyph map to a default value, that is, the connection weights of the connections between each reference character and the decomposed glyph it belongs to in the character glyph map are the same; for example, the connection weights of the connections between each reference character and the decomposed glyph it belongs to in the character glyph map are all set to 1.

[0069] Furthermore, the intelligent device generates a character glyph map carrying connection weights according to the connection weights of the connections between each reference character and the decomposed glyph that each reference character belongs to (that is, the generated character glyph map includes the connection relationships between each reference character and the decomposed glyph that each reference character belongs to, and the connection weights of the connections between each reference character and the decomposed glyph that each reference character belongs to). The intelligent device determines each reference character and the decomposed character as the characters to be recognized, and the characters to be recognized include the reference character and the non-decomposable characters in the decomposed glyph.

[0070] S404. The intelligent device determines the associated text sequence of the reference character i according to the character glyph map.

[0071] It should be noted that if the step of determining each reference character and the decomposed glyph it belongs to as the connection nodes in the character glyph map is not executed in step S403, the intelligent device determines each reference character and the decomposed glyph it belongs to as the connection nodes in the character glyph map; otherwise, the step of determining each reference character and the decomposed glyph it belongs to as the connection nodes in the character glyph map is no longer executed.

[0072] In one embodiment, the intelligent device obtains at least two candidate associated text sequences of the reference character i from the character glyph graph; each candidate associated text sequence of the reference character i includes M connection nodes connected to the reference character i in sequence, where M is a positive integer less than or equal to the total number of connection nodes in the character glyph graph. The intelligent device respectively obtains the sum of the edge weights corresponding to each candidate text association sequence according to the edge weights of the edges associated with each candidate associated text sequence; and determines the candidate text association sequence with the largest sum of the edge weights among the at least two candidate text association sequences as the associated text sequence of the reference character i. Figure 5a FIG. is a flowchart of a method for determining an associated text sequence provided by an embodiment of the present application. As Figure 5a shown, assume M = 3, and the candidate associated text sequences of the reference character "fan" include: fan → yu → zhai; fan → yu → ling; fan → yu → xu. From Figure 5a the edge weights of each edge in it, it can be calculated that the sum of the edge weights corresponding to the candidate associated text sequence 1: "fan → yu → zhai" is 7, the sum of the edge weights corresponding to the candidate associated text sequence 2: "fan → yu → ling" is 9, and the sum of the edge weights corresponding to the candidate associated text sequence 3: "fan → yu → xu" is 11. Then the intelligent device determines the candidate associated text sequence 3: "fan → yu → xu" as the associated text sequence of the reference character "fan".

[0073] It should be noted that the larger the value of M, the greater the feature information carried by each candidate associated text sequence of the reference character i, and the greater the computational complexity, that is, the amount of information of the feature information carried by each candidate associated text sequence of the reference character i is proportional to the computational complexity.

[0074] Further, the intelligent device trains the initial model based on the associated text sequences of each reference character (that is, optimizes and trains the parameters in the initial model through the associated text sequences of each reference character) to obtain a character glyph model. In one embodiment, reference characters for training and their character decomposition results are collected in advance, and the edge weights of the edges between the reference characters and their decomposed glyphs are calculated using the formula in step S403; then an unsupervised word vector algorithm (such as the node2vec algorithm, word2vec algorithm, etc.) is used to obtain the glyph vectors of each reference character from the constructed character glyph graph, and the glyph vectors of each reference character are used to optimize and train the parameters in the initial model to obtain a character glyph model. After obtaining the character glyph model, the text to be recognized is used as the input of the character glyph model, and the character feature vectors of each text to be recognized output by the character glyph model are obtained. Figure 5b FIG. is a schematic diagram of a process for processing text to be recognized by a character glyph model provided by an embodiment of the present application. As Figure 5bAs shown, the text to be recognized is input into the input layer of the text glyph model. After the input layer of the text glyph model obtains the input data, it processes (feature extraction) the data (text to be recognized) input to the input layer through the hidden layer, and finally outputs the processing result (text feature vector of the text to be recognized) through the output layer.

[0075] S405. The intelligent device determines the text similarity between each text to be recognized according to the text feature vector of each text to be recognized.

[0076] For the specific implementation manner of step S405, reference can be made to Figure 2 the implementation manner of step S204 in , which will not be elaborated here.

[0077] S406. The intelligent device obtains the title of the text to be corrected.

[0078] The title of the text to be corrected is uploaded from the terminal device to the intelligent device or input into the intelligent device by the user; the title of the text to be corrected is obtained after preprocessing the title to be corrected (such as deleting the punctuation marks, spaces, etc. in the title).

[0079] S407. The intelligent device obtains the similar texts of the texts to be recognized included in the title of the text to be corrected according to the text similarity between each text to be recognized.

[0080] In one implementation manner, the intelligent device takes the texts to be recognized included in the title of the text to be corrected as the input of the text glyph model, obtains the feature vectors of each text to be recognized included in the title of the text to be corrected output by the text glyph model, and obtains the similar texts of the texts to be recognized included in the title of the text to be corrected based on the feature vectors. Specifically, the intelligent device determines the text similarity between each text to be recognized included in the title of the text to be corrected according to the feature vectors of each text to be recognized included in the title of the text to be corrected, and determines the texts with similarity higher than the threshold as the similar texts of the texts to be recognized included in the title of the text to be corrected.

[0081] S408. The intelligent device corrects the title of the text to be corrected according to the similar texts to obtain the corrected title of the text.

[0082] In one implementation manner, the intelligent device screens out the correct texts from the similar texts of the texts to be recognized included in the title of the text to be corrected, and replaces the wrong texts in the title of the text to be corrected with the correct texts to obtain the corrected title of the text. Specifically, the intelligent device takes the feature vectors of each text to be recognized included in the title of the text to be corrected output by the text glyph model as the input of the text error correction model, and obtains the corrected title of the text output by the text error correction model; wherein, the text error correction model is obtained by training the neural network model with the feature vectors of multiple dimensions such as the pronunciation, strokes, and structure of the text.

[0083] Figure 5c This is a schematic diagram of page correction for a page with an incorrect text title provided by an embodiment of the present application. As Figure 5c shown, after the user triggers the "Search" function in the function page 501, they are redirected to the search page 502. The search page 502 includes a search input box 5021, and the user enters the content to be searched in the search input box 5021 to search for the content of interest. As shown in the search page 502, if the text titles in the network are not processed, the search results may contain titles with typos, such as "XXXX Quanwang Video Membership", "Dragon Object Video", "Funny Video", "Car Introduction", etc. These titles containing typos will greatly affect the user experience and even become a way for illegal users to perform illegal operations (such as posting false advertisements). The optimized search page after optimizing the text titles through the above steps S401 - S408 is shown in the search page 503. As shown in the search page 503, by correcting the text titles, the user search experience can be improved, and it can assist the content optimization model to optimize the search results and filter out illegal information in the search results.

[0084] In the embodiment of the present application, on the basis of Figure 2 the embodiment, the edge weights of each edge in the text glyph map are calculated based on the text attributes of the reference text in the text glyph map and the glyph attributes of the split glyphs to which they belong, and the initial model is trained based on the text glyph map carrying the edge weights to obtain a trained text glyph model. In addition to predicting the similarity between each text according to the text feature vector of each text, improving the efficiency and accuracy of text similarity recognition, the text glyph model can also be used to assist in text (similar characters) error correction and text recognition (determining the target text as the text with the highest similarity to the text in the image to be recognized according to the similarity), thereby improving the convenience of the text processing process.

[0085] The above details the method of the embodiment of the present application. To facilitate better implementation of the above solution of the embodiment of the present application, correspondingly, the device of the embodiment of the present application is provided below.

[0086] Please refer to Figure 6 , Figure 6 This is a schematic structural diagram of a text processing device provided by an embodiment of the present application. The text processing device can be mounted on the intelligent device in the above method embodiment. The intelligent device can specifically be Figure 1 the terminal device 101 in Figure 6 shown, or the server 102. Figure 2 and Figure 4 The text processing device shown can be used to execute some or all of the functions described in the above

[0087] An acquisition unit 601, configured to acquire at least two reference characters, split the at least two reference characters, and obtain split glyphs of the at least two reference characters; the split glyphs of the at least two reference characters include split characters.

[0088] A processing unit 602, configured to construct a character glyph graph based on the at least two reference characters and the split glyphs of the at least two reference characters; the character glyph graph includes an association relationship between each reference character and its corresponding split glyph.

[0089] And configured to determine each reference character and the split character as a character to be recognized, and generate a character feature vector for each character to be recognized based on the character glyph graph.

[0090] And configured to determine a character similarity between each character to be recognized according to the character feature vector of each character to be recognized.

[0091] In one embodiment, the at least two reference characters include a reference character i, where i is a positive integer less than or equal to the total number of the at least two reference characters; the processing unit 602 is configured to construct a character glyph graph based on the at least two reference characters and the split glyphs of the at least two reference characters, specifically:

[0092] Establish an edge connection between the reference character i and the split glyph of the reference character i.

[0093] Determine an edge weight of the edge connection between the reference character i and the split glyph of the reference character i according to the character attribute information of the reference character i and the glyph attribute information of the split glyph of the reference character i.

[0094] Generate the character glyph graph according to the edge weight of the edge connection between the reference character i and the split glyph of the reference character i.

[0095] In one embodiment, the processing unit 602 is configured to generate a character feature vector for each character to be recognized based on the character glyph graph, specifically:

[0096] Determine each reference character and the split glyph of each reference character as connection nodes in the character glyph graph.

[0097] Determine an associated text sequence of the reference character i according to the character glyph graph; the associated text sequence includes M connection nodes that are sequentially connected to the reference character i, where M is a positive integer less than or equal to the total number of connection nodes in the character glyph graph.

[0098] Train an initial model based on the associated text sequence to obtain a character glyph model.

[0099] Generate a character feature vector for each character to be recognized based on the character glyph model.

[0100] In one embodiment, the processing unit 602 is configured to determine an associated text sequence of the reference character i according to the character glyph map, specifically:

[0101] Obtain at least two candidate associated text sequences of the reference character i from the character glyph map;

[0102] According to the edge weights of the edges associated with each candidate associated text sequence, respectively obtain the sum of the edge weights corresponding to each candidate text associated sequence;

[0103] Determine the candidate text associated sequence with the largest sum of edge weights among the at least two candidate text associated sequences as the associated text sequence of the reference character i.

[0104] In one embodiment, the processing unit 602 is configured to determine the edge weight of the edge between the reference character i and its split glyph according to the character attribute information of the reference character i and the glyph attribute information of the split glyph of the reference character i, specifically:

[0105] According to the character attribute information and the glyph attribute information, obtain a first edge weight of the reference character i for the split glyph to which it belongs, and obtain a second edge weight of the split glyph of the reference character i for the reference character i;

[0106] Determine the first edge weight and the second edge weight as the edge weight of the edge between the reference character i and its split glyph.

[0107] In one embodiment, the processing unit 602 is configured to obtain a first edge weight of the reference character i for the split glyph to which it belongs and a second edge weight of the split glyph of the reference character i for the reference character i according to the character attribute information and the glyph attribute information, specifically:

[0108] Determine each reference character and the split glyph of each reference character as connection nodes in the character glyph map;

[0109] Obtain the first stroke information of the reference character i and the first number of nodes of the adjacent connection nodes of the reference character i in the character glyph map, and determine the first stroke information and the first number of nodes as the character attribute information;

[0110] Obtain the second stroke information of the split glyph of the reference character i and the second number of nodes of the adjacent connection nodes of the split glyph of the reference character i in the character glyph map, and determine the second stroke information and the second number of nodes as the glyph attribute information;

[0111] Determine the first edge connection weight according to the first stroke information and the glyph attribute information in the character attribute information, and determine the second edge connection weight according to the second stroke information in the glyph attribute information and the character attribute information.

[0112] In one embodiment, each character to be recognized includes a first character to be recognized and a second character to be recognized; the processing unit 602 is configured to determine the character similarity between each character to be recognized according to the character feature vector of each character to be recognized, specifically:

[0113] Calculate the vector distance between the character feature vector of the first character to be recognized and the character feature vector of the second character to be recognized;

[0114] Determine the character similarity between the first character to be recognized and the second character to be recognized according to the vector distance.

[0115] In one embodiment, the processing unit 602 is further configured to:

[0116] Obtain the title of the text to be corrected;

[0117] Obtain the similar characters of the characters to be recognized included in the title of the text to be corrected according to the character similarity between each character to be recognized;

[0118] Correct the title of the text to be corrected according to the similar characters to obtain the corrected title of the text.

[0119] According to an embodiment of the present application, Figure 2 and Figure 4 Some of the steps involved in the text processing method shown can be executed by each unit in the Figure 6 text processing device shown. For example, Figure 2 the step S201 shown in Figure 6 can be executed by the obtaining unit 601 shown, and the steps S202-S204 can be executed by the Figure 6 processing unit 602 shown. Figure 4 The steps S401 and S406 shown in Figure 6 can be executed by the obtaining unit 601 shown, and the steps S402-S405, the steps S407 and S408 can be executed by the Figure 6 processing unit 602 shown. Figure 6Each unit in the described word processing device can be separately or all combined into one or several other units to form, or a certain one (or some) of the units can be further split into multiple smaller units with more specific functions to form. This can achieve the same operations without affecting the realization of the technical effects of the embodiments of this application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of this application, the word processing device can also include other units. In practical applications, these functions can also be assisted by other units and can be realized through the cooperation of multiple units.

[0120] According to another embodiment of this application, it can be achieved by running a computer program (including program code) that can execute the respective steps involved in the corresponding methods shown in Figure 2 and Figure 4 on a general computing device such as a computer that includes processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), to construct a word processing device as shown in Figure 6 and to implement the word processing method of the embodiments of this application. The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.

[0121] Based on the same inventive concept, the principle and beneficial effects of the word processing device provided in the embodiments of this application for solving problems are similar to those of the word processing device in the method embodiments of this application. The principle and beneficial effects of the method implementation can be referred to. For the sake of brevity, they will not be elaborated here.

[0122] Please refer to Figure 7 , Figure 7A schematic structural diagram of an intelligent device provided by an embodiment of the present application. The intelligent device at least includes a processor 701, a communication interface 702, and a memory 703. Among them, the processor 701, the communication interface 702, and the memory 703 can be connected through a bus or other means. Among them, the processor 701 (or the Central Processing Unit (CPU)) is the computing core and control core of the terminal. It can parse various instructions in the terminal and process various data of the terminal. For example, the CPU can be used to parse the power-on and power-off instructions sent by the user to the terminal and control the terminal to perform power-on and power-off operations. Another example is that the CPU can transmit various interactive data between the internal structures of the terminal, and so on. The communication interface 702 may optionally include a standard wired interface, a wireless interface (such as WI-FI, a mobile communication interface, etc.). Controlled by the processor 701, it can be used to send and receive data. The communication interface 702 can also be used for the transmission and interaction of internal data of the terminal. The memory 703 (Memory) is a memory device in the terminal, used to store programs and data. It can be understood that the memory 703 here can include both the built-in memory of the terminal and, of course, the extended memory supported by the terminal. The memory 703 provides a storage space, and this storage space stores the operating system of the terminal, which may include but is not limited to: Android system, iOS system, Windows Phone system, etc. The present application does not make any limitations on this.

[0123] In an embodiment of the present application, the processor 701 is configured to perform the following operations by running the executable program code in the memory 703:

[0124] Obtain at least two reference characters through the communication interface 702, split the at least two reference characters, and obtain the split glyphs of the at least two reference characters; the split glyphs of the at least two reference characters include split characters;

[0125] Based on the at least two reference characters and the split glyphs of the at least two reference characters, construct a character glyph map; the character glyph map includes the association relationship between each reference character and its affiliated split glyph;

[0126] Determine each reference character and the split character as the characters to be recognized, and generate a character feature vector for each character to be recognized based on the character glyph map;

[0127] According to the character feature vectors of each character to be recognized, determine the character similarity between each character to be recognized.

[0128] As an alternative embodiment, the at least two reference characters include a reference character i, where i is a positive integer less than or equal to the total number of the at least two reference characters; a specific embodiment in which the processor 701 constructs a character glyph map based on the at least two reference characters and the split glyphs of the at least two reference characters is as follows:

[0129] Establish an edge connection between the reference character i and the split glyph of the reference character i;

[0130] Determine the edge weight of the edge connection between the reference character i and the split glyph to which the reference character i belongs according to the character attribute information of the reference character i and the glyph attribute information of the split glyph of the reference character i;

[0131] Generate the character glyph map according to the edge weight of the edge connection between the reference character i and the split glyph to which the reference character i belongs.

[0132] As an alternative embodiment, a specific embodiment in which the processor 701 generates a character feature vector for each character to be recognized based on the character glyph map is as follows:

[0133] Determine each reference character and the split glyph of each reference character as connection nodes in the character glyph map;

[0134] Determine the associated text sequence of the reference character i according to the character glyph map; the associated text sequence includes M connection nodes that are sequentially connected to the reference character i, where M is a positive integer less than or equal to the total number of connection nodes in the character glyph map;

[0135] Train an initial model based on the associated text sequence to obtain a character glyph model;

[0136] Generate a character feature vector for each character to be recognized based on the character glyph model.

[0137] As an alternative embodiment, a specific embodiment in which the processor 701 determines the associated text sequence of the reference character i according to the character glyph map is as follows:

[0138] Obtain at least two candidate associated text sequences of the reference character i from the character glyph map;

[0139] Obtain the sum of the edge weights corresponding to each candidate text associated sequence respectively according to the edge weights of the edges associated with each candidate associated text sequence;

[0140] Determine the candidate text associated sequence with the largest sum of the edge weights among the at least two candidate text associated sequences as the associated text sequence of the reference character i.

[0141] As an alternative embodiment, a specific embodiment in which the processor 701 determines the edge weight of the edge between the reference character i and its split glyph according to the character attribute information of the reference character i and the glyph attribute information of the split glyph of the reference character i is as follows:

[0142] According to the character attribute information and the glyph attribute information, obtain the first edge weight of the reference character i with respect to its split glyph, and obtain the second edge weight of the split glyph of the reference character i with respect to the reference character i;

[0143] Determine the first edge weight and the second edge weight as the edge weight of the edge between the reference character i and its split glyph.

[0144] As an alternative embodiment, a specific embodiment in which the processor 701 obtains the first edge weight of the reference character i with respect to its split glyph and the second edge weight of the split glyph of the reference character i with respect to the reference character i according to the character attribute information and the glyph attribute information is as follows:

[0145] Determine each reference character and the split glyph of each reference character as connection nodes in the character glyph graph;

[0146] Obtain the first stroke information of the reference character i and the first number of adjacent connection nodes of the reference character i in the character glyph graph, and determine the first stroke information and the first number of nodes as the character attribute information;

[0147] Obtain the second stroke information of the split glyph of the reference character i and the second number of adjacent connection nodes of the split glyph of the reference character i in the character glyph graph, and determine the second stroke information and the second number of nodes as the glyph attribute information;

[0148] Determine the first edge weight according to the first stroke information in the character attribute information and the glyph attribute information, and determine the second edge weight according to the second stroke information in the glyph attribute information and the character attribute information.

[0149] As an alternative embodiment, each text to be recognized includes a first text to be recognized and a second text to be recognized; a specific embodiment in which the processor 701 determines the text similarity between each text to be recognized according to the text feature vector of each text to be recognized is as follows:

[0150] Calculate the vector distance between the text feature vector of the first text to be recognized and the text feature vector of the second text to be recognized;

[0151] Determine the text similarity between the first text to be recognized and the second text to be recognized according to the vector distance.

[0152] As an optional embodiment, the processor 701 is further configured to:

[0153] Obtain the title of the text to be corrected;

[0154] According to the text similarity between each text to be recognized, obtain the similar texts of the texts to be recognized included in the title of the text to be corrected;

[0155] Correct the title of the text to be corrected according to the similar texts to obtain the corrected title of the text.

[0156] Based on the same inventive concept, the principle of problem-solving and the beneficial effects of the intelligent device provided in the embodiments of the present application are similar to the principle of problem-solving and the beneficial effects of the text processing method in the method embodiments of the present application. For the principle and beneficial effects of the method implementation, refer to the relevant content. For the sake of brevity, it will not be elaborated here.

[0157] The embodiments of the present application further provide a computer-readable storage medium, in which one or more instructions are stored, and the one or more instructions are adapted to be loaded and executed by a processor to perform the text processing method described in the above method embodiments.

[0158] The embodiments of the present application further provide a computer program product containing instructions, which, when running on a computer, causes the computer to execute the text processing method described in the above method embodiments.

[0159] The embodiments of the present application further provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above text processing method.

[0160] The steps in the method embodiments of the present application can be adjusted, combined, and deleted according to actual needs.

[0161] The modules in the device embodiments of the present application can be combined, divided, and deleted according to actual needs.

[0162] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The readable storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, etc.

[0163] The above-disclosed is only a preferred embodiment of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present application still fall within the scope covered by the invention.

Claims

1. A word processing method, characterized in that, The method includes: Obtain at least two reference characters, split the at least two reference characters to obtain the split glyphs of the at least two reference characters; the split glyphs of the at least two reference characters include split characters; Based on the at least two reference characters and the split glyphs of the at least two reference characters, construct a character glyph graph; the character glyph graph includes the association relationship between each reference character and its affiliated split glyph; Determine each reference character and the split character as a character to be recognized, and generate a character feature vector for each character to be recognized based on the character glyph graph; Determine the character similarity between each character to be recognized according to the character feature vector of each character to be recognized.

2. The method according to claim 1, wherein The at least two reference characters include reference character i, where i is a positive integer less than or equal to the total number of the at least two reference characters; The constructing a character glyph graph based on the at least two reference characters and the split glyphs of the at least two reference characters includes: Establish an edge connection between reference character i and the split glyph of reference character i; Determine the edge weight of the edge connection between reference character i and its affiliated split glyph according to the character attribute information of reference character i and the glyph attribute information of the split glyph of reference character i; Generate the character glyph graph according to the edge weight of the edge connection between reference character i and its affiliated split glyph.

3. The method according to claim 2, wherein The generating a character feature vector for each character to be recognized based on the character glyph graph includes: Determine each reference character and the split glyph of each reference character as connection nodes in the character glyph graph; Determine the associated text sequence of reference character i according to the character glyph graph; the associated text sequence includes M connection nodes that are sequentially connected to reference character i, where M is a positive integer less than or equal to the total number of connection nodes in the character glyph graph; Train an initial model based on the associated text sequence to obtain a character glyph model; Generate a character feature vector for each character to be recognized based on the character glyph model.

4. The method according to claim 3, wherein The determining the associated text sequence of reference character i according to the character glyph graph includes: Obtain at least two candidate associated text sequences of reference character i from the character glyph graph; Respectively obtain the sum of the edge weights corresponding to each candidate associated text sequence according to the edge weights of the edges associated with each candidate associated text sequence; Determine the candidate associated text sequence with the largest sum of the edge weights among the at least two candidate associated text sequences as the associated text sequence of reference character i.

5. The method according to claim 2, wherein The determining the edge weight of the edge connection between reference character i and its affiliated split glyph according to the character attribute information of reference character i and the glyph attribute information of the split glyph of reference character i includes: According to the character attribute information and the glyph attribute information, obtain the first edge weight of reference character i for its affiliated split glyph, and obtain the second edge weight of the split glyph of reference character i for reference character i; Determine the first edge weight and the second edge weight as the edge weight of the edge between the reference character i and its split glyph.

6. The method according to claim 5, wherein The obtaining the first edge weight of the reference character i for its split glyph according to the character attribute information and the glyph attribute information, and obtaining the second edge weight of the split glyph of the reference character i for the reference character i includes: Determine each reference character and the split glyph of each reference character as connection nodes in the character glyph graph; Obtain the first stroke information of the reference character i and the first node number of the adjacent connection nodes of the reference character i in the character glyph graph, and determine the first stroke information and the first node number as the character attribute information; Obtain the second stroke information of the split glyph of the reference character i and the second node number of the adjacent connection nodes of the split glyph of the reference character i in the character glyph graph, and determine the second stroke information and the second node number as the glyph attribute information; Determine the first edge weight according to the first stroke information in the character attribute information and the glyph attribute information, and determine the second edge weight according to the second stroke information in the glyph attribute information and the character attribute information.

7. The method according to claim 1, wherein Each of the to-be-recognized characters includes a first to-be-recognized character and a second to-be-recognized character; the determining the text similarity between each of the to-be-recognized characters according to the text feature vectors of each of the to-be-recognized characters includes: Calculate the vector distance between the text feature vector of the first to-be-recognized character and the text feature vector of the second to-be-recognized character; Determine the text similarity between the first to-be-recognized character and the second to-be-recognized character according to the vector distance.

8. The method according to claim 1, wherein The method further includes: Obtain the title of the text to be corrected; Obtain the similar characters of the to-be-recognized characters included in the title of the text to be corrected according to the text similarity between each of the to-be-recognized characters; Correct the title of the text to be corrected according to the similar characters to obtain the corrected title of the text.

9. An intelligent device, characterized in that, Includes: A storage device and a processor; A computer program is stored in the storage device; The processor executes the computer program to implement the text processing method according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the text processing method according to any one of claims 1-8 is implemented.

Citation Information

Patent Citations

  • Method and device for determining characters with similar forms in search engine

    CN103927330A

  • Text error correction method and device, and related equipment

    CN108874174A