Text processing method, device, electronic device and storage medium

By generating a set of associated characters and using the co-occurrence frequency in the corpus to determine the replacement characters, the problem of deep learning models having difficulty identifying variant connections is solved, the accuracy of text recognition is improved, and advertising traffic and network environment management are supported.

CN119721016BActive Publication Date: 2025-09-19BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411855632.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-09-19
Estimated Expiration
2044-12-16

AI Technical Summary

Technical Problem

Existing deep learning models find it difficult to effectively identify variant forms of contact information, resulting in low recognition accuracy, which affects advertising traffic and network environment governance.

Method used

By generating a set of associated characters, using the co-occurrence frequency in the corpus to determine the replacement characters, generating replacement text, and using a deep learning model to identify the connection information in the replacement text.

Benefits of technology

The accuracy of text recognition has been improved, and it can effectively identify variant forms of contact information, supporting advertising traffic and network environment management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119721016B_ABST
    Figure CN119721016B_ABST
Patent Text Reader

Abstract

The present disclosure provides a text processing method, apparatus, electronic device, and storage medium, relating to the fields of artificial intelligence technology, particularly natural language processing and deep learning technology. A specific implementation scheme comprises: for each original character in a text to be processed, determining a set of associated characters for the original character based on the original character and at least one associated character generated based on the original character; determining replacement characters for each of the multiple original characters in the text to be processed based on co-occurrence relationships between multiple elements from the multiple associated character sets; and generating replacement text based on the replacement characters for each of the multiple original characters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the fields of natural language processing and deep learning technology. More specifically, the present disclosure provides a text processing method, apparatus, electronic device, storage medium, and computer program product. Background Art

[0002] With the continuous development of computer and internet technologies, a vast amount of information exists online. Some information often contains text variations. Identifying information containing text variations to determine whether it contains sensitive information, such as contact information, is a key measure for combating ad fraud and ad misdirection, and maintaining a healthy online environment. Summary of the Invention

[0003] The present disclosure provides a text processing method, apparatus, electronic device, storage medium, and computer program product.

[0004] According to a first aspect, a text processing method is provided, the method comprising: for each original character in a text to be processed, determining a set of associated characters for the original character based on the original character and at least one associated character generated based on the original character; determining replacement characters for each of the multiple original characters in the text to be processed based on a co-occurrence relationship between multiple elements respectively from the multiple associated character sets; and generating replacement text based on the replacement characters for each of the multiple original characters.

[0005] According to a second aspect, a text processing device is provided, which includes: an associated character set determination module for determining, for each original character in a text to be processed, an associated character set of the original character based on the original character and at least one associated character generated based on the original character; a replacement character determination module for determining, based on a co-occurrence relationship between multiple elements respectively from multiple associated character sets, a replacement character for each of multiple original characters in the text to be processed; and a replacement text determination module for generating replacement text based on the replacement characters for each of the multiple original characters.

[0006] According to a third aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided according to the present disclosure.

[0007] According to a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method provided according to the present disclosure.

[0008] According to a fifth aspect, a computer program product is provided, comprising a computer program stored on at least one of a readable storage medium and an electronic device, wherein the computer program implements the method provided according to the present disclosure when executed by a processor.

[0009] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.

[0011] Figure 1 is a schematic diagram of an exemplary system architecture to which a text processing method and apparatus can be applied according to an embodiment of the present disclosure;

[0012] Figure 2 is a flowchart of a text processing method according to an embodiment of the present disclosure;

[0013] Figure 3 is a schematic diagram of an associated character linked list according to an embodiment of the present disclosure;

[0014] Figure 4 is a schematic diagram of a text processing method according to an embodiment of the present disclosure;

[0015] Figure 5 is a schematic diagram of a text processing method according to another embodiment of the present disclosure;

[0016] Figure 6 is a block diagram of a text processing apparatus according to an embodiment of the present disclosure; and

[0017] Figure 7 FIG. 4 is a block diagram of an electronic device according to a text processing method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0018] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0019] Identifying contact information in text is a key technology for managing ad traffic. Examples of contact information include phone numbers, email addresses, and social media accounts. However, text containing contact information often uses variations to avoid recognition, making it difficult to identify.

[0020] Currently, it's possible to identify variant text by training a deep learning model to learn its characteristics. For example, a contact information recognition model can be trained using samples containing variant contact information. However, the context of the variant contact information is normal text, so the model cannot effectively learn the variant characteristics during training, resulting in poor recognition accuracy.

[0021] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0022] In the technical solution disclosed herein, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.

[0023] Figure 1 This is a schematic diagram of an exemplary system architecture to which the text processing method and apparatus can be applied according to an embodiment of the present disclosure. It should be noted that: Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure may not be used in other devices, systems, environments or scenarios.

[0024] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0025] Users can use terminal devices 101, 102, 103 to interact with server 105 via network 104 to receive or send messages, etc. Terminal devices 101, 102, 103 can be various electronic devices, including but not limited to smartphones, tablet computers, laptop computers, etc.

[0026] Server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using terminal devices 101, 102, and 103. The background management server may analyze and process received data such as user requests and feedback the processing results to the terminal device.

[0027] The text processing method provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the text processing device provided by the embodiments of the present disclosure can generally be disposed in the server 105. The text processing method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103, and / or the server 105. Correspondingly, the text processing device provided by the embodiments of the present disclosure can also be disposed in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103, and / or the server 105.

[0028] Figure 2 is a flowchart of a text processing method according to an embodiment of the present disclosure.

[0029] As Figure 2 shown, the text processing method 200 includes operations S210 to S230.

[0030] In operation S210, for each original character in the text to be processed, an associated character set of the original character is determined according to the original character and at least one associated character generated based on the original character.

[0031] The text to be processed can be text obtained from published information, articles, comments, push messages, etc. on the network. The text to be processed may contain characters in variant forms.

[0032] Each character in the text to be processed can be used as an original character first, and then for each original character, multiple variant forms of the original character can be expanded to obtain multiple associated characters of the original character. The original character can be a normal character without drainage cheating behavior or a variant character with drainage cheating behavior. For example, if the text to be processed is "clip skirt chat", the original characters "clip" and "skirt" are variant characters, and "chat" is a normal character.

[0033] For the original character "clip", multiple associated characters can be generated through various variant forms such as homophonic characters, homomorphic characters, homophonic symbols, etc. For example, associated characters such as "add", "home", "+", and an emoji of a house (representing home) can be generated.

[0034] The original character "clip" and the associated characters "add", "home", "+", and an emoji of a house (representing home) can form the associated character set of the original character "clip".

[0035] In operation S220, according to the co-occurrence relationship between multiple elements from multiple associated character sets, replacement characters for each of the multiple original characters in the text to be processed are determined.

[0036] For example, multiple sets of associated characters may include a set of associated characters for the original character "夹", a set of associated characters for the original character "裙", and a set of associated characters for the original character "聊".

[0037] Any combination can be made for the elements respectively from different sets of associated characters to obtain combined phrases. For example, the combined phrases may include "夹裙聊", "加裙聊", "加群聊", etc. For each combined phrase, the co-occurrence frequency of multiple characters (such as 3 characters) in the phrase in the corpus can be calculated, and the multiple characters in the phrase with the highest co-occurrence frequency are respectively determined as the replacement characters for the multiple original characters in the text to be processed.

[0038] For example, the three characters "加", "群", "聊" in the phrase "加群聊" have the highest co-occurrence frequency in the corpus. Therefore, the three characters "加", "群", "聊" can be respectively used as the replacement characters for the three characters "夹", "裙", "聊" in the text to be processed.

[0039] In operation S230, replacement text is generated according to the replacement characters of the multiple original characters respectively.

[0040] For example, the original characters "夹裙聊" in the text to be processed can be replaced with "加群聊", and "加群聊" is the replacement text. The replacement text can restore the true information of the text to be processed, facilitating subsequent processing such as sensitive information recognition and contact information recognition for anti-advertising drainage and cheating of the text.

[0041] According to an embodiment of the present disclosure, for each original character in the text to be processed, at least one associated character of each original character is generated, the original character and the at least one associated character are combined into a set of associated characters, and according to the co-occurrence relationship of multiple elements respectively from multiple sets of associated characters in the corpus, the multiple elements with the highest co-occurrence frequency are determined as the replacement characters for the multiple original characters in the text to be processed, the original characters are replaced with the replacement characters to obtain replacement text, and the replacement text can restore the true information of the text to be processed and improve the accuracy of subsequent text recognition.

[0042] Figure 3 It is a schematic diagram of an associated character linked list according to an embodiment of the present disclosure.

[0043] According to an embodiment of the present disclosure, for each original character, at least one associated character of the original character is generated according to at least one variant type, where the variant type includes at least one of homophonic characters, homonymic characters, homomorphic characters, homophonic numbers, and homophonic symbols; and according to the original character and the at least one associated character, a set of associated characters is determined.

[0044] Such as Figure 3As shown, taking the original character "夹" as an example, according to various variant types such as homophonic characters, homophonic characters, homomorphic characters, homophonic numbers, homophonic symbols, etc., multiple associated characters such as "家", "加", "+", "jia", and the graphic or expression of "家" can be generated. The original character "夹" and each associated character can form an associated character set 310.

[0045] According to an embodiment of the present disclosure, for each associated character set, determine the frequency of occurrence of each associated character in the corpus for at least one associated character in the associated character set; and taking the original character in the associated character set as the starting node and each associated character in the associated character set as the associated node, sort the associated nodes according to the frequency of occurrence of each associated character in the corpus to obtain an associated character linked list.

[0046] For example, taking the original character "夹" as the starting node and each associated character as the associated node, calculate the frequency of occurrence of each associated character in the corpus, and sort the associated nodes in descending order according to the frequency of occurrence of the associated character in the corpus to obtain an associated character linked list 320.

[0047] For example, the frequencies of occurrence of associated characters such as "加", "家", …, "+" in the corpus decrease in turn. Therefore, they can be arranged in turn behind the original character "夹" to obtain an associated character linked list 320.

[0048] According to an embodiment of the present disclosure, arranging the associated characters in descending order of the frequency of occurrence in the corpus facilitates quickly combining the characters with the highest co-occurrence frequency during subsequent character combination, thereby improving the efficiency of determining replacement characters.

[0049] The following is combined with Figures 4 and 5 to illustrate the determination of replacement characters in an embodiment of the present disclosure.

[0050] Figure 4 is a schematic diagram of a text processing method according to an embodiment of the present disclosure.

[0051] As Figure 4 shown, the associated character linked lists 410, 420, and 430 are the associated character linked lists of the original characters "夹", "裙", and "聊" respectively. The associated characters in each associated character linked list are arranged in descending order according to the frequency of occurrence in the corpus.

[0052] According to an embodiment of the present disclosure, combine multiple elements respectively from multiple associated character sets to obtain multiple first phrases; for each first phrase, calculate the co-occurrence frequency of the multiple elements in the corpus; and determine the multiple elements in the first phrase with the highest co-occurrence frequency as the replacement characters for the multiple original characters in the text to be processed respectively.

[0053] For example, elements from the associated character linked list 410, elements from the associated character linked list 420, and elements from the associated character linked list 430 can be combined arbitrarily to obtain multiple first phrases. The first phrases include, for example, "pin group chat", "pin group", "join group chat", etc. For each first phrase, the co-occurrence frequency of multiple characters in the group in the corpus can be calculated, and the elements in the first phrase with the highest co-occurrence frequency are determined as the replacement characters for the original characters.

[0054] For example, to improve the efficiency of determining replacement characters, the associated character linked list can be divided into multiple character blocks according to the number of tasks that can be executed in parallel each time, and the combinations are carried out in sequence according to the order of the character blocks. During the combination of characters in each character block, the co-occurrence frequency of the characters in the first phrase combined in real time can be calculated.

[0055] For example, if the number of tasks that can be executed in parallel each time is 3, then "pin", "join", "home" in the associated character linked list 410, "skirt", "group", "qun" in the associated character linked list 420, and "chat", "le", "liao" in the associated character linked list 430 can be divided into the first character block. Use "pin", "join", "home" in the associated character linked list 410 to be combined with "skirt", "group", "qun" in the associated character linked list 420 and "chat", "le", "liao" in the associated character linked list 430 in parallel to obtain multiple first phrases.

[0056] During the process of combining the characters in the first character block, the co-occurrence frequency of multiple elements in the first phrase combined in real time in the corpus can be calculated. A threshold for the co-occurrence frequency can be set (for example, 70%). Among all the character combination results of the first character block, the first phrases with the co-occurrence frequency reaching the threshold can be used as candidate phrases, and then the first phrase with the highest co-occurrence frequency among the candidate phrases can be selected as the target first phrase. After obtaining the target first phrase, the character combination of the subsequent character blocks can no longer be carried out, but the elements in the target first phrase are used as the replacement characters for the original characters. This is because the associated characters in the associated character linked list are arranged in descending order of the frequency of appearance in the corpus. Therefore, the first phrase with the highest co-occurrence frequency is likely to be generated in the character combination of the character block with a higher ranking. Therefore, the first phrase that has exceeded the threshold and has the highest co-occurrence frequency can be used as the first target phrase, which can improve the efficiency of determining the first target phrase. The elements in the first target phrase are the replacement characters for the original characters, so the efficiency of determining the replacement characters can be improved.

[0057] For example, "join group chat" is the combination result of characters in the first character block, which was generated relatively early, has a co-occurrence frequency exceeding the threshold, and has the highest co-occurrence frequency among all character combination results in the first character block. Therefore, "join group chat" can be determined as the replacement text for the text to be processed "clip skirt chat".

[0058] In an embodiment of the present disclosure, the associated characters in the associated character linked list are arranged in descending order of the frequency of occurrence in the corpus. By combining according to the associated character linked list, multiple first word groups are obtained, and the co-occurrence frequency of multiple elements in the first word group in the corpus is calculated. The elements with the highest co-occurrence frequency are determined as the replacement characters of the original characters, which can improve the determination efficiency of the replacement characters and the replacement text.

[0059] Figure 5 It is a schematic diagram of a text processing method according to another embodiment of the present disclosure.

[0060] According to an example of the present disclosure, for two adjacent associated character sets corresponding to two adjacent original characters respectively, according to the co-occurrence relationship between two elements respectively from the two adjacent associated character sets, the replacement characters of the two adjacent original characters are determined respectively.

[0061] As Figure 5 shown, in this embodiment, for every two adjacent original characters, the character combination is performed according to two adjacent associated character linked lists and the replacement characters are determined. For the text to be processed with a relatively large number of characters, the frequency of multiple characters co-occurring may be low. Therefore, for every two adjacent original characters, the replacement characters can be determined in turn according to the co-occurrence frequency.

[0062] According to an embodiment of the present disclosure, each element in the first associated character set is combined with each element in the second associated character set respectively to obtain multiple second word groups; for each second word group, the co-occurrence frequency of the two elements in the second word group in the corpus is calculated; and the two elements in the second word group with the highest co-occurrence frequency are determined as the replacement characters of the two adjacent original characters respectively.

[0063] For example, the first associated character set can be the associated character linked list 510, and the second associated character set can be the associated character linked list 520. The second word group "join group" can be determined as the target second word group in a manner similar to the character combination and replacement character determination with the above three associated character linked lists, that is, the element "join" is the replacement character of the original character "clip", and the element "group" is the replacement character of the original character "skirt".

[0064] Next, the first set of associated characters can be determined as the linked list of associated characters 520, and the second set of associated characters can be determined as the linked list of associated characters 530. Since it has been determined that "group" in the linked list of associated characters 520 is the replacement character for the original character "skirt", therefore, "group" in the linked list of associated characters 520 can be directly combined with each element in the linked list of associated characters 530, and the replacement character for the original character "chat" can be determined according to the co-occurrence frequency.

[0065] Alternatively, each element in the linked list of associated characters 520 can be combined with each element in the linked list of associated characters 530 to obtain multiple second word groups. At the same time, the co-occurrence frequency of the second word groups containing "group" is weighted, and then the second word group with the highest co-occurrence frequency is determined from all the second word groups as the second target word group. The two elements in the second target word group are used as the replacement characters for the character "skirt" and the character "chat" respectively.

[0066] For example, the target second word group "group chat" is the replacement character for the original character "skirt chat". Combining with the replacement character "add" for the original character "clip", the replacement text "add group chat" can be obtained.

[0067] The embodiments of the present disclosure can restore the real text information for long texts and improve the determination efficiency of replacement characters and replacement texts by performing character combination and determination of replacement characters through the linked lists of associated characters for every two adjacent original characters.

[0068] According to the embodiments of the present disclosure, sensitive information recognition is performed on the replacement text to obtain a recognition result, including: using a deep learning model to perform contact information recognition on the replacement text to obtain a recognition result indicating whether the replacement text contains contact information.

[0069] After converting the text to be processed into a replacement text, the replacement text restores the real information of the text to be processed. Therefore, the replacement text can be recognized to identify whether it contains sensitive or illegal information.

[0070] In one example, a deep learning model can be used to recognize the contact information in the replacement text, so as to identify whether the replacement text contains contact information. The contact information can include phone numbers, email addresses, social accounts, etc. If it is recognized that the replacement text contains contact information, it can be determined that the text to be processed belongs to text for advertising drainage or cheating.

[0071] The above-mentioned deep learning model can be obtained by training using contact sample text containing normal text. Compared with the model obtained by training using contact sample text containing variant text in the related art, it is difficult to learn variant features, resulting in poor text recognition accuracy. The embodiment of the present disclosure uses a deep learning model trained using contact sample text containing normal text to identify replacement text containing normal text, which can improve the text recognition accuracy.

[0072] According to an embodiment of the present disclosure, the present disclosure also provides a text processing device.

[0073] Figure 6 is a block diagram of a text processing apparatus according to an embodiment of the present disclosure.

[0074] like Figure 6 As shown, the text processing apparatus 600 includes an associated character set determination module 610 , a replacement character determination module 620 , and a replacement text determination module 630 .

[0075] The associated character set determination module 610 is configured to determine, for each original character in the text to be processed, an associated character set of the original character according to the original character and at least one associated character generated based on the original character.

[0076] The replacement character determination module 620 is configured to determine replacement characters for respective original characters in the to-be-processed text based on co-occurrence relationships between a plurality of elements respectively from a plurality of associated character sets.

[0077] The replacement text determination module 630 is configured to generate replacement text according to the replacement characters of the respective original characters.

[0078] The replacement character determination module 620 includes a first combining unit, a first co-occurrence frequency determining unit, and a first replacement character determining unit.

[0079] The first combining unit is used to combine multiple elements from multiple associated character sets to obtain multiple first phrases.

[0080] The first co-occurrence frequency determination unit is configured to calculate, for each first phrase, the co-occurrence frequency of a plurality of elements in the first phrase in the corpus.

[0081] The first replacement character determination unit is configured to determine replacement characters for each of a plurality of original characters in the to-be-processed text according to the first phrase with the highest co-occurrence frequency.

[0082] The first replacement character determination unit is configured to determine the multiple elements in the first phrase with the highest co-occurrence frequency as replacement characters for the multiple original characters in the text to be processed.

[0083] The replacement character determination module 620 is further configured to determine, for each of the two adjacent associated character sets corresponding to the two adjacent original characters, a replacement character based on a co-occurrence relationship between two elements from the two adjacent associated character sets.

[0084] The two adjacent associated character sets include a first associated character set and a second associated character set. The replacement character determination module 620 includes a second combining unit, a second co-occurrence frequency determination unit, and a second replacement character determination unit.

[0085] The second combining unit is used to combine each element in the first associated character set with each element in the second associated character set to obtain a plurality of second phrases.

[0086] The second co-occurrence frequency determination unit is configured to calculate, for each second phrase, a co-occurrence frequency of two elements in the second phrase in the corpus.

[0087] The second replacement character determination unit is configured to determine replacement characters for each of two adjacent original characters according to the second phrase with the highest co-occurrence frequency.

[0088] The second replacement character determination unit is configured to determine two elements in the second phrase with the highest co-occurrence frequency as replacement characters for two adjacent original characters, respectively.

[0089] The associated character set determination module 610 includes an associated character generation unit and an associated character set determination unit.

[0090] The associated character generating unit is used to generate at least one associated character of each original character according to at least one variant type, wherein the variant type includes at least one of homophones, homophones, homographs, homophone numbers, and homophone symbols.

[0091] The associated character set determining unit is configured to determine an associated character set according to the original character and at least one associated character.

[0092] According to an embodiment of the present disclosure, the text processing apparatus 600 further includes a frequency calculation module and an associated character list determination module.

[0093] The frequency calculation module is used to determine, for each associated character set, a frequency of occurrence of each associated character in at least one associated character in the associated character set in the corpus.

[0094] The associated character list determination module is used to take the original character in the associated character set as the starting node, take each associated character in the associated character set as the associated node, sort the associated nodes according to the frequency of each associated character appearing in the corpus, and obtain the associated character list.

[0095] The text processing apparatus 600 further includes a recognition module.

[0096] The recognition module is used to identify sensitive information in the replacement text and obtain a recognition result.

[0097] The recognition module is also used to use a deep learning model to identify contact information for the replacement text, and obtain an identification result indicating whether the replacement text contains contact information.

[0098] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0099] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0100] like Figure 7 As shown, device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. RAM 703 may also store various programs and data required for the operation of device 700. Computing unit 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.

[0101] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0102] The computing unit 701 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as the text processing method. For example, in some embodiments, the text processing method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the text processing method described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the text processing method by any other suitable means (e.g., via firmware).

[0103] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0104] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0105] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0106] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0107] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0108] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.

[0109] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0110] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A text processing method, comprising: For each original character in the text to be processed, determining a character set associated with the original character according to the original character and at least one associated character generated based on the original character; determining, based on co-occurrence relationships between a plurality of elements from a plurality of associated character sets, a plurality of elements with the highest co-occurrence frequencies as replacement characters for a plurality of original characters in the to-be-processed text; as well as generating a replacement text according to the replacement characters of the respective original characters; The step of determining the associated character set of the original character includes: For each original character, generating at least one associated character of the original character according to at least one variant type, wherein the variant type includes at least one of homophones, homophones, homographs, homophone numbers, and homophone symbols; and The associated character set is determined according to the original character and the at least one associated character.

2. The method according to claim 1, wherein The step of determining, based on the co-occurrence relationship between the multiple elements from the multiple associated character sets, the replacement characters for the multiple original characters in the to-be-processed text comprises: Combining multiple elements from the multiple associated character sets to obtain multiple first phrases; For each first phrase, calculating the co-occurrence frequency of multiple elements in the first phrase in the corpus; and Determine, according to the first phrase with the highest co-occurrence frequency, replacement characters for each of the plurality of original characters in the to-be-processed text.

3. The method according to claim 2, wherein: The step of determining, based on the first phrase with the highest co-occurrence frequency, replacement characters for each of the plurality of original characters in the to-be-processed text comprises: The multiple elements in the first phrase with the highest co-occurrence frequency are respectively determined as replacement characters for the multiple original characters in the text to be processed.

4. The method according to claim 1, wherein The step of determining, based on the co-occurrence relationship between the multiple elements from the multiple associated character sets, the replacement characters for the multiple original characters in the to-be-processed text comprises: For two adjacent associated character sets corresponding to two adjacent original characters respectively, replacement characters for the two adjacent original characters are determined according to a co-occurrence relationship between two elements from the two adjacent associated character sets.

5. The method according to claim 4, wherein The two adjacent associated character sets include a first associated character set and a second associated character set; and determining the replacement characters for the two adjacent original characters based on a co-occurrence relationship between two elements from the two adjacent associated character sets includes: Combining each element in the first associated character set with each element in the second associated character set to obtain a plurality of second phrases; For each second word group, calculating the co-occurrence frequency of two elements in the second word group in the corpus; and According to the second phrase with the highest co-occurrence frequency, replacement characters for each of the two adjacent original characters are determined.

6. The method according to claim 5, wherein: The step of determining the replacement characters for the two adjacent original characters based on the second phrase with the highest co-occurrence frequency includes: The two elements in the second phrase with the highest co-occurrence frequency are respectively determined as replacement characters for the two adjacent original characters.

7. The method according to claim 1, further comprising: For each associated character set, determining a frequency of occurrence of each associated character in the at least one associated character in the associated character set in the corpus; as well as The original character in the associated character set is used as a starting node, each associated character in the associated character set is used as an associated node, and the associated nodes are sorted according to the frequency of each associated character appearing in the corpus to obtain an associated character linked list.

8. The method according to claim 1, further comprising: Sensitive information is identified on the replacement text to obtain an identification result.

9. The method according to claim 8, wherein The sensitive information identification of the replacement text to obtain the identification result includes: A deep learning model is used to identify contact information for the replacement text to obtain an identification result indicating whether the replacement text contains contact information.

10. A text processing device comprising: an associated character set determining module, configured to determine, for each original character in the text to be processed, an associated character set of the original character based on the original character and at least one associated character generated based on the original character; a replacement character determination module, configured to determine, based on a co-occurrence relationship between a plurality of elements respectively from a plurality of associated character sets, a plurality of elements with the highest co-occurrence frequency as replacement characters for a plurality of original characters in the to-be-processed text; as well as a replacement text determination module, configured to generate a replacement text based on the replacement characters of the plurality of original characters; The associated character set determination module includes: an associated character generating unit, configured to generate, for each original character, at least one associated character of the original character according to at least one variant type, wherein the variant type includes at least one of homophones, homophones, homographs, homophone numbers, and homophone symbols; and The associated character set determining unit is configured to determine the associated character set according to the original character and the at least one associated character.

11. The device according to claim 10, wherein The replacement character determination module includes: A first combining unit is configured to combine a plurality of elements from the plurality of associated character sets to obtain a plurality of first phrases; A first co-occurrence frequency determination unit is configured to calculate, for each first phrase, the co-occurrence frequency of a plurality of elements in the first phrase in the corpus; and The first replacement character determination unit is configured to determine replacement characters for each of the plurality of original characters in the to-be-processed text according to the first phrase with the highest co-occurrence frequency.

12. The device according to claim 11, wherein The first replacement character determination unit is configured to determine the multiple elements in the first phrase with the highest co-occurrence frequency as replacement characters for the multiple original characters in the to-be-processed text.

13. The device according to claim 10, wherein The replacement character determination module is further configured to determine, for each of the two adjacent associated character sets corresponding to the two adjacent original characters, a replacement character for each of the two adjacent original characters based on a co-occurrence relationship between two elements from the two adjacent associated character sets.

14. The device according to claim 13, wherein The two adjacent associated character sets include a first associated character set and a second associated character set; The replacement character determination module includes: a second combining unit, configured to combine each element in the first associated character set with each element in the second associated character set to obtain a plurality of second phrases; A second co-occurrence frequency determination unit is configured to calculate, for each second phrase, a co-occurrence frequency of two elements in the second phrase in the corpus; and The second replacement character determination unit is configured to determine replacement characters for each of the two adjacent original characters according to the second phrase with the highest co-occurrence frequency.

15. The device according to claim 14, wherein The second replacement character determination unit is configured to determine two elements in the second phrase with the highest co-occurrence frequency as replacement characters for the two adjacent original characters, respectively.

16. The apparatus according to claim 10, further comprising: a frequency calculation module, configured to determine, for each associated character set, a frequency of occurrence of each associated character in the at least one associated character in the associated character set in the corpus; as well as The associated character list determination module is used to take the original character in the associated character set as the starting node, take each associated character in the associated character set as the associated node, sort the associated nodes according to the frequency of each associated character appearing in the corpus, and obtain the associated character list.

17. The apparatus according to claim 10, further comprising: The recognition module is used to identify sensitive information in the replacement text and obtain a recognition result.

18. The device according to claim 17, wherein The recognition module is used to use a deep learning model to identify the contact information of the replacement text, and obtain an identification result indicating whether the replacement text contains the contact information.

19. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.

20. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.

21. A computer program product, comprising a computer program, wherein the computer program is stored on at least one of a readable storage medium and an electronic device, and when the computer program is executed by a processor, implements the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Text error correction method and device

    CN111079412A

  • Text error correction method and device

    CN113919326A