Method and apparatus for recognizing text

By hashmapping the original words merged after text word segmentation and identity identifiers, the existing sensitive word filtering algorithm is solved, and efficient and accurate text recognition is achieved.

CN113779980BActive Publication Date: 2025-05-16BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110127068.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-29
Publication Date
2025-05-16
Estimated Expiration
2041-01-29

AI Technical Summary

Technical Problem

The existing sensitive word filtering algorithms are inefficient when identifying sensitive words in text and are difficult to accurately identify according to user-defined rules.

Method used

By obtaining the identity identifier and the original text to be identified, the word segmentation strategy is used to divide the text word, the original word and the identity identifier are merged, and the merged original word is mapped to the corresponding position of the binary array by using hash mapping, and whether the original word is the target word is determined by the state of the array position.

Benefits of technology

It improves the computing efficiency of text recognition, can customize recognition rules based on user identity identification, and enhances the accurate recognition ability of sensitive words in text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113779980B_ABST
    Figure CN113779980B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure disclose a method and device for identifying text. A specific implementation of the method includes: obtaining an identity identifier and an original text to be identified; based on a preset segmentation strategy, segmenting the original text to obtain an original word set; for each original word in the original word set, performing the following steps: merging the original word with the identity identifier to obtain a merged original word; using a first hash map to determine the first hash value of the merged original word; based on the first hash value of the merged original word, determining the corresponding subscript of the merged original word in a pre-generated first binary array; in response to determining that the state of the array position pointed to by the corresponding subscript of the merged original word in the first binary array is 1, determining that the original word is a target word. By determining whether the original word is a target word through the state of the array position, the text can be identified according to the user's identity identifier and the computational efficiency of text recognition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, specifically to the field of text recognition, and more particularly to a method and device for recognizing text. Background Art

[0002] All business scenarios on the Internet require UGC (User Generated Content) content for content risk management. Among them, the sensitive word filtering algorithm is widely used in various types of risk control systems due to its simple operation, high recognition efficiency and low technical difficulty. The risk control strategy based on sensitive word filtering can effectively identify illegal content (such as pornography, politics, terrorism, fraud, harassment, pornography, etc.).

[0003] In the related art, there are two types of sensitive word filtering algorithms: one is a sensitive word matching algorithm based on regular expressions, the principle of which is to compose a regular expression with fixed rules into a sensitive word library, and perform regular matching on the text content that needs to be identified; the other is a sensitive word algorithm based on DFA (Deterministic Finite Automaton), including various finite automaton algorithms, the principle of which is to use a hash table structure in local memory to store the starting characters of each sensitive word, and use a tree structure to store non-starting characters of sensitive words. In the filtering stage, hash search is performed by character, and if the character hits, the tree query is continued to identify whether there are sensitive words in the text. Summary of the invention

[0004] Embodiments of the present disclosure provide a method and apparatus for recognizing text.

[0005] In a first aspect, an embodiment of the present disclosure provides a method for recognizing text, the method comprising: obtaining an identity identifier and an original text to be recognized; based on a preset segmentation strategy, segmenting the original text to obtain an original word set; for each original word in the original word set, performing the following recognition steps: merging the original word with the identity identifier to obtain a merged original word; using a first hash map to determine a first hash value of the merged original word; based on the first hash value of the merged original word, determining a corresponding subscript of the merged original word in a pre-generated first binary array; in response to determining that the state of the array position pointed to by the subscript corresponding to the merged original word in the first binary array is 1, determining that the original word is a target word.

[0006] In some embodiments, the first hash map includes a first preset number of first hash functions, each first hash function corresponds to a first hash value; and, in response to determining that the state of the array position pointed to by the subscript corresponding to the merged original word in the first binary array is 1, determining that the original word is the target word, including: in response to the states of the first preset number of array positions pointed to by the subscript corresponding to the merged original word in the first binary array are all 1, determining that the original word is the target word.

[0007] In some embodiments, the first binary array is generated via the following steps: obtaining a set of target words and an identity identifier; constructing a first initial binary array; merging each target word with the identity identifier to obtain each merged target word; using a first hash map to determine a first hash value for each merged target word; based on the first hash value of each merged target word, determining a corresponding subscript for each merged target word in the first initial binary array; setting the state of the array position pointed to by the corresponding subscript in each first initial binary array to 1 to obtain a first binary array.

[0008] In some embodiments, the method also includes: in response to a request instruction to add a target word, obtaining an identity indicated by the request instruction and a newly added target word to be added; merging the newly added target word with the identity indicated by the request instruction to obtain a merged newly added target word; using a first hash map to determine a first hash value of the merged newly added target word; based on the first hash value of the merged newly added target word, determining a corresponding subscript of the merged newly added target word in the first binary array; and setting the state of the array position pointed to by the corresponding subscript of the merged newly added target word in the first binary array to 1.

[0009] In some embodiments, the method further includes: in response to a request instruction to delete a target word, obtaining an identity indicated by the request instruction and a target word to be deleted; constructing a second initial binary array; merging the target word to be deleted with the identity indicated by the request instruction to obtain a merged target word to be deleted; using a second hash map to respectively determine a second hash value of each merged target word to be deleted; based on the second hash value of each merged target word to be deleted, respectively determine the corresponding subscript of each merged target word to be deleted in the second initial binary array; setting the state of the array position pointed to by the corresponding subscript in each second initial binary array to 1 to obtain a second binary array; and, in response to determining that the state of the array position pointed to by the subscript corresponding to the merged original word in the first binary array is 1, determining that the original word is the target word, including: in response to determining that the state of the array position pointed to by the target subscript is 1, using a second hash map to determine the second hash value of the merged original word; based on the second hash value of the merged original word, determining the corresponding subscript of the merged original word in the second binary array; in response to determining that the state of the array position pointed to by the subscript corresponding to the merged original word in the second binary array is 0, determining that the original word is the target word.

[0010] In some embodiments, the second hash map includes a second preset number of second hash functions, each second hash function corresponds to a second hash value; and, in response to determining that the state of the array position pointed to by the subscript corresponding to the original word in the second binary array is 0, determining that the original word is the target word includes: in response to determining that the states of the second preset number of array positions pointed to by the subscript corresponding to the original word in the second binary array are not all 1, determining that the original word is the target word.

[0011] In a second aspect, an embodiment of the present disclosure provides a device for recognizing text, the device comprising: an acquisition unit, configured to acquire an identity identifier and an original text to be recognized; a word segmentation unit, configured to segment the original text based on a preset word segmentation strategy to obtain an original word set; an identification unit, configured to perform the following identification steps for each original word in the original word set: merge the original word with the identity identifier to obtain a merged original word; use a first hash map to determine a first hash value of the merged original word; based on the first hash value of the merged original word, determine the subscript corresponding to the merged original word in a pre-generated first binary array; in response to determining that the state of the array position pointed to by the subscript corresponding to the merged original word in the first binary array is 1, determine that the original word is a target word.

[0012] In some embodiments, the first hash map includes a first preset number of first hash functions, each first hash function corresponds to a first hash value; and the recognition unit is further configured to: in response to the states of a first preset number of array positions pointed to by the subscripts corresponding to the merged original word in the first binary array being all 1, determine that the original word is the target word.

[0013] In some embodiments, the device further includes a first array generation unit configured to: obtain a target word set and an identity identifier; construct a first initial binary array; merge each target word with the identity identifier to obtain each merged target word; use a first hash map to determine the first hash value of each merged target word; based on the first hash value of each merged target word, determine the corresponding subscript of each merged target word in the first initial binary array; set the state of the array position pointed to by the corresponding subscript in each first initial binary array to 1, and obtain the first binary array

[0014] In some embodiments, the device also includes a target word adding unit, which is configured to: in response to a request instruction to add a target word, obtain an identity indicated by the request instruction and a newly added target word to be added; merge the newly added target word with the identity indicated by the request instruction to obtain a merged newly added target word; use a first hash mapping to determine a first hash value of the merged newly added target word; based on the first hash value of the merged newly added target word, determine a corresponding subscript of the merged newly added target word in the first binary array; set the state of the array position pointed to by the corresponding subscript of the merged newly added target word in the first binary array to 1.

[0015] In some embodiments, the device also includes a target word deletion unit, which is configured to: in response to a request instruction to delete the target word, obtain the identity indicated by the request instruction and the target word to be deleted; construct a second initial binary array; merge the target word to be deleted with the identity indicated by the request instruction to obtain a merged target word to be deleted; use a second hash map to determine the second hash value of each merged target word to be deleted; based on the second hash value of each merged target word to be deleted, determine the corresponding subscript of each merged target word to be deleted in the second initial binary array; set the state of the array position pointed to by the corresponding subscript in each second initial binary array to 1 to obtain a second binary array; and the recognition unit is further configured to: in response to determining that the state of the array position pointed to by the target subscript is 1, use the second hash map to determine the second hash value of the merged original word; based on the second hash value of the merged original word, determine the corresponding subscript of the merged original word in the second binary array; in response to determining that the state of the array position pointed to by the subscript corresponding to the original word in the second binary array is 0, determine that the original word is the target word.

[0016] In some embodiments, the second hash map includes a second preset number of second hash functions, each second hash function corresponds to a second hash value; and the recognition unit is further configured to: in response to determining that the states of the second preset number of array positions pointed to by the subscripts corresponding to the merged original word in the second binary array are not all 1, determine that the original word is the target word.

[0017] The method and device for recognizing text provided by the embodiments of the present disclosure merge the original words obtained after word segmentation of the original text with the identity identifier, map the merged original words to the corresponding positions of the binary array through hash mapping, determine whether the original words are the target words through the state of the array position, and thus recognize the text according to the user's identity identifier and improve the computational efficiency of text recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Other features, objects and advantages of the present disclosure will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0019] Figure 1 is an exemplary system architecture diagram in which some embodiments of the present disclosure may be applied;

[0020] Figure 2 is a flow chart of an embodiment of a method for recognizing text according to the present disclosure;

[0021] Figure 3 yes Figure 2 A schematic diagram of an application scenario of an embodiment of a method for recognizing text is shown;

[0022] Figure 4 is a flowchart of another embodiment of a method for recognizing text according to the present disclosure;

[0023] Figure 5 is a structural schematic diagram of an embodiment of a device for recognizing text according to the present disclosure;

[0024] Figure 6 It is a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION

[0025] The present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It is understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It is also necessary to explain that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.

[0026] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0027] Figure 1 An exemplary system architecture 100 is shown to which a method for recognizing text or an apparatus for recognizing text according to an embodiment of the present disclosure can be applied.

[0028] like Figure 1 As shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0029] The user can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. For example, various interactive applications (such as shopping applications, social applications, or video live broadcast applications) can be loaded on the terminal device, and the user can send information to the server through the interactive application and receive information from the server.

[0030] The user can also send the original text to be identified and the identity identifier to the server through the terminal device, and the server will identify the original text to determine whether the target word (for example, a sensitive word or a keyword) exists in the original text.

[0031] Terminal devices 101, 102, 103 can be hardware or software. When terminal devices 101, 102, 103 are hardware, they can be electronic devices with communication functions, including but not limited to smart phones, tablet computers, e-book readers, laptop computers, desktop computers, etc. When terminal devices 101, 102, 103 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software or software modules for providing distributed services, or they can be implemented as a single software or software module. No specific limitation is made here.

[0032] The server 105 may be a server that provides various services, such as a background information server that interacts with user information uploaded by the terminal devices 101, 102, and 103 (e.g., replies to messages sent by users). The background data server may also identify the received user information to determine whether the target word exists in the message sent by the user.

[0033] It should be noted that the method for recognizing text provided in the embodiment of the present disclosure can be executed by the terminal devices 101, 102, 103, or by the server 105. Accordingly, the apparatus for recognizing text can be set in the terminal devices 101, 102, 103, or in the server 105. No specific limitation is made here.

[0034] It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules for providing distributed services, or it can be implemented as a single software or software module. No specific limitation is made here.

[0035] Continue to refer Figure 2 , shows a process 200 of an embodiment of a method for recognizing text according to the present disclosure. The method for recognizing text comprises the following steps:

[0036] Step 201, obtaining an identity identifier and an original text to be identified.

[0037] In this embodiment, the identity tag represents the user's custom rule for the target word. For example, the same word may be identified as a target word by user A, but may be identified as a non-target word by user B. The identity tag can be a string of characters. The original text can be a short sentence, a phrase or a text paragraph, and can be a real-time message or a historical text, which is not limited in this application.

[0038] Taking the online customer service system of an e-commerce platform as an example, the execution subject can be the business server of the e-commerce platform, and the user interacts with the execution subject through the client of the e-commerce platform loaded on the terminal. The execution subject can obtain the text message sent by the user in real time as the original text to be identified. The identity identifier at this time can be the identity identifier of the e-commerce platform itself, and the target word is defined by the operator of the e-commerce platform.

[0039] For another example, in a live video broadcast scenario, the execution subject may be the service server of the live broadcast platform, and users watching the live broadcast can interact with the execution subject through the client of the live broadcast platform. The execution subject can obtain the text message sent by the user in real time as the original text to be identified. The identity identifier at this time may be the identity identifier of the live broadcast room, indicating that the target word in this scenario is customized by the operator of the live broadcast room.

[0040] For another example, the execution subject can also be Figure 1In the terminal device shown in FIG. 1 , the user can define the target word according to his / her needs and set an identity for the target word rule. Then the user can send the original text to be recognized and the device identity to the terminal device.

[0041] Step 202: segment the original text based on a preset segmentation strategy to obtain an original word set.

[0042] In this embodiment, the word segmentation strategy can be determined based on a preset target word.

[0043] As an example, if the target word set includes "paid brushing orders", the execution entity can set a higher word segmentation weight for "paid brushing orders" in the text segmenter. For example, if the original text obtained by the execution entity in step 201 is "Please search for paid brushing orders", the original word set obtained by the execution entity after inputting the original text into the text segmenter is "please", "search", and "paid brushing orders". If the target word set includes "brushing orders", the execution entity can set a higher word segmentation weight for "brushing orders" in the text segmenter. In this way, the original word set obtained by the execution entity from the text segmenter is "please", "search", "paid" and "brushing orders".

[0044] After the execution subject obtains the original word set, steps 203 to 206 are executed for each original word in the original word set.

[0045] Step 203: merge the original word with the identity identifier to obtain a merged original word.

[0046] As an example, the execution subject may connect the original word and the identity identifier into a string, which is the combined original word.

[0047] Step 204: Use the first hash map to determine the first hash value of the merged original word.

[0048] In this embodiment, the execution entity may input the merged original word into a preset hash function to obtain a first hash value of the merged original word.

[0049] Step 205 : determining the subscript corresponding to the merged original word in the pre-generated first binary array based on the first hash value of the merged original word.

[0050] In this embodiment, the execution entity may perform an “&” operation based on the array length of the first binary array and the first hash value to determine the subscript corresponding to the merged original word in the first binary array.

[0051] As an example, the execution entity can convert the array length of the first binary array and the first hash value of the merged original word into binary numbers, and then perform an "&" operation on the first hash value of the merged original word and the array length of the first binary array. The resulting number is the subscript corresponding to the merged original word in the first binary array.

[0052] Step 206 , in response to determining that the state of the array position pointed to by the subscript corresponding to the merged original word in the first binary array is 1, determine that the original word is the target word.

[0053] In this embodiment, the state of the array position is used to characterize whether the word corresponding to the array position is a target word. For example, if the state of the array position is 1, it means that the word corresponding to the position is a target word, and if the state of the array position is 0, it means that the word corresponding to the position is a non-target word.

[0054] In some optional implementations of the present embodiment, the first hash map includes a first preset number of first hash functions, each first hash function corresponds to a first hash value; and, in response to determining that the state of the array position pointed to by the subscript corresponding to the merged original word in the first binary array is 1, determining that the original word is the target word, including: in response to the states of the first preset number of array positions pointed to by the subscript corresponding to the merged original word in the first binary array are all 1, determining that the original word is the target word.

[0055] Generally, there is a certain collision probability in hash mapping, that is, the hash values ​​obtained by different words through hash mapping may be the same. In order to reduce the collision probability and improve the accuracy of text recognition, the first hash mapping in this implementation is a multi-level hash mapping.

[0056] As an example, the first hash map includes three hash functions, and the execution subject can obtain three first hash values ​​of the merged original word after step 204, and then obtain three subscripts after step 205. After that, the execution subject can respectively obtain the states of the array positions pointed to by the three subscripts in the first binary array. If the states of the three array positions are all 1, it means that the original word is the target word. If the states of the three array positions are not all 1, it means that the original word is a non-target word.

[0057] In some optional implementations of the present embodiment, the first binary array is generated through the following steps: obtaining a set of target words and an identity identifier; constructing a first initial binary array; merging each target word with the identity identifier to obtain each merged target word; using a first hash map to determine the first hash value of each merged target word; based on the first hash value of each merged target word, determining the corresponding subscript of each merged target word in the first initial binary array; setting the state of the array position pointed to by the corresponding subscript in each first initial binary array to 1, to obtain a first binary array.

[0058] In this implementation, the execution subject merges the user's identity and the pre-selected target word, and then maps the merged target word to the corresponding array position in the first binary array through the first hash map, and represents the target word through the state of the array position, so as to realize the conversion of the target word set into the state of the array position in the first binary array. On the one hand, the storage space of the target word set can be effectively reduced. Taking a binary array of 2 billion bits as an example, the memory consumed by a binary array of this length is about 120M. If each target word is mapped to 8 bits in the array, the target words that can be stored in the array can reach a theoretical storage capacity of 250 million. On the other hand, the target word can be coupled with the user's identity. After the same word is merged with different identity identifiers, it is mapped to different array positions in the first binary array. There is no association between the two, so that the business isolation of the target word set is realized.

[0059] In a specific example, the execution subject is a business server of a live broadcast platform. The execution subject can receive the target words (for example, sensitive words: A, B and C) selected by the operator of the live broadcast room through the network, and use the ID (Identity document) of the live broadcast room as the identity identifier (the live broadcast room ID can be 123, for example). Afterwards, the execution subject can construct a first initial binary array, and the length of the first initial binary array can be determined based on the number of target words, for example, it can be 3. The execution subject merges each target word with the live broadcast room ID respectively, and the merged target words obtained are: 123A, 123B and 123C. Then, the execution subject uses the first hash map to determine the first hash value of the merged target word respectively. Then, the execution subject determines the array subscripts corresponding to the three merged target words based on the first hash value and the array length of the first initial binary array. Finally, the states of the array positions pointed to by the three array subscripts are set to 1, and the first binary array is obtained.

[0060] In this way, the execution subject can combine the target words customized by the live broadcast room with the live broadcast room ID and map them to different array positions respectively, so that the target word sets of each live broadcast room do not interfere with each other. When the execution subject identifies the real-time messages in each live broadcast room in combination with the identity of the live broadcast room, as an example, the live broadcast room with the live broadcast room ID of 456 has customized target words B and C, then the message containing A will be identified as the target word in the live broadcast room with ID 123, but will not be identified as the target word in the live broadcast room with ID 456.

[0061] It should be noted that the first hash mapping in this implementation can also adopt multi-level mapping to reduce the collision probability in the process of constructing the first binary array. In conjunction with the above example, if the first hash mapping is a three-level mapping, the array length of the first initial binary array is the product of the number of mapping levels and the number of target words: 9. In addition, each merged target word corresponds to 3 first hash values ​​and 3 array positions.

[0062] In some optional implementations of the present embodiment, the method further includes: in response to a request instruction for adding a target word, obtaining an identity indicated by the request instruction and a newly added target word to be added; merging the newly added target word with the identity indicated by the request instruction to obtain a merged newly added target word; using a first hash map to determine a first hash value of the merged newly added target word; based on the first hash value of the merged newly added target word, determining a corresponding subscript of the merged newly added target word in the first binary array; and setting the state of the array position pointed to by the corresponding subscript of the merged newly added target word in the first binary array to 1.

[0063] In this implementation, the user can add a target word by adding a request instruction, which can improve the timeliness and flexibility of the text recognition provided by the embodiment of the present disclosure.

[0064] It should be noted that the first hash mapping in this implementation may also adopt multi-level mapping to reduce the collision probability in the process of adding target words.

[0065] Continue to see Figure 3 , Figure 3 2 is a flow chart of an embodiment of the method. Figure 3In the scenario shown, a user can interact with the business server 302 of the e-commerce platform through the e-commerce platform client loaded on the smart phone 301. For example, the user can enter the information "Do you need help with the store to brush orders?" in the customer service system of the e-commerce platform. The server 302 can obtain the information in real time and identify it as the original text to determine whether the information contains sensitive words. The server 302 obtains the original text 303 to be identified and the identity "customer service" corresponding to the current dialogue scene. After that, the server 302 inputs the original text 303 into the text word separator to obtain the original word set 305, including: "need", "help", "shop", "brushing orders". The original words are then merged with the identity 304 to obtain the merged original word set, including: "customer service needs", "customer service help", "customer service shop", "customer service brushing orders". Then, the server 302 uses the first hash mapping to map the merged original words to the array positions in the first binary array 307. For example, the array position corresponding to "customer service needs" is the 1st, the array position corresponding to "customer service help" is the 3rd, the array position corresponding to "customer service store" is the 5th, and the array position corresponding to "customer service brush order" is the 7th. Among them, only the status of the 5th and 7th bits in the first binary array is 1, thereby determining that "store" and "brushing order" in the original text are sensitive words.

[0066] The method and device for recognizing text provided by the embodiments of the present disclosure merge the original words obtained after word segmentation of the original text with the identity identifier, map the merged original words to the corresponding positions of the binary array through hash mapping, determine whether the original words are the target words through the state of the array position, and thus recognize the text according to the user's identity identifier and improve the computational efficiency of text recognition.

[0067] Further references Figure 4 , which shows a process 400 of another embodiment of a method for recognizing text. The process 400 of the method for recognizing text includes the following steps:

[0068] Step 401, obtaining the original text to be identified and the identity identifier.

[0069] Step 402: segment the original text based on a preset segmentation strategy to obtain an original word set.

[0070] Step 403: merge the original word with the identity identifier to obtain a merged original word.

[0071] Step 404: Use the first hash map to determine the first hash value of the merged original word.

[0072] Step 405 : determining the subscript corresponding to the merged original word in the pre-generated first binary array based on the first hash value of the merged original word.

[0073] Steps 401 to 405 correspond to the aforementioned steps 201 to 205 and will not be repeated here.

[0074] Step 406 , in response to determining that the state of the array position pointed to by the target index is 1, using a second hash map to determine a second hash value of the merged original word.

[0075] In this embodiment, the method for recognizing text also includes the following step of deleting target words: in response to a request instruction to delete the target word, obtaining an identity indicated by the request instruction and the target word to be deleted; constructing a second initial binary array; merging the target word to be deleted with the identity indicated by the request instruction to obtain a merged target word to be deleted; using a second hash map to respectively determine the second hash value of each merged target word to be deleted; based on the second hash value of each merged target word to be deleted, respectively determine the corresponding subscript of each merged target word to be deleted in the second initial binary array; setting the state of the array position pointed to by the corresponding subscript in each second initial binary array to 1 to obtain a second binary array.

[0076] Generally, the state of the array position in the first binary array can only be set to 1, not 0. Therefore, the execution subject in this embodiment can represent the deleted target word through the state of the array position in the second binary array. For example, if the state of the array position in the second binary array is 1, it means that the word corresponding to the array position is the deleted target word.

[0077] The steps of constructing the second binary array are similar to the steps of constructing the first binary array, and will not be repeated here.

[0078] Step 407: Determine the subscript corresponding to the original word in the second binary array based on the second hash value of the merged original word.

[0079] The execution entity may determine the subscript corresponding to the merged original word in the second binary array based on the second hash value of the merged original word and the length of the second binary array.

[0080] Step 408 , in response to determining that the state of the array position pointed to by the subscript corresponding to the merged original word in the second binary array is 0, determine that the original word is the target word.

[0081] In this embodiment, the array position in the second binary array is 0, indicating that the target word mapped to the array position has not been deleted. At this time, it can be determined that the original word is the target word.

[0082] In some optional implementations of the present embodiment, the second hash map includes a second preset number of second hash functions, each second hash function corresponds to a second hash value; and, in response to determining that the state of the array position pointed to by the subscript corresponding to the original word in the second binary array is 0, determining that the original word is the target word includes: in response to determining that the states of the second preset number of array positions pointed to by the subscript corresponding to the original word in the second binary array are not all 1, determining that the original word is the target word.

[0083] In this implementation, the second hash mapping adopts a multi-level mapping, which can reduce the collision probability of the hash mapping in the process of deleting the target word and determining whether the target word is deleted, thereby improving the accuracy of text recognition.

[0084] As an example, the execution entity may use three first hash functions as hash functions of the first hash map, and five second hash functions as hash functions of the second hash map. In this way, each merged original word corresponds to three first hash values ​​and three array positions in the three first binary values. If the states of the three array positions of the merged original word in the first binary value are all 1, the second hash map is used to determine the five second hash values ​​of the merged original word and the five array positions in the second binary array. If the states of the five array positions in the second binary array are not all 0, the original word is determined to be the target word.

[0085] from Figure 4 It can be seen that process 400 of the method for recognizing text in this embodiment reflects the step of determining whether the target word matched by the original word is deleted through the state of the array position in the second binary array during the text recognition process. The use of the second binary array to represent the target word to be deleted can improve the flexibility of user-defined target words.

[0086] In some optional implementations of the above embodiment, the method may further include the following steps of updating the word segmentation strategy: in response to a request instruction to add a target word, updating the word segmentation strategy based on the newly added target word; in response to a request instruction to delete a target word, updating the word segmentation strategy based on the target word to be deleted.

[0087] In this implementation, the word segmentation strategy can be updated synchronously based on the update instruction of the target word set, which can improve the accuracy of word segmentation and further improve the accuracy of text recognition.

[0088] As an example, when the execution subject receives a request instruction to add a target word, a higher segmentation weight can be set for the newly added target word in the text segmentation; when the execution subject receives a request instruction to delete a target word, a lower segmentation weight can be set for the deleted target word in the text segmenter.

[0089] Further references Figure 5 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a device for recognizing text. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0090] like Figure 5 As shown, the device 500 for recognizing text in this embodiment includes: an acquisition unit 501, configured to acquire an identity identifier and an original text to be recognized; a word segmentation unit 502, configured to segment the original text based on a preset word segmentation strategy to obtain an original word set; an identification unit 503, configured to perform the following identification steps for each original word in the original word set: merge the original word with the identity identifier to obtain a merged original word; use a first hash map to determine a first hash value of the merged original word; based on the first hash value of the merged original word, determine the subscript corresponding to the merged original word in a pre-generated first binary array; in response to determining that the state of the array position pointed to by the subscript corresponding to the merged original word in the first binary array is 1, determine that the original word is a target word.

[0091] In this embodiment, the first hash map includes a first preset number of first hash functions, each first hash function corresponds to a first hash value; and the recognition unit 503 is further configured to: in response to the states of the first preset number of array positions pointed to by the subscripts corresponding to the merged original word in the first binary array are all 1, determine that the original word is the target word.

[0092] In this embodiment, the device 500 also includes a first array generation unit, which is configured to: obtain a target word set and an identity identifier; construct a first initial binary array; merge each target word with the identity identifier to obtain each merged target word; use a first hash map to determine the first hash value of each merged target word; based on the first hash value of each merged target word, determine the corresponding subscript of each merged target word in the first initial binary array; set the state of the array position pointed to by the corresponding subscript in each first initial binary array to 1, and obtain the first binary array

[0093] In this embodiment, the device 500 also includes a target word adding unit, which is configured to: in response to a request instruction for adding a target word, obtain the identity indicated by the request instruction and the newly added target word to be added; merge the newly added target word with the identity indicated by the request instruction to obtain a merged newly added target word; use a first hash mapping to determine the first hash value of the merged newly added target word; based on the first hash value of the merged newly added target word, determine the subscript corresponding to the merged newly added target word in the first binary array; set the state of the array position pointed to by the subscript corresponding to the merged newly added target word in the first binary array to 1.

[0094] In this embodiment, the device 500 also includes a target word deletion unit, which is configured to: in response to a request instruction to delete the target word, obtain the identity indicated by the request instruction and the target word to be deleted; construct a second initial binary array; merge the target word to be deleted with the identity indicated by the request instruction to obtain a merged target word to be deleted; use a second hash map to determine the second hash value of each merged target word to be deleted; based on the second hash value of each merged target word to be deleted, determine the corresponding subscript of each merged target word to be deleted in the second initial binary array; set the state of the array position pointed to by the corresponding subscript in each second initial binary array to 1 to obtain a second binary array; and the identification unit 503 is further configured to: in response to determining that the state of the array position pointed to by the target subscript is 1, use the second hash map to determine the second hash value of the merged original word; based on the second hash value of the merged original word, determine the corresponding subscript of the merged original word in the second binary array; in response to determining that the state of the array position pointed to by the subscript corresponding to the merged original word in the second binary array is 0, determine that the original word is the target word.

[0095] In this embodiment, the second hash map includes a second preset number of second hash functions, each second hash function corresponds to a second hash value; and the recognition unit 503 is further configured to: in response to determining that the states of the second preset number of array positions pointed to by the subscripts corresponding to the merged original word in the second binary array are not all 1, determine that the original word is the target word.

[0096] Reference below Figure 6 , which shows an electronic device (eg, Figure 1 The terminal device in the embodiment of the present disclosure may include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6The terminal device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0097] like Figure 6 As shown, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0098] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 6 Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0099] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, the above functions defined in the method of the embodiment of the present disclosure are executed. It should be noted that the computer-readable medium described in the embodiment of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection with one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In an embodiment of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, an apparatus, or a device. In an embodiment of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in combination with an instruction execution system, an apparatus, or a device. The program code embodied on the computer readable medium may be transmitted using any appropriate medium, including but not limited to: wire, optical cable, RF (radio frequency), etc., or any suitable combination of the foregoing.

[0100] The computer-readable medium may be included in the electronic device; or it may exist independently without being assembled into the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains the identity identifier and the original text to be identified; based on the preset word segmentation strategy, the original text is segmented to obtain the original word set; for each original word in the original word set, the following recognition steps are performed: the original word is merged with the identity identifier to obtain the merged original word; the first hash value of the merged original word is determined by using the first hash map; based on the first hash value of the merged original word, the corresponding subscript of the merged original word in the pre-generated first binary array is determined; in response to determining that the state of the array position pointed to by the corresponding subscript of the merged original word in the first binary array is 1, the original word is determined to be the target word.

[0101] Computer program code for performing the operations of embodiments of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on a user's computer, partially on a user's computer, as a separate software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0102] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0103] The units involved in the embodiments described in the present disclosure may be implemented by software or by hardware. The units described may also be provided in a processor, for example, may be described as: a processor includes an acquisition unit, a word segmentation unit, and a recognition unit. The names of these units do not constitute a limitation on the units themselves in certain circumstances, for example, the acquisition unit may also be described as a "unit for acquiring an identity identifier and an original text to be recognized".

[0104] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) to form a technical solution.

Claims

1. A method for recognizing text, wherein: include: Obtaining the identity identifier and the original text to be identified; Based on a preset word segmentation strategy, segment the original text to obtain an original word set; For each original word in the original word set, the following identification steps are performed: merging the original word with the identity identifier to obtain a merged original word; using a first hash map to determine a first hash value of the merged original word; Based on the first hash value of the merged original word, determining a subscript corresponding to the merged original word in a pre-generated first binary array; In response to determining that the state of the array position pointed to by the subscript corresponding to the merged original word in the first binary array is 1, the original word is determined to be a target word, and the first binary array is generated through the following steps: obtaining a target word set and the identity identifier; constructing a first initial binary array; merging each of the target words with the identity identifier to obtain each merged target word; Using a first hash map, respectively determine a first hash value of each of the merged target words; Based on the first hash value of each of the merged target words, the corresponding subscripts of each of the merged target words in the first initial binary array are determined respectively; the states of the array positions pointed to by the corresponding subscripts in each of the first initial binary arrays are set to 1, so as to obtain the first binary array.

2. The method according to claim 1, wherein: The first hash map includes a first preset number of first hash functions, each of the first hash functions corresponds to a first hash value; as well as, In response to determining that the state of the array position pointed to by the subscript corresponding to the merged original word in the first binary array is 1, determining that the original word is a target word includes: In response to the fact that the states of a first preset number of array positions pointed to by the subscripts corresponding to the merged original word in the first binary array are all 1, the original word is determined to be a target word.

3. The method according to claim 1, further comprising: In response to a request instruction for adding a target word, obtaining an identity indicated by the request instruction and a new target word to be added; Merging the newly added target word with the identity indicated by the request instruction to obtain a merged newly added target word; Using a first hash map, determining a first hash value of the merged newly added target word; Determining, based on the first hash value of the merged newly added target word, a subscript corresponding to the merged newly added target word in the first binary array; The state of the array position pointed to by the subscript corresponding to the merged newly added target word in the first binary array is set to 1.

4. The method according to claim 1, wherein: The method further comprises: In response to a request instruction to delete a target word, obtain an identity indicated by the request instruction and a target word to be deleted; construct a second initial binary array; merge the target word to be deleted with the identity indicated by the request instruction to obtain a merged target word to be deleted; use a second hash map to respectively determine a second hash value of each merged target word to be deleted; based on the second hash value of each merged target word to be deleted, respectively determine a corresponding subscript of each merged target word to be deleted in the second initial binary array; set the state of the array position pointed to by the corresponding subscript in each second initial binary array to 1 to obtain a second binary array; And, in response to determining that the state of the array position pointed to by the subscript corresponding to the merged original word in the first binary array is 1, determining that the original word is a target word includes: In response to determining that the state of the array position pointed to by the subscript corresponding to the merged original word in the first binary array is 1, a second hash value of the merged original word is determined using a second hash map; based on the second hash value of the merged original word, the subscript corresponding to the merged original word in the second binary array is determined; in response to determining that the state of the array position pointed to by the subscript corresponding to the merged original word in the second binary array is 0, the original word is determined to be a target word.

5. The method according to claim 4, wherein: The second hash map includes a second preset number of second hash functions, each of the second hash functions corresponds to a second hash value; And, in response to determining that the state of the array position pointed to by the subscript corresponding to the original word in the second binary array is 0, determining that the original word is a target word includes: In response to determining that the states of the second preset number of array positions pointed to by the subscripts corresponding to the merged original word in the second binary array are not all 1, the original word is determined to be a target word.

6. A device for recognizing text, wherein: include: An acquisition unit, configured to acquire an identity identifier and an original text to be recognized; A word segmentation unit is configured to segment the original text based on a preset word segmentation strategy to obtain an original word set; The recognition unit is configured to perform the following recognition steps for each original word in the original word set: merge the original word with the identity identifier to obtain a merged original word; use a first hash map to determine a first hash value of the merged original word; Based on the first hash value of the merged original word, determining a subscript corresponding to the merged original word in a pre-generated first binary array; In response to determining that the state of the array position pointed to by the subscript corresponding to the merged original word in the first binary array is 1, determining that the original word is a target word; The first array generation unit is configured to: obtain a target word set and the identity identifier; constructing a first initial binary array; Merging each of the target words with the identity identifier to obtain each merged target word; Using a first hash map, respectively determine a first hash value of each of the merged target words; Based on the first hash value of each of the merged target words, the corresponding subscripts of each of the merged target words in the first initial binary array are determined respectively; the states of the array positions pointed to by the corresponding subscripts in each of the first initial binary arrays are set to 1, so as to obtain the first binary array.

7. The device according to claim 6, wherein: The first hash map includes a first preset number of first hash functions, each of the first hash functions corresponds to a first hash value; and, The recognition unit is further configured to: in response to the states of a first preset number of array positions pointed to by the subscripts corresponding to the merged original word in the first binary array all being 1, determine that the original word is a target word.

8. The device according to claim 6, further comprising a target word adding unit configured to: In response to a request instruction for adding a target word, obtaining an identity indicated by the request instruction and a new target word to be added; Merging the newly added target word with the identity indicated by the request instruction to obtain a merged newly added target word; Using a first hash map, determining a first hash value of the merged newly added target word; Determining, based on the first hash value of the merged newly added target word, a subscript corresponding to the merged newly added target word in the first binary array; The state of the array position pointed to by the subscript corresponding to the merged newly added target word in the first binary array is set to 1.

9. The device according to claim 6, further comprising a target word deletion unit configured to: in response to a request instruction to delete the target word, obtain the identity indicated by the request instruction and the target word to be deleted; construct a second initial binary array; Merging the target word to be deleted with the identity indicated by the request instruction to obtain a merged target word to be deleted; using a second hash mapping to respectively determine a second hash value of each of the merged target words to be deleted; Based on the second hash values ​​of the merged target words to be deleted, respectively determine the subscripts corresponding to the merged target words to be deleted in the second initial binary array; set the states of the array positions pointed to by the corresponding subscripts in the second initial binary arrays to 1, to obtain a second binary array; And, the identification unit is further configured to: In response to determining that the state of the array position pointed to by the subscript corresponding to the merged original word in the first binary array is 1, using a second hash map to determine a second hash value of the merged original word; Determine, based on the second hash value of the merged original word, a subscript corresponding to the merged original word in the second binary array; In response to determining that the state of the array position pointed to by the subscript corresponding to the merged original word in the second binary array is 0, the original word is determined to be the target word.

10. The device according to claim 9, wherein: The second hash map includes a second preset number of second hash functions, each of the second hash functions corresponds to a second hash value; And, the recognition unit is further configured to: in response to determining that the states of a second preset number of array positions pointed to by the subscripts corresponding to the merged original word in the second binary array are not all 1, determine that the original word is a target word.

11. An electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.

12. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • File processing method and device, electronic equipment and readable storage medium

    CN110858191A

  • Information processing method and device, server and storage medium

    CN111488529A