Text recognition method and device, equipment, intelligent agent, program product and medium
By combining multiple identification methods, including semantic coherence and pronunciation violation detection, the problem of identifying homophonic variant content on online platforms has been solved, achieving efficient homophonic variant filtering and ensuring the security of the online environment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-14
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies are insufficient to effectively identify and filter homophonic variant content, leading to the spread of negative impacts on online platforms.
By combining multiple identification methods such as semantic coherence recognition, pronunciation violation recognition, and violation classification recognition, the pronunciation violation recognition mode is adjusted using the results of semantic coherence recognition to improve recognition accuracy and targeting. By combining semantic coherence and pronunciation violation recognition, homophonic variant adversarial content can be identified and filtered.
It improves the accuracy and efficiency of identifying homophone variants of content, ensuring the security of the network environment and preventing the spread of harmful information.
Smart Images

Figure CN121859916A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, particularly to the fields of natural language processing, deep learning, and model training, and specifically to text recognition methods, devices, electronic devices, intelligent agents, program products, and storage media. Background Technology
[0002] With the rapid development of internet technology, users can communicate with friends on social media platforms and with merchants on e-commerce platforms via text, thereby reducing communication costs and improving the efficiency of interaction. However, some potential harms have also arisen, such as the transmission of illegal or harmful content through online platforms. Therefore, identifying illegal information has become an important research direction in order to maintain a healthy online environment. Summary of the Invention
[0003] This disclosure provides a text recognition method, apparatus, electronic device, intelligent agent, program product, and storage medium.
[0004] According to one aspect of this disclosure, a text recognition method is provided, comprising: performing semantic coherence recognition on the text to be recognized to obtain a first recognition result; performing pronunciation violation recognition on the text to be recognized based on predetermined violation information according to a target recognition pattern that matches the first recognition result to obtain a second recognition result, wherein the pronunciation violation recognition is used to determine whether there is text content in the text to be recognized that matches the pronunciation of the predetermined violation information; and when the second recognition result indicates that there is no text content in the text to be recognized that matches the pronunciation of the predetermined violation information, performing violation classification recognition on the text to be recognized to obtain a target recognition result.
[0005] According to another aspect of this disclosure, a text recognition device is provided, comprising: a semantic recognition module for performing semantic coherence recognition on the text to be recognized to obtain a first recognition result; a pronunciation recognition module for performing pronunciation violation recognition on the text to be recognized according to a target recognition pattern matching the first recognition result and based on predetermined violation information to obtain a second recognition result, wherein the pronunciation violation recognition is used to determine whether there is text content in the text to be recognized that matches the pronunciation of the predetermined violation information; and a classification module for performing violation classification recognition on the text to be recognized when the second recognition result indicates that there is no text content in the text to be recognized that matches the pronunciation of the predetermined violation information to obtain a target recognition result.
[0006] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described above.
[0007] According to another aspect of this disclosure, an artificial intelligence-based intelligent agent is provided, comprising: an input module for receiving input information; a processing module for determining a target task based on the input information received by the input module, determining a large model based on the target task, and obtaining output information by calling the large model to execute the method described above; and an output module for outputting the output information obtained by the processing module.
[0008] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the method described above.
[0009] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method described above.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0012] Figure 1 This illustration schematically shows an exemplary system architecture to which text recognition methods and apparatus can be applied according to embodiments of the present disclosure;
[0013] Figure 2 A flowchart illustrating a text recognition method according to an embodiment of the present disclosure is shown schematically.
[0014] Figure 3A The illustration shows a schematic diagram of pronunciation violation recognition of a text segment recognition pattern according to an embodiment of the present disclosure;
[0015] Figure 3B The illustration shows a schematic diagram of pronunciation violation recognition of a byte recognition pattern according to another embodiment of the present disclosure;
[0016] Figure 4 A schematic diagram illustrating the determination of a perplexity threshold according to an embodiment of the present disclosure is shown.
[0017] Figure 5 A flowchart illustrating a text recognition method according to an embodiment of the present disclosure is shown schematically.
[0018] Figure 6 A flowchart illustrating model training according to an embodiment of the present disclosure is shown schematically;
[0019] Figure 7 A block diagram of a text recognition device according to an embodiment of the present disclosure is shown schematically;
[0020] Figure 8 A schematic diagram illustrating the structure of an intelligent agent of artificial intelligence according to embodiments of the present disclosure; and
[0021] Figure 9 A block diagram of an electronic device suitable for implementing a text recognition method according to an embodiment of the present disclosure is shown schematically. Detailed Implementation
[0022] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0023] Homophone variant adversarial content can be understood as content that differs in character form but has the same pronunciation. Once homophone variant adversarial content is sent to online platforms, such as social media, forums, online comments, and live streaming of digital humans, it will have a negative impact.
[0024] It can identify the text content to be sent to the network platform, filter out homophonic variants of anti-content, and prevent it from spreading through the Internet and causing adverse effects.
[0025] In view of this, this disclosure provides a text recognition method, comprising: performing semantic coherence recognition on the text to be recognized to obtain a first recognition result; performing pronunciation violation recognition on the text to be recognized based on predetermined violation information according to a target recognition pattern that matches the first recognition result to obtain a second recognition result, wherein the pronunciation violation recognition is used to determine whether there is text content in the text to be recognized that matches the pronunciation of the predetermined violation information; and, if the second recognition result indicates that there is no text content in the text to be recognized that matches the pronunciation of the violation information, performing violation classification recognition on the text to be recognized to obtain a target recognition result.
[0026] According to embodiments of this disclosure, multiple recognition methods of different types, such as semantic coherence recognition, pronunciation violation recognition, and violation classification recognition, are combined to improve the accuracy of identifying adversarial content with homophone variants. Furthermore, based on the semantic coherence recognition results, the recognition mode for subsequent pronunciation violation recognition is adaptively adjusted, thereby improving the targeting of pronunciation violation recognition and ultimately enhancing the effectiveness of the target recognition results.
[0027] Figure 1 The illustration schematically shows an exemplary system architecture to which text recognition methods and apparatus can be applied according to embodiments of the present disclosure.
[0028] It is important to note that Figure 1 The examples shown are merely examples of system architectures that can be applied to embodiments of this disclosure, intended to help those skilled in the art understand the technical content of this disclosure. However, they do not imply that embodiments of this disclosure cannot be used in other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which the text recognition method and apparatus can be applied may include a terminal device. However, the terminal device may implement the text recognition method and apparatus provided by embodiments of this disclosure without interacting with a server.
[0029] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0030] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).
[0031] Terminal devices 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0032] Server 105 can be a server that provides various services, such as a backend management server that supports the content browsed by users using terminal devices 101, 102, and 103 (for example only). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0033] It should be noted that the text recognition method provided in this embodiment can generally be executed by terminal devices 101, 102, or 103. Accordingly, the text recognition device provided in this embodiment can also be disposed in terminal devices 101, 102, or 103.
[0034] Alternatively, the text recognition method provided in this embodiment can generally be executed by server 105. Correspondingly, the text recognition device provided in this embodiment can generally be located in server 105. The text recognition method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the text recognition device provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.
[0035] For example, terminal devices 101, 102, and 103 can acquire the text to be identified that a user has pre-published on a webpage, and then send the acquired text to server 105. Server 105 performs semantic coherence recognition on the text to be identified, obtaining a first recognition result. Based on predetermined violation information and a target recognition pattern matching the first recognition result, the text to be identified undergoes pronunciation violation recognition, obtaining a second recognition result. Pronunciation violation recognition is used to determine whether there is text content in the text to be identified that matches the pronunciation of the predetermined violation information. If the second recognition result indicates that there is no text content in the text to be identified that matches the pronunciation of the violation information, the text to be identified undergoes violation classification recognition, obtaining a target recognition result. Alternatively, a server or server cluster capable of communicating with terminal devices 101, 102, and 103 and / or server 105 can identify the text to be identified and ultimately determine the target recognition result.
[0036] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0037] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure, and application of users' personal information comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and there is no violation of public order and good morals.
[0038] In the technical solution disclosed herein, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.
[0039] It should be noted that the serial numbers of each operation in the following method are only used as representations of the operation for description, and should not be regarded as indicating the execution order of each operation. Unless explicitly stated, the method does not need to be executed exactly in the order shown.
[0040] Figure 2 The flowchart of the text recognition method according to an embodiment of the present disclosure is schematically shown.
[0041] As Figure 2 shown, the method includes operations S210 to S230.
[0042] In operation S210, semantic coherence recognition is performed on the text to be recognized to obtain a first recognition result.
[0043] In operation S220, according to the target recognition mode matching the first recognition result, based on the predetermined violation information, pronunciation violation recognition is performed on the text to be recognized to obtain a second recognition result.
[0044] In operation S230, when the second recognition result indicates that there is no text content in the text to be recognized that matches the pronunciation of the predetermined violation information, violation classification recognition is performed on the text to be recognized to obtain a target recognition result.
[0045] Pronunciation violation can be understood as that there is content in the text that has the same pronunciation as the predetermined violation information. This same pronunciation can be understood as the same pinyin. For example, "傻子" and "沙子" both have the pinyin "sha zi", but it is not limited to this, and it can also be understood as homophonic, such as "混沌" and "混蛋".
[0046] When there is a pronunciation violation in the text to be recognized, because some text content in the text to be recognized is homophonically replaced by the violation information, there is a semantic coherence problem in the semantics expressed by the whole text to be recognized. For example, if "泥" in "泥土的芬芳" is replaced by "尼", then the semantic coherence of "尼土的芬芳" has a problem.
[0047] Semantic coherence recognition can be performed on the text to be recognized to obtain a first recognition result. Thus, using semantic coherence recognition as a primary recognition method, based on the first recognition result obtained by the primary recognition method, a target recognition mode that matches the first recognition result is determined from multiple recognition modes. Thereby, the flexible adjustment and pertinence of pronunciation violation recognition are improved, and the recognition accuracy is further improved.
[0048] For example, a first recognition mode may include matching the pinyin of a predetermined violation information with the pinyin of the text to be recognized. A second recognition mode may include inputting the predetermined violation information and the text to be recognized into a pronunciation violation recognition model for recognition. A third recognition mode may include converting the predetermined violation information and the text to be recognized into audio information respectively, and then matching them using audio features.
[0049] If the first recognition result indicates semantic coherence of the text to be recognized, the first recognition pattern is used as the target recognition pattern. If the first recognition result indicates semantic incoherence of the text to be recognized, the second recognition pattern is used as the target recognition pattern. However, this is not the only possibility. Alternatively, if the first recognition result indicates semantic coherence of the text to be recognized is at the first level, the first recognition pattern can be used as the target recognition pattern. If the first recognition result indicates semantic coherence of the text to be recognized is at the second level, the second recognition pattern can be used as the target recognition pattern. If the first recognition result indicates semantic coherence of the text to be recognized is at the third level, the third recognition pattern can be used as the target recognition pattern. The degree of semantic coherence represented by the first, second, and third levels decreases progressively.
[0050] Pronunciation violation identification is used to determine whether there is text content in the text to be identified that matches the pronunciation of a predetermined violation. After obtaining a second identification result based on pronunciation violation identification, if the second identification result indicates that there is text content in the text to be identified that matches the pronunciation of the violation, a target identification result indicating that the text to be identified has a pronunciation violation can be determined based on the second identification result.
[0051] If the second recognition result indicates that the text to be recognized does not contain any text content that matches the pronunciation of the predetermined violation information, it can only be determined that the text to be recognized does not contain any text content that matches the pronunciation of the predetermined violation information, but it cannot be determined that the text to be recognized does not contain any pronunciation violation information or other types of violation information. Therefore, violation classification and recognition can be performed on the text to be recognized to obtain the target recognition result.
[0052] According to embodiments of this disclosure, multiple recognition methods of different types, such as semantic coherence recognition, pronunciation violation recognition, and violation classification recognition, are combined to improve the accuracy of identifying homophonic variant adversarial content. Furthermore, based on the semantic coherence recognition results, the recognition mode for subsequent pronunciation violation recognition is adaptively adjusted, thereby improving the targeting of pronunciation violation recognition and thus enhancing the effectiveness of the target recognition results. Additionally, when the text to be recognized contains violation content different from the predetermined violation information, violation classification recognition of the text to be recognized can supplement the pronunciation violation recognition, compensating for the real-time updates and changes in online slang, thereby improving the recognition scope and effectiveness.
[0053] According to embodiments of this disclosure, for example, Figure 2 Following the operation S220 shown, the text recognition method further includes: if the second recognition result indicates that the text to be recognized contains text content that matches the pronunciation of a predetermined violation information, then based on the second recognition result, determining the target recognition result.
[0054] If the second recognition result indicates that the text to be recognized contains text content that matches the pronunciation of the predetermined violation information, it can be determined that the target recognition result indicates that the text to be recognized contains text content that matches the pronunciation of the predetermined violation information.
[0055] The text to be identified can be anonymized, such as by masking or deleting, so that the online content disseminated on the Internet platform is normal, compliant, and non-violation-compliant content.
[0056] According to embodiments of this disclosure, if it is determined that the text to be identified contains a pronunciation violation based on pronunciation violation recognition, then the target recognition result is determined, which can improve recognition efficiency while ensuring recognition accuracy.
[0057] The above provides an overall overview of text recognition methods. The following section will explain how to perform pronunciation violation recognition.
[0058] The first recognition result indicates whether the semantics of the text to be recognized are coherent. If the first recognition result indicates that the text to be recognized is semantically coherent, it means that the probability of the text to be recognized having a pronunciation violation is low. Conversely, if the first recognition result indicates that the text to be recognized is not semantically coherent, it means that the probability of the text to be recognized having a pronunciation violation is high.
[0059] Two different recognition modes can be set. If the first recognition result indicates that the text to be recognized is semantically coherent, the text segment recognition mode is used as the target recognition mode to simplify the recognition operation, reduce recognition difficulty, and improve recognition efficiency. If the first recognition result indicates that the text to be recognized is not semantically coherent, the byte recognition mode is used as the target recognition mode to improve recognition accuracy.
[0060] Figure 3A The illustration shows a schematic diagram of pronunciation violation recognition of a text fragment recognition pattern according to an embodiment of the present disclosure.
[0061] like Figure 3A As shown, if the first recognition result indicates that the text to be recognized is semantically coherent, the text to be recognized can be segmented to obtain multiple text fragments.
[0062] The pronunciation information of multiple text segments can be determined. The pronunciation information of each text segment is then matched with the pronunciation information of a predetermined violation to obtain a second recognition result.
[0063] like Figure 3A As shown, the pinyin information of each text segment can be used as pronunciation information. The pinyin information of each of the multiple text segments is matched with the pinyin information of each of the multiple predetermined violations to obtain the second recognition result.
[0064] If the pinyin information of a text segment matches the pinyin information of one of the predetermined violations, the second recognition result indicates that the text to be recognized has a pronunciation violation. If the pinyin information of multiple text segments does not match the pinyin information of all predetermined violations, the second recognition result indicates that the text to be recognized does not have a pronunciation violation.
[0065] According to another embodiment of this disclosure, pronunciation violation identification can also be performed according to a byte identification pattern. For example, if the first identification result indicates that the text to be identified is not semantically coherent, the pronunciation information of each byte in the text to be identified is matched with the pronunciation information of a predetermined violation to obtain a second identification result.
[0066] Figure 3B A schematic diagram illustrating pronunciation violation recognition of a byte recognition pattern according to another embodiment of the present disclosure is shown.
[0067] like Figure 3B As shown, the pinyin information of each byte (token) of the text to be identified is determined. The pinyin information of each of the multiple predetermined violations is matched one by one with the pinyin information of each byte in the text to be identified, for example, by sliding the matching according to the description order of the text to be identified (see the dotted arrow), to obtain the second identification result.
[0068] like Figure 3B As shown, the predetermined violation information can be multiple bytes. A sliding match can be performed between the multiple bytes of pinyin information of the predetermined violation information (represented by shaded boxes) and the adjacent multiple bytes of pinyin information of the text to be recognized (represented by boxes).
[0069] If a group of adjacent bytes in the text to be identified matches the pinyin information of one of the predetermined violation information, the second identification result indicates that the text to be identified has a pronunciation violation. If any group of adjacent bytes in the text to be identified does not match the pinyin information of any of the predetermined violation information, the second identification result indicates that the text to be identified does not have a pronunciation violation.
[0070] The above section provided a detailed explanation of how to identify pronunciation violations. The following section will explain how to identify semantic coherence.
[0071] For example Figure 2The operation S210 shown, which performs semantic coherence recognition on the text to be recognized to obtain a first recognition result, may include: performing perplexity recognition on the text to be recognized to obtain a perplexity level. The perplexity level represents the degree of semantic coherence of the text to be recognized. Using a perplexity level threshold as reference information, a first recognition result characterizing whether the text to be recognized is semantically coherent is determined based on the perplexity level of the text to be recognized.
[0072] According to embodiments of this disclosure, perplexity identification is performed on the text to be identified to quantify the semantic coherence of the text through perplexity level. This improves the accuracy and interpretability of semantic coherence identification and reduces identification errors. Furthermore, perplexity identification is a simple and mature method, reducing the difficulty of identification.
[0073] As mentioned above, perplexity identification is relatively easy to perform and readily quantifiable. However, how to correlate the perplexity level of the quantified representation with the semantic coherence of the text to be identified, with the perplexity threshold serving as a reference, becomes a crucial influencing factor.
[0074] The following will explain how to determine the confusion threshold.
[0075] According to embodiments of this disclosure, when performing such Figure 2 Prior to the operation S210 shown, the text recognition method may further include the operation of determining a perplexity threshold from a plurality of predetermined perplexity thresholds.
[0076] Determining a perplexity threshold from multiple predetermined perplexity thresholds may include: evaluating multiple predetermined perplexity thresholds using multiple evaluation samples to obtain evaluation results for each predetermined perplexity threshold; the evaluation samples include evaluation text and evaluation tags characterizing whether the evaluation text is semantically coherent; and determining the perplexity threshold from the multiple predetermined perplexity thresholds based on the evaluation results for each predetermined perplexity threshold.
[0077] By using an evaluation method to filter multiple predetermined perplexity thresholds, the adaptability of the perplexity thresholds to the actual content to be processed can be improved, thereby improving the recognition accuracy of whether the semantic coherence of the text to be identified is correct.
[0078] The evaluation text can be semantically coherent and free of pronunciation violations, but it is not limited to this; it can also be semantically incoherent and contain pronunciation violations. As long as each evaluation text includes a genuine evaluation label, the resulting evaluation sample will contain both the evaluation text and the evaluation label. Multiple evaluation samples can be set to accurately and effectively filter for the perplexity threshold.
[0079] The following section will provide a detailed explanation of how to determine the evaluation results.
[0080] According to embodiments of this disclosure, multiple evaluation samples are used to evaluate multiple predetermined perplexity thresholds to obtain evaluation results for each of the predetermined perplexity thresholds. This may include: performing perplexity identification on multiple evaluation texts to obtain multiple evaluation perplexities; using the predetermined perplexity thresholds as reference information, determining the evaluation identification results of whether the semantic representations of the multiple evaluation texts are coherent based on the multiple evaluation perplexities; and determining the evaluation result of the predetermined perplexity threshold based on the evaluation labels of the multiple evaluation texts and the evaluation identification results of the multiple evaluation texts.
[0081] Each predetermined perplexity threshold can be used to perform semantic coherence identification on multiple evaluation texts, determining whether the semantic representations of each text are coherent. The evaluation labels representing true semantic coherence can be used as a reference to determine the number or proportion of correctly identified evaluation results for each predetermined perplexity threshold, which serves as the evaluation result for that threshold. Correctly identified evaluation results can be understood as those representing the same content as the evaluation label.
[0082] Figure 4 A schematic diagram illustrating the determination of a perplexity threshold according to an embodiment of the present disclosure is shown.
[0083] like Figure 4 As shown, a perplexity threshold evaluation curve is generated with the identifier of the predetermined perplexity threshold as the horizontal axis and the evaluation result of the predetermined perplexity threshold as the vertical axis.
[0084] like Figure 4 As shown, since the evaluation result can characterize the number or proportion of correctly identified evaluation results, the highest point on the vertical axis of the confusion threshold evaluation curve can be determined as the confusion threshold, such as a predetermined confusion threshold identified as threshold 5.
[0085] According to embodiments of this disclosure, by using evaluation labels, which represent whether the semantics are truly coherent, as a reference to evaluate multiple perplexity thresholds, the accuracy and effectiveness of the evaluation can be improved.
[0086] The preceding text provided a detailed explanation of how to perform semantic coherence identification. The following text will explain how to perform violation classification identification.
[0087] According to embodiments of this disclosure, for example, Figure 2The operation S230 shown involves classifying the text to be identified as a violation to obtain a target recognition result. This includes: performing primary classification on the text features of the text to be identified to obtain a primary classification result; performing secondary classification on the text features if the primary classification result indicates that the text to be identified violates a predetermined rule to obtain a secondary classification result, where the secondary classification result represents the rule category violated by the text to be identified; and obtaining the target recognition result based on the secondary classification result.
[0088] When performing primary classification on the text features of the text to be identified, it can be preliminarily determined whether the text violates predetermined rules. In some embodiments, primary classification can be implemented using a pre-trained primary classifier, which can preliminarily determine whether the text violates predetermined rules based on text features. For example, the primary classifier can preliminarily identify whether the text contains any risk of violating rules.
[0089] If the primary classification result indicates that the text to be identified violates a predetermined rule, secondary classification can be performed on the text features to further identify the specific rule category violated by the text. Accordingly, secondary classification can also be accomplished using a pre-trained secondary classifier, which can more finely divide the rule categories violated by the text to be identified, such as subdividing the violated rule categories into specific rule categories involving illegality, negativity, and violence.
[0090] According to embodiments of this disclosure, a final target recognition result can be obtained based on the secondary classification result. This target recognition result can not only indicate whether the text to be recognized violates a predetermined rule, but also specifically point out the category of the violated rule, thereby improving the precision of the recognition.
[0091] According to embodiments of this disclosure, for example, Figure 2 The operation S230 shown further includes: determining the target recognition result based on the first-level classification result, provided that the first-level classification result indicates that the text to be recognized conforms to a predetermined rule.
[0092] When the primary classification result indicates that the text to be identified conforms to the predetermined rules, it means that the text to be identified is preliminarily judged to have no illegal content. At this time, the target identification result can be determined directly based on the primary classification result, that is, the text to be identified is determined to be non-illegal content that meets the standard. There is no need to perform subsequent secondary classification identification and other operations. This can improve the overall identification efficiency while ensuring the accuracy of identification.
[0093] Figure 5 A flowchart illustrating a text recognition method according to an embodiment of the present disclosure is shown schematically.
[0094] For example Figure 5 As shown, the identification method includes the following operations.
[0095] The initial text is preprocessed, such as removing extra spaces, special symbols, invisible characters, zero-width spaces, etc., to obtain the text to be recognized.
[0096] The text to be identified is subjected to perplexity identification to obtain the perplexity level.
[0097] Determine if the perplexity is less than the perplexity threshold T. If the perplexity is less than the perplexity threshold T, determine that the first recognition result indicates semantic coherence of the text to be recognized, and perform word segmentation and fragment matching operations. Otherwise, determine that the first recognition result indicates that the text to be recognized is not semantically coherent, and perform byte matching operations.
[0098] Word segmentation can include: segmenting the text to be recognized into multiple text fragments. Fragment matching can include: matching the pronunciation information of each of the multiple text fragments with the pronunciation information of a predetermined violation information to obtain a second recognition result.
[0099] The byte matching operation may include: matching the pronunciation information of each byte in the text to be identified with the pronunciation information of a predetermined violation information to obtain a second identification result.
[0100] Determine whether the second recognition result indicates that the text to be recognized contains text content that matches the pronunciation of the predetermined violation information. If the second recognition result indicates that the text to be recognized does not contain text content that matches the pronunciation of the predetermined violation information, perform the violation classification and recognition operation; otherwise, determine the target recognition result.
[0101] The violation classification identification operation may include: performing primary classification identification on the text features of the text to be identified, and obtaining the primary classification result.
[0102] Determine whether the primary classification result indicates that the text to be identified violates the predetermined rules. If the primary classification result indicates that the text to be identified violates the predetermined rules, perform secondary classification recognition; otherwise, determine the target recognition result.
[0103] The text features are subjected to secondary classification to obtain the secondary classification results. Based on the secondary classification results, the target recognition result is determined.
[0104] The preceding text has provided a detailed explanation of how to perform violation classification and identification. In this embodiment, a violation classification model can be used to classify and identify violations in the text to be identified. The training process of the violation classification model will be described below.
[0105] The violation classification model may include a feature extraction module for feature extraction, a first-level classifier for determining whether a predetermined rule has been violated, and a second-level classifier for determining the category of the violated rule.
[0106] The violation classification model is trained as follows: The initial feature extraction module's parameters are adjusted using the original sample text and enhanced sample text to obtain the feature extraction module. The enhanced sample text is obtained by replacing the original sample text with content of the same pronunciation. Using the original sample text and the primary classification labels, the parameters of the feature extraction module are frozen, and the parameters of the initial primary classifier are adjusted to obtain the primary classifier. The primary classification labels characterize whether the original sample text truly violates the predetermined rules.
[0107] According to embodiments of this disclosure, the parameters of the feature extraction module can be frozen using the original sample text and secondary classification labels, and the parameters of the initial secondary classifier can be adjusted to obtain a secondary classifier, wherein the secondary classification labels represent the true rule categories violated by the original sample text.
[0108] The violation classification model can be trained from the initial violation classification model. Accordingly, the initial violation classification model may include an initial feature extraction module, an initial first-level classifier, and an initial second-level classifier. In some embodiments, the base of the initial violation classification model may be a pre-trained model.
[0109] When training the initial violation classification model, the parameters of the initial feature extraction module can be adjusted first to obtain the feature extraction module. Then, the parameters of the feature extraction module can be frozen, and the parameters of the initial first-level classifier and the initial second-level classifier can be adjusted respectively to obtain the first-level classifier and the second-level classifier, thus obtaining the violation classification model.
[0110] The initial first-level classifier performs binary classification on the text features output by the feature extraction module, obtaining the first-level classification result of the samples. For example, the first-level classification result can include 0 and 1, where a first-level classification result of 0 indicates that the original sample text violates a predetermined rule, and a first-level classification result of 1 indicates that the original sample text does not violate the predetermined rule. Correspondingly, the first-level classification label can also include 0 and 1, where a first-level classification label of 0 indicates that the original sample text truly violates the predetermined rule, and a first-level classification label of 1 indicates that the original sample text does not truly violate the predetermined rule.
[0111] The initial secondary classifier performs multi-classification on the text features output by the feature extraction module, obtaining the sample secondary classification results. These results can include the rule categories violated by the original sample text, such as illegal, negative, and violent categories. Correspondingly, the secondary classification labels can also include the actual violation of the original sample text, such as illegal, negative, and violent categories.
[0112] To improve the violation classification model's ability to identify homophonic content, the parameters of the initial feature extraction module can be adjusted by simultaneously using the original sample text and the enhanced sample text obtained by replacing the original sample text with homophonic content. This allows the feature extraction module to better capture the pronunciation feature information in the text, thereby improving the violation classification model's accuracy in identifying homophonic violations.
[0113] When constructing enhanced sample text, a homophonic content library can be used to replace core elements such as nouns and verbs in the original sample text with homophonic content. Each original sample text can generate multiple enhanced sample texts, and the primary classification label of the enhanced sample text can be consistent with the primary classification label of the original sample text.
[0114] In some embodiments, the original sample text and its enhanced sample text can be input into the initial feature extraction module to obtain the original sample text features and the enhanced sample text features. Then, a first loss function is used to process the original sample text features and the enhanced sample text features to obtain a first loss value. The parameters of the initial feature extraction module are adjusted using the first loss value to obtain the feature extraction module.
[0115] The first loss function can be represented by the following formula (1):
[0116] Loss1=diff(x,x') formula (1);
[0117] Where Loss1 represents the first loss value, and diff(x,x') represents the correlation between the original sample text feature x and the enhanced sample text feature x'.
[0118] In other embodiments, the original sample text can be input into the initial feature extraction module to obtain the original sample text features, and the original sample text features can be input into the initial first-level classifier and the initial second-level classifier respectively to obtain the original sample first-level classification result and the original sample second-level classification result. Based on the cross-entropy loss value between the original sample first-level classification result and the first-level classification label, the cross-entropy loss value between the original sample second-level classification result and the second-level classification label, and the first loss value obtained by processing the original sample text features and the enhanced sample text features using the first loss function, a second loss value is obtained. The parameters of the initial feature extraction module are then adjusted using the second loss value to obtain the feature extraction module. The second loss function can be shown in the following formula (2):
[0119] Loss2=CEbin1(x)+CEmulti1(x)+λdiff(x,x') Formula (2);
[0120] Where Loss2 represents the second loss value, CEbin(x) represents the cross-entropy loss value between the original sample's first-level classification result and the first-level classification label, CEmulti(x) represents the cross-entropy loss value between the original sample's second-level classification result and the second-level classification label, λ is an adjustment parameter, and diff(x,x') represents the first loss value.
[0121] After obtaining the feature extraction module, the parameters of the feature extraction module can be frozen, and the parameters of the initial first-level classifier and the initial second-level classifier can be adjusted respectively.
[0122] When adjusting the parameters of the initial first-level classifier, the original sample text can be input into the feature extraction module to obtain sample text features. The sample text features are then input into the initial first-level classifier to obtain the sample first-level classification result. Then, the third loss function is used to determine the cross-entropy loss value between the sample first-level classification result and the first-level classification label of the original sample text, thus obtaining the third loss value. The parameters of the initial first-level classifier are then adjusted using the third loss value to obtain the first-level classifier. The third loss function can be shown in the following formula (3):
[0123] Loss3=CEbin2(x) formula (3);
[0124] Where Loss3 represents the third loss value, and CEbin2(x) represents the cross-entropy loss value between the primary classification result and the primary classification label of the sample.
[0125] Similar to adjusting the parameters of the initial first-level classifier, the sample text features can be input into the initial second-level classifier to obtain the sample second-level classification result. Then, the fourth loss function is used to determine the cross-entropy loss value between the sample second-level classification result and the second-level classification label of the original sample text, thus obtaining the fourth loss value. The parameters of the initial second-level classifier are then adjusted using the fourth loss value to obtain the second-level classifier. The fourth loss function can be shown in the following formula (4):
[0126] Loss4=CEmulti2(x) formula (4);
[0127] Where Loss4 represents the fourth loss value, and CEmulti2(x) represents the cross-entropy loss value between the sample's secondary classification result and the secondary classification label.
[0128] According to embodiments of this disclosure, adjusting the parameters of the initial feature extraction module using enhanced sample text can improve the recognition accuracy of the violation classification model for homophonic violations. Furthermore, by freezing the parameters of the feature extraction module and adjusting the parameters of the initial first-level and initial second-level classifiers, not only is the classification accuracy of the violation classification model improved, but the training efficiency of the violation classification model is also increased.
[0129] In embodiments of this disclosure, the feature extraction module may include one or more of the following: encoder-decoder, attention mechanism, long short-term memory network, convolutional module, etc., as long as it is a network structure capable of text feature extraction. The first-level classifier and the second-level classifier may include one or more of the following: activation function, fully connected layer, etc., as long as they are capable of classification.
[0130] Figure 6 A flowchart illustrating model training according to an embodiment of the present disclosure is shown schematically.
[0131] like Figure 6 As shown, model training includes the following operations.
[0132] Enhanced sample text is obtained by replacing the original sample text with content that has the same pronunciation.
[0133] The parameters of the initial feature extraction module are adjusted using the original sample text and the enhanced sample text to obtain the feature extraction module.
[0134] Using the original sample text and primary classification labels, the parameters of the feature extraction module are frozen, and the parameters of the initial primary classifier are adjusted to obtain the primary classifier.
[0135] Using the original sample text and secondary classification labels, the parameters of the feature extraction module are frozen, and the parameters of the initial secondary classifier are adjusted to obtain the secondary classifier.
[0136] Figure 7 A block diagram of a text recognition device according to an embodiment of the present disclosure is shown schematically.
[0137] like Figure 7 As shown, the text recognition device 700 includes a semantic recognition module 710, a pronunciation recognition module 720, and a classification module 730.
[0138] The semantic recognition module 710 is used to perform semantic coherence recognition on the text to be recognized and obtain the first recognition result.
[0139] The pronunciation recognition module 720 is used to perform pronunciation violation recognition on the text to be recognized according to a target recognition pattern that matches the first recognition result and based on predetermined violation information, to obtain a second recognition result. The pronunciation violation recognition is used to determine whether there is text content in the text to be recognized that matches the pronunciation of the predetermined violation information.
[0140] The classification module 730 is used to classify and identify the text to be identified when the second recognition result indicates that there is no text content that matches the pronunciation of the predetermined violation information, so as to obtain the target recognition result.
[0141] According to embodiments of this disclosure, the text recognition device further includes a target determination module.
[0142] The target determination module is used to determine the target recognition result based on the second recognition result, when the second recognition result indicates that the text to be recognized contains text content that matches the pronunciation of the predetermined violation information.
[0143] According to embodiments of this disclosure, the pronunciation recognition module includes: a word segmentation submodule and a first matching submodule.
[0144] The word segmentation submodule is used to segment the text to be identified into multiple text fragments, provided that the first recognition result indicates that the text to be identified is semantically coherent.
[0145] The first matching submodule is used to match the pronunciation information of each of the multiple text segments with the pronunciation information of the predetermined violation information to obtain the second recognition result.
[0146] According to embodiments of this disclosure, the pronunciation recognition module includes a second matching submodule.
[0147] The second matching submodule is used to match the pronunciation information of each byte in the text to be identified with the pronunciation information of the predetermined violation information when the first recognition result indicates that the text to be identified is not semantically coherent, so as to obtain the second recognition result.
[0148] According to embodiments of this disclosure, the semantic recognition module includes a perplexity recognition submodule and a judgment submodule.
[0149] The perplexity identification submodule is used to identify perplexity in the text to be identified and obtain the perplexity score, which represents the degree of semantic coherence of the text to be identified.
[0150] The judgment submodule is used to determine the first recognition result, which represents whether the text to be recognized is semantically coherent, based on the perplexity threshold as reference information and the perplexity of the text to be recognized.
[0151] According to embodiments of this disclosure, the semantic recognition module further includes an evaluation submodule and a perplexity determination submodule.
[0152] The evaluation submodule is used to evaluate multiple predetermined perplexity thresholds using multiple evaluation samples, and obtain the evaluation results for each of the multiple predetermined perplexity thresholds. The evaluation samples include evaluation text and evaluation labels that characterize whether the evaluation text is semantically coherent.
[0153] The perplexity determination submodule is used to determine a perplexity threshold from multiple predetermined perplexity thresholds based on the evaluation results of each of the multiple predetermined perplexity thresholds.
[0154] According to embodiments of this disclosure, the evaluation submodule includes a confusion evaluation unit, a coherence evaluation unit, and a threshold evaluation unit.
[0155] The perplexity assessment unit is used to identify perplexity in multiple assessment texts and obtain multiple perplexity assessments.
[0156] The coherence evaluation unit is used to determine the evaluation and recognition results of whether the semantic representations of multiple evaluation texts are coherent, based on multiple evaluation perplexity levels and using a predetermined perplexity threshold as reference information.
[0157] The threshold evaluation unit is used to determine the evaluation result of a predetermined perplexity threshold based on the evaluation labels of multiple evaluation texts and the evaluation recognition results of multiple evaluation texts.
[0158] According to embodiments of this disclosure, the classification module includes a primary classification submodule, a secondary classification submodule, and a first classification result determination submodule.
[0159] The first-level classification submodule is used to perform first-level classification recognition on the text features of the text to be recognized, and obtain the first-level classification result.
[0160] The secondary classification submodule is used to perform secondary classification on the text features when the primary classification result indicates that the text to be identified violates a predetermined rule, and obtain the secondary classification result, where the secondary classification result represents the rule category violated by the text to be identified.
[0161] The first classification result determination submodule is used to obtain the target recognition result based on the second-level classification result.
[0162] According to embodiments of this disclosure, the classification module further includes a second classification result submodule.
[0163] The second classification result submodule is used to determine the target recognition result based on the first-level classification result, provided that the text to be recognized conforms to the predetermined rules.
[0164] According to embodiments of this disclosure, a violation classification model is used to classify and identify the text to be identified. The violation classification model includes a feature extraction module for feature extraction and a first-level classifier for determining whether a predetermined rule is violated.
[0165] The violation classification model was trained using the following modules:
[0166] The first parameter adjustment module is used to adjust the parameters of the initial feature extraction module using the original sample text and the enhanced sample text to obtain the feature extraction module. The enhanced sample text is obtained by replacing the original sample text with content that has the same pronunciation.
[0167] The second parameter adjustment module is used to freeze the parameters of the feature extraction module using the original sample text and the first-level classification label, and adjust the parameters of the initial first-level classifier to obtain the first-level classifier. The first-level classification label represents whether the original sample text truly violates the predetermined rules.
[0168] According to embodiments of this disclosure, the violation classification model further includes a secondary classifier for determining the category of the violated rule.
[0169] The training of the violation classification model also includes the following modules:
[0170] The third parameter adjustment module is used to freeze the parameters of the target feature extraction module using the original sample text and the secondary classification labels, and adjust the parameters of the initial secondary classifier to obtain the secondary classifier. The secondary classification labels represent the true rule categories violated by the original sample text.
[0171] Figure 8 A schematic block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is shown.
[0172] In embodiments of this disclosure, such as Figure 8 As shown, the intelligent agent 800 may include an input module 810, a processing module 820, and an output module 830.
[0173] Input module 810 is used to receive input information.
[0174] The processing module 820 is used to determine the target task based on the input information received by the input module, determine the large model based on the target task, and obtain output information by calling the large model to execute the text recognition method provided according to the embodiments of this disclosure.
[0175] Output module 830 is used to output the output information obtained by the processing module.
[0176] According to embodiments of this disclosure, the input module 810 is responsible for receiving or sensing information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or the external environment), and converting it into a format that the intelligent agent 800 can understand and process. The input module 810 is the primary link for the intelligent agent 800 to interact with the outside world, enabling the intelligent agent 800 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.
[0177] In the example, the input module 810 can input the text to be recognized, as described above.
[0178] In the example, processing module 820 is the core support for the ability of agent 800 to handle complex tasks. Processing module 820 can execute the text recognition method described above.
[0179] In the example, the performance of processing module 820 is closely related to the large model on which agent 800 is based. To fully leverage the capabilities of the large model, the internal structure of processing module 820 can be designed to be highly configurable and scalable to handle various types of tasks and requirements in real-world scenarios.
[0180] In the example, after the agent 800 acquires the text to be recognized, the processing module 820 can use a large model to perform semantic coherence recognition on the text to be recognized, obtaining a first recognition result. Based on predetermined violation information and a target recognition pattern matching the first recognition result, the processing module 820 performs pronunciation violation recognition on the text to be recognized, obtaining a second recognition result. The pronunciation violation recognition is used to determine whether there is text content in the text to be recognized whose pronunciation matches the predetermined violation information. If the second recognition result indicates that there is no text content in the text to be recognized whose pronunciation matches the violation information, the processing module 820 performs violation classification recognition on the text to be recognized, obtaining a target recognition result. This target recognition result is then passed to the output module 830.
[0181] Understandably, while large models possess excellent language understanding and generation capabilities, like humans, their ability to solve tasks is limited without the aid of any tools. When Agent 800 is given the ability to invoke tools, it can perform tasks such as using a calculator to perform mathematical calculations, using Python to perform data analysis, and using a search engine to create weather forecasts.
[0182] In the example, output module 830 can output the target recognition results described above.
[0183] The intelligent agent 800 according to the embodiments of this disclosure can simply and effectively improve the level of intelligence, and enhance flexibility and versatility.
[0184] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0185] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described above.
[0186] According to embodiments of the present disclosure, a non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to cause a computer to perform the method described above.
[0187] According to an embodiment of this disclosure, a computer program product includes a computer program that, when executed by a processor, implements the method described above.
[0188] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0189] like Figure 9 As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded from storage unit 908 into random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.
[0190] Multiple components in device 900 are connected to input / output (I / O) interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0191] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as text recognition methods. For example, in some embodiments, the text recognition method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the text recognition method described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to perform text recognition methods by any other suitable means (e.g., by means of firmware).
[0192] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0193] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0194] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0195] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0196] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0197] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0198] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0199] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A text recognition method, comprising: Semantic coherence recognition is performed on the text to be recognized to obtain the first recognition result; According to the target recognition pattern that matches the first recognition result, based on the predetermined violation information, the text to be recognized is subjected to pronunciation violation recognition to obtain a second recognition result, wherein the pronunciation violation recognition is used to determine whether there is text content in the text to be recognized that matches the pronunciation of the predetermined violation information; as well as If the second recognition result indicates that there is no text content in the text to be recognized that matches the pronunciation of the predetermined violation information, the text to be recognized is classified as a violation to obtain the target recognition result.
2. The method according to claim 1, further comprising: If the second recognition result indicates that the text to be recognized contains text content that matches the pronunciation of the predetermined violation information, the target recognition result is determined based on the second recognition result.
3. The method according to claim 1 or 2, wherein, The step of identifying pronunciation violations in the text to be identified according to a target identification pattern that matches the first identification result, based on predetermined violation information, to obtain a second identification result, includes: If the first recognition result indicates that the text to be recognized is semantically coherent, then the text to be recognized is segmented into multiple text fragments; and The second recognition result is obtained by matching the pronunciation information of each of the multiple text segments with the pronunciation information of the predetermined violation information.
4. The method according to any one of claims 1 to 3, wherein, The step of identifying pronunciation violations in the text to be identified according to a target identification pattern that matches the first identification result, based on predetermined violation information, to obtain a second identification result, includes: If the first recognition result indicates that the text to be recognized is not semantically coherent, the pronunciation information of each byte in the text to be recognized is matched with the pronunciation information of the predetermined violation information to obtain the second recognition result.
5. The method according to any one of claims 1 to 4, wherein, The process of performing semantic coherence recognition on the text to be recognized to obtain a first recognition result includes: The text to be identified is subjected to perplexity identification to obtain a perplexity score, where the perplexity score represents the semantic coherence of the text to be identified; and Using a perplexity threshold as reference information, and based on the perplexity of the text to be identified, a first identification result characterizing whether the text to be identified is semantically coherent is determined.
6. The method according to claim 5, further comprising: Multiple evaluation samples are used to evaluate multiple predetermined perplexity thresholds to obtain evaluation results for each of the predetermined perplexity thresholds. The evaluation samples include evaluation text and evaluation tags that characterize whether the evaluation text is semantically coherent. as well as The perplexity threshold is determined from the plurality of predetermined perplexity thresholds based on the evaluation results of each of the predetermined perplexity thresholds.
7. The method according to claim 6, wherein, The process of evaluating multiple predetermined perplexity thresholds using multiple evaluation samples to obtain evaluation results for each of the predetermined perplexity thresholds includes: Multiple evaluation texts are subjected to perplexity identification to obtain multiple evaluation perplexity levels; Using the predetermined perplexity threshold as reference information, and based on multiple evaluation perplexities, determine the evaluation and recognition results to determine whether the semantic representations of the multiple evaluation texts are coherent; and The evaluation result of the predetermined perplexity threshold is determined based on the evaluation labels of each of the multiple evaluation texts and the evaluation recognition results of each of the multiple evaluation texts.
8. The method according to any one of claims 1 to 7, wherein, The step of classifying and identifying violations in the text to be identified to obtain the target identification result includes: The text features of the text to be identified are subjected to primary classification recognition to obtain the primary classification result; If the primary classification result indicates that the text to be identified violates a predetermined rule, the text features are then subjected to secondary classification to obtain a secondary classification result, wherein the secondary classification result indicates the rule category violated by the text to be identified. Based on the secondary classification results, the target recognition results are obtained.
9. The method according to claim 8, further comprising: If the primary classification result indicates that the text to be identified conforms to a predetermined rule, the target identification result is determined based on the primary classification result.
10. The method according to any one of claims 1 to 9, wherein, The text to be identified is classified and identified using a violation classification model, which includes a feature extraction module for feature extraction and a first-level classifier for determining whether a predetermined rule is violated. The violation classification model was trained using the following method: The initial feature extraction module is obtained by adjusting the parameters of the original sample text and the enhanced sample text. The enhanced sample text is obtained by replacing the original sample text with content that has the same pronunciation. Using the original sample text and the primary classification label, the parameters of the feature extraction module are frozen, and the parameters of the initial primary classifier are adjusted to obtain the primary classifier. The primary classification label represents whether the original sample text truly violates the predetermined rules.
11. The method according to claim 10, wherein, The violation classification model also includes a secondary classifier, which is used to determine the category of the violated rule; The training operation of the violation classification model also includes: Using the original sample text and the secondary classification labels, the parameters of the target feature extraction module are frozen, and the parameters of the initial secondary classifier are adjusted to obtain the secondary classifier, wherein the secondary classification labels represent the true rule categories violated by the original sample text.
12. A text recognition device, comprising: The semantic recognition module is used to perform semantic coherence recognition on the text to be recognized and obtain the first recognition result; The pronunciation recognition module is used to perform pronunciation violation recognition on the text to be recognized according to a target recognition pattern that matches the first recognition result and based on predetermined violation information, to obtain a second recognition result, wherein the pronunciation violation recognition is used to determine whether there is text content in the text to be recognized that matches the pronunciation of the predetermined violation information; as well as The classification module is used to classify and identify the text to be identified as a violation when the second recognition result indicates that there is no text content that matches the pronunciation of the predetermined violation information, thereby obtaining the target recognition result.
13. An intelligent agent based on artificial intelligence, comprising: The input module is used to receive input information; The processing module is configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and obtain output information by calling the large model to execute the method of any one of claims 1 to 11. An output module is used to output the output information obtained by the processing module.
14. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 11.
15. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 11.
16. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 11.