Text processing method and device based on artificial intelligence, computer equipment and medium
The AI-driven method enhances Chinese speech synthesis by using multiple classifiers to process and aggregate multiple-tone character predictions, addressing inefficiencies and inaccuracies in rule-based systems, thereby improving processing efficiency and accuracy.
Patent Information
- Application Number
- CN202510374230.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-15
AI Technical Summary
The existing rule-driven system has problems of low processing efficiency and low accuracy when dealing with Chinese polyphonics, and it is difficult to adapt to rapid changes in languages and new phenomena.
Using an artificial intelligence-based method, after receiving the multi-tone text sequence input by the user, adjusting and semantic feature extraction, multiple classifiers (such as fully connected network, BLSTM and Transformer blocks) are used to predict and aggregate processing to generate target pronunciation results.
It improves the processing efficiency and accuracy of polyphonic pronunciation recognition, can better adapt to language changes, and improves the flexibility and accuracy of the system.
Smart Images

Figure CN120317243A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and can be applied to fields such as fintech and digital healthcare. In particular, it relates to a text processing method, device, computer device, and storage medium based on artificial intelligence. Background Art
[0002] In a Chinese Mandarin speech synthesis system, accurately converting Chinese characters into corresponding phonemes is a key step to ensure the quality of speech synthesis. This conversion process is particularly complex because there are a large number of polyphonic characters in Chinese, and these characters may have different pronunciations in different contexts. Therefore, how to effectively handle the problem of polyphonic characters has become an important challenge in the field of Chinese speech synthesis.
[0003] Currently, the industry mainly relies on rule-driven systems to handle the problem of polyphonic characters. The working principle of such systems is based on a series of rules summarized by speech experts and converted into a computer-processable form. These rules are designed to assign the correct pronunciation to Chinese characters according to their context. However, although this method has achieved certain results to a certain extent, its inherent limitations have become increasingly prominent.
[0004] First, the development cost of rule-driven systems is high. Due to the need for the extensive participation of speech experts and the in-depth accumulation of language knowledge, building and improving such a rule system not only requires a large amount of human and time resources, but may also lead to low development efficiency. Second, as the number of rules continues to increase, the problem of rule conflicts has become increasingly serious. Due to the complexity and variability of the Chinese context, some polyphonic characters may match multiple rules in different contexts, which makes it difficult for the system to select the correct pronunciation, thus reducing the accuracy of polyphonic character pronunciation recognition. In addition, rule-driven systems have poor adaptability to newly emerging language phenomena. With the continuous development of society and the continuous evolution of language, new words, phrases, and expressions emerge in an endless stream. However, since rule-driven systems rely on a pre-defined rule library, it is difficult to quickly adapt to these new language phenomena, thus limiting their flexibility and universality in practical applications.
[0005] In summary, the existing rule-driven systems have problems of low processing efficiency and low accuracy in dealing with Chinese polyphonic characters. Therefore, there is an urgent need in the industry for a more efficient and accurate method for dealing with polyphonic characters to promote the further development of Chinese Mandarin speech synthesis technology. Summary of the Invention
[0006] The purpose of the embodiments of this application is to propose a text processing method, device, computer device, and storage medium based on artificial intelligence to solve the technical problems of low processing efficiency and low accuracy existing in the existing rule-driven systems when dealing with Chinese polyphonic characters.
[0007] In a first aspect, there is provided an artificial intelligence-based text processing method, including:
[0008] Receiving an original text sequence containing polyphonic characters input by a user;
[0009] Performing adjustment processing on the original text sequence to obtain a corresponding target text sequence;
[0010] Performing semantic feature extraction processing on the target text sequence to obtain corresponding semantic features;
[0011] Obtaining the data scale of the original text sequence and determining whether the data scale is greater than a preset scale;
[0012] If so, calling a target number of classifiers and respectively performing prediction processing on the semantic features based on each of the classifiers to obtain corresponding multiple pronunciation results;
[0013] Performing aggregation processing on all the pronunciation results to obtain a corresponding target pronunciation result;
[0014] Performing output processing on the target pronunciation result.
[0015] In a second aspect, there is provided an artificial intelligence-based text processing apparatus, including:
[0016] A receiving module, configured to receive an original text sequence containing polyphonic characters input by a user;
[0017] An adjustment module, configured to perform adjustment processing on the original text sequence to obtain a corresponding target text sequence;
[0018] An extraction module, configured to perform semantic feature extraction processing on the target text sequence to obtain corresponding semantic features;
[0019] A judgment module, configured to obtain the data scale of the original text sequence and determine whether the data scale is greater than a preset scale;
[0020] A prediction module, configured to, if so, call a target number of classifiers and respectively perform prediction processing on the semantic features based on each of the classifiers to obtain corresponding multiple pronunciation results;
[0021] An aggregation module, configured to perform aggregation processing on all the pronunciation results to obtain a corresponding target pronunciation result;
[0022] A first output module, configured to perform output processing on the target pronunciation result.
[0023] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned text processing method based on artificial intelligence are implemented.
[0024] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned text processing method based on artificial intelligence are implemented.
[0025] In the solution implemented by the above-mentioned text processing method, device, computer device, and storage medium based on artificial intelligence, first, an original text sequence containing polyphonic characters input by a user is received; and the original text sequence is preprocessed and adjusted to obtain a corresponding target text sequence; then, semantic feature extraction processing is performed on the target text sequence to obtain corresponding semantic features; after that, the data scale of the original text sequence is obtained, and it is determined whether the data scale is greater than a preset scale; if so, a target number of classifiers are called, and the semantic features are respectively predicted based on each of the classifiers to obtain corresponding multiple pronunciation results; subsequently, all the pronunciation results are aggregated to obtain a corresponding target pronunciation result; finally, output processing is performed on the target pronunciation result. In this application, the original text sequence containing polyphonic characters input by the user is adjusted to obtain a target text sequence, then semantic feature extraction processing is performed on the target text sequence to obtain semantic features, after that, the data scale of the original text sequence is obtained, and when it is detected that the data scale of the original text sequence is greater than the preset scale, the target number of classifiers are intelligently called, and the semantic features are respectively predicted based on each classifier to obtain multiple pronunciation results, subsequently, all the pronunciation results are aggregated to obtain a target pronunciation result, and finally, output processing is performed on the target pronunciation result. Different from the existing processing method for identifying the pronunciation of polyphonic characters based on rules, this application combines the use of multiple classifiers to predict the pronunciation of polyphonic characters and aggregate them to generate the corresponding target pronunciation result, effectively improving the processing efficiency and accuracy of identifying the pronunciation of polyphonic characters, and effectively improving the accuracy of the generated target pronunciation result. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the solutions in this application, the following will briefly introduce the drawings required for the description of the embodiments of this application. Obviously, the following drawings are some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0027] Figure 1 It is an exemplary system architecture diagram to which this application can be applied;
[0028] Figure 2 is a flowchart of an embodiment of an artificial intelligence-based text processing method according to the present application;
[0029] Figure 3 is a schematic structural diagram of an embodiment of an artificial intelligence-based text processing apparatus according to the present application;
[0030] Figure 4 is a schematic structural diagram of an embodiment of a computer device according to the present application. Detailed implementation manners
[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.
[0032] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0033] To enable those skilled in the technical field to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0034] As Figure 1 shown, the system architecture 100 may include a terminal device 101, a network 102, and a server 103. The terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0035] The user may use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications may be installed on the terminal device 101, such as a web browser application, a shopping application, a search application, an instant messaging tool, an email client, a social platform software, etc.
[0036] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop 1011, tablet computer 1012, or mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer, a desktop computer, and so on.
[0037] The server 103 can be a server that provides various services, such as a background server that supports the pages displayed on the terminal device 101.
[0038] It should be noted that the artificial intelligence-based text processing method provided by the embodiments of the present application is generally executed by the server / terminal device. Correspondingly, the artificial intelligence-based text processing device is generally set in the server / terminal device.
[0039] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in
[0040] Continuing to refer to Figure 2 , a flowchart of an embodiment of the artificial intelligence-based text processing method according to the present application is shown. According to different requirements, the order of the steps in this flowchart can be changed, and some steps can be omitted. The artificial intelligence-based text processing method provided by the embodiments of the present application can be applied to any scenario that requires polyphonic character recognition. Then, the artificial intelligence-based text processing method can be applied to the products in these scenarios, such as polyphonic character recognition in the financial field or the medical field. The artificial intelligence-based text processing method includes the following steps:
[0041] Step S201, receiving an original text sequence containing polyphonic characters input by a user.
[0042] In this embodiment, the electronic device on which the artificial intelligence-based text processing method runs (such as Figure 1The server / terminal device shown in the figure) can obtain the original text sequence of the polyphonic character by wired connection or wireless connection. It should be noted that the above wireless connection method may include but is not limited to 3G / 4G / 5G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other wireless connection methods known now or developed in the future. The execution subject of this application is a text processing system, which can be referred to as the system. The above original text sequence can be a string and can contain multiple polyphonic characters and ordinary Chinese characters. This application can be applied to business processing scenarios of polyphonic character recognition in the financial field and the medical field. Exemplarily, in the financial field, the content of the original text sequence can be: "Our bank will release the latest financial report next week, which will detail our assets and liabilities and profit growth rate." Polyphonic character description: In this example, "行" is a polyphonic character, which is usually pronounced as "xíng" in "我行" (in the context of "our bank" or "I'm going to do something"), but the specific pronunciation still needs to be judged in combination with the context and the semantics of the entire sentence. However, here we assume that it is an abbreviation for "bank" and should be pronounced "háng".
[0043] Alternatively, in the financial field, the original text sequence may also read: "The stock market has been volatile recently. Investors should be cautious and should not blindly follow the trend." Note on polyphones: Although there are no obvious polyphones in this example, it is worth noting that some Chinese characters in financial terms may have different interpretations or pronunciation tendencies in different contexts (although they may not necessarily constitute polyphones in the strict sense). For example, "涨" in the stock market usually refers to price increases and is pronounced "zhǎng", but it may also have the pronunciation of "zhàng" in other contexts (such as "泡涨"). However, here we focus on polyphones in the traditional sense.
[0044] In the medical field, the content of the original text sequence may be: "After receiving chemotherapy, patients often experience adverse reactions such as nausea and vomiting." Explanation of polyphones: In this example, "恶" is a polyphone. It is pronounced as "ě" in "难" (indicating a feeling of discomfort and nausea), and is pronounced as "è" in other contexts (such as "malicious" and "nasty").
[0045] Alternatively, in the medical field, the original text sequence may also read: "The doctor advises the patient to monitor blood pressure regularly, so as to pay attention to changes in blood pressure in time and prevent cardiovascular diseases such as hypertension." Polyphone explanation: Although there are no obvious polyphones in this example, the word "blood" has different pronunciation tendencies in medical terms. Usually, "blood" is pronounced as "xuè" (referring to blood) when used alone, but it may also be pronounced as a light "xie" (without tone) in some compound words or idioms (such as "blood sweat" and "blood tears"). However, in professional texts in the medical field, "blood" is almost always pronounced as "xuè". In addition, although the word "pressure" is not a polyphone, it has a specific meaning and pronunciation in medical terms (such as "pressure" in "blood pressure" is pronounced as "yā").
[0046] In this embodiment, illustratively, in the business scenario of financial insurance product push, the business tracking data may include transaction data, payment data, business data, and the like.
[0047] Step S202: adjusting the original text sequence to obtain a corresponding target text sequence.
[0048] In this embodiment, the above-mentioned specific implementation process of adjusting the original text sequence to obtain the corresponding target text sequence will be further described in detail in subsequent specific embodiments of the present application, and will not be elaborated on here.
[0049] Step S203: performing semantic feature extraction processing on the target text sequence to obtain corresponding semantic features.
[0050] In the present embodiment, the above-mentioned semantic feature extraction processing is performed on the target text sequence to obtain the specific implementation process of the corresponding semantic features. This application will provide further details on this in subsequent specific embodiments, and will not be elaborated on here.
[0051] Step S204, obtaining the data size of the original text sequence, and determining whether the data size is greater than a preset size.
[0052] In the present embodiment, the above-mentioned data scale refers to the number of polyphones contained in the above-mentioned original text sequence. There is no specific limitation on the selection of the numerical value of the above-mentioned preset scale, which can be determined according to actual business needs. Exemplarily, the preset scale can be set to 3. If the data scale of the original text sequence, that is, the number of polyphones contained in the original text sequence is greater than 3, it indicates that the data scale of the original text sequence is large. If the number of polyphones contained in the original text sequence is less than 3, it indicates that the data scale of the original text sequence is small.
[0053] Step S205, if so, call a classifier with a target number, and respectively perform prediction processing on the semantic features based on each of the classifiers to obtain a corresponding plurality of pronunciation results.
[0054] In this embodiment, the above target number is 3, and the above plurality of classifiers specifically include three neural network-based classifiers, namely a fully connected network classifier, a bidirectional long short-term memory network (BLSTM) classifier, and a classifier based on a transformer block. Among them, (1) Fully connected network classifier: Process the speech features of polyphonic characters through a simple two-layer fully connected network, and predict its pronunciation through a separate output layer. (2) Bidirectional long short-term memory network (BLSTM) classifier: Use BLSTM to capture the dynamic dependence information of the context of polyphonic characters, especially the influence of the adjacent context, to provide more detailed context modeling for pronunciation prediction. (3) Classifier based on a transformer block: Model the long-distance context information through the transformer structure to ensure that the long-distance context information can also participate in the prediction decision.
[0055] Specifically, for the fully connected network classifier, the fully connected network classifier receives the semantic features, performs feature transformation through a two-layer fully connected network (i.e., a linear layer + an activation function), calculates the probability distribution of each pronunciation label using the softmax function, and takes the label with the highest probability as the prediction result, that is, the pronunciation result. For the bidirectional long short-term memory network (BLSTM) classifier, through the semantic features, and then transfer the features to the BLSTM layer to capture the dynamic dependence information of the context. Subsequently, a fully connected layer and the softmax function are applied to the output of the BLSTM layer to predict the pronunciation label to obtain the corresponding pronunciation result. For the classifier based on a Transformer block: Receive the semantic features output by the BERT model, and then use an additional Transformer block (or self-attention mechanism) to further process the feature sequence to capture the long-distance context information. Subsequently, a fully connected layer and the softmax function are applied to the output of the Transformer block to predict the pronunciation label to obtain the corresponding pronunciation result.
[0056] In addition, the process of training the classifier includes: training the fully-connected network classifier, the BLSTM classifier, and the Transformer-block-based classifier separately. The model is trained using the training set, and the hyperparameters such as the learning rate, batch size, number of network layers, etc. are adjusted using the validation set. Then, the performance of each classifier on the validation set, such as accuracy, F1 score, etc., is recorded. After that, the performance of the ensemble model is evaluated using the test set, and the training and evaluation processes are repeated until satisfactory performance is achieved, thereby obtaining the trained classifier. Subsequently, the trained ensemble model is deployed to the production environment. And the performance of the classifier in actual applications, such as accuracy, response time, etc., is monitored. Among them, the change of the data distribution can be regularly checked, and whether the classifier needs to be retrained to adapt to the new data. Also, when updating the classifier, methods such as incremental learning and online learning can be considered to efficiently utilize the new data.
[0057] In addition, the system designs an independent output layer for each polyphonic character separately to avoid the confusion problem between the pronunciation labels of different polyphonic characters. This design enables the system to flexibly expand to handle newly added polyphonic characters, only by adding a new output layer, which significantly improves the scalability of the system.
[0058] Step S206: Aggregate all the pronunciation results to obtain the corresponding target pronunciation result.
[0059] In this embodiment, the specific implementation process of aggregating all the pronunciation results to obtain the corresponding target pronunciation result will be further described in detail in the subsequent specific embodiments of this application, and will not be elaborated here too much.
[0060] Step S207: Perform output processing on the target pronunciation result.
[0061] In this embodiment, the specific implementation process of performing output processing on the target pronunciation result will be further described in detail in the subsequent specific embodiments of this application, and will not be elaborated here too much.
[0062] This application first receives the original text sequence containing polyphonic characters input by the user; then adjusts and processes the original text sequence to obtain the corresponding target text sequence; then performs semantic feature extraction processing on the target text sequence to obtain the corresponding semantic features; then obtains the data scale of the original text sequence and determines whether the data scale is greater than the preset scale; if so, calls the target number of classifiers and predicts the semantic features based on each classifier respectively to obtain the corresponding multiple pronunciation results; subsequently, aggregates all the pronunciation results to obtain the corresponding target pronunciation result; and finally performs output processing on the target pronunciation result. This application adjusts and processes the original text sequence containing polyphonic characters input by the user to obtain the target text sequence, then performs semantic feature extraction processing on the target text sequence to obtain semantic features, then obtains the data scale of the original text sequence, and when it detects that the data scale of the original text sequence is greater than the preset scale, it will intelligently call the target number of classifiers and predict the semantic features based on each classifier respectively to obtain multiple pronunciation results, then aggregates all the pronunciation results to obtain the target pronunciation result, and finally performs output processing on the target pronunciation result. Different from the existing processing method for identifying the pronunciation of polyphonic characters based on rules, this application combines the use of multiple classifiers to predict the pronunciation of polyphonic characters and aggregates them to generate the corresponding target pronunciation result, effectively improving the processing efficiency and accuracy of identifying the pronunciation of polyphonic characters, and effectively improving the accuracy of the generated target pronunciation result.
[0063] In some optional implementation manners of this embodiment, step S202 includes the following steps:
[0064] Perform cleaning processing on the original text sequence to obtain the corresponding first text sequence.
[0065] In this embodiment, the above cleaning processing refers to removing irrelevant characters, including deleting non-Chinese characters such as punctuation marks and numbers in the text.
[0066] Perform word segmentation processing on the first text sequence to obtain the corresponding second text sequence.
[0067] In this embodiment, the processing method for the above word segmentation processing is not specifically limited, and character-level word segmentation can be used, or more complex word segmentation methods can be used, such as dictionary-based word segmentation or deep learning-based word segmentation.
[0068] Perform standardization processing on the second text sequence to obtain the corresponding third text sequence.
[0069] In this embodiment, the above standardization processing refers to converting the text into a unified format, such as converting full-width characters to half-width characters.
[0070] Use the third text sequence as the target text sequence.
[0071] In this application, the original text sequence is cleaned to obtain the corresponding first text sequence; then the first text sequence is tokenized to obtain the corresponding second text sequence; afterwards, the second text sequence is normalized to obtain the corresponding third text sequence; subsequently, the third text sequence is used as the target text sequence. By cleaning, tokenizing, and normalizing the original text sequence, this application can automatically and accurately complete the adjustment process of the original text sequence, effectively ensuring the accuracy and standardization of the generated target text sequence.
[0072] In some alternative implementation manners, step S203 includes the following steps:
[0073] Perform format conversion processing on the target text sequence to obtain the corresponding fourth text sequence.
[0074] In this embodiment, the above format conversion processing refers to converting the input target text sequence into the input format required by the BERT model, including processing such as adding special tokens (such as [CLS] and [SEP]).
[0075] Invoke the pre-trained BERT model.
[0076] In this embodiment, BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model based on the Transformer architecture, aiming to generate deep language representations through bidirectional context understanding. The BERT model has made breakthrough progress in the field of natural language processing through its unique bidirectional architecture and pre-training-fine-tuning paradigm, becoming one of the current state-of-the-art language models. Its applications are extensive and far-reaching, promoting the rapid development of NLP technology.
[0077] By using the pre-trained BERT model, semantic information can be learned from large-scale unsupervised data. The BERT model is based on the transformer structure, and its bidirectional context awareness ability can effectively capture the context information of polyphonic characters. Compared with traditional methods, the semantic features generated by the BERT model are more comprehensive and accurate, and can better support the polyphonic pronunciation prediction task.
[0078] Perform encoding processing on the fourth text sequence based on the BERT model to obtain the embedding representation corresponding to the fourth text sequence.
[0079] In this embodiment, the above fourth text sequence can be passed to the BERT model to obtain the embedding representation of each character corresponding to the fourth text sequence. These embedding representations are the semantic features for subsequent classification tasks.
[0080] Use the embedding representation as the semantic feature.
[0081] This application performs format conversion processing on the target text sequence to obtain the corresponding fourth text sequence; then calls the pre-trained BERT model; then encodes the fourth text sequence based on the BERT model to obtain the embedding representation corresponding to the fourth text sequence; subsequently, the embedding representation is used as the semantic feature. By performing format conversion processing on the target text sequence to obtain the corresponding fourth text sequence, and then encoding the fourth text sequence based on the use of the pre-trained BERT model, this application can efficiently and accurately extract the semantic features corresponding to the target text sequence, improve the extraction efficiency of semantic features, and ensure the accuracy of the obtained semantic features.
[0082] In some optional implementation manners, step S206 includes the following steps:
[0083] Obtain a preset ensemble learning strategy.
[0084] In this embodiment, there is no specific limitation on the selection of the above ensemble learning strategy, which can be determined according to actual business requirements. Specifically, the ensemble learning strategy can be selected as the voting method (such as majority voting, weighted average voting), the stacking method (Stacking), or the bagging method (Bagging), etc. For the voting method, each classifier independently predicts the test samples, and then summarizes the prediction results, and selects the label with the most occurrences as the final prediction. For the stacking method, a meta-classifier (such as logistic regression, decision tree, etc.) can be trained, and its input is the prediction results (usually probability distributions) of each base classifier, and the output is the final prediction label. For the bagging method, the bagging method is used to improve the stability of the model. By randomly sampling the training set multiple times and training multiple base classifiers, and then integrating their prediction results.
[0085] Aggregate all the pronunciation results based on the ensemble learning strategy to obtain the corresponding aggregation result.
[0086] In this embodiment, all the pronunciation results output by multiple classifiers can be aggregated according to the processing steps of the selected ensemble learning strategy to obtain the corresponding aggregation result. Among them, each classifier will output the probability distribution of one or more pronunciation labels.
[0087] Determine the corresponding target pronunciation label based on the aggregation result.
[0088] In this embodiment, according to the aggregation result, a correct pronunciation label can be determined for each polyphonic character and used as the target pronunciation label. Specifically, the pronunciation label with the highest probability can be selected as the final prediction result, or the prediction results of multiple classifiers can be weighted and averaged, and then the label with the highest probability can be selected as the final prediction result, that is, the target pronunciation label.
[0089] Use the target pronunciation label as the target pronunciation result.
[0090] This application obtains a preset ensemble learning strategy; then aggregates all the pronunciation results based on the ensemble learning strategy to obtain a corresponding aggregation result; then determines a corresponding target pronunciation label based on the aggregation result; subsequently, uses the target pronunciation label as the target pronunciation result. By aggregating all the pronunciation results based on the use of the ensemble learning strategy to obtain a corresponding aggregation result, and then determining the target pronunciation label based on the aggregation result and using it as the corresponding target pronunciation result, this application combines different prediction results through the ensemble learning strategy, can make full use of the advantages of multiple classifiers, and effectively improves the accuracy and robustness of polyphonic character recognition.
[0091] In some alternative implementation manners, step S207 includes the following steps:
[0092] Align the target pronunciation result with the original text sequence to obtain a processed fifth text sequence.
[0093] In this embodiment, the target pronunciation result can be aligned with the original text sequence so that the user can clearly see the pronunciation of each Chinese character. Specifically, it can be achieved by inserting the target pronunciation result into the original text or generating a new text sequence.
[0094] Obtain a preset variety of output manners.
[0095] In this embodiment, the above-mentioned variety of output manners may at least include pinyin, phonetic symbols, or other forms of pronunciation representations.
[0096] Determine a target output manner from all the output manners.
[0097] In this embodiment, there is no specific limitation on the determination method of the above-mentioned target output manner, and it can be selected according to actual business requirements, and any one of pinyin, phonetic symbols, or other forms of pronunciation representations can be selected.
[0098] Based on the target output manner, perform output processing on the fifth text sequence.
[0099] In this embodiment, according to the selected target output manner, the fifth text sequence can be output in a user-friendly manner corresponding thereto.
[0100] This application performs alignment processing on the target pronunciation result and the original text sequence to obtain a processed fifth text sequence; then obtains a variety of preset output methods; then determines a target output method from all the output methods; subsequently, based on the target output method, performs output processing on the fifth text sequence. By performing alignment processing on the target pronunciation result and the original text sequence to obtain a processed fifth text sequence, then determining a target output method from a variety of preset output methods, and further performing output processing on the fifth text sequence based on the use of the target output method, this application can output the target text sequence in a user-friendly manner, which is beneficial to improving the output intelligence of the target text sequence and further enhancing the user experience.
[0101] In some optional implementation manners of this embodiment, after step S203, the above electronic device may further perform the following steps:
[0102] If the data scale is smaller than the preset scale, then screen out a specified classifier from all the classifiers.
[0103] In this embodiment, for the specific implementation process of screening out the specified classifier from all the classifiers, this application will further describe the details in subsequent specific embodiments and will not elaborate too much here.
[0104] Perform prediction processing on the semantic feature based on the specified classifier to obtain a corresponding specified pronunciation result.
[0105] In this embodiment, the above semantic feature can be passed to the above specified classifier, so that the specified classifier can perform classification prediction according to the semantic feature to determine the correct pronunciation label for each polyphonic character and output the corresponding specified pronunciation result.
[0106] Perform output processing on the specified pronunciation result.
[0107] In this embodiment, for the processing process of performing output processing on the specified pronunciation result, reference can be made to the specific implementation process of performing output processing on the target pronunciation result above, and details will not be elaborated here.
[0108] If the data rule is detected to be smaller than the preset scale in this application, a specified classifier is screened out from all the classifiers; then, based on the specified classifier, prediction processing is performed on the semantic features to obtain a corresponding specified pronunciation result; subsequently, output processing is performed on the specified pronunciation result. When it is detected that the data scale of the original text sequence is smaller than the preset scale, this application will automatically and intelligently screen out a specified classifier from all the classifiers, then perform prediction processing on the semantic features based on the specified classifier to obtain a corresponding specified pronunciation result, and further perform output processing on the specified pronunciation result, thereby realizing the selection of a single classifier for pronunciation prediction processing when the data scale of the original text sequence is small, which can effectively improve the processing efficiency of pronunciation prediction processing, reduce the processing workload and energy consumption of pronunciation prediction processing.
[0109] In some optional implementation manners of this embodiment, the screening out of a specified classifier from all the classifiers includes the following steps:
[0110] Obtain the processing efficiency of each classifier.
[0111] In this embodiment, text containing polyphonic characters collected in advance can be obtained as test data, and then the test data is used to perform test processing on each classifier, and the processing time of each classifier on the test data is recorded. Then, the reciprocal of the processing time is calculated and used as the processing efficiency of the classifier.
[0112] Obtain the processing accuracy rate of each classifier.
[0113] In this embodiment, each classifier can be tested using the above test data, and the accuracy rate of each classifier on the test data is recorded as the above processing accuracy rate.
[0114] Generate a processing score for each classifier based on the processing efficiency and processing accuracy rate of each classifier.
[0115] In this embodiment, the processing efficiency and processing accuracy rate of each classifier can be calculated using a weighted summation algorithm to generate a processing score for each classifier. Among them, the numerical selection of the first preset weight for the processing efficiency and the second preset weight for the processing accuracy rate is not specifically limited and can be selected according to actual business requirements.
[0116] Screen out the target classifier with the highest processing score from all the classifiers.
[0117] In this embodiment, the processing scores of the generated classifiers can be numerically compared, and then the target classifier with the highest processing score can be screened out from all the classifiers according to the numerical comparison result.
[0118] Use the target classifier as the specified classifier.
[0119] In this application, the processing efficiency of each classifier is obtained; the processing accuracy rate of each classifier is obtained; then, based on the processing efficiency and processing accuracy rate of each classifier, the processing score of each classifier is generated; after that, the target classifier with the highest processing score is selected from all the classifiers; subsequently, the target classifier is used as the specified classifier. By obtaining the processing efficiency and processing accuracy rate of each classifier, and then generating the processing score of each classifier based on the processing efficiency and processing accuracy rate of each classifier, this application further screens out the target classifier with the highest processing score from all the classifiers and uses it as the specified classifier, so as to determine the specified classifier that meets the requirements by considering the performance of the processing efficiency and processing accuracy rate of each classifier. Since the specified classifier has the optimal comprehensive performance corresponding to accuracy and efficiency, the accuracy and intelligence of the determination of the specified classifier are effectively ensured, enabling the subsequent use of the specified classifier to perform predictive processing on semantic features, which can effectively improve the processing efficiency and processing accuracy of polyphonic character recognition processing for the original text sequence.
[0120] In some alternative implementation manners, the obtained user information has obtained the consent of the user and complies with the provisions of relevant laws and relevant policies.
[0121] In addition, the non-company software tools or components that appear in the embodiments of this application are only introduced by way of example and do not represent actual use.
[0122] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0123] It should be emphasized that to further ensure the privacy and security of the above target pronunciation result, the above target pronunciation result can also be stored in a node of a blockchain.
[0124] The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.
[0125] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0126] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0127] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, Read-Only Memory (ROM), or a Random Access Memory (RAM), etc.
[0128] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0129] Further reference Figure 3 to Figure 2 As an implementation of the method shown above, the present application provides an embodiment of a text processing device based on artificial intelligence. This device embodiment corresponds to the method embodiment shown in Figure 2 and this device can be specifically applied to various electronic devices.
[0130] Such as Figure 3As shown in the figure, the artificial intelligence-based text processing device 300 described in this embodiment includes: a receiving module 301, an adjustment module 302, an extraction module 303, a judgment module 304, a prediction module 305, an aggregation module 306, and a first output module 307. Among them:
[0131] The receiving module 301 is configured to receive an original text sequence containing polyphonic characters input by a user;
[0132] The adjustment module 302 is configured to perform adjustment processing on the original text sequence to obtain a corresponding target text sequence;
[0133] The extraction module 303 is configured to perform semantic feature extraction processing on the target text sequence to obtain corresponding semantic features;
[0134] The judgment module 304 is configured to obtain the data scale of the original text sequence and judge whether the data scale is greater than a preset scale;
[0135] The prediction module 305 is configured to, if so, call a target number of classifiers and perform prediction processing on the semantic features based on each of the classifiers to obtain corresponding multiple pronunciation results;
[0136] The aggregation module 306 is configured to perform aggregation processing on all the pronunciation results to obtain a corresponding target pronunciation result;
[0137] The first output module 307 is configured to perform output processing on the target pronunciation result.
[0138] In this embodiment, the operations respectively performed by the above modules or units correspond one by one to the steps of the artificial intelligence-based text processing method in the foregoing embodiment, and will not be elaborated herein.
[0139] In some optional implementation manners of this embodiment, the adjustment processing module 302 includes:
[0140] The first processing sub-module is configured to perform cleaning processing on the original text sequence to obtain a corresponding first text sequence;
[0141] The second processing sub-module is configured to perform word segmentation processing on the first text sequence to obtain a corresponding second text sequence;
[0142] The third processing sub-module is configured to perform standardization processing on the second text sequence to obtain a corresponding third text sequence;
[0143] The first determination sub-module is configured to use the third text sequence as the target text sequence.
[0144] In this embodiment, the operations respectively performed by the above-mentioned modules or units correspond one by one to the steps of the text processing method based on artificial intelligence in the foregoing embodiment, and will not be elaborated herein.
[0145] In some optional implementation manners of this embodiment, the extraction module 303 includes:
[0146] A fourth processing sub-module, configured to perform format conversion processing on the target text sequence to obtain a corresponding fourth text sequence;
[0147] A calling sub-module, configured to call a pre-trained BERT model;
[0148] An encoding sub-module, configured to perform encoding processing on the fourth text sequence based on the BERT model to obtain an embedding representation corresponding to the fourth text sequence;
[0149] A second determination sub-module, configured to use the embedding representation as the semantic feature.
[0150] In this embodiment, the operations respectively performed by the above-mentioned modules or units correspond one by one to the steps of the text processing method based on artificial intelligence in the foregoing embodiment, and will not be elaborated herein.
[0151] In some optional implementation manners of this embodiment, the aggregation module 306 includes:
[0152] A first acquisition sub-module, configured to acquire a preset ensemble learning strategy;
[0153] An aggregation sub-module, configured to perform aggregation processing on all the pronunciation results based on the ensemble learning strategy to obtain a corresponding aggregation result;
[0154] A third determination sub-module, configured to determine a corresponding target pronunciation label based on the aggregation result;
[0155] A fourth determination sub-module, configured to use the target pronunciation label as the target pronunciation result.
[0156] In this embodiment, the operations respectively performed by the above-mentioned modules or units correspond one by one to the steps of the text processing method based on artificial intelligence in the foregoing embodiment, and will not be elaborated herein.
[0157] In some optional implementation manners of this embodiment, the first output module 307 includes:
[0158] An alignment sub-module, configured to perform alignment processing on the target pronunciation result and the original text sequence to obtain a processed fifth text sequence;
[0159] A second acquisition sub-module, configured to acquire a preset variety of output manners;
[0160] A fifth determination sub-module, configured to determine a target output mode from all the output modes;
[0161] An output sub-module, configured to perform output processing on the fifth text sequence based on the target output mode.
[0162] In this embodiment, the operations respectively performed by the above modules or units correspond one by one to the steps of the artificial intelligence-based text processing method in the foregoing embodiment, and will not be elaborated herein.
[0163] In some optional implementation manners of this embodiment, the artificial intelligence-based text processing apparatus further includes:
[0164] A screening module, configured to screen out a specified classifier from all the classifiers if the data scale is smaller than the preset scale;
[0165] A processing module, configured to perform prediction processing on the semantic features based on the specified classifier to obtain a corresponding specified pronunciation result;
[0166] A second output module, configured to perform output processing on the specified pronunciation result.
[0167] In this embodiment, the operations respectively performed by the above modules or units correspond one by one to the steps of the artificial intelligence-based text processing method in the foregoing embodiment, and will not be elaborated herein.
[0168] In some optional implementation manners of this embodiment, the screening module includes:
[0169] A third acquisition sub-module, configured to acquire the processing efficiency of each classifier;
[0170] A fourth acquisition sub-module, configured to acquire the processing accuracy rate of each classifier;
[0171] A generation sub-module, configured to generate a processing score for each classifier based on the processing efficiency and processing accuracy rate of each classifier;
[0172] A screening sub-module, configured to screen out a target classifier with the highest processing score from all the classifiers;
[0173] A sixth determination sub-module, configured to use the target classifier as the specified classifier. In this embodiment, the operations respectively performed by the above modules or units correspond one by one to the steps of the artificial intelligence-based text processing method in the foregoing embodiment, and will not be elaborated herein.
[0174] To solve the above technical problems, an embodiment of the present application further provides a computer device. For details, please refer to Figure 4 , Figure 4This is the basic structural block diagram of the computer device in this embodiment.
[0175] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that communicate with each other through a system bus. It should be noted that only the computer device 4 with components 41 - 43 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of this technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0176] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can interact with the user through a keyboard, a mouse, a remote control, a touchpad, or a voice control device, etc.
[0177] The memory 41 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 4. Of course, the memory 41 can also include both the internal storage unit and the external storage device of the computer device 4. In this embodiment, the memory 41 is usually used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for the text processing method based on artificial intelligence. In addition, the memory 41 can also be used to temporarily store various data that have been output or will be output.
[0178] In some embodiments, the processor 42 may be a Central Processing Unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run the computer-readable instructions stored in the memory 41 or process data, such as running the computer-readable instructions of the artificial intelligence-based text processing method.
[0179] The network interface 43 may include a wireless network interface or a wired network interface, and the network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.
[0180] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects:
[0181] In the embodiments of the present application, the present application first receives an original text sequence containing polyphonic characters input by a user; and performs adjustment processing on the original text sequence to obtain a corresponding target text sequence; then performs semantic feature extraction processing on the target text sequence to obtain corresponding semantic features; then obtains the data scale of the original text sequence, and determines whether the data scale is greater than a preset scale; if so, invokes a target number of classifiers, and respectively performs prediction processing on the semantic features based on each classifier to obtain corresponding multiple pronunciation results; subsequently performs aggregation processing on all the pronunciation results to obtain a corresponding target pronunciation result; and finally performs output processing on the target pronunciation result. The present application obtains a target text sequence by performing adjustment processing on the received original text sequence containing polyphonic characters input by the user, then performs semantic feature extraction processing on the target text sequence to obtain semantic features, then obtains the data scale of the original text sequence, and when it is detected that the data scale of the original text sequence is greater than the preset scale, intelligently invokes a target number of classifiers, and respectively performs prediction processing on the semantic features based on each classifier to obtain multiple pronunciation results, subsequently performs aggregation processing on all the pronunciation results to obtain a target pronunciation result, and finally performs output processing on the target pronunciation result. Different from the existing processing method of identifying the pronunciation of polyphonic characters based on rules, the present application performs prediction processing on the pronunciation of polyphonic characters by combining the use of multiple classifiers and aggregates to generate corresponding target pronunciation results, effectively improving the processing efficiency and processing accuracy of identifying the pronunciation of polyphonic characters, and effectively improving the accuracy of the generated target pronunciation results.
[0182] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor, so that the at least one processor executes the steps of the above-mentioned artificial intelligence-based text processing method.
[0183] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects:
[0184] In the embodiments of the present application, the present application first receives an original text sequence containing polyphonic characters input by a user; and performs adjustment processing on the original text sequence to obtain a corresponding target text sequence; then performs semantic feature extraction processing on the target text sequence to obtain corresponding semantic features; then obtains the data scale of the original text sequence, and determines whether the data scale is greater than a preset scale; if so, invokes a target number of classifiers, and respectively performs prediction processing on the semantic features based on each classifier to obtain corresponding multiple pronunciation results; subsequently performs aggregation processing on all the pronunciation results to obtain a corresponding target pronunciation result; and finally performs output processing on the target pronunciation result. The present application obtains a target text sequence by performing adjustment processing on the received original text sequence containing polyphonic characters input by the user, then performs semantic feature extraction processing on the target text sequence to obtain semantic features, then obtains the data scale of the original text sequence, and when it is detected that the data scale of the original text sequence is greater than the preset scale, intelligently invokes a target number of classifiers, and respectively performs prediction processing on the semantic features based on each classifier to obtain multiple pronunciation results, subsequently performs aggregation processing on all the pronunciation results to obtain a target pronunciation result, and finally performs output processing on the target pronunciation result. Different from the existing processing manner of identifying the pronunciation of polyphonic characters based on rules, the present application combines the use of multiple classifiers to perform prediction processing on the pronunciation of polyphonic characters and aggregates to generate a corresponding target pronunciation result, effectively improving the processing efficiency and processing accuracy of identifying the pronunciation of polyphonic characters, and effectively improving the accuracy of the generated target pronunciation result.
[0185] Through the description of the above implementation manners, those skilled in the art can clearly understand that the method of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present application.
[0186] Obviously, the embodiments described above are only a part of the embodiments of this application, rather than all of them. The preferred embodiments of this application are shown in the accompanying drawings, but they do not limit the patent scope of this application. This application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of this application more thorough and comprehensive. Although this application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure that makes use of the content of the specification and drawings of this application, directly or indirectly applied in other related technical fields, is equally within the scope of patent protection of this application.
Claims
1. A text processing method based on artificial intelligence, characterized in that, It includes the following steps: Receive the original text sequence containing polyphonic characters input by the user; Perform adjustment processing on the original text sequence to obtain the corresponding target text sequence; Perform semantic feature extraction processing on the target text sequence to obtain the corresponding semantic features; Obtain the data scale of the original text sequence and determine whether the data scale is greater than the preset scale; If so, call the target number of classifiers and perform prediction processing on the semantic features based on each of the classifiers to obtain the corresponding multiple pronunciation results; Perform aggregation processing on all the pronunciation results to obtain the corresponding target pronunciation result; Perform output processing on the target pronunciation result.
2. The text processing method based on artificial intelligence according to claim 1, wherein The step of performing adjustment processing on the original text sequence to obtain the corresponding target text sequence specifically includes: Perform cleaning processing on the original text sequence to obtain the corresponding first text sequence; Perform word segmentation processing on the first text sequence to obtain the corresponding second text sequence; Perform standardization processing on the second text sequence to obtain the corresponding third text sequence; Use the third text sequence as the target text sequence.
3. The text processing method based on artificial intelligence according to claim 1, wherein The step of performing semantic feature extraction processing on the target text sequence to obtain the corresponding semantic features specifically includes: Perform format conversion processing on the target text sequence to obtain the corresponding fourth text sequence; Call the pre-trained BERT model; Perform encoding processing on the fourth text sequence based on the BERT model to obtain the embedding representation corresponding to the fourth text sequence; Use the embedding representation as the semantic features.
4. The text processing method based on artificial intelligence according to claim 1, characterized in that The step of performing aggregation processing on all the pronunciation results to obtain the corresponding target pronunciation result specifically includes: Obtain the preset ensemble learning strategy; Perform aggregation processing on all the pronunciation results based on the ensemble learning strategy to obtain the corresponding aggregation result; Determine the corresponding target pronunciation label based on the aggregation result; Use the target pronunciation label as the target pronunciation result.
5. The text processing method based on artificial intelligence according to claim 1, wherein The step of performing output processing on the target pronunciation result specifically includes: Perform alignment processing on the target pronunciation result and the original text sequence to obtain the processed fifth text sequence; Obtain the preset multiple output methods; Determine the target output method from all the output methods; Perform output processing on the fifth text sequence based on the target output method.
6. The text processing method based on artificial intelligence according to claim 1, characterized in that, After the step of obtaining the data scale of the original text sequence and determining whether the data scale is greater than the preset scale, it further includes: If the data scale is less than the preset scale, then screen out the specified classifier from all the classifiers; Perform prediction processing on the semantic features based on the specified classifier to obtain the corresponding specified pronunciation result; Perform output processing on the specified pronunciation result.
7. The text processing method based on artificial intelligence according to claim 6, wherein, The step of screening out the specified classifier from all the classifiers specifically includes: Obtain the processing efficiency of each classifier; Obtain the processing accuracy of each classifier; Generate the processing scores of each classifier based on the processing efficiency and processing accuracy of each classifier; Select the target classifier with the highest processing score from all the classifiers; Use the target classifier as the specified classifier.
8. An artificial intelligence-based text processing device, characterized in that, Comprising: A receiving module, configured to receive an original text sequence containing polyphonic characters input by a user; An adjustment module, configured to perform adjustment processing on the original text sequence to obtain a corresponding target text sequence; An extraction module, configured to perform semantic feature extraction processing on the target text sequence to obtain corresponding semantic features; A judgment module, configured to obtain the data scale of the original text sequence and judge whether the data scale is greater than a preset scale; A prediction module, configured to, if so, call a target number of classifiers and perform prediction processing on the semantic features respectively based on each of the classifiers to obtain corresponding multiple pronunciation results; An aggregation module, configured to perform aggregation processing on all the pronunciation results to obtain a corresponding target pronunciation result; A first output module, configured to perform output processing on the target pronunciation result.
9. A computer device, characterized in that, Comprising a memory and a processor, wherein computer-readable instructions are stored in the memory, and when the processor executes the computer-readable instructions, the steps of the text processing method based on artificial intelligence according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that, Computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by a processor, the steps of the text processing method based on artificial intelligence according to any one of claims 1 to 7 are implemented.