Text detection method and device, electronic equipment and readable storage medium
By performing word segmentation, clustering and short-term prediction type annotation on applied text, combined with pre-trained short-term prediction recognition model, the problem of inaccurate understanding of text semantics in the prior art is solved, and the accuracy of text detection is improved.
Patent Information
- Application Number
- CN202311677991.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art cannot accurately understand text semantics when detecting whether text is compliant in applications, resulting in a decrease in detection accuracy.
By performing word segmentation, clustering and short-term prediction type annotation of the detected text, short-term prediction recognition values are obtained, and a pre-trained short-term prediction recognition model is used to determine whether there is short-term prediction content in the text.
The accuracy of whether text detection is compliant is improved, and it is possible to identify relatively certainly whether there is short-term predicted content in the text to be detected.
Smart Images

Figure CN120144766A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and particularly to a text detection method, apparatus, electronic device, and readable storage medium. Background Art
[0002] When a user uses a service provided by an application software (referred to as an application), such as a financial product recommendation service, etc., the application can transmit information that the user needs to pay attention to to the user in the form of voice broadcast or text display. Generally, to implement the function of voice broadcast or text display of information, the application developer needs to pre-formulate the text content in advance. When the application runs, the voice broadcast engine or the text rendering engine can broadcast the text content to the user in voice, or display the text content to the user in the form of text.
[0003] To ensure the compliance of the application, the operator of the application will detect the text in the application. In the related art, the text in the application can be compared with the pre-obtained compliant text or non-compliant text through a text difference comparison algorithm, such as the Mayers algorithm, or a text similarity or distance algorithm, such as Euclidean distance, cosine similarity, minimum edit distance, etc., and it is determined whether the text in the application is compliant according to the result of the difference comparison. However, since there is a problem in the related art that the semantics of the text in the application cannot be accurately understood, it is impossible to accurately identify whether the text in the application is compliant, reducing the accuracy of detecting whether the text is compliant. Summary of the Invention
[0004] To solve the problems in the related art, embodiments of the present disclosure provide a text detection method, apparatus, electronic device, and readable storage medium.
[0005] In a first aspect, an embodiment of the present disclosure provides a text detection method, including:
[0006] Obtain a text to be detected, and perform word segmentation on the text to be detected to obtain the word segmentation of the text to be detected;
[0007] Perform clustering on the word segmentation of the text to be detected, and label the clustered word segmentation with a short-term prediction type;
[0008] Based on the short-term prediction type of each clustered word segmentation, obtain a first short-term prediction recognition value of the text to be detected;
[0009] Obtain a pre-trained short-term prediction recognition model, and use the text to be detected, the first short-term prediction recognition value, and the current time as inputs to obtain a second short-term prediction recognition value output by the short-term prediction recognition model;
[0010] Determine whether the text to be detected has short-term prediction content according to the second short-term prediction recognition value.
[0011] In one embodiment of the present disclosure, obtaining a first short-term prediction recognition value of the text to be detected based on the short-term prediction types of each clustering word segmentation includes:
[0012] Determining a weight value corresponding to each short-term prediction type according to the current time;
[0013] Obtaining a short-term prediction recognition value corresponding to each clustering word segmentation according to the weight value corresponding to each short-term prediction type and the word order of each clustering word segmentation in the text to be detected;
[0014] Obtaining the first short-term prediction recognition value according to the short-term prediction recognition value corresponding to each clustering word segmentation.
[0015] In one embodiment of the present disclosure, the short-term prediction types include positive short-term prediction, negative short-term prediction, first-quarter prediction, second-quarter prediction, third-quarter prediction, and fourth-quarter prediction.
[0016] In one embodiment of the present disclosure, obtaining a pre-trained short-term prediction recognition model and using the text to be detected, the first short-term prediction recognition value, and the current time as inputs to obtain a second short-term prediction recognition value output by the short-term prediction recognition model includes:
[0017] Converting the text to be detected, the first short-term prediction recognition value, and the current time into texts respectively, and concatenating them into a text input item in sequence;
[0018] Inputting the text input item into the short-term prediction recognition model to obtain the second short-term prediction recognition value.
[0019] In one embodiment of the present disclosure, the short-term prediction recognition model includes an output gate. The input item of the activation function in the output gate is x. When x > 0, the value σ of the activation function in the output gate is obtained by σ = x / (1 + exp(-βx)). When x ≤ 0, the value σ of the activation function in the output gate is obtained by σ = α(exp(x) - 1), where β is the first short-term prediction recognition value and α is a fixed parameter. In one embodiment of the present disclosure, the method further includes:
[0020] Labeling the misleading description types for the clustering word segmentation;
[0021] Obtaining a first misleading description recognition value of the text to be detected based on the misleading description types of each clustering word segmentation;
[0022] Obtaining a pre-trained misleading description recognition model, obtaining the input of the misleading description recognition model according to the text to be detected, and inputting it into the misleading description recognition model to obtain a second misleading description recognition value output by the misleading description recognition model;
[0023] Determine whether there is misleading description content in the text to be detected according to the value identified by the first misleading description and the value identified by the second misleading description.
[0024] In an embodiment of the present disclosure, obtaining the input of the misleading description recognition model according to the text to be detected and the value of the first misleading description recognition includes:
[0025] Based on the grammatical structure of the text to be detected, determine at least one subject token in the tokenization of the text to be detected, and determine the target subject token among the at least one subject token, and perform numerical conversion on the target subject token;
[0026] Successively splice the text to be detected and the numerical conversion result of the target subject token to obtain the input of the misleading description recognition model.
[0027] In a second aspect, an embodiment of the present disclosure provides a text detection device, including:
[0028] A text tokenization module, configured to obtain the text to be detected and tokenize the text to be detected to obtain the tokenization of the text to be detected;
[0029] A token clustering module, configured to cluster the tokens of the text to be detected and label the short-term prediction type for the clustered tokens;
[0030] A first recognition module, configured to obtain the first short-term prediction recognition value of the text to be detected based on the short-term prediction type of each clustered token;
[0031] A second recognition module, configured to obtain a pre-trained short-term prediction recognition model, and use the text to be detected, the first short-term prediction recognition value, and the current time as inputs to obtain the second short-term prediction recognition value output by the short-term prediction recognition model;
[0032] A short-term prediction recognition module, configured to determine whether there is short-term prediction content in the text to be detected according to the second short-term prediction recognition value.
[0033] In a third aspect, an embodiment of the present disclosure provides an electronic device, including a memory and a processor, wherein the memory is used to store one or more computer instructions, and wherein the one or more computer instructions are executed by the processor to implement the method according to any one of the first aspects.
[0034] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the method according to the first aspect is implemented.
[0035] According to the technical solution provided by the embodiments of the present disclosure, by obtaining the text to be detected, segmenting the text to be detected to obtain the word segmentation of the text to be detected; clustering the word segmentation of the text to be detected, and labeling the clustered word segmentation with short-term prediction types; based on the short-term prediction types of each clustered word segmentation, obtaining the first short-term prediction recognition value of the text to be detected, where the first short-term prediction recognition value can reflect the number of words with short-term prediction in the semantic sense in the text to be detected; obtaining a pre-trained short-term prediction recognition model, and using the text to be detected, the first short-term prediction recognition value, and the current time as inputs to obtain the second short-term prediction recognition value output by the short-term prediction recognition model, where the second short-term prediction recognition value can reflect the recognition result of the pre-trained short-term prediction recognition model at the current time and with the introduction of the second short-term prediction recognition value as a reference for the number of words with short-term prediction in the text to be detected. Therefore, according to the second short-term prediction recognition value, it can be more determined whether there is short-term prediction content in the text to be detected, which helps to accurately identify whether the text to be detected is compliant and improves the accuracy of detecting whether the text is compliant.
[0036] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings
[0037] In combination with the drawings, through the following detailed description of non-limiting embodiments, other features, objects, and advantages of the present disclosure will become more obvious. In the drawings:
[0038] Figure 1 The flowchart of the text detection method according to the embodiments of the present disclosure is shown.
[0039] Figure 2 The structural block diagram of the text detection device according to the embodiments of the present disclosure is shown.
[0040] Figure 3 The structural block diagram of the electronic device according to the embodiments of the present disclosure is shown.
[0041] Figure 4 The structural schematic diagram of the computer system suitable for implementing the method according to the embodiments of the present disclosure is shown. Detailed Embodiments
[0042] Hereinafter, the exemplary embodiments of the present disclosure will be described in detail with reference to the drawings, so that those skilled in the art can easily implement them. In addition, for the sake of clarity, parts irrelevant to the description of the exemplary embodiments are omitted in the drawings.
[0043] In the present disclosure, it should be understood that terms such as "including" or "having" are intended to indicate the presence of features, numbers, steps, actions, components, parts, or combinations thereof disclosed in this specification, and are not intended to exclude the possibility of the presence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0044] In addition, it should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0045] In the present disclosure, when it comes to operations of obtaining user information or user data or operations of presenting user information or user data to others, such operations are all operations authorized, confirmed by the user, or actively selected by the user.
[0046] When a user uses the services provided by an application software (referred to as an application) (such as a financial product recommendation service, etc.), the application can transmit information that the user needs to pay attention to to the user in the form of voice broadcast or text display.
[0047] Exemplarily, taking a financial service application as an example for illustration, the developer of the financial service application can pre - formulate the text content of financial product recommendations. When the user uses the financial product recommendation service, the financial service application can use the speaker on the user terminal and, based on the corresponding text content of financial product recommendations, present the financial product recommendation content for the corresponding financial product to the user in the form of voice broadcast. The financial service application can also use the display screen on the user terminal and, based on the corresponding text content of financial product recommendations, present the financial product recommendation content for the corresponding financial product to the user in the form of text display.
[0048] Compliance management is an inherent requirement for the stable operation of an enterprise, and it is also a basic prerequisite for preventing compliance risks. It is a part that every company must manage and a powerful weapon to protect its own interests. The improvement of the internal control system is inseparable from compliance management and operations, so as to maximize the effectiveness of the enterprise's internal control system. For financial institutions, compliance requirements are even the bottom line to ensure the normal operation of the company. In recent years, there are still financial institutions that have been punished due to compliance issues, and the punishment measures also highlight the regulatory authorities' emphasis on financial compliance. To ensure the compliance of an application, the operator of the application will detect the text in the application.
[0049] In the related art, through a text difference comparison algorithm, such as the Mayers algorithm, or a text similarity or distance algorithm, the text in the application can be compared with the pre-obtained compliant text or non-compliant text for differences, such as Euclidean distance, cosine similarity, minimum edit distance, etc., and it is determined whether the text in the application is compliant according to the result of the difference comparison. In the related art, through a text difference comparison algorithm, such as the Mayers algorithm, or a text similarity or distance algorithm for difference comparison, such as Euclidean distance, cosine similarity, minimum edit distance, etc.
[0050] However, the inventor found that the above-mentioned solutions can all be understood as physically calculating the difference or similarity between the two texts to be detected, and do not understand the difference between the two texts semantically. Exemplarily, if the semantics of the text to be detected is the same as that of the non-compliant text, but their description methods are different, then based on the above-mentioned solutions, a detection result indicating that the difference between the text to be detected and the non-compliant text is small may be obtained, resulting in the inability to accurately identify whether the corresponding text is compliant, and reducing the accuracy of detecting whether the text is compliant.
[0051] To solve the above problems, the present disclosure provides a text detection method, apparatus, electronic device, and readable storage medium.
[0052] According to the technical solution provided by the embodiment of the present disclosure, by obtaining the text to be detected, segmenting the text to be detected to obtain the segmented words of the text to be detected; clustering the segmented words of the text to be detected, and annotating the short-term prediction type for the clustered segmented words; based on the short-term prediction type of each clustered segmented word, obtaining the first short-term prediction recognition value of the text to be detected, where the first short-term prediction recognition value can reflect the number of words with short-term prediction in the semantics in the text to be detected; obtaining a pre-trained short-term prediction recognition model, and taking the text to be detected, the first short-term prediction recognition value, and the current time as inputs to obtain the second short-term prediction recognition value output by the short-term prediction recognition model, where the second short-term prediction recognition value can reflect the recognition result of the pre-trained short-term prediction recognition model at the current time and introducing the above-mentioned second short-term prediction recognition value as a reference for the number of words with short-term prediction in the text to be detected. Therefore, according to the second short-term prediction recognition value, it can be more certain whether there is short-term prediction content in the text to be detected, which helps to accurately identify whether the text to be detected is compliant and improves the accuracy of detecting whether the text is compliant.
[0053] Figure 1 The flowchart of the text detection method according to an embodiment of the present disclosure is shown. As Figure 1 shown, the text detection method includes the following steps S101-S105:
[0054] In step S101, obtain the text to be detected, and perform word segmentation on the text to be detected to obtain the word segmentation of the text to be detected.
[0055] In step S102, perform clustering on the word segmentation of the text to be detected, and label the clustered word segmentation with short-term prediction types.
[0056] In step S103, based on the short-term prediction types of each clustered word segmentation, obtain the first short-term prediction recognition value of the text to be detected.
[0057] In step S104, obtain the pre-trained short-term prediction recognition model, and use the text to be detected, the first short-term prediction recognition value, and the current time as inputs to obtain the second short-term prediction recognition value output by the short-term prediction recognition model.
[0058] In step S105, determine whether the text to be detected has short-term prediction content according to the second short-term prediction recognition value.
[0059] In an implementation manner of the present disclosure, performing word segmentation on the text to be detected can be understood as performing word segmentation on the text to be detected according to a pre-obtained word segmentation algorithm or word segmentation model, etc., where the word segmentation is an indivisible word or character.
[0060] In an implementation manner of the present disclosure, performing clustering on the word segmentation of the text to be detected can be understood as vectorizing each word segmentation based on the One-Hot method; after obtaining the vectorized representation of each word segmentation, use a clustering algorithm to perform unsupervised clustering on the collected word segmentation.
[0061] In an implementation manner of the present disclosure, performing clustering on the word segmentation of the text to be detected can also be understood as clustering the word segmentation of the text to be detected based on the K-means algorithm.
[0062] Exemplarily, the center c_kz of the k-th clustered word segmentation C_k can be obtained based on c_kz = (1 / |C_k|)*Σ(x_i), where x_i is the sample word segmentation belonging to the clustered word segmentation C_k, |C_k| is the number of word segmentations in the clustered word segmentation C_k, and Σ(x_i) is the sum of all word segmentations belonging to C_k.
[0063] For each sample word segmentation x_i, obtain the cluster cluster(x_i) to which each sample word segmentation x_i belongs based on the formula cluster(x_i) = argmin||x_i - c_k||^2, where argmin represents obtaining the parameter value when minimizing the objective function, and ||x_i - c_k||^2 represents the square of the Euclidean distance between the sample word segmentation x_i and the cluster c_kz.
[0064] Based on the above formula, through multiple iterations, continuously update the center of each clustering word segment and the cluster to which each word segment belongs until the center of the clustering word segment no longer changes or reaches a predetermined number of iterations, thereby obtaining the final word segment clustering result.
[0065] In one implementation of the present disclosure, the short-term prediction type can be understood as the type used to indicate the information that may be associated with short-term prediction expressed by the corresponding clustering word segment.
[0066] In one implementation of the present disclosure, the short-term prediction type may include positive short-term prediction, negative short-term prediction, first-quarter prediction, second-quarter prediction, third-quarter prediction, and fourth-quarter prediction.
[0067] Exemplarily, if the clustering word segments are word segments with strong association with short-term prediction content such as "next half year", "next quarter", "within a few weeks", "recent market conditions", etc., the short-term prediction type of this clustering word segment can be labeled as positive short-term prediction;
[0068] If the clustering word segments are word segments with relatively weak association with short-term prediction content such as "after half a year", "quarter-on-quarter", "past", "last year", etc., the short-term prediction type of this clustering word segment can be labeled as negative short-term prediction;
[0069] If the clustering word segments are word segments with strong association with the first quarter such as "beginning of the year", "January", "Spring Festival", "spring", etc., the short-term prediction type of this clustering word segment can be labeled as first-quarter prediction;
[0070] If the clustering word segments are word segments with strong association with the second quarter such as "May Day", "May", "summer", "Tomb-Sweeping Festival", etc., the short-term prediction type of this clustering word segment can be labeled as second-quarter prediction;
[0071] If the clustering word segments are word segments with strong association with the third quarter such as "third quarter", "September", "autumn", "summer vacation", etc., the short-term prediction type of this clustering word segment can be labeled as third-quarter prediction;
[0072] If the clustering word segments are word segments with strong association with the fourth quarter such as "end of the year", "December", "winter", "Double Eleven", etc., the short-term prediction type of this clustering word segment can be labeled as fourth-quarter prediction.
[0073] In one implementation of the present disclosure, labeling the short-term prediction type for the clustering word segment can be understood as matching the clustering word segment with a dictionary obtained in advance for indicating words under each short-term prediction type, and determining the short-term prediction type to which the clustering word segment belongs according to the matching result. It can also be understood as substituting the clustering word segment into a pre-obtained short-term prediction type labeling algorithm for calculation, and determining the short-term prediction type to which the clustering word segment belongs according to the calculation result.
[0074] In an implementation of the present disclosure, obtaining a first short-term prediction recognition value of the text to be detected based on the short-term prediction type of each clustering word segmentation can be understood as looking up based on the short-term prediction recognition value form obtained in advance according to the short-term prediction type of each clustering word segmentation to obtain the corresponding first short-term prediction recognition value; it can also be understood as calculating based on the short-term prediction type of each clustering word segmentation and the weight corresponding to each short-term prediction type to obtain the first short-term prediction recognition value of the text to be detected.
[0075] In an implementation of the present disclosure, the short-term prediction recognition model can be understood as a Hidden Markov Model (HMM), a conditional random field (CRF) model, a Long Short-Term Memory (LSTM), a Bi-directional Long Short-Term Memory (Bi-LSTM) model, a Bidirectional Encoder Representations from Transformers (BERT) model, etc.
[0076] In an implementation of the present disclosure, determining whether the text to be detected has short-term prediction content according to the second short-term prediction recognition value can be understood as when the second short-term prediction recognition value meets the short-term prediction condition (for example, is greater than or equal to the short-term prediction threshold, or belongs to the short-term prediction interval), it can be determined that the text to be detected has short-term prediction content, otherwise it can be determined that there is no short-term prediction content.
[0077] According to the technical solution provided by the embodiments of the present disclosure, by obtaining the text to be detected, segmenting the text to be detected to obtain the word segmentation of the text to be detected; clustering the word segmentation of the text to be detected and labeling the clustered word segmentation with short-term prediction types; based on the short-term prediction types of each clustered word segmentation, obtaining the first short-term prediction recognition value of the text to be detected, where the first short-term prediction recognition value can reflect the number of words with short-term prediction in semantics in the text to be detected; obtaining a pre-trained short-term prediction recognition model, and using the text to be detected, the first short-term prediction recognition value, and the current time as inputs to obtain the second short-term prediction recognition value output by the short-term prediction recognition model, where the second short-term prediction recognition value can reflect the recognition result of the pre-trained short-term prediction recognition model at the current time and with the introduction of the second short-term prediction recognition value as a reference for the number of words with short-term prediction in the text to be detected. Therefore, according to the second short-term prediction recognition value, it can be more determined whether there is short-term prediction content in the text to be detected, which helps to accurately identify whether the text to be detected is compliant and improves the accuracy of detecting whether the text is compliant.
[0078] In an implementation manner of the present disclosure, obtaining the first short-term prediction recognition value of the text to be detected based on the short-term prediction types of each clustered word segmentation includes:
[0079] Determining the weight value corresponding to each short-term prediction type according to the current time;
[0080] Obtaining the short-term prediction recognition value corresponding to each clustered word segmentation according to the weight value corresponding to each short-term prediction type and the word order of each clustered word segmentation in the text to be detected;
[0081] Obtaining the first short-term prediction recognition value according to the short-term prediction recognition value corresponding to each clustered word segmentation.
[0082] In an implementation manner of the present disclosure, determining the weight value corresponding to each short-term prediction type according to the current time can be understood as querying based on a pre-obtained weight value form according to the current time to obtain the weight value corresponding to each short-term prediction type; or it can also be understood as substituting the current time into a pre-obtained weight value algorithm for calculation to obtain the weight value corresponding to each short-term prediction type, where the weight value corresponding to the short-term prediction type is positively correlated with the short-term prediction recognition value corresponding to the clustered word segmentation belonging to the short-term prediction type.
[0083] Exemplarily, if the current time belongs to the third quarter, then the predicted values for the third quarter and the fourth quarter are more likely to be short-term predictions. Therefore, for the positive short-term prediction, the third-quarter prediction, and the fourth-quarter prediction in the short-term prediction type, the corresponding weight values α_1, α_5, and α_6 are all positive, while for the negative short-term prediction, the first-quarter prediction, and the second-quarter prediction, the corresponding weight values α_2, α_3, and α_4 are all negative.
[0084] In an implementation manner of the present disclosure, the word order of the clustering and word segmentation in the text to be detected can be understood as the positional relationship between the clustering and word segmentation corresponding to the short-term prediction type and the clustering and word segmentation of other short-term prediction types in the text to be detected.
[0085] Exemplarily, if the current time belongs to the third quarter, then the predicted values for the third quarter and the fourth quarter are more likely to be short-term predictions. Therefore, in the text to be detected, if a clustering and word segmentation belonging to the positive short-term prediction is adjacent to a clustering and word segmentation belonging to the third-quarter prediction or the fourth-quarter prediction, and another clustering and word segmentation belonging to the positive short-term prediction is adjacent to a clustering and word segmentation belonging to the first-quarter prediction or the second-quarter prediction, then the short-term prediction recognition value corresponding to the former clustering and word segmentation should be greater than the short-term prediction recognition value corresponding to the latter clustering and word segmentation.
[0086] According to the technical solution provided by the embodiments of the present disclosure, by determining the weight value corresponding to each short-term prediction type according to the current time; obtaining the short-term prediction recognition value corresponding to each clustering and word segmentation according to the weight value corresponding to each short-term prediction type and the word order of each clustering and word segmentation in the text to be detected; and obtaining the first short-term prediction recognition value according to the short-term prediction recognition value corresponding to each clustering and word segmentation, it can be ensured that the first short-term prediction recognition value can more comprehensively reflect whether there is short-term prediction description content in the text to be detected.
[0087] In an implementation manner of the present disclosure, obtaining a pre-trained short-term prediction recognition model, and using the text to be detected, the first short-term prediction recognition value, and the current time as inputs to obtain the second short-term prediction recognition value output by the short-term prediction recognition model includes:
[0088] Converting the text to be detected, the first short-term prediction recognition value, and the current time into texts respectively, and sequentially concatenating them into a text input item;
[0089] Inputting the text input item into the short-term prediction recognition model to obtain the second short-term prediction recognition value.
[0090] Exemplarily, the text input items can be combined by means of data splicing. That is, [CLS] is used to represent the start of the text to be detected, and [SEP] is used to represent the separation between two adjacent types of content. The text to be detected is: "I think there may be a significant market decline in July this year", the current time is "June 28th", and the first short-term prediction recognition value is σ. Then the text input item can be [CLS] "I think there may be a significant market decline in July this year" + [SEP] + σ + [SEP] + the time stamp corresponding to "June 28th".
[0091] According to the technical solution provided by the embodiment of the present disclosure, by converting the text to be detected, the first short-term prediction recognition value, and the current time into texts respectively, and splicing them into a text input item in sequence; and inputting the text input item into the short-term prediction recognition model to obtain the second short-term prediction recognition value, it can enable the model to understand the current time while understanding the semantics, which helps the model to dynamically perceive the timeliness in the text to be detected, enables the model to better distinguish the prediction timeliness while recognizing and predicting the semantics, and strengthens the discrimination effect of the model.
[0092] In an implementation manner of the present disclosure, the short-term prediction recognition model can be a bidirectional long short-term memory network model (Bi-directional Long Short-Term Memory, Bi-LSTM). This short-term prediction recognition model includes a forgetting gate, an input gate, and an output gate. The input of the forgetting gate is the input of the short-term prediction recognition model, that is, the text input item formed by converting the text to be detected, the first short-term prediction recognition value, and the current time into texts respectively and splicing them in sequence. The output of the output gate is the output of the short-term prediction recognition model. Among them, the output x of the input gate is the input item of the activation function in the output gate. When x > 0, the value σ of the activation function in the output gate can be obtained by σ = x / (1 + exp(-βx)); when x ≤ 0, the value σ of the activation function in the output gate can be obtained by σ = α(exp(x) - 1), where β is the first short-term prediction recognition value; α is a fixed parameter, and the value of α can be between 0 and 2, preferably 1.67.
[0093] According to the technical solution provided by the embodiment of the present disclosure, by setting the activation function in the output gate of the short-term prediction recognition model such that when x > 0, the value σ of the activation function in the output gate is obtained by σ = x / (1 + exp(-βx)), and when x ≤ 0, the value σ of the activation function in the output gate is obtained by σ = α(exp(x) - 1), it can be ensured that the larger the value of β, the more obvious the short-term time is, and the function is closer to 1. On the contrary, when the value of β is smaller, the more obvious the non-short-term time is, and the function will be farther away from 1. Thus, the output mean of the model is close to 0, which can help the model accelerate learning towards 0.
[0094] In one embodiment of the present disclosure, the method further includes:
[0095] Labeling the misleading description type for the clustered word segmentation;
[0096] Based on the misleading description type of each clustered word segmentation, obtaining the first misleading description recognition value of the text to be detected;
[0097] Obtaining a pre-trained misleading description recognition model, obtaining the input of the misleading description recognition model according to the text to be detected, and inputting it into the misleading description recognition model to obtain the second misleading description recognition value output by the misleading description recognition model;
[0098] Determining whether there is misleading description content in the text to be detected according to the first misleading description recognition value and the second misleading description recognition value.
[0099] In one implementation manner of the present disclosure, the misleading description type may include at least one of master (for example, the corresponding clustered word segmentations are treasure fund managers, goddesses, etc.), risk (for example, the corresponding clustered word segmentations are risk aversion optimization, almost no risk, etc.), drawdown (for example, the corresponding clustered word segmentations are small drawdown degree, strict drawdown control, etc.), opportunity (for example, the corresponding clustered word segmentations are not to be missed, layout opportunity, etc.), return (for example, the corresponding clustered word segmentations are obvious excess return, extremely high rate of return, etc.), experience (for example, the corresponding clustered word segmentations are better investment experience, happier, etc.), cost performance (for example, the corresponding clustered word segmentations are extremely cost-effective, constructing a high cost-effective portfolio, etc.), performance (for example, the corresponding clustered word segmentations are performance witnessing glory, ranking among the top, etc.), most (for example, the corresponding clustered word segmentations are the most powerful, the most cutting-edge, etc.), excellent (for example, the corresponding clustered word segmentations are high-light moments, setting new highs repeatedly, etc.).
[0100] In one implementation manner of the present disclosure, labeling the misleading description type for the clustered word segmentation can be understood as matching the clustered word segmentation with a pre-obtained dictionary for indicating words under each misleading description type, and determining the misleading description type to which the clustered word segmentation belongs according to the matching result. It can also be understood as substituting the clustered word segmentation into a pre-obtained misleading description type labeling algorithm for calculation, and determining the misleading description type to which the clustered word segmentation belongs according to the calculation result.
[0101] In one implementation manner of the present disclosure, obtaining the second misleading description recognition value of the text to be detected based on the misleading description type of each clustered word segmentation can be understood as looking up based on the pre-obtained misleading description recognition value form according to the misleading description type of each clustered word segmentation to obtain the corresponding first misleading description recognition value; it can also be understood as calculating based on the misleading description type of each clustered word segmentation and the weight corresponding to each misleading description type to obtain the first misleading description recognition value of the text to be detected.
[0102] In one implementation of the present disclosure, the misleading description recognition model can be understood as a Hidden Markov Model (HMM), a conditional random field (CRF) model, a Long Short-Term Memory (LSTM), a Bi-directional Long Short-Term Memory (Bi-LSTM) model, a Bidirectional Encoder Representations from Transformers (BERT) model, etc.
[0103] In one implementation of the present disclosure, determining whether there is misleading description content in the text to be detected according to the first misleading description recognition value and the second misleading description recognition value can be understood as when at least one of the first misleading description recognition value and the second misleading description recognition value meets the misleading description condition (for example, greater than or equal to the misleading description threshold, or belonging to the misleading description interval), it can be determined that there is misleading description content in the text to be detected, otherwise it can be determined that there is no misleading description content.
[0104] According to the technical solution provided by the embodiments of the present disclosure, by clustering word segmentation to label the misleading description type; based on the misleading description type of each clustering word segmentation, obtaining the first misleading description recognition value of the text to be detected; obtaining the pre-trained misleading description recognition model, obtaining the input of the misleading description recognition model according to the text to be detected, and inputting it into the misleading description recognition model to obtain the second misleading description recognition value output by the misleading description recognition model; determining whether there is misleading description content in the text to be detected according to the first misleading description recognition value and the second misleading description recognition value, it can be more accurately determined whether there is misleading description content in the text to be detected, which helps to further identify whether the text to be detected is compliant and improves the accuracy of detecting whether the text is compliant.
[0105] In one implementation of the present disclosure, obtaining the input of the misleading description recognition model according to the text to be detected and the first misleading description recognition value includes:
[0106] Based on the grammatical structure of the text to be detected, determining at least one subject word segmentation in the word segmentation of the text to be detected, and determining the target subject word segmentation in the at least one subject word segmentation, and performing numerical conversion on the target word segmentation;
[0107] The numerical conversion results of the text to be detected and the target subject participles are concatenated in sequence to obtain the input of the misleading description recognition model.
[0108] In one implementation manner of the present disclosure, the target subject participle can be understood as the main object in a specific industry that the user is more concerned about. For example, if the specific industry is the fund industry, the target subject participle may include participles corresponding to personal names, participles corresponding to fund products, participles corresponding to fund types, etc.
[0109] Exemplarily, if the text to be detected is "A is a newly emerged fund goddess, a master of winning rate, and has a leading advantage in the B fund product track", then A and B therein can be understood as the target subject participles; if the text to be detected is "The C hybrid fund has a very strong income ability recently and is worthy of customers' trust", then C therein can be understood as the target subject participle; if the text to be detected is "D is the champion of this year's E competition and has very strong abilities", then D and E therein can be understood as the target subject participles.
[0110] According to the technical solution provided by the embodiment of the present disclosure, at least one subject participle is determined from the participles of the text to be detected based on the grammatical structure of the text to be detected, and the target subject participle is determined from the at least one subject participle, and numerical conversion is performed on the target participle; the numerical conversion results of the text to be detected and the target subject participle are concatenated in sequence to obtain the input of the misleading description recognition model, so as to obtain the input of the misleading description recognition model, which can help the model notice whether there are differences in the main bodies described in the text to be detected, so that the model can more accurately judge whether there is a misleading description in the text to be detected.
[0111] Figure 2 The structural block diagram of the text detection device according to the embodiment of the present disclosure is shown. Among them, the device can be implemented as part or all of an electronic device through software, hardware, or a combination of both.
[0112] As Figure 2 shown, the text detection device 200 includes:
[0113] A text participle module, configured to obtain the text to be detected and perform participle on the text to be detected to obtain the participles of the text to be detected;
[0114] A participle clustering module, configured to cluster the participles of the text to be detected and label the short-term prediction types for the clustered participles;
[0115] A first recognition module, configured to obtain the first short-term prediction recognition value of the text to be detected based on the short-term prediction types of each clustered participle;
[0116] A second recognition module, configured to obtain a pre-trained short-term prediction recognition model, and use the text to be detected, the first short-term prediction recognition value, and the current time as inputs to obtain a second short-term prediction recognition value output by the short-term prediction recognition model;
[0117] A short-term prediction recognition module, configured to determine whether there is short-term prediction content in the text to be detected according to the second short-term prediction recognition value.
[0118] According to the technical solution provided by the embodiments of the present disclosure, by obtaining the text to be detected, segmenting the text to be detected to obtain the segmented words of the text to be detected; clustering the segmented words of the text to be detected, and labeling the clustered segmented words with short-term prediction types; based on the short-term prediction types of each clustered segmented word, obtaining the first short-term prediction recognition value of the text to be detected, where the first short-term prediction recognition value can reflect the number of words with short-term prediction in the semantic sense in the text to be detected; obtaining a pre-trained short-term prediction recognition model, and using the text to be detected, the first short-term prediction recognition value, and the current time as inputs to obtain a second short-term prediction recognition value output by the short-term prediction recognition model, where the second short-term prediction recognition value can reflect the recognition result of the pre-trained short-term prediction recognition model at the current time and introducing the above-mentioned second short-term prediction recognition value as a reference for the number of words with short-term prediction in the text to be detected. Therefore, according to the second short-term prediction recognition value, it can be relatively determined whether there is short-term prediction content in the text to be detected, which helps to accurately identify whether the text to be detected is compliant and improves the accuracy of detecting whether the text is compliant.
[0119] The present disclosure also discloses an electronic device, Figure 3 showing a structural block diagram of an electronic device according to an embodiment of the present disclosure.
[0120] As Figure 3 shown, the electronic device includes a memory and a processor, where the memory is used to store one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the method according to the embodiments of the present disclosure.
[0121] An embodiment of the present disclosure provides a text detection method, including:
[0122] Obtaining the text to be detected, and segmenting the text to be detected to obtain the segmented words of the text to be detected;
[0123] Clustering the segmented words of the text to be detected, and labeling the clustered segmented words with short-term prediction types;
[0124] Obtaining the first short-term prediction recognition value of the text to be detected based on the short-term prediction types of each clustered segmented word;
[0125] Obtain the pre-trained short-term prediction recognition model, and use the text to be detected, the first short-term prediction recognition value, and the current time as inputs to obtain the second short-term prediction recognition value output by the short-term prediction recognition model;
[0126] Determine whether there is short-term prediction content in the text to be detected according to the second short-term prediction recognition value.
[0127] In an implementation manner of the present disclosure, obtaining the first short-term prediction recognition value of the text to be detected based on the short-term prediction type of each clustered word segment includes:
[0128] Determine the weight value corresponding to each short-term prediction type according to the current time;
[0129] Obtain the short-term prediction recognition value corresponding to each clustered word segment according to the weight value corresponding to each short-term prediction type and the word order of each clustered word segment in the text to be detected;
[0130] Obtain the first short-term prediction recognition value according to the short-term prediction recognition value corresponding to each clustered word segment.
[0131] In an implementation manner of the present disclosure, the short-term prediction types include positive short-term prediction, negative short-term prediction, first-quarter prediction, second-quarter prediction, third-quarter prediction, and fourth-quarter prediction.
[0132] In an implementation manner of the present disclosure, obtaining the pre-trained short-term prediction recognition model, and using the text to be detected, the first short-term prediction recognition value, and the current time as inputs to obtain the second short-term prediction recognition value output by the short-term prediction recognition model includes:
[0133] Convert the text to be detected, the first short-term prediction recognition value, and the current time into texts respectively, and splice them into a text input item in sequence;
[0134] Input the text input item into the short-term prediction recognition model to obtain the second short-term prediction recognition value.
[0135] In an implementation manner of the present disclosure, the short-term prediction recognition model includes an output gate. The input item of the activation function in the output gate is x. When x>0, the value σ of the activation function in the output gate is obtained by σ = x / (1 + exp(-βx)). When x≤0, the value σ of the activation function in the output gate is obtained by σ = α(exp(x) - 1), where β is the first short-term prediction recognition value and α is a fixed parameter. In an implementation manner of the present disclosure, the method further includes:
[0136] Label the misleading description type for the clustered word segment;
[0137] Obtain the first misleading description recognition value of the text to be detected based on the misleading description type of each clustered word segment;
[0138] Obtain a pre-trained misleading description recognition model, obtain the input of the misleading description recognition model according to the text to be detected, and input it into the misleading description recognition model to obtain the second misleading description recognition value output by the misleading description recognition model;
[0139] Determine whether there is misleading description content in the text to be detected according to the first misleading description recognition value and the second misleading description recognition value.
[0140] In an implementation manner of the present disclosure, obtaining the input of the misleading description recognition model according to the text to be detected and the first misleading description recognition value includes;
[0141] Based on the grammatical structure of the text to be detected, determine at least one subject token in the word segmentation of the text to be detected, determine the target subject token in the at least one subject token, and perform numerical conversion on the target subject token;
[0142] Successively splice the text to be detected and the numerical conversion result of the target subject token to obtain the input of the misleading description recognition model.
[0143] Figure 4 Show a schematic structural diagram of a computer system suitable for implementing the method according to the embodiments of the present disclosure.
[0144] As Figure 4 shown, the computer system includes a processing unit, which can execute various methods in the above embodiments according to the program stored in the read-only memory (ROM) or the program loaded from the storage part into the random access memory (RAM). In the RAM, various programs and data required for the operation of the computer system are also stored. The processing unit, ROM, and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus.
[0145] The following components are connected to the I / O interface: an input part including a keyboard, a mouse, etc.; an output part including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage part including a hard disk, etc.; and a communication part including a network interface card such as a LAN card, a modem, etc. The communication part performs a communication process via a network such as the Internet. The drive is also connected to the I / O interface as needed. Removable media, such as magnetic disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on the drive as needed, so that the computer programs read from them can be installed into the storage part as needed. Among them, the processing unit can be implemented as a processing unit such as a CPU, GPU, TPU, FPGA, NPU, etc.
[0146] In particular, according to an embodiment of the present disclosure, the method described above can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing the above method. In such an embodiment, the computer program can be downloaded and installed from a network via a communication section, and / or installed from a removable medium.
[0147] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0148] The units or modules involved in the embodiments described in the present disclosure can be implemented in software or in programmable hardware. The units or modules described can also be provided in a processor, and the names of these units or modules do not in some cases constitute a limitation on the units or modules themselves.
[0149] As another aspect, the present disclosure also provides a computer-readable storage medium, which can be the computer-readable storage medium included in the electronic device or computer system in the above embodiments; or can exist separately and be a computer-readable storage medium not assembled into the device. The computer-readable storage medium stores one or more programs, and the one or more programs are used by one or more processors to execute the methods described in the present disclosure.
[0150] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the present disclosure that have similar functions.
Claims
1. A text detection method, characterized in that, it includes: Obtain the text to be detected, and segment the text to be detected to obtain the segmented words of the text to be detected; Cluster the segmented words of the text to be detected, and label the clustered segmented words with short-term prediction types; Based on the short-term prediction types of each clustered segmented word, obtain the first short-term prediction recognition value of the text to be detected; Obtain a pre-trained short-term prediction recognition model, and use the text to be detected, the first short-term prediction recognition value, and the current time as inputs to obtain the second short-term prediction recognition value output by the short-term prediction recognition model; Determine whether there is short-term prediction content in the text to be detected according to the second short-term prediction recognition value.
2. The text detection method according to claim 1, characterized in that, The step of obtaining the first short-term prediction recognition value of the text to be detected based on the short-term prediction types of each clustered segmented word includes: Determine the weight value corresponding to each short-term prediction type according to the current time; According to the weight value corresponding to each short-term prediction type and the word order of each clustered segmented word in the text to be detected, obtain the short-term prediction recognition value corresponding to each clustered segmented word; Obtain the first short-term prediction recognition value according to the short-term prediction recognition value corresponding to each clustered segmented word.
3. The text detection method according to claim 1, characterized in that, The short-term prediction types include positive short-term prediction, negative short-term prediction, first-quarter prediction, second-quarter prediction, third-quarter prediction, and fourth-quarter prediction.
4. The text detection method according to claim 1, characterized in that, The step of obtaining a pre-trained short-term prediction recognition model and using the text to be detected, the first short-term prediction recognition value, and the current time as inputs to obtain the second short-term prediction recognition value output by the short-term prediction recognition model includes: Convert the text to be detected, the first short-term prediction recognition value, and the current time into texts respectively, and splice them into a text input item in sequence; Input the text input item into the short-term prediction recognition model to obtain the second short-term prediction recognition value.
5. The text detection method according to claim 1, characterized in that, The short-term prediction recognition model includes an output gate, where the input item of the activation function in the output gate is x. When x > 0, the value σ of the activation function in the output gate is obtained by σ = x / (1 + exp(-βx)); when x ≤ 0, the value σ of the activation function in the output gate is obtained by σ = α(exp(x) - 1), where β is the first short-term prediction recognition value and α is a fixed parameter.
6. The text detection method according to any one of claims 1-5, characterized in that, The method further includes: Label the clustered segmented words with misleading description types; Based on the misleading description types of each clustered segmented word, obtain the first misleading description recognition value of the text to be detected; Obtain a pre-trained misleading description recognition model, obtain the input of the misleading description recognition model according to the text to be detected, and input it into the misleading description recognition model to obtain the second misleading description recognition value output by the misleading description recognition model; Determine whether there is misleading description content in the text to be detected according to the first misleading description recognition value and the second misleading description recognition value.
7. The text detection method according to claim 6, characterized in that The obtaining of the input of the misleading description recognition model according to the text to be detected and the first misleading description recognition value includes; Based on the grammatical structure of the text to be detected, determine at least one subject token in the tokens of the text to be detected, and determine a target subject token among the at least one subject token, and perform numerical conversion on the target subject token; Successively splice the text to be detected and the numerical conversion result of the target subject token to obtain the input of the misleading description recognition model.
8. A text detection device, characterized in that including: A text tokenization module configured to obtain a text to be detected and tokenize the text to be detected to obtain the tokens of the text to be detected; A token clustering module configured to cluster the tokens of the text to be detected and label the clustered tokens with short-term prediction types; A first recognition module configured to obtain the first short-term prediction recognition value of the text to be detected based on the short-term prediction types of each clustered token; A second recognition module configured to obtain a pre-trained short-term prediction recognition model, and use the text to be detected, the first short-term prediction recognition value, and the current time as inputs to obtain the second short-term prediction recognition value output by the short-term prediction recognition model; A short-term prediction recognition module configured to determine whether there is short-term prediction content in the text to be detected according to the second short-term prediction recognition value.
9. An electronic device, characterized in that It includes a memory and a processor; wherein, the memory is used to store one or more computer instructions, and wherein, the one or more computer instructions are executed by the processor to implement the method steps described in any one of claims 1-8.
10. A computer-readable storage medium, on which computer instructions are stored, characterized in that When the computer instructions are executed by a processor, the method steps described in any one of claims 1-8 are implemented.