Text type detection method and apparatus

By performing preliminary checks on long texts and selecting sentences with a large number of keywords, and combining this with a pre-trained model to determine the identifying sentences, the problem of low efficiency in long text type detection in existing technologies is solved, and efficient and accurate text type recognition is achieved.

CN115827867BActive Publication Date: 2026-04-28BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2022-12-12
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

The lack of effective methods for long text type detection in existing technologies leads to low efficiency in text type recognition for internet push services.

Method used

After preliminary examination of the text to be tested, sentences with a large number of keywords are selected as target sentences after sentence segmentation. A pre-trained model is used to identify the identifying sentences, and the text type is determined when the preset conditions are met.

Benefits of technology

It simplifies the detection process and improves the accuracy and efficiency of long text type recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115827867B_ABST
    Figure CN115827867B_ABST
Patent Text Reader

Abstract

The present disclosure provides a text type detection method and device, relates to the technical field of data processing, and particularly relates to natural language processing. The implementation scheme is as follows: performing preliminary inspection on a to-be-detected text; in response to the to-be-detected text being preliminarily inspected as a suspected preset type of text, performing sentence segmentation on the to-be-detected text to obtain a sentence set containing a plurality of sentences; determining, for each sentence in the sentence set, a number of predetermined keywords contained in the sentence; selecting a plurality of target sentences from the plurality of sentences according to the number of keywords contained in the plurality of sentences; respectively determining whether each target sentence in the plurality of target sentences is an identification sentence; and in response to the number of identification sentences satisfying a preset condition, determining that the to-be-detected text is a preset type of text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and more particularly to natural language processing, specifically to a method and apparatus for detecting text types, electronic devices, computer-readable storage media, and computer program products. Background Technology

[0002] The internet contains a vast amount of text information, which may include various types, such as sports-related text and lifestyle-related text. When the internet provides text push services to users, it needs to push text of relevant types based on user preferences. Therefore, the internet currently requires the servers responsible for push services to be able to quickly and effectively identify the text type to facilitate subsequent push services. However, most existing text type detection methods are for short texts, and there are currently no methods for detecting the text type of long texts.

[0003] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention

[0004] This disclosure provides a method and apparatus for detecting text types, an electronic device, a computer-readable storage medium, and a computer program product.

[0005] According to one aspect of this disclosure, a method for detecting text types is provided, comprising: performing a preliminary examination on the text to be detected; in response to the preliminary examination indicating that the text to be detected is suspected to be text of a preset type, segmenting the text to be detected into sentences to obtain a set of sentences containing multiple sentences; for each sentence in the set of sentences, determining the number of predetermined keywords contained in the sentence, wherein the keywords are associated with a preset type; selecting multiple target sentences from the multiple sentences based on the number of keywords contained in the multiple sentences in the set of sentences; determining whether each target sentence in the multiple target sentences is an identifier sentence; and in response to determining that the number of identifier sentences meets a preset condition, determining that the text to be detected is text of a preset type.

[0006] According to another aspect of this disclosure, a text type detection apparatus is provided, comprising: an inspection unit configured to perform a preliminary inspection on a text to be detected; a sentence segmentation unit configured to, in response to the text to be detected being preliminarily identified as text of a suspected preset type, segment the text to be detected into sentences to obtain a sentence set containing multiple sentences; a first determination unit configured to, for each sentence in the sentence set, determine the number of predetermined keywords contained in that sentence, wherein the keywords are associated with a preset type; a selection unit configured to, based on the number of keywords contained in the multiple sentences in the sentence set, select multiple target sentences from the multiple sentences; a second determination unit configured to, respectively, determine whether each target sentence among the multiple target sentences is an identifier sentence; and a third determination unit configured to, in response to determining that the number of identifier sentences meets a preset condition, determine that the text to be detected is text of a preset type.

[0007] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the methods described above.

[0008] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the methods described above.

[0009] According to another aspect of this disclosure, a computer program product is provided, including a computer program, wherein the computer program implements the above-described method when executed by a processor.

[0010] According to one or more embodiments of this disclosure, the text to be detected can be segmented into multiple sentences, and a portion of the target sentences that are likely to be the identifying sentences can be selected from the multiple sentences for inspection, thereby avoiding the detection of the entire text, simplifying the detection process, and improving the accuracy of detection.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0013] Figure 1A flowchart of a text type detection method according to an embodiment of the present disclosure is shown;

[0014] Figure 2 A flowchart illustrating a method for selecting a target statement from a plurality of statements according to an embodiment of the present disclosure is shown;

[0015] Figure 3 A flowchart illustrating a method for determining keywords from multiple preset types of text detected in historical data according to an embodiment of the present disclosure is shown;

[0016] Figure 4 A flowchart illustrating a method for segmenting text to be detected according to an embodiment of the present disclosure is shown;

[0017] Figure 5 A structural block diagram of a text detection device of a preset type according to an embodiment of the present disclosure is shown;

[0018] Figure 6 A structural block diagram of a text detection device of a preset type according to another embodiment of the present disclosure is shown;

[0019] Figure 7 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0020] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0021] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.

[0022] The terminology used in the description of the various examples in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.

[0023] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings. Figure 1 A flowchart of a text type detection method 100 according to an embodiment of this disclosure is shown. This detection method 100 can be used with servers related to text processing. These servers may be servers for receiving user-uploaded documents, such as servers for cloud storage, servers for websites for publishing articles, etc. These servers may also be servers for providing text push services to relevant users, such as servers related to social media, advertising, or cloud providers, etc. The aforementioned preset text types include, but are not limited to, sports-related text, music-related text, etc.

[0024] like Figure 1 As shown, the method 100 includes:

[0025] Step 110: Perform a preliminary inspection of the text to be tested;

[0026] Step 120: In response to the text to be detected being preliminarily identified as a text of a pre-defined type, the text to be detected is segmented into sentences to obtain a set of sentences containing multiple sentences.

[0027] Step 130: For each statement in the statement set, determine the number of predetermined keywords contained in the statement, wherein the keywords are associated with a preset type;

[0028] Step 140: Select multiple target statements from multiple statements based on the number of keywords contained in the multiple statements in the statement set;

[0029] Step 150: Determine whether each of the multiple target statements is an identifier statement; and

[0030] Step 160: In response to the fact that the number of identification statements meets the preset conditions, determine that the text to be detected is text of a preset type.

[0031] The detection method of one or more embodiments of this disclosure segments the text to be detected into multiple sentences, selects a portion of the target sentences that are likely to be the identifying sentences from the multiple sentences for inspection, thereby avoiding detection of the entire text, simplifying the detection process, and improving the accuracy of detection.

[0032] The preset types mentioned above are text classification types, including but not limited to: music type, sports type, lifestyle type, travel type and military type, etc. The text content in text belonging to the preset type is related to that type. For example, text belonging to the music type mainly describes content related to music.

[0033] In step 110, a preliminary check is first performed on the text to be detected. Due to the massive amount of text information on the server, only a small portion of this text may be of the preset type, with the majority being text of a different type. Therefore, a preliminary check is performed before subsequent precise checks to filter out the vast majority of text that is clearly not of the preset type. In some embodiments, to improve the accuracy of the preliminary check, a large amount of text to be detected can be input into a text classification model for prediction, and the prediction results can be used to determine whether it is text of a suspected preset type. In other embodiments, other methods can be used for the preliminary check, such as extracting high-frequency words from the text and then determining whether these high-frequency words are domain-specific terms associated with the preset type. If the high-frequency words are domain-specific terms, the text is determined to be text of a suspected preset type.

[0034] In step 120, the text to be detected is segmented into sentences to facilitate accurate subsequent verification. In some embodiments, to improve the efficiency of sentence segmentation, the text to be detected can be segmented into sentences at preset character intervals. In other embodiments, to ensure the accuracy of sentence segmentation, punctuation marks and stop words in the text can be identified first, and then sentences can be segmented using punctuation marks or stop words as dividing points. The multiple sentences obtained through the above sentence segmentation will constitute a sentence set.

[0035] In step 130, the keyword indicates that a statement containing that keyword is likely a relevant identifying statement. Keywords are keywords detected in certain fields and associated with the aforementioned preset types. For example, when detecting text about music, keywords include, but are not limited to, words such as "song," "singer," and "performance"; for example, when detecting text about sports, keywords include, but are not limited to, words such as "football," "basketball," and "racing." These keywords can be predetermined, and a keyword table can be pre-stored in the server executing the detection method. During detection, this keyword table can be queried to obtain relevant keywords. Generally speaking, the more keywords a statement contains, the more likely it is to be an identifying statement.

[0036] In step 140, statements with a high number of keywords can be selected from the statement set as target statements, and these target statements will be the focus of subsequent detection. In some embodiments, statements containing more than a threshold number of keywords can be selected as target statements. In other embodiments, statements containing keywords that constitute a percentage of the total word count of the statement can be selected as target statements.

[0037] In step 150, each target sentence can be input into a sentence recognition model for prediction, and the prediction result of the model can be used to determine whether the target sentence is an identifier sentence. In other embodiments, the identification of a target sentence can also be determined in other ways, such as classifying the keywords appearing in the sentence. Words with a higher degree of relevance to the domain are assigned a higher level; for example, in sports-related text, the keyword "football" has a higher level than the keyword "fitness." Then, by comprehensively analyzing the levels of keywords in each sentence, the identification of the target sentence can be determined. For example, if a sentence contains three or more keywords with a preset level, it is determined to be an identifier sentence.

[0038] It is understood that if the text to be detected contains only a small number of identifier statements, it cannot constitute text of the preset type. Therefore, in step 160, it can be determined whether the text to be detected is text of the preset type by judging whether the number of identifier statements meets a preset condition. The preset condition can be that the proportion of identifier statements among multiple target statements exceeds a threshold proportion. If the proportion exceeds the threshold proportion, it indicates that the text contains a large number of identifier statements, and the text to be detected is determined to be text of the preset type. The threshold proportion can be 50%, 60%, 70%, etc. In some other embodiments, the judgment can also be made by the absolute number of identifier statements, and the preset condition is not limited here.

[0039] Figure 2 A flowchart of a method 200 for selecting a target statement from a plurality of statements, as shown in the fundamental disclosure embodiment, is provided. Figure 2 As shown, the method 200 includes:

[0040] Step 210: Sort the multiple statements according to the number of keywords they each contain, in descending order of the number of keywords; and

[0041] Step 220: Select multiple statements with a preset proportion that appear first in the sorting results as multiple target statements.

[0042] It is understood that the more keywords a target statement contains, the greater the likelihood that it is an identifying statement, and target statements ranking higher in the sorting results are more likely to be identifying statements. In this embodiment, sorting can accurately locate target statements, thereby improving the accuracy of subsequent text detection.

[0043] In some embodiments, each target sentence can be fed into a pre-trained ALBERT model, and the model's prediction results can be used to determine whether the target sentence is a flagged sentence. The pre-trained model can be one previously trained by developers to solve similar problems. When solving similar pre-defined text detection problems, it is not necessary to train a new model from scratch; a model trained on similar problems can be retrained. The pre-trained ALBERT model has already been trained on a large corpus of historical data, which can come from real internet data. Therefore, compared to other text classification models, the pre-trained ALBERT model has more accurate predictions and can accurately predict whether a target sentence is a flagged sentence.

[0044] In some embodiments, after acquiring the text to be detected, preprocessing operations can be performed on the text. Preprocessing operations include one or more of the following: text type conversion, replacing consecutive numbers with preset characters, and removing non-text symbols. In this embodiment, preprocessing operations can remove interference from numbers and non-text characters, improving the accuracy of subsequent detection.

[0045] In some embodiments, preliminary verification of the text to be detected includes: inputting the text to be detected into a text classification model, and determining whether the text to be detected belongs to a suspected preset type based on the prediction results of the text classification model. The aforementioned text classification model can be, for example, the fastText model. fastText is a word vector calculation and text classification tool that determines the category of the text to be detected based on the Euclidean distance between the word vectors of identified high-frequency words in the text in the vector space. Using a text classification model such as the fastText model for preliminary verification allows the fastText model to predict the entire text without requiring sentence segmentation. The fastText model predicts faster, but its accuracy is lower; it can be used for preliminary verification to improve text detection efficiency. As mentioned above, due to the massive amount of text information on the server, and the fact that the vast majority of this text information is not of a preset type, a preliminary verification is performed before subsequent precise verification to filter out the vast majority of text that clearly does not belong to the preset type. Subsequent detection operations are only performed if the fastText model predicts that the text is a suspected preset type; if the fastText model predicts that the text is not a suspected preset type, subsequent operations such as sentence segmentation and target sentence selection are not performed.

[0046] The training process of the above text classification model includes: training the text classification model using both positive samples and negative samples at the same time. Among them, the sample input of the positive sample is text suspected of a preset type containing at least one keyword, and the sample input of the negative sample is text not suspected of a preset type containing at least one keyword. In this embodiment, the text classification model can be trained using both positive and negative samples at the same time, so as to accurately adjust the parameters of the model and improve the accuracy of subsequent model prediction.

[0047] In some embodiments, the process of determining keywords includes: determining multiple keywords from multiple texts of preset types detected from historical data. Using the texts of preset types that have been determined in historical data to determine keywords makes the types of determined keywords more accurate and complete. The multiple texts of preset types detected from historical data can be the same texts of preset types previously detected by the server for detecting the text to be detected. In addition, the texts of preset types in these historical data can also be texts of preset types marked manually. For example, when detecting texts of the music type, the historical data can be other texts stored in the relevant server that have been determined to be of the music type.

[0048] Figure 3 The flowchart of method 300 for determining keywords from multiple texts of preset types detected from historical data according to an embodiment of the present disclosure is shown, as Figure 3 shown, the method 300 includes:

[0049] Step 310, segmenting the multiple texts of preset types to obtain multiple candidate keywords;

[0050] Step 320, determining the inverse document frequency of each candidate keyword among the multiple candidate keywords, where the inverse document frequency is determined according to the number of multiple texts of preset types and the number of texts of preset types containing the candidate keyword among the multiple texts of preset types; and

[0051] Step 330, determining multiple keywords from the multiple candidate keywords according to the inverse document frequencies of the multiple candidate keywords.

[0052] In step 310, words or phrases in the texts of preset types are mined using cohesion, mutual information, etc. First, the texts of preset types are segmented according to n-gram (for example, when it is 2-gram, that is, every two characters in the original sentence are combined). For example, "hello world" can be segmented into "hello, good world, world". The cohesion (i.e., point-wise mutual information) can be used as the basis for segmentation. If the cohesion is less than a certain threshold, it is separated to obtain candidate keywords. The calculation formula of cohesion is as follows:

[0053]

[0054] Where p(x,y) represents the frequency of simultaneous occurrence of characters x and y in the text, and p(x) and p(y) represent the individual frequencies of occurrence of characters x and y in the text. In other embodiments, word segmentation tools such as HMM (Hidden Markov Model) can also be used to segment sentences in historically annotated text of a preset type.

[0055] In step 320, the inverse text frequency (IDF) of each candidate keyword obtained in step 310 is calculated. Inverse text frequency represents the frequency of a candidate keyword appearing in text of a preset type. The fewer texts of a preset type containing candidate keyword t, the larger the IDF, indicating that candidate keyword t has good category discrimination ability. For example, in the field of detecting sports-related text, candidate keyword t is a keyword that distinguishes sports-related content. If the number of texts containing candidate keyword t in a certain type of text is m, and the number of texts containing t in other types of text is k, then the total number of texts containing t is obviously n = m + k. When m is large, n is also large, and the IDF value obtained according to the IDF formula will be relatively small, indicating that candidate keyword t has weak category discrimination ability and is not a keyword used for detecting text of a preset type. The formula for calculating inverse text frequency is as follows:

[0056]

[0057] Where, n d This represents the total number of texts of the preset type, and df(d,t) represents the number of texts containing the candidate keyword.

[0058] In step 330, candidate keywords with inverse text frequencies greater than a frequency threshold can be selected from multiple candidate keywords to form the final set of keywords. These keywords are then used to select target sentences from the text to be detected, as described in step 140 of method 100. Before performing step 330, the candidate keywords selected in step 310 can also be manually screened to remove obvious non-keywords (such as words with high IDF values ​​like "walking," "breathing," and "exercise" in sports articles).

[0059] In this embodiment, target keywords can be determined from multiple candidate keywords using inverse text frequency. These determined target keywords have excellent category discrimination capabilities. For example, in the field of detecting sports-related text, these determined keywords have a good function of distinguishing whether a statement is a sports-related identifier. The method in this embodiment improves the accuracy of keyword determination.

[0060] Figure 4A flowchart of a method 400 for segmenting text to be detected according to an embodiment of the present disclosure is shown, such as... Figure 4 As shown, the method 400 includes:

[0061] Step 410: The text to be detected is segmented into sentences at preset intervals of a certain number of characters to obtain multiple sentences; and

[0062] Step 420: Delete incomplete statements from multiple statements.

[0063] In step 410, the character count can be obtained empirically; for example, after experimentation, the character count can be set to 256. In step 420, the short text after sentence segmentation is preprocessed, which requires removing incomplete sentences from multiple sentences. Specifically, the rules for removing incomplete sentences may include:

[0064] a) If there is a punctuation mark at the beginning or end of the split statement (less than 5 characters), delete the punctuation mark and the characters before or after it to ensure that the split statement contains as much complete meaning as possible.

[0065] b) If the number of non-Chinese characters in the segmented sentence is less than the threshold, then delete it. For example, if the number of non-Chinese characters is less than 10, then discard the sentence to avoid interference from non-textual information such as phone numbers.

[0066] Figure 5 A structural block diagram of a text detection device 500 of a preset type according to an embodiment of the present disclosure is shown. Figure 5 As shown, the device 500 includes: an inspection unit 510 configured to perform a preliminary inspection on the text to be inspected; a sentence segmentation unit 520 configured to segment the text to be inspected into sentences in response to the preliminary inspection finding the text to be inspected to be suspected of being of a preset type, thereby obtaining a sentence set containing multiple sentences; a first determination unit 530 configured to determine, for each sentence in the sentence set, the number of predetermined keywords contained in that sentence, wherein the keywords are associated with a preset type; a selection unit 540 configured to select multiple target sentences from the multiple sentences based on the number of keywords contained in the multiple sentences in the sentence set; a second determination unit 550 configured to determine whether each of the multiple target sentences is an identifier sentence; and a third determination unit 560 configured to determine that the text to be inspected is text of a preset type in response to the determination that the number of identifier sentences meets a preset condition.

[0067] Figure 6 A structural block diagram of a text detection device 600 of a preset type according to another embodiment of the present disclosure is shown. Figure 6As shown, in some embodiments, the selection unit 640 includes: a sorting module 641, configured to sort multiple statements in descending order of the number of keywords contained in each statement; and a first determining module 642, configured to determine multiple statements with a preset proportion of the sorted results as multiple target statements.

[0068] In some embodiments, the second determining unit 650 is further configured to: input each target statement into a pre-trained ALBERT model, and determine whether the target statement is an identifier statement based on the model's prediction results.

[0069] In some embodiments, the above-described apparatus 600 further includes a preprocessing unit 670 configured to perform preprocessing operations on the text to be detected, wherein the preprocessing operations include one or more of the following operations: text type conversion, replacing consecutive numbers with preset characters, and removing non-text symbols.

[0070] In some embodiments, the verification unit 610 is further configured to: input the text to be detected into a text classification model, and determine whether the text to be detected is a text of a suspected preset type based on the prediction result of the text classification model.

[0071] In some embodiments, the above-described apparatus 600 further includes a keyword determination unit 680, configured to determine multiple keywords from multiple preset types of text detected in historical data.

[0072] In some embodiments, the keyword determination unit 680 includes: a word segmentation module 681 configured to segment multiple preset types of text to obtain multiple candidate keywords; a second determination module 682 configured to determine the inverse text frequency of each candidate keyword among the multiple candidate keywords, wherein the inverse text frequency is determined based on the number of multiple preset types of text and the number of preset types of text containing the candidate keyword; and a third determination module 683 configured to determine multiple keywords from the multiple candidate keywords based on the inverse text frequency of the multiple candidate keywords.

[0073] It should be understood that Figure 5 Each unit of the device 500 shown can be connected to a reference. Figure 1 The steps in the described method 100 correspond to each other. Figure 6 The various units and modules of the device 600 shown can be compared with the reference. Figures 2 to 4 The steps described in methods 200-400 correspond to each other. Therefore, the operations, features, and advantages described above for methods 100-400 also apply to apparatus 500, apparatus 600, and their constituent units and modules. For the sake of brevity, some operations, features, and advantages will not be repeated here.

[0074] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0075] According to embodiments of this disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.

[0076] refer to Figure 7 The present invention describes a structural block diagram of an electronic device 700 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0077] like Figure 7 As shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. The RAM 703 may also store various programs and data required for the operation of the electronic device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0078] Multiple components in electronic device 700 are connected to I / O interface 705, including: input unit 706, output unit 707, storage unit 708, and communication unit 709. Input unit 706 can be any type of device capable of inputting information to electronic device 700. Input unit 706 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 707 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 708 may include, but is not limited to, hard disk and optical disk. Communication unit 709 allows electronic device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0079] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as text type detection methods. For example, in some embodiments, the text type detection method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the text type detection method described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform text type detection methods by any other suitable means (e.g., by means of firmware).

[0080] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0081] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0082] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0083] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0084] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.

[0085] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0086] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0087] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.

Claims

1. A method for detecting text types, comprising: The preliminary examination of the text to be detected includes: inputting the text to be detected into a text classification model, and determining whether the text to be detected is a text of a suspected preset type based on the prediction results of the text classification model. The text classification model is the fastText model. In response to the text to be detected being preliminarily identified as text of the preset type, subsequent detection operations are performed on the text to be detected, including: The text to be detected is segmented into sentences to obtain a set of sentences containing multiple sentences, wherein segmenting the text to be detected into sentences includes: The text to be detected is segmented into sentences at preset intervals of a certain number of characters to obtain multiple sentences; and Delete incomplete statements from the given statements; For each statement in the statement set, determine the number of predetermined keywords contained in that statement, wherein the keywords are associated with the preset type; Based on the number of keywords contained in the multiple statements in the statement set, select multiple target statements from the multiple statements, including: Based on the number of keywords contained in each of the multiple statements, the statements are sorted in descending order of the number of keywords; and The multiple statements that are ranked first in the sorting results according to a predetermined proportion are identified as the multiple target statements. Determining whether each of the plurality of target statements is an identifier statement includes: inputting each target statement into a pre-trained ALBERT model, and determining whether the target statement is an identifier statement based on the model's prediction result; and In response to determining that the number of the identified statements among the plurality of target statements meets a preset condition, the text to be detected is determined to be text of the preset type.

2. The method according to claim 1, further comprising: The text to be detected is preprocessed, wherein the preprocessing operation includes one or more of the following operations: text type conversion, replacing consecutive numbers with preset characters, and removing non-text symbols.

3. The method according to claim 1, wherein, The training process of the text classification model includes: The text classification model is trained using both positive and negative samples. The input of the positive samples is text of the pre-defined type that contains at least one keyword, and the input of the negative samples is text of the non-pre-defined type that contains at least one keyword.

4. The method according to any one of claims 1-3, wherein, The process of determining the keywords includes: Multiple keywords are identified from multiple preset types of text detected from historical data.

5. The method according to claim 4, wherein, The determination of multiple keywords from multiple preset types of text detected from historical data includes: The multiple preset text types are segmented to obtain multiple candidate keywords; Determine the inverse text frequency (IMR) of each candidate keyword from the plurality of candidate keywords, wherein the IMR is determined based on the number of texts of the plurality of preset types and the number of texts of the preset type containing the candidate keyword; and The multiple keywords are determined from the multiple candidate keywords based on their inverse text frequencies.

6. A text type detection device, comprising: The inspection unit is configured to perform a preliminary inspection on the text to be detected. The inspection unit is further configured to: input the text to be detected into a text classification model, and determine whether the text to be detected is a text of a suspected preset type based on the prediction result of the text classification model. The text classification model is the fastText model. The detection device is configured to perform subsequent detection operations on the text to be detected in response to the text being initially identified as potentially belonging to the preset type, and the detection device further includes: The sentence segmentation unit is configured to segment the text to be detected into sentences to obtain a sentence set containing multiple sentences, wherein the segmentation of the text to be detected into sentences includes: The text to be detected is segmented into sentences at preset intervals of a certain number of characters to obtain multiple sentences; and Delete incomplete statements from the given statements; The first determining unit is configured to determine, for each statement in the statement set, the number of predetermined keywords contained in that statement, wherein the keywords are associated with the predetermined type; A selection unit is configured to select multiple target statements from the multiple statements in the statement set based on the number of keywords contained therein. The selection unit includes: The sorting module is configured to sort the multiple statements in descending order of the number of keywords contained in each statement; and The first determining module is configured to determine a plurality of statements that are ranked first in the sorting results according to a preset proportion as the plurality of target statements. The second determining unit is configured to determine whether each of the plurality of target statements is an identifier statement, wherein the second determining unit is further configured to: input each target statement into a pre-trained ALBERT model, and determine whether the target statement is the identifier statement based on the model's prediction result; and The third determining unit is configured to determine the text to be detected as text of the preset type in response to determining that the number of the identifier statements among the plurality of target statements meets a preset condition.

7. The apparatus according to claim 6, further comprising: The preprocessing unit is configured to perform preprocessing operations on the text to be detected, wherein the preprocessing operations include one or more of the following operations: text type conversion, replacing consecutive numbers with preset characters, and removing non-text symbols.

8. The apparatus according to claim 6 or 7, further comprising: The keyword identification unit is configured to identify multiple keywords from multiple preset types of text detected in historical data.

9. The apparatus according to claim 8, wherein, The keyword determination unit includes: The word segmentation module is configured to segment text of the multiple preset types to obtain multiple candidate keywords; The second determining module is configured to determine the inverse text frequency of each candidate keyword among the plurality of candidate keywords, wherein the inverse text frequency is determined based on the number of texts of the plurality of preset types and the number of texts of the preset type containing the candidate keyword; and The third determining module is configured to determine the multiple keywords from the multiple candidate keywords based on the inverse text frequency of the multiple candidate keywords.

10. An electronic device, comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.

11. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-5.

12. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the method of any one of claims 1-5.

Citation Information

Patent Citations

  • Scientific and technological achievement classification method and device, equipment and medium

    CN111177372A

  • Text core content extraction method and device

    CN111767393A