Statement extraction method, apparatus, electronic device, and computer-readable storage medium

By extracting core statements first by chunking and then filtering, the problem that smart devices cannot accurately understand user intentions is solved, and high-quality service provision is achieved, which is highly universal.

CN114298030BActive Publication Date: 2025-06-13CLOUDMINDS SHANGHAI ROBOTICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111528935.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-14
Publication Date
2025-06-13
Estimated Expiration
2041-12-14

AI Technical Summary

Technical Problem

Due to factors such as external noise and VAD truncation, the text information received by the natural language understanding module is messy and cannot accurately understand the user's true intentions, resulting in the services provided by the smart device being unable to meet the user's actual needs.

Method used

The core statement is extracted by blocking first and then filtering, and the text information to be processed is divided into chunking through the pre-trained statement blocking model to obtain several candidate statements, and then deduplication and denoising filtering are performed according to the preset filtering rules, and the target statement representing the user's true intention is output.

Benefits of technology

Scientifically, reasonably and accurately extract core statements that can represent the user's true intentions, improving the accuracy of smart devices in understanding user's intentions, thereby providing high-quality services that meet users' actual needs, which is highly universal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114298030B_ABST
    Figure CN114298030B_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to the technical field of natural language processing, and disclose a sentence extraction method, apparatus, electronic device, and computer-readable storage medium. The sentence extraction method includes: obtaining text information to be processed; wherein the text information to be processed includes text information generated by converting obtained speech information; inputting the text information to be processed into a pre-trained sentence chunking model to obtain a number of candidate sentences; filtering the number of candidate sentences according to a preset filtering rule to obtain and output a target sentence; wherein the target sentence is the sentence remaining after filtering among the number of candidate sentences. The sentence extraction method provided by the embodiments of the present application can scientifically, reasonably, and accurately extract the core sentence that can represent the true intention of the user from the text information, so as to provide high-quality services that meet the actual needs of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of natural language processing, and in particular, to a method and apparatus for extracting sentences, an electronic device, and a computer-readable storage medium. Background Art

[0002] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It is a technology for realizing effective communication between humans and computers through natural language, and is a science integrating linguistics, computer science, and mathematics. The research content of natural language processing includes, but is not limited to, text classification, information extraction, automatic summarization, intelligent dialogue, topic recommendation, machine translation, subject term recognition, knowledge base construction, deep text representation, named entity recognition, text generation, etc. The rise of natural language processing technology has greatly changed people's lives and led people into the era of intelligent life.

[0003] In people's daily intelligent life, people often interact with various intelligent devices in the form of conversations. The words spoken by the user are first sent to the speech recognition module of the intelligent device to be converted into text form, and then input to the natural language understanding module for understanding and analysis, so that the intelligent device can understand the user's intention and provide corresponding services to the user. Therefore, the result recognized by the speech recognition module determines the quality of understanding of the natural language understanding module.

[0004] However, in actual use, due to the influence of external noise and factors such as Voice Activity Detect (VAD) truncation, the text information received by the natural language understanding module is very messy and unsatisfactory, that is, the text information output by the speech recognition module may not necessarily represent the true intention of the user, which makes the natural language understanding module unable to accurately understand the true intention of the user, and ultimately results in the services provided by the intelligent device not meeting the actual needs of the user, bringing a bad user experience to the user. Summary of the Invention

[0005] The purpose of the embodiments of the present application is to provide a method and apparatus for extracting sentences, an electronic device, and a computer-readable storage medium, which can scientifically, reasonably, and accurately extract the core sentences that can represent the true intention of the user from the text information, so as to provide high-quality services that meet the actual needs of the user.

[0006] To solve the above technical problems, an embodiment of the present application provides a sentence extraction method, including the following steps: obtaining text information to be processed; wherein, the text information to be processed includes text information converted from the obtained voice information; inputting the text information to be processed into a pre-trained sentence chunking model to obtain a number of candidate sentences; filtering the number of candidate sentences according to a preset filtering rule to obtain and output a target sentence; wherein, the target sentence is the sentence remaining after filtering among the number of candidate sentences.

[0007] An embodiment of the present application also provides a sentence extraction device, including: an acquisition module, configured to obtain text information to be processed; wherein, the text information to be processed includes text information converted from the obtained voice information; a chunking module, configured to input the text information to be processed into a pre-trained sentence chunking model to obtain a number of candidate sentences; a filtering module, configured to filter the number of candidate sentences according to a preset filtering rule to obtain and output a target sentence; wherein, the target sentence is the sentence remaining after filtering among the number of candidate sentences.

[0008] An embodiment of the present application also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the above sentence extraction method.

[0009] An embodiment of the present application also provides a computer-readable storage medium, storing a computer program, and when the computer program is executed by a processor, the above sentence extraction method is implemented.

[0010] The sentence extraction method, device, electronic device, and computer-readable storage medium provided by the embodiments of the present application. The server first obtains the text information to be processed including the text information generated by converting the obtained voice information. Subsequently, the obtained text information to be processed is input into a pre-trained sentence chunking model for chunking to obtain several candidate sentences. Then, the several candidate sentences are filtered according to a preset filtering rule, and the sentences remaining after filtering among the several candidate sentences are output as target sentences. Considering that whether it is rewriting the user's question, extracting by regular expressions, or the technical solution of downstream task generalization, the extracted core sentences are not accurate and reasonable enough, and may not necessarily represent the true intention of the user. Moreover, these solutions may only be applicable in some specific scenarios, with very poor generalization ability and no universality. However, the embodiments of the present application adopt the method of first chunking and then filtering to extract the core sentences, which can scientifically, reasonably, and accurately extract the core sentences that can represent the true intention of the user from the text information to be processed, so that the intelligent device can perform various downstream tasks according to the extracted core sentences, thereby providing high-quality services that meet the actual needs of the user. At the same time, the embodiments of the present application can adapt to various application scenarios by adjusting the filtering rule, and have strong universality.

[0011] In addition, the filtering rule includes a duplicate removal filtering rule. The step of filtering the several candidate sentences according to the preset filtering rule to obtain and output the target sentence includes: detecting whether there are multiple candidate sentences that are exactly the same among the several candidate sentences; if there are multiple candidate sentences that are exactly the same among the several candidate sentences, only retain any one of the multiple candidate sentences that are exactly the same; output the retained candidate sentence as the target sentence. Considering that the user may say repeated words in some cases. For example, the user thinks that what he said the first time was not clear, and then immediately repeats this sentence in a louder voice. The intentions of these two sentences are the same. If not filtered, the intelligent device is very likely to execute the same task twice according to these two sentences, which is not what the user expects. Therefore, this embodiment can remove duplicates from the repeated words, that is, multiple candidate sentences that are exactly the same, in the text information to be processed, further meeting the actual needs of the user.

[0012] In addition, the filtering rules include denoising filtering rules. Filtering the plurality of candidate statements according to the preset filtering rules to obtain and output a target statement includes: inputting the plurality of candidate statements into a preset perplexity calculation model to obtain the perplexity of each candidate statement; respectively determining whether the perplexity of each candidate statement is less than a preset threshold, and retaining the candidate statements whose perplexity is less than the preset threshold; outputting the retained candidate statements as the target statement. Due to the existence of factors such as environmental noise, the text information to be processed is very likely to include unfinished sentences, and intelligent devices cannot understand unfinished sentences. Perplexity can well measure whether a candidate statement is a complete sentence. In the embodiments of the present application, only candidate statements with a perplexity less than the preset threshold are retained to ensure that the extracted statements are complete and can clearly represent the user's intention.

[0013] In addition, after respectively determining whether the perplexity of each candidate statement is less than the preset threshold, it includes: if the perplexity of each candidate statement is greater than or equal to the preset threshold, generating a re-acquisition instruction, where the re-acquisition instruction is used to re-acquire the text information to be processed. When the perplexity of each candidate statement is greater than or equal to the preset threshold, it indicates that the noise has caused a very serious impact on the text information to be processed and cannot be used. The server generates a re-acquisition instruction to re-acquire the text information to be processed until there are candidate statements with a perplexity less than the preset threshold.

[0014] In addition, if there are a plurality of the target statements, the outputting the target statement includes: determining the positions of the plurality of target statements in the text information to be processed; outputting the target statement with the last position. Generally speaking, the last sentence spoken by the user can best represent the user's current intention. In the embodiments of the present application, when there are multiple target statements left after filtering, only the target statement with the last position is output, which is more in line with the user's true intention and provides high-quality services that best meet the actual needs of the user.

[0015] In addition, the pre-trained sentence chunking model is trained through the following steps: obtaining training samples; where each of the training samples contains a plurality of short sentences; labeling the first word of each short sentence in the training sample with a first label, and labeling the other words except the first word of each short sentence with a second label; iteratively training the sentence chunking model according to the training samples, the first label, and the second label, which can enable the sentence chunking model to obtain accurate sentence chunking capabilities.

[0016] In addition, the training samples include first-class training samples, second-class training samples, and third-class training samples. Obtaining the training samples includes: randomly splicing n short sentences among a number of obtained short sentences to obtain the first-class training samples, where n is an integer greater than 1; randomly generating a number of strings and splicing the strings at both ends of any one of the obtained short sentences to obtain the second-class training samples; obtaining third-class training samples from the target text, where the target text includes at least novels, logs, and news. By obtaining different types of training samples in various ways to train the sentence chunking model, the sentence chunking ability of the model can be effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings, and these exemplary illustrations do not limit the embodiments.

[0018] Figure 1 is a flowchart of a sentence extraction method according to an embodiment of the present application;

[0019] Figure 2 is a process of filtering a number of candidate sentences according to a preset filtering rule to obtain and output a target sentence in an embodiment of the present application Figure 1 ;

[0020] Figure 3 is a process of filtering a number of candidate sentences according to a preset filtering rule to obtain and output a target sentence in another embodiment of the present application Figure 2 ;

[0021] Figure 4 is a flowchart of training a sentence chunking model in an embodiment of the present application;

[0022] Figure 5 is a schematic diagram of a sentence extraction device according to another embodiment of the present application;

[0023] Figure 6 is a schematic diagram of the structure of an electronic device according to another embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will elaborate on each embodiment of this application with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of this application, many technical details are presented to help readers better understand this application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in this application can still be implemented. The division of the following embodiments is for convenience of description and should not impose any limitation on the specific implementation of this application. The various embodiments can be combined and cross-referenced with each other on the premise of no conflict.

[0025] In people's daily intelligent life, people often interact with various intelligent devices in the form of conversations. The words spoken by the user are first sent to the speech recognition module of the intelligent device to be converted into text form, and then input into the natural language understanding module for understanding and analysis, enabling the intelligent device to understand the user's intention and provide corresponding services to the user. However, due to the influence of external noise and VAD truncation and other factors, the text information received by the natural language understanding module is very messy and unsatisfactory. It is necessary to accurately extract the core sentences representing the user's true intention from the text information. Only based on the core sentences can the services required by the user be provided. Therefore, it is very important to accurately extract the core sentences from the text information. Rewriting the user's question to rewrite the core sentence, extracting the core sentence based on regular expressions, and extracting the core sentence based on the generalization of downstream tasks are common methods for extracting core sentences from the text.

[0026] Rewriting the user's question means modifying the second sentence according to the first sentence of two consecutive sentences, such as removing the first or last word of the second sentence, etc. However, rewriting the user's question is likely to rewrite incorrect questions and change the user's true intention; the generalization ability of extracting core sentences based on regular expressions is very poor and can only be used in some specific scenarios, and only some problems can be solved; when extracting core sentences based on the generalization of downstream tasks, it is necessary to generalize according to the downstream tasks. To be applicable to different downstream tasks, different generalizations are required, and the cost is too high.

[0027] To solve the technical problems that the core sentences extracted by the above sentence extraction methods are not accurate and reasonable enough, may not necessarily represent the user's true intention, and these solutions may only be applicable in some specific scenarios, have very poor generalization ability, and lack universality, an embodiment of this application provides a sentence extraction method, which is applied to an electronic device. Herein, the electronic device can be a terminal or a server. In this embodiment and the following embodiments, the electronic device is taken as an example of a server for illustration. The following will specifically describe the implementation details of the sentence extraction method of this embodiment. The following content is only implementation details provided for convenience of understanding and is not necessary for implementing this solution.

[0028] The flowchart of the statement extraction method in this embodiment can be as Figure 1 shown, including:

[0029] Step 101, obtain the text information to be processed.

[0030] Specifically, when performing statement extraction, the server first obtains the text information to be processed. The text information to be processed obtained by the server includes the text information generated by converting the obtained voice information.

[0031] In one example, the text information to be processed obtained by the server can also be the text information input by the user.

[0032] In a specific implementation, the server can monitor the sounds in the surrounding environment in real time. When detecting the user's voice, the server can obtain the voice information spoken by the user and convert the voice information into text form to obtain the text information to be processed.

[0033] Step 102, input the text information to be processed into a pre-trained statement chunking model to obtain several candidate statements.

[0034] In a specific implementation, after obtaining the text information to be processed, the server can input the text information to be processed into the pre-trained statement chunking model to obtain several candidate statements output by the statement chunking model. Among them, the statement chunking model can split a long sentence into several short sentences, and there is at most one subject, one predicate, and one object in each short sentence.

[0035] In one example, the text information to be processed obtained by the server is: "Please keep a safe distance from me. What's the weather like today?", and the server inputs the text information to be processed into the statement chunking model to obtain candidate statements "Please keep a safe distance from me" and "What's the weather like today?".

[0036] In one example, the text information to be processed obtained by the server is: "Can you dance? Can you dance?", and the server inputs the text information to be processed into the statement chunking model to obtain candidate statements "Can you dance?" and "Can you dance?".

[0037] In one example, the text information to be processed obtained by the server is: "Nearby, where can I go to check my luggage?", and the server inputs the text information to be processed into the statement chunking model to obtain candidate statements "Nearby" and "Where can I go to check my luggage nearby?".

[0038] Step 103, filter the several candidate statements according to the preset filtering rules to obtain and output the target statements.

[0039] In a specific implementation, after the server obtains a number of candidate statements output by the statement chunking model, it can filter the number of candidate statements according to a preset filtering rule, and output the statements remaining after filtering among the number of candidate statements as target statements. Among them, the preset filtering rule can be set by those skilled in the art according to actual needs.

[0040] In one example, if the number of target statements obtained by the server is several, the server can first determine the positions of the several target statements in the text information to be processed, and then output the target statement with the last position. Generally speaking, the last sentence spoken by the user best represents the user's current intention. In the embodiments of the present application, when there are multiple target statements remaining after filtering, only the last target statement is output, which is more in line with the user's true intention and provides high-quality services that best meet the actual needs of the user.

[0041] For example: The text information to be processed obtained by the server is: "Please keep a safe distance from me. How's the weather today?" The target statements obtained by the server are "Please keep a safe distance from me" and "How's the weather today?" Among them, the target statement "How's the weather today" is in the last position in the text information to be processed, and the server then outputs "How's the weather today".

[0042] In another example, if the number of target statements obtained by the server is several, the server can first determine the positions of the several target statements in the text information to be processed, and then output the target statement with the first position.

[0043] In another example, if the number of target statements obtained by the server is several, the server can first determine the positions of the several target statements in the text information to be processed, and then output a target statement at a specified position.

[0044] In this embodiment, compared with technical solutions such as rewriting the core statement by rewriting the user's question, extracting the core statement based on regular expressions, and extracting the core statement based on downstream task generalization, in the embodiment of the present application, the server first obtains the text information to be processed including the text information generated by converting the obtained voice information, and then inputs the obtained text information to be processed into a pre-trained statement chunking model for chunking to obtain a number of candidate statements, and then filters the number of candidate statements according to a preset filtering rule, and outputs the statements remaining after filtering in the number of candidate statements as the target statement. Considering that whether it is rewriting the user's question, extracting by regular expressions, or the technical solution of downstream task generalization, the extracted core statements are not accurate and reasonable enough, and may not necessarily represent the true intention of the user, and these solutions may only be applicable in some specific scenarios, with very poor generalization ability and no universality. However, the embodiment of the present application extracts the core statement by first chunking and then filtering, which can scientifically, reasonably, and accurately extract the core statement that can represent the true intention of the user from the text information to be processed, so that the intelligent device can perform various downstream tasks according to the extracted core statement, thereby providing high-quality services that meet the actual needs of the user. At the same time, the embodiment of the present application can adapt to various application scenarios by adjusting the filtering rule, and has strong universality.

[0045] In one embodiment, the preset filtering rule includes a duplicate removal filtering rule. The server filters the number of candidate statements according to the preset filtering rule to obtain and output the target statement, which can be implemented through the following steps as shown in Figure 2 : Specifically including:

[0046] Step 201, detect whether there are multiple candidate statements that are exactly the same among the number of candidate statements.

[0047] In a specific implementation, the filtering performed by the server on the number of candidate statements includes duplicate removal filtering. The server can detect whether there are multiple candidate statements that are exactly the same among the number of candidate statements, and the exactly the same candidate statements are the statements that appear repeatedly.

[0048] In an example, the text information to be processed obtained by the server is "Can you dance? Can you dance? What's the weather like today?", and the candidate statements obtained by the server include candidate statement A "Can you dance?", candidate statement B "Can you dance?", and candidate statement C "What's the weather like today?". Among them, the server detects that candidate statement A and candidate statement B are exactly the same, and candidate statement A and candidate statement B are the repeated statements.

[0049] Step 202, if there are multiple candidate statements that are exactly the same among the number of candidate statements, then only retain any one of the multiple candidate statements that are exactly the same.

[0050] Step 203: Output the retained candidate statement as the target statement.

[0051] In a specific implementation, if the server detects that there are multiple identical candidate statements among several candidate statements, the server can only retain any one of the multiple identical candidate statements and discard the other candidate statements among the multiple identical candidate statements. Finally, only the retained candidate statement is output as the target statement.

[0052] In an example, the candidate statements obtained by the server include candidate statement A "Can you dance?", candidate statement B "Can you dance?", and candidate statement C "What's the weather like today?". The server detects that candidate statement A and candidate statement B are identical, retains only candidate statement A, and discards candidate statement B. The server detects that there is no candidate statement identical to candidate statement C, and then retains candidate statement C. Finally, the output target statements include "Can you dance?" and "What's the weather like today?".

[0053] In this embodiment, the filtering rule includes a deduplication filtering rule. Filtering the several candidate statements according to the preset filtering rule to obtain and output the target statement includes: detecting whether there are multiple identical candidate statements among the several candidate statements; if there are multiple identical candidate statements among the several candidate statements, only retain any one of the multiple identical candidate statements; output the retained candidate statement as the target statement. Considering that the user may say repeated words in some cases, for example, the user thinks that what he said for the first time was not clear, and then immediately repeated this sentence in a louder voice. The intentions of these two sentences are the same. If not filtered, the intelligent device is very likely to execute the repeated task twice according to these two sentences, which is not what the user expects. Therefore, this embodiment can deduplicate the repeated words, that is, multiple identical candidate statements, in the text information to be processed, further meeting the actual needs of the user.

[0054] In another embodiment, the preset filtering rule includes a denoising filtering rule. The server filters several candidate statements according to the preset filtering rule to obtain and output the target statement, which can be implemented through the following steps: Figure 3 as shown below, specifically including:

[0055] Step 301: Input the several candidate statements into a preset perplexity calculation model to obtain the perplexity of each candidate statement.

[0056] In a specific implementation, the filtering performed by the server on a number of candidate statements includes denoising filtering. The server can input the number of candidate statements into a preset perplexity calculation model to obtain the perplexity of each candidate statement output by the perplexity calculation model. The perplexity calculation model is used to calculate the perplexity of the candidate statement. The higher the perplexity of the candidate statement, the more abnormal the candidate statement is, and it may be a statement with noise or an incorrect statement.

[0057] Step 302: Determine whether the perplexity of each candidate statement is less than a preset threshold, and retain the candidate statements whose perplexity is less than the preset threshold.

[0058] Step 303: Output the retained candidate statements as target statements.

[0059] In a specific implementation, after the server calculates the perplexity of each candidate statement through the perplexity calculation model, it can determine whether the perplexity of each candidate statement is less than a preset threshold, retain only the candidate statements whose perplexity is less than the preset threshold, and discard the candidate statements whose perplexity is greater than or equal to the preset threshold, that is, filter out the statements with noise and incorrect statements, and finally output the retained candidate statements as target statements. Among them, the preset threshold can be set by those skilled in the art according to actual needs, and the embodiments of the present application do not make specific limitations on this.

[0060] In an example, if the server detects that the perplexity of each candidate statement is greater than or equal to the preset threshold, it can generate a re-acquisition instruction to re-acquire the text information to be processed. Considering that when the perplexity of each candidate statement is greater than or equal to the preset threshold, it means that the noise has a very serious impact on the text information to be processed and cannot be used. The server generates a re-acquisition instruction to re-acquire the text information to be processed until there are candidate statements with a perplexity less than the preset threshold.

[0061] In this embodiment, the filtering rule includes a denoising filtering rule. Filtering the number of candidate statements according to the preset filtering rule to obtain and output target statements includes: inputting the number of candidate statements into a preset perplexity calculation model to obtain the perplexity of each candidate statement; respectively determining whether the perplexity of each candidate statement is less than a preset threshold, and retaining the candidate statements whose perplexity is less than the preset threshold; outputting the retained candidate statements as target statements. Due to the existence of factors such as environmental noise, the text information to be processed is very likely to include unfinished sentences, and intelligent devices cannot understand unfinished sentences. Perplexity can well measure whether a candidate statement is a complete sentence. The embodiments of the present application only retain candidate statements with a perplexity less than the preset threshold to ensure that the extracted statements are complete and can clearly represent the user's intention.

[0062] In one embodiment, the server can, for exampleFigure 4 Each of the steps shown, for the sentence chunking model, specifically includes:

[0063] Step 401: Obtain training samples, where each training sample contains several short sentences.

[0064] In one example, the training samples include the first type of training samples, the second type of training samples, and the third type of training samples. When the server obtains the training samples, it can first obtain several short sentences, randomly splice n short sentences among the obtained several short sentences to obtain the first type of training samples, where n is an integer greater than 1; the server can also randomly generate several strings and splice the several strings at both ends of any one of the obtained several short sentences to obtain the second type of training samples; the server can also obtain the third type of training samples from target texts such as novels, logs, and news. By obtaining different types of training samples in various ways to train the sentence chunking model, the sentence chunking ability of the model can be effectively improved.

[0065] Step 402: Label the first character of each short sentence in the training sample with a first label, and label the other characters of each short sentence except the first character with a second label.

[0066] In one example, the server can use the sequence annotation BI annotation to annotate the training samples, label the first character of each short sentence in the training sample with a first label, that is, "[B_sen]", and label the other characters of each short sentence except the first character with a second label, that is, "[I_sen]".

[0067] For example, the training sample obtained by the server is "What's the weather like tomorrow? Can you dance?", and this training sample contains two short sentences, namely "What's the weather like tomorrow?" and "Can you dance?". The server labels the first labels "[B_sen]" for "tomorrow" and "you", and labels the second label "[I_sen]" for the other characters of the training sample. The label of this training sample can be expressed as "[B_sen][I_sen][I_sen][I_sen][I_sen][I_sen][I_sen][I_sen][B_sen][I_sen][I_sen][I_sen][I_sen]".

[0068] Step 403: Iteratively train the sentence chunking model according to the training samples, the first label, and the second label.

[0069] In specific implementation, after the server annotates the labels for the training samples, it can iteratively train the sentence chunking model according to the training samples, the first label, and the second label.

[0070] In one example, the sentence chunking model can be built based on the "bert + bilstm + crf" structure.

[0071] In this embodiment, the pre-trained sentence chunking model is trained through the following steps: obtaining training samples; wherein, each of the training samples contains a number of short sentences; labeling the first word of each short sentence in the training samples with a first label, and labeling the other words of each short sentence except the first word with a second label; iteratively training the sentence chunking model according to the training samples, the first label, and the second label, so that the sentence chunking model can obtain accurate sentence chunking ability.

[0072] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are within the protection scope of this patent; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but not changing the core design of its algorithm and process are within the protection scope of this patent.

[0073] Another embodiment of the present application relates to a sentence extraction device. The implementation details of the sentence extraction device in this embodiment will be specifically described below. The following content is only the implementation details provided for convenient understanding and is not necessary for implementing this solution. The schematic diagram of the sentence extraction device in this embodiment can be as Figure 5 shown, including:

[0074] An acquisition module 501, configured to acquire text information to be processed, wherein the text information to be processed includes text information converted from the acquired voice information.

[0075] A chunking module 502, configured to input the text information to be processed into the pre-trained sentence chunking model to obtain a number of candidate sentences.

[0076] A filtering module 503, configured to filter a number of candidate sentences according to a preset filtering rule to obtain and output a target sentence, wherein the target sentence is the remaining sentence after filtering among the number of candidate sentences.

[0077] It is worth mentioning that each module involved in this embodiment is a logical module. In actual applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, in order to highlight the innovative part of the present application, units not closely related to solving the technical problems proposed by the present application are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.

[0078] Another embodiment of the present application relates to an electronic device, such asFigure 6 As shown, it includes: at least one processor 601; and a memory 602 communicatively connected to the at least one processor 601; wherein, the memory 602 stores instructions executable by the at least one processor 601, and when the instructions are executed by the at least one processor 601, the at least one processor 601 is enabled to execute the statement extraction method in the above embodiments.

[0079] Among them, the memory and the processor are connected by a bus. The bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and the memory together. The bus may also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver may be a component or multiple components, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted over the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.

[0080] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. And the memory can be used to store the data used by the processor when executing operations.

[0081] Another embodiment of the present application relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the method embodiments described above are implemented.

[0082] That is, those skilled in the art can understand that all or part of the steps of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium and includes several instructions for enabling a device (which may be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0083] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application, and in practical applications, various changes can be made in form and details without departing from the spirit and scope of the present application.

Claims

1. A method for extracting sentences, characterized in that, it includes: Obtain the text information to be processed; wherein, the text information to be processed includes the text information converted from the obtained voice information; Input the text information to be processed into a pre-trained sentence chunking model to obtain a number of candidate sentences; Filter the number of candidate sentences according to a preset filtering rule to obtain and output the target sentence; wherein, the target sentence is the sentence remaining after filtering among the number of candidate sentences; The pre-trained sentence chunking model is trained through the following steps: Obtain training samples; wherein, each of the training samples contains a number of short sentences; Label the first character of each short sentence in the training sample with a first label, and label the other characters of each short sentence except the first character with a second label; Iteratively train the sentence chunking model according to the training sample, the first label, and the second label.

2. The sentence extraction method according to claim 1, characterized in that, The filtering rule includes a duplicate removal filtering rule. Filtering the number of candidate sentences according to the preset filtering rule to obtain and output the target sentence includes: Detect whether there are multiple candidate sentences that are exactly the same among the number of candidate sentences; If there are multiple candidate sentences that are exactly the same among the number of candidate sentences, only retain any one of the multiple candidate sentences that are exactly the same; Output the retained candidate sentence as the target sentence.

3. The sentence extraction method according to any one of claims 1 or 2, characterized in that, The filtering rule includes a denoising filtering rule. Filtering the number of candidate sentences according to the preset filtering rule to obtain and output the target sentence includes: Input the number of candidate sentences into a preset perplexity calculation model to obtain the perplexity of each candidate sentence; Respectively determine whether the perplexity of each candidate sentence is less than a preset threshold, and retain the candidate sentences whose perplexity is less than the preset threshold; Output the retained candidate sentence as the target sentence.

4. The sentence extraction method according to claim 3, characterized in that, After respectively determining whether the perplexity of each candidate sentence is less than the preset threshold, it includes: If the perplexity of each candidate sentence is greater than or equal to the preset threshold, generate a re-acquisition instruction, wherein the re-acquisition instruction is used to re-acquire the text information to be processed.

5. The sentence extraction method according to any one of claims 1 to 3, characterized in that, If there are a number of target sentences, the output of the target sentence includes: Determine the positions of the number of target sentences in the text information to be processed; Output the target sentence with the last position.

6. The sentence extraction method according to claim 1, characterized in that, The training samples include first-class training samples, second-class training samples, and third-class training samples. Obtaining the training samples includes: Randomly splice n short sentences among the obtained number of short sentences to obtain the first-class training samples; wherein, the n is an integer greater than 1; Randomly generate a number of strings, and splice the number of strings at both ends of any one of the several short sentences obtained, to obtain the second type of training sample; Obtain the third type of training sample from the target text; wherein, the target text includes at least novels, logs and news.

7. A sentence extraction device, characterized in that, comprising: an acquisition module, configured to acquire text information to be processed; wherein, the text information to be processed includes text information converted from acquired voice information; a chunking module, configured to input the text information to be processed into a pre-trained sentence chunking model to obtain a number of candidate sentences; a filtering module, configured to filter the number of candidate sentences according to a preset filtering rule, to obtain and output a target sentence; wherein, the target sentence is the sentence remaining after filtering among the number of candidate sentences; The pre-trained sentence chunking model is trained through the following steps: Obtain training samples; wherein, each of the training samples contains a number of short sentences; Label a first label for the first character of each short sentence in the training sample, and label a second label for other characters of each short sentence except the first character; Iteratively train the sentence chunking model according to the training sample, the first label and the second label.

8. An electronic device, characterized in that, comprising: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute the sentence extraction method according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements the sentence extraction method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Statement rationality judgment method and device based on semantic parsing and computer equipment

    CN109992769A

  • Conference summary processing method and device, equipment and medium

    CN113011169A