Method and device for detecting text abstract generated by large model, medium and equipment

By dividing the text summaries generated by the large model into time intervals and performing content matching detection, the problems of low detection accuracy and high manual cost in the existing technology are solved, and efficient and automated text summary authenticity detection is achieved.

CN121996785APending Publication Date: 2026-05-08BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-11-06
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing text summarization schemes generated by large detection models suffer from problems such as low detection accuracy or excessive manual costs, especially due to their reliance on human subjective judgment and the high quality of reference summaries.

Method used

By acquiring the modalities of the target data, the target text is extracted and divided into multiple time intervals. A large model is used to generate content summaries, and a second large model is used to determine whether the text content of each time interval matches the summary, thus automatically detecting the authenticity of the text summary.

Benefits of technology

It improves the accuracy and efficiency of text summarization generated by large detection models, reduces manpower and time costs, can automatically and accurately identify non-real content, and reduces computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121996785A_ABST
    Figure CN121996785A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and device for detecting a text abstract generated by a large model, a medium and equipment. The method comprises the steps of obtaining target data, and determining a data mode of the target data; if the data modality is a video or audio modality, extracting a target text corresponding to the target data, and generating a first text abstract of the target text through a first large model, the first text abstract comprising a plurality of time intervals divided according to the target text and respective content abstracts of the plurality of time intervals; according to the plurality of time intervals, extracting respective text contents of the plurality of time intervals from the target text; determining whether the text content of each time interval is matched with the content abstract of the time interval or not through a second large model; and determining whether unreal content exists in the first text abstract or not according to whether the text content of each time interval is matched with the content abstract or not. Through the method, whether unreal content exists in the text abstract generated by detecting the large model or not can be efficiently and accurately determined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a method, apparatus, medium, and device for detecting text summaries generated by large models. Background Technology

[0002] Large models refer to artificial intelligence models with hundreds of millions or more parameters, pre-trained on massive datasets. Text summarization typically refers to the process of extracting key information from a long text to form a concise and accurate summary. With the development of large model technology, generating text summaries of long texts using large models has become a common method for obtaining text summaries. However, due to the "illusion" problem of large models—that is, the content generated by large models can contain non-realistic content—it is often necessary to detect whether there is non-realistic content in the text summaries generated by large models. However, existing methods for detecting text summaries generated by large models still suffer from problems such as low detection accuracy or excessive manual costs. Summary of the Invention

[0003] This disclosure describes a method, apparatus, medium, and device for detecting text summaries generated by a large model.

[0004] According to the first aspect, a method for detecting text summaries generated by large models is provided, including:

[0005] Acquire target data and determine the data modality of the target data; if the data modality is a video modality or an audio modality, extract the target text corresponding to the target data, and generate a first text summary corresponding to the target text through a first major model. The first text summary includes multiple time intervals divided according to the target text and content summaries corresponding to each of the multiple time intervals; extract the text content corresponding to each of the multiple time intervals from the target text according to the multiple time intervals.

[0006] The second major model determines whether the text content of each time interval in the multiple time intervals matches the content summary of the time interval; based on whether the text content of each time interval matches the content summary of the time interval, it is determined whether there is any false content in the first text summary.

[0007] According to the second aspect, an apparatus for detecting text summaries generated by a large model is provided, comprising:

[0008] The acquisition unit is configured to: acquire target data; determine the data modality of the target data; if the data modality is a video modality or an audio modality; extract the target text corresponding to the target data; generate a first text summary corresponding to the target text through a first large model; the first text summary includes multiple time intervals divided according to the target text, and content summaries corresponding to each of the multiple time intervals; and extract the text content corresponding to each of the multiple time intervals from the target text according to the multiple time intervals.

[0009] The judgment unit is configured to, through the second major model, determine whether the text content of each time interval in the plurality of time intervals matches the content summary of the time interval; and, based on whether the text content of each time interval matches the content summary of the time interval, determine whether there is any false content in the first text summary.

[0010] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed in a computer, causes the computer to perform the method of the first aspect.

[0011] According to a fourth aspect, an electronic device is provided, including a memory and a processor, wherein the memory stores executable code, and the processor executes the executable code to implement the method of the first aspect.

[0012] The present disclosure provides a method, apparatus, device, and medium for detecting text summaries generated by a large detection model. First, target data is acquired, and the data modality of the target data is determined. If the data modality is video or audio, target text corresponding to the target data is extracted. A first text summary corresponding to the target text is generated using a first large detection model. The first text summary includes multiple time intervals divided according to the target text, and content summaries for each of the multiple time intervals. Text content for each of the multiple time intervals is extracted from the target text based on the multiple time intervals. Then, a second large detection model is used to determine whether the text content of each time interval matches the content summary of that time interval. Based on whether the text content of each time interval matches the content summary, it is determined whether there is any false content in the first text summary. This method can automatically, efficiently, and accurately determine whether there is any false content in the text summaries generated by the large detection model. Attached Figure Description

[0013] Figure 1 This diagram illustrates the manual detection of text summaries generated by a large model.

[0014] Figure 2 This diagram illustrates the detection of text summaries generated by a large model using reference summaries.

[0015] Figure 3 A schematic diagram of a method for detecting text summaries generated by a large model according to an embodiment of the present disclosure is shown;

[0016] Figure 4 A flowchart illustrating a method for detecting text summaries generated by a large model according to an embodiment of the present disclosure is shown.

[0017] Figure 5 A schematic diagram of a method for detecting text summaries generated by a large model according to another embodiment of the present disclosure is shown;

[0018] Figure 6 A schematic block diagram of an apparatus for detecting text summaries generated by a large model according to an embodiment of the present disclosure is shown;

[0019] Figure 7 A schematic diagram of the structure of an electronic device suitable for implementing embodiments of the present disclosure is shown;

[0020] Figure 8 A schematic diagram of the structure of a storage medium suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation

[0021] The technical solutions provided in this specification will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the relevant invention and not intended to limit the invention. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0022] In the description of the implementations disclosed herein, the term "comprising" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one / an implementation" or "the implementation" should be understood as "at least one / an implementation". The term "some implementations" should be understood as "at least some implementations". Other explicit and implicit definitions may also be included below.

[0023] As mentioned earlier, text summarization typically refers to processing long texts to extract key information and form a concise and accurate summary. Large models refer to artificial intelligence models with hundreds of millions or more parameters, pre-trained on massive datasets. With the development of large model technology, generating text summaries of long texts using large models is an efficient way to obtain text summaries. However, due to the "illusion" problem of large models—that is, the content generated by large models can contain inaccurate information—it is often necessary to detect whether inaccurate content exists in the text summaries generated by large models. Existing methods for detecting text summaries generated by large models typically rely on manual verification or verification based on reference summaries.

[0024] Figure 1 This diagram illustrates the manual detection of text summaries generated by a large model. For example... Figure 1 As shown, the target text to be summarized can be input into a large model to obtain a generated text summary. Then, users can detect whether the text summary contains non-realistic content, or "illusionary" content. However, manual detection relies on the subjective judgment of the detectors; different users may have different detection standards and preferences, leading to inconsistent detection results for the same summary. Furthermore, manual detection requires significant manpower.

[0025] Figure 2 This diagram illustrates the detection of text summaries generated by a large model using reference summaries. For example... Figure 2 As shown, after generating a text summary of the target text using a large model, a pre-defined term frequency algorithm (such as an n-gram algorithm, i.e., an N-gram algorithm) can be used to determine the similarity between the text summary and a pre-obtained reference summary. Then, the accuracy of the text summary is determined based on the similarity. However, this approach also has the following problems: First, the detection results heavily depend on the quality of the reference summary; inaccurate or incomplete reference summaries can lead to inaccurate detection results. In real-world production scenarios, it is often difficult to obtain high-quality reference summaries of the text being detected, or obtaining them requires high costs. Second, existing term frequency algorithms for confirming the similarity between text summaries and reference summaries typically determine the similarity based on the frequency of occurrence of words or word sequences in both. This method usually cannot measure whether the two truly match semantically. Therefore, the accuracy and reasonableness of detection results obtained based on term frequency algorithms are often not ideal.

[0026] To address the aforementioned technical problems, this disclosure provides a flowchart illustrating a method for detecting text summaries generated by a large model. Figure 3 A schematic diagram of a method for detecting text summaries generated by a large model according to an embodiment of the present disclosure is shown. Figure 3As shown, in some embodiments, such as when the target data is video or audio data, corresponding target text can be generated based on the target data. A first major model is used to obtain multiple time intervals of the target data, divided according to the target text, and a content summary for each time interval. Based on the multiple time intervals, the text content of each time interval is extracted from the target text. Then, a second major model is used to determine whether the text content and content summary of each time interval match, thereby determining whether there is any non-real content in the text summary of the target text. In other embodiments, such as when the target data is text data, a text summary of the target data can also be obtained using the first major model. Then, a third major model is used to extract multiple events from the second text summary, and the third major model is used to determine whether each event can be inferred from the target data, thereby determining whether there is any non-real content in the text summary of the target data.

[0027] The advantages of this method are as follows: First, compared to manually detecting the accuracy of text summaries generated by large models, this method can efficiently and automatically determine whether there is false content in the text summaries generated by large models. This not only avoids the influence of subjective judgment by human users, improving detection accuracy, but also significantly reduces the time and manpower costs of detection. Second, this method detects the quality of text summaries generated by a large model (e.g., the first large model) based on an evaluation large model (e.g., the second large model). Compared to detection schemes based on word frequency algorithms, it can obtain more accurate and reasonable detection results based on the semantics of the detected text (e.g., the target text) and the text summaries. Third, in some embodiments, the text corresponding to video / audio data can be divided into multiple time intervals, and the text content of each time interval can be determined to match the content summary, thereby determining whether there is false content in the entire text summary of the target text. This approach reduces the computational complexity of verifying video / audio data summaries and improves the accuracy of the summary verification results. Fourth, in some embodiments, multiple events can be extracted from the text summary of the text data, and the presence of false content in the entire text summary of the target text can be determined based on whether each event is inferred separately from the text data. This reduces the computational complexity of verifying text data summaries and improves the accuracy of the summary verification results.

[0028] The following describes the detailed process of this method.

[0029] Figure 4 A flowchart illustrating a method for detecting text summaries generated by a large model according to an embodiment of this disclosure is shown. Figure 4 As shown, the method includes at least the following steps:

[0030] Step S401: Obtain target data and determine the data modality of the target data; if the data modality is a video modality or an audio modality, extract the target text corresponding to the target data, and generate a first text summary corresponding to the target text through a first large model. The first text summary includes multiple time intervals divided according to the target text and content summaries corresponding to each of the multiple time intervals; extract the text content corresponding to each of the multiple time intervals from the target text according to the multiple time intervals.

[0031] Step S403: Using the second major model, determine whether the text content of each time interval in the plurality of time intervals matches the content summary of the time interval; based on whether the text content of each time interval matches the content summary of the time interval, determine whether there is any false content in the first text summary.

[0032] First, in step S301, target data is acquired, and the data modality of the target data is determined. If the data modality is a video modality or an audio modality, the target text corresponding to the target data is extracted. Modality refers to the form or type of data. In different embodiments, the target data may be acquired through different specific methods and used for different specific business or purposes; this specification does not impose any limitations on this. In this step, the data modality of the target data can be determined. If the data modality of the target data is a video modality or an audio modality, the target text corresponding to the target data can be extracted.

[0033] In different specific embodiments, the specific methods for extracting target text can differ depending on whether the target data is in video or audio modality. In one embodiment, if the data modality is video, a first text corresponding to the target data can be extracted using an optical character recognition (OCR) tool, and a second text corresponding to the target data can be extracted using an automatic speech recognition (ASR) tool; the first text and the second text are used as the target text. Optical character recognition (OCR) tools can be used to extract text information from images or video frames. In different specific embodiments, different types of OCR tools can be used to extract text information from the target data as the first text. Automatic speech recognition (ASR) tools can be used to convert human speech in audio data into corresponding text information. Since video data often includes audio content, ASR tools can be used to extract text information corresponding to the audio content of the target data as the second text. In different specific embodiments, the ASR tool used to extract the second text can be of different types, and this specification does not limit this. In this way, when the target data is in video modality, the text information corresponding to the target data can be obtained efficiently and accurately.

[0034] In another embodiment, if the data modality is audio, an automatic speech recognition tool is used to extract the second text corresponding to the target data; this second text is then used as the target text. This method allows for efficient and accurate acquisition of the text information corresponding to the target data when the target data is in an audio modality.

[0035] After obtaining the target text, a first text summary corresponding to the target text can be generated using the first large model. Furthermore, text content corresponding to multiple time intervals can be extracted from the target text based on multiple time intervals. A large model typically refers to an artificial intelligence model with hundreds of millions or more parameters, pre-trained on large-scale data. In different embodiments, the first large model can be a large model of different specific types or with different neural network structures. The second and third large models mentioned later can also be large models of different specific types or with different neural network structures.

[0036] Because video or audio data is typically time-series, the target text usually includes a large amount of detailed information marked by timestamps. These timestamps are often difficult for users to understand directly, and the details are often numerous and overly detailed, making it difficult for users to quickly form a summary of the entire video or audio. Therefore, the first text summary can include multiple time intervals divided according to the target text, as well as content summaries for each time interval. Using the first main model, the target text is divided into multiple time intervals based on the timestamps of the numerous detailed information, and content summaries for each time interval are obtained. Essentially, this involves dividing the target text into multiple paragraphs or chapters based on the time dimension and obtaining content summaries for each paragraph or chapter. Thus, users can quickly form a holistic understanding of the video or audio through the various time intervals and their content summaries.

[0037] Specifically, in one embodiment, a prompt word can be constructed to instruct the generation of a text summary corresponding to the target text. This prompt word is then input into a first large model to obtain a first text summary. The prompt word refers to the text information input into the large model, with the purpose of guiding the model to generate corresponding content based on the prompt word. In different embodiments, the specific form of the prompt word and the specific form of the obtained first text summary may differ, and this specification does not impose any limitations on this.

[0038] Then, in step S403, the second model can be used to determine whether the text content of each time interval in the multiple time intervals matches the content summary of that time interval; based on whether the text content of each time interval matches the content summary of that time interval, it can be determined whether there is any false content in the first text summary.

[0039] In different embodiments, the specific methods for determining whether there is false content in the first text summary can vary. In one embodiment, the presence of false content in the first text summary can be determined as follows: if a target interval exists among the plurality of time intervals, and the text content of the target interval does not match the content summary of the target interval, then it is determined that there is false content in the text summary; if no target interval exists among the plurality of time intervals, then it is determined that there is no false content in the first text summary. This method allows for efficient and accurate determination of whether there is false content in text summaries of video or audio data generated using large models.

[0040] In different embodiments, the specific methods for determining whether the text content of each time interval matches the content summary of that time interval can also differ. In one embodiment, the following method can be used to determine whether the text content of each time interval matches the content summary of that time interval: For each time interval, if the length of the text summary of the time interval is greater than a predetermined threshold, then multiple first events are extracted from the text summary using a third major model; a second major model is used to determine whether each first event is inferred from the text content of the time interval; if there is a first event among the multiple first events that is not inferred from the text content of the time interval, then it is determined that the text content of the time interval does not match the content summary of the time interval; if there is no first event among the multiple first events that is not inferred from the text content of the time interval, then it is determined that the text content of the time interval matches the content summary of the time interval. Specifically, for time intervals where the length of the text summary is greater than the predetermined threshold, multiple first events of that time interval can be obtained by inputting prompts indicating the extraction of multiple events from the text summary into the third major model. In this way, the computational complexity of determining whether the text content of a specific time interval matches the content summary of that time interval can be reduced when the length of the text summary of a specific time interval is greater than the predetermined threshold.

[0041] An event typically refers to an independent knowledge point or fact. In different specific embodiments, the specific form of the prompt words input into the third model and the first event obtained can vary. The predetermined threshold can also vary in different specific embodiments. In one specific embodiment, the predetermined threshold can be 100 characters.

[0042] In another embodiment, it can also be determined whether the text content of each time interval matches the content summary of that time interval in the following way: For each time interval, if the length of the text summary of the time interval is not greater than a predetermined threshold, the second large model determines whether the content summary of the time interval can be inferred from the text content of the time interval; if the content summary of the time interval can be inferred from the text content of the time interval, then it is determined that the text content of the time interval matches the content summary of the time interval; if the content summary of the time interval cannot be inferred from the text content of the time interval, then it is determined that the text content of the time interval does not match the content summary of the time interval. Specifically, for time intervals where the length of the text summary is not greater than the predetermined threshold, a prompt word indicating whether the content summary of the time interval can be inferred from the text content of the time interval can be input into the second large model to obtain the result of whether the content summary can be inferred from the text content of the time interval. In different specific embodiments, the specific form of the prompt word input into the second large model can also be different. In this way, the efficiency of determining whether the text content of a specific time interval matches the content summary of that time interval can be improved when the length of the text summary of a specific time interval is not greater than the predetermined threshold.

[0043] In some scenarios, the target data itself can be text data, meaning the data modality of the target data is text. Therefore, a text summary of the target data can be directly generated using a large model. To test the accuracy of this type of text summary, in one embodiment, if the data modality determined in step S401 is text, a second text summary corresponding to the target data can be generated using a first large model, and multiple second events can be extracted from the second text summary using a third large model. The second large model is then used to determine whether each second event is inferred from the target data. Based on whether each second event is inferred from the target data, it is determined whether there is any non-realistic content in the second text summary, such as... Figure 5 As shown, this method allows for the efficient and automated verification of the accuracy of text summaries of text data. Furthermore, this method essentially transforms the problem of determining whether there is false content in the overall text summary of text data into determining whether individual events can be inferred from the text data, thereby reducing computational complexity and improving the accuracy of the summary verification results.

[0044] In one embodiment, if any of the plurality of second events is not inferred from the target data, then it is determined that the second text summary contains non-real content; if none of the plurality of second events is inferred from the target data, then it is determined that the second text summary does not contain non-real content. This method can efficiently and accurately detect the absence of non-real content in text summaries generated from text data using a large model. Furthermore, this method can be used to reasonably determine whether non-real content exists in text summaries of text data.

[0045] Figure 6 A schematic block diagram of an apparatus for detecting text summaries generated by a large model according to an embodiment of the present disclosure is provided. The apparatus is used to perform, for example... Figure 4 The method shown. (As shown) Figure 6 As shown, the device 600 includes:

[0046] Acquisition unit 601 is configured to acquire target data and determine the data modality of the target data; if the data modality is a video modality or an audio modality, extract the target text corresponding to the target data, and generate a first text summary corresponding to the target text through a first large model. The first text summary includes multiple time intervals divided according to the target text and content summaries corresponding to each of the multiple time intervals; and extract the text content corresponding to each of the multiple time intervals from the target text according to the multiple time intervals.

[0047] The judgment unit 602 is configured to determine, through the second major model, whether the text content of each time interval in the plurality of time intervals matches the content summary of the time interval; and to determine whether there is false content in the first text summary based on whether the text content of each time interval matches the content summary of the time interval.

[0048] This disclosure also provides an electronic device, including a memory and a processor. The memory stores executable code, and when the processor executes the executable code, it implements, for example... Figure 4 The method shown.

[0049] The following can also be referenced Figure 7 It shows a schematic diagram of the structure of an electronic device 700 suitable for implementing embodiments of the present disclosure. Figure 7 The illustrated electronic device 700 is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0050] like Figure 7As shown, the electronic device 700 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701. The aforementioned processing device 701 may be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and may further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) 702 or a program loaded from storage device 708 into random access memory (RAM) 703. RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing device 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0051] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 7 Each box shown can represent a device or multiple devices as needed.

[0052] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the method for detecting text summaries generated from large models provided by embodiments of this disclosure.

[0053] This disclosure also provides a computer-readable storage medium storing a computer program thereon, which, when executed in a computer, causes the computer to perform the functions provided in this disclosure. Figure 4 The image shows a method for detecting text summaries generated by a large model. Figure 8 A schematic diagram illustrating a storage medium for implementing an embodiment of this disclosure. For example, such as... Figure 8 As shown, the storage medium 800 can be a non-transitory computer-readable storage medium used to store non-transitory computer-executable instructions 801. When the non-transitory computer-executable instructions 801 are executed by a processor, they can implement a method for detecting text summaries generated by large models provided in this disclosure. For example, when the non-transitory computer-executable instructions 801 are executed by a processor, one or more steps in the method for detecting text summaries generated by large models provided in this disclosure can be performed. For example, the storage medium 800 can be applied in the above-mentioned electronic device. For example, the storage medium 800 can include a memory in the electronic device. The description of the storage medium 800 can be found in the description of the memory in the embodiments of the electronic device, and will not be repeated here. The specific functions and technical effects of the storage medium 800 can be found in the description of a method for detecting text summaries generated by large models provided in this disclosure, and will not be repeated here.

[0054] It should be noted that the computer-readable medium in the embodiments of this disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a memory card of a smartphone, a storage component of a tablet computer, a portable computer disk, a hard disk of a personal computer, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, wherein computer-readable program code is carried. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (Radio Frequency), etc., or any suitable combination thereof.

[0055] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs, which, when executed by the server, cause the electronic device to implement the method for detecting text summaries generated by a large model provided in the embodiments of this disclosure.

[0056] Computer program code for performing the operations of embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0057] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. The units described in the embodiments of the present disclosure may be implemented in software or hardware. The names of the units do not necessarily constitute a limitation on the unit itself. The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), Systems-on-Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0058] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for storage media and computing devices are basically similar to the method embodiments, so they are described more simply; relevant parts can be referred to the descriptions of the method embodiments.

[0059] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in this disclosure. Furthermore, although the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of a single embodiment may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0060] The above detailed embodiments further illustrate the purpose, technical solution, and beneficial effects of the embodiments of the present invention. Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. It should be understood that the above are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solutions of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for detecting text summaries generated by a large model, comprising: Acquire target data and determine the data modality of the target data; If the data modality is a video modality or an audio modality, the target text corresponding to the target data is extracted, and a first text summary corresponding to the target text is generated through a first large model. The first text summary includes multiple time intervals divided according to the target text, and content summaries corresponding to each of the multiple time intervals. Based on the multiple time intervals, the text content corresponding to each of the multiple time intervals is extracted from the target text. The second major model determines whether the text content of each time interval in the multiple time intervals matches the content summary of the time interval; based on whether the text content of each time interval matches the content summary of the time interval, it is determined whether there is any false content in the first text summary.

2. The method of claim 1, wherein, Based on whether the text content of each time interval matches the content summary of the time interval, determine whether there is any false content in the first text summary, including: If a target interval exists among the multiple time intervals, and the text content of the target interval does not match the content summary of the target interval, then it is determined that there is false content in the text summary; If the target interval does not exist among the multiple time intervals, then it is determined that there is no false content in the first text summary.

3. The method according to claim 2, wherein, Determining whether the text content of each time interval within the plurality of time intervals matches the content summary of the time interval includes: For each time interval If the length of the text summary of the time interval is greater than a predetermined threshold, then multiple first events are extracted from the text summary using the third major model; then, using the second major model, it is determined whether each first event can be inferred from the text content of the time interval. If any of the plurality of first events is not inferred from the text content of the time interval, then it is determined that the text content of the time interval does not match the content summary of the time interval; if none of the plurality of first events is inferred from the text content of the time interval, then it is determined that the text content of the time interval matches the content summary of the time interval.

4. The method according to claim 3, wherein, Determining whether the text content of each time interval in the plurality of time intervals matches the content summary of the time interval further includes: For each time interval If the length of the text summary of the time interval is not greater than a predetermined threshold, the second major model is used to determine whether the content summary of the time interval can be inferred based on the text content of the time interval. If a content summary of the time interval is inferred from the text content of the time interval, then it is determined that the text content of the time interval matches the content summary of the time interval; if a content summary of the time interval cannot be inferred from the text content of the time interval, then it is determined that the text content of the time interval does not match the content summary of the time interval.

5. The method according to claim 1, wherein, If the data modality is a video modality or an audio modality, extract the target text corresponding to the target data, including: If the data modality is a video modality, the first text corresponding to the target data is extracted using an optical character recognition tool, and the second text corresponding to the target data is extracted using an automatic speech recognition tool; the first text and the second text are used as the target text.

6. The method according to claim 1, wherein, If the data modality is a video modality or an audio modality, extract the target text corresponding to the target data, including: If the data modality is an audio modality, the second text corresponding to the target data is extracted using an automatic speech recognition tool; the second text is then used as the target text.

7. The method according to claim 1, further comprising: If the data modality is a text modality, a second text summary corresponding to the target data is generated through the first major model, and multiple second events are extracted from the second text summary through the third major model; The second major model is used to determine whether each second event can be inferred from the target data, and based on whether each second event can be inferred from the target data, it is determined whether there is any non-real content in the second text summary.

8. The method according to claim 7, wherein, Based on whether each of the second events was inferred from the target data, it is determined whether there is any non-true content in the second text summary, including: If any of the plurality of second events is not inferred from the target data, then it is determined that there is false content in the second text summary; if none of the plurality of second events is inferred from the target data, then it is determined that there is no false content in the second text summary.

9. An apparatus for detecting text summaries generated by a large model, comprising: The acquisition unit is configured to acquire target data and determine the data modality of the target data. If the data modality is a video modality or an audio modality, the target text corresponding to the target data is extracted, and a first text summary corresponding to the target text is generated through a first large model. The first text summary includes multiple time intervals divided according to the target text, and content summaries corresponding to each of the multiple time intervals. Based on the multiple time intervals, the text content corresponding to each of the multiple time intervals is extracted from the target text. The judgment unit is configured to, through the second major model, determine whether the text content of each time interval in the plurality of time intervals matches the content summary of the time interval; and, based on whether the text content of each time interval matches the content summary of the time interval, determine whether there is any false content in the first text summary.

10. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method of any one of claims 1-8.

11. An electronic device comprising a memory and a processor, wherein the memory stores executable code, and the processor, when executing the executable code, implements the method of any one of claims 1-8.