Text report generation device, generation system and generation method

CN122575369APending Publication Date: 2026-08-14GETAC TECH CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

然而,这种方式不仅容易因人为疏失漏记了影音内容中的关键信息,且对报告制作者来说由于需紧盯影音内容,不仅耗时更耗费心力

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575369A_ABST
    Figure CN122575369A_ABST
Patent Text Reader

Abstract

A text report generation apparatus, system, and method are disclosed. The method generates a text report from an audio-visual recording of an event. First, the audio recordings in the audio-visual recording are sequentially converted into a first text record based on a time sequence. Next, an image record of the audio-visual recording under a timestamp in the time sequence is analyzed to generate a second text record about the content of this image record. Then, a third text record is generated based on the identification of the image record content using a key object. Finally, the first, second, and third text records are merged based on the timestamp to form a text report. This improves the efficiency and accuracy of data integration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a report processing system and method, and more particularly to a text report generation device, generation system and generation method. Background Technology

[0002] Traditionally, to facilitate subsequent review or searching of the scene of an event, corresponding text reports are generated and saved for each recorded audio and video recording of the scene. Especially for audio and video recordings of major events, in order to accurately present the scene, in addition to transcribing the recorded audio content word for word, text descriptions of specific video content are also required.

[0003] However, existing technologies primarily rely on manual review and recording of audio-visual content to generate corresponding text reports. This method is prone to human error in omitting crucial information from the audio-visual content, and it is also time-consuming and labor-intensive for report creators who must constantly monitor the content.

[0004] Therefore, there is an urgent need for a text report generation device, generation system, and generation method that can solve the above-mentioned technical problems. Summary of the Invention

[0005] One embodiment of this invention provides a text report generation method for generating a text report from an audio-visual recording made during an event, comprising: sequentially converting the audio recording in the audio-visual recording into a first text record according to a time sequence of the audio-visual recording; generating a second text record about the image content of the image record based on an image record in the audio-visual recording at a timestamp in the time sequence; identifying the image content of the image record based on at least one key object to generate a third text record; and merging the first text record, the second text record, and the third text record to form the text report.

[0006] In some implementations, the text report generation method further includes: converting the audio recording in the audio-visual recording into a verbatim transcript; and analyzing the content of the verbatim transcript, identifying the corresponding speaker in the verbatim transcript, and forming the first text record.

[0007] In some implementations, a large language model is used to analyze the transcript to identify the corresponding speaker in the transcript.

[0008] In some implementations, the second text record describes an event presented by the image content recorded in the image record.

[0009] In some implementations, an attention-based deep learning model is used to analyze the image record and the event that the image content is presented to generate the second text record.

[0010] In some implementations, the third text record describes whether the image content of the image record contains the at least one key object.

[0011] In some implementations, at least one critical object is an illegitimate object.

[0012] In some implementations, a zero-shot object detection model is used to detect the image content of the image record to generate the third text record.

[0013] In some implementations, the second and third text records are written into the first text record at the corresponding timestamp, so as to merge the first, second, and third text records to form the text report.

[0014] Another embodiment of this invention provides a text report generation system, comprising an audio / video capture device and a text report generation device. The audio / video capture device is used to record audio / video data of an event to generate an audio / video record. The text report generation device is coupled to the audio / video capture device and is used to receive the audio / video record to generate a text report. The text report generation device further includes a storage device and a processor. The storage device stores at least one instruction. The processor, coupled to the storage device, is used to access the at least one instruction to execute: sequentially converting the audio recording in the audio / video record into a first text record according to a time sequence of the audio / video record; generating a second text record about the image content of the image record according to an image record at a timestamp in the time sequence; identifying the image content of the image record according to at least one key object to generate a third text record; and merging the first text record, the second text record, and the third text record to form the text report.

[0015] In some implementations, the processor is also configured to access the at least one instruction to execute: transcribing the speech record in the audio-visual recording into a transcript; and analyzing the content of the transcript, identifying the corresponding speaker in the transcript, and forming the first text record.

[0016] In some implementations, the processor is also used to access the at least one instruction to perform a large language model analysis of the verbatim text, identifying the corresponding speaker in the verbatim text.

[0017] In some implementations, the second text record describes an event that occurred within the image content recorded by the image record.

[0018] In some implementations, the processor is also configured to access the at least one instruction to execute: an attention mechanism deep learning model analyzes the events occurring in the image recording to generate the second text record.

[0019] In some implementations, at least one critical object is an illegitimate object.

[0020] In some implementations, the processor is also configured to access the at least one instruction to execute: a zero-shot object detection model detects the image content of the image record to generate a third text record of whether the image content contains the at least one key object.

[0021] In some implementations, the processor is also configured to access the at least one instruction to execute: writing the second text record and the third text record into the first text record at the corresponding timestamp, so as to merge the first text record, the second text record and the third text record to form the text report.

[0022] Another embodiment of this invention provides a text report generation apparatus for receiving an audio-visual recording from an audio-visual capture device to generate a text report. The text report generation apparatus includes a storage device and a processor. The storage device stores at least one instruction. The processor, coupled to the storage device, accesses the at least one instruction to execute: sequentially converting the audio recording in the audio-visual recording into a first text record according to a time sequence of the audio-visual recording; generating a second text record related to the image content of the image record based on an image record at a timestamp in the time sequence; identifying the image content of the image record based on at least one key object to generate a third text record; and merging the first text record, the second text record, and the third text record to form the text report.

[0023] The text report generation device, system, and method in this case can convert the audio recording in the audio-visual recording into a verbatim transcript, then use a large-scale language model to analyze the transcript, identify the corresponding speaker, and mark the speaker to form a first text record. It also uses video event analysis technology to infer the content of the video recording at a time stamp, generating a second text record describing the event in the video recording. Simultaneously, it integrates object detection technology to identify the video recording at a time stamp and generate a third text record indicating whether the video recording contains an illegal object. Finally, the first, second, and third text records are merged using timestamps to form an accurate and complete text report. Besides significantly improving the efficiency and accuracy of data integration, this method also avoids omissions or errors that may occur during manual recording of audio-visual content due to human error, ensuring the consistency and completeness of the record. Attached Figure Description

[0024] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this case and, together with the specification, serve to explain the technical solutions of the embodiments of this case.

[0025] Figure 1 This is a schematic diagram illustrating a text report generation system according to some embodiments of the present invention.

[0026] Figure 2 This is a flowchart illustrating a text report generation method according to some embodiments of the present invention.

[0027] The reference numerals in the attached figures are explained as follows:

[0028] 100: Text Report Generation System

[0029] 110: Audio and video capture device

[0030] 111: Storage element

[0031] 112: Audio and video recording

[0032] 113: Voice Recording

[0033] 114: Video Recording

[0034] 115: Transcript

[0035] 116: First text record

[0036] 117: Second Text Record

[0037] 118: Third Text Record

[0038] 119: Written Report

[0039] 120: Text Report Generation Device

[0040] 121: Storage device

[0041] 122: Processor

[0042] 130: Network

[0043] 200: Methods for Generating Text Reports

[0044] 210-250: Steps Detailed Implementation

[0045] The concept of this invention will be clearly explained below with reference to the accompanying drawings and detailed description. Any person skilled in the art who understands the embodiments of this invention may make changes and modifications based on the technology taught in this invention without departing from the concept and scope of this invention.

[0046] The terminology used herein is for the purpose of describing specific embodiments only and is not intended to limit the scope of this work. Singular forms such as “a,” “this,” “this,” “the,” and “the” as used herein also include plural forms.

[0047] The terms "coupled" or "connected" as used in this article can refer to two or more components or devices making direct physical contact with each other, or making indirect physical contact with each other, or to two or more components or devices operating or acting on each other.

[0048] The terms “include,” “including,” “have,” “contain,” etc., used in this article are all open-ended terms, meaning that they include but are not limited to.

[0049] The term "and / or" as used herein includes any or all of the things mentioned.

[0050] Unless otherwise specified, the terms used herein generally have their ordinary meaning in the context of the art, the subject matter, and the specific content of this case. Certain terms used to describe this case will be discussed below or elsewhere in this specification to provide additional guidance to those skilled in the art in describing the case.

[0051] Traditionally, on-site audio-visual content is manually reviewed, and corresponding written reports are generated through manual recording. However, this method is prone to human error in omitting key information from the audio-visual content, and it is also time-consuming and labor-intensive for the report creator, who must constantly monitor the content. Therefore, this invention provides a written report generation device, system, and method. It uses a speech recognition model to convert the speech recording from an on-site audio-visual recording into a first written record. It also uses a deep learning model to identify a timestamped image record within the audio-visual content to generate a second written record, and a specific object detection model to identify this image record to generate a third written record. Finally, it integrates the first, second, and third written records based on the timestamps to generate a written report. This significantly improves the efficiency and accuracy of data integration, avoids omitting key information from the audio-visual content due to human error, and ensures the consistency and completeness of the record.

[0052] Figure 1 This is a schematic diagram illustrating a text report generation system according to some embodiments of the present invention. Please refer to... Figure 1 The text report generation system 100 includes an audio / video capture device 110 and a text report generation device 120. The audio / video capture device 110 is coupled to the text report generation device 120 via a network 130.

[0053] The audio-visual capture device 110 is used to record audio-visual events of a crime scene to generate an audio-visual record 112, which is stored in the storage element 111 of the audio-visual capture device 110. In some embodiments, the audio-visual capture device 110 is a personal audio-visual recorder carried by the recorder, for example, a police officer carrying a covert recorder to record audio-visual events occurring at a crime scene, but this case is not limited to this. The text report generation device 120 is coupled to the audio-visual capture device 110 via a network 130 and receives the audio-visual record 112 uploaded by the audio-visual capture device 110 via the network 130. In some embodiments, the text report generation device 120 executes a text report generation method to convert the audio-visual record 112 into a text report 119 for output.

[0054] In some embodiments, the text report generation apparatus 120 includes at least a storage device 121 and a processor 122, wherein the storage device 121 is communicatively coupled or electrically coupled to the processor 122. In some embodiments, the processor 122 may include, but is not limited to, a single processing unit or an integration of multiple microprocessors, which may be electrically coupled to the storage device 121, wherein the storage device 121 may be internal or external memory, including volatile or non-volatile memory. In this embodiment, the processor 122 may access and execute at least one instruction from the storage device 121 to further implement, for example... Figure 2 The text report generation method shown.

[0055] Figure 2 This is a schematic diagram illustrating a text report generation method according to some embodiments of the present invention. In some embodiments, the text report generation method 200 may be... Figure 1 The text report generation system 100 shown in the embodiment is executed to identify and integrate the captured audio-visual recordings 112 to generate a text report 119. It should be understood that, unless otherwise specified, the order of operations in the text report generation method 200 mentioned in this embodiment can be adjusted according to actual needs, and may even be executed simultaneously or partially simultaneously. Furthermore, in different embodiments, these operations can be adaptively added, replaced, and / or omitted. Please also refer to... Figure 1 as well as Figure 2 .

[0056] First, in step 210, an audio-visual record is generated. In some embodiments, the audio-visual record 112 used to form the text report 119 is formed by recording a live event using the audio-visual capture device 110. Therefore, in this case, the audio-visual capture device 110 first captures an audio-visual record 112 of a live event and uploads the audio-visual record 112 to the text report generation device 120 via the network 130 to generate the corresponding text report 119. In some embodiments, the audio-visual record 112 consists of voice and video, therefore the audio-visual record 112 includes voice recording 113 and video recording 114. In this case, the voice recording 113 and video recording 114 are processed separately to form their respective text records, and then the respective text records are integrated to generate the text report 119.

[0057] The speech recording 113 in the audio-visual recording 112 is processed to form a corresponding text recording. Accordingly, in step 220, the speech recording is converted word-for-word into a verbatim transcript. In some embodiments, the speech recording 113 in the audio-visual recording 112 is converted into a verbatim transcript 115 sequentially according to a time sequence of the audio-visual recording 112. In some embodiments, a known Whisper transcription technique can be used to transcribe the audio file of the speech recording 113 into a verbatim transcript 115. However, it is worth noting that the above is only one embodiment, and other transcription techniques can also be used to convert the speech recording into a verbatim transcript.

[0058] In step 222, based on the content of the transcript 115, the corresponding speaker is identified in the transcript 115 to form a first text record 116. In some embodiments, a large language model is used to analyze the transcript 115, identify the corresponding speaker in the transcript 115, and mark the speaker to form the first text record 116. In some embodiments, if the audio-visual capture device 110 is a covert recorder carried by a police officer, the transcript 115 can be analyzed using a large language model to further identify whether the speech record in the transcript 115 comes from a police officer or a party involved and mark it to form the first text record 116. In some embodiments, known large language models such as ChaptGPT or Llama 3.1 can be used to analyze the transcript 115 and mark the corresponding speaker. However, it is worth noting that the above is only one embodiment, and other types of large language models can also be used to analyze the transcript 115. In some implementations, the voice recordings 113 each have a timestamp indicating the recording time, so the converted first text recordings 116 not only record spoken sentences but also have corresponding timestamps.

[0059] The video recording 114 in the audio-visual recording 112 is processed to form a corresponding text record. Accordingly, in step 230, a second text record 117 related to the video content of the video recording 114 is generated. In some embodiments, since the converted first text record 116 contains not only spoken sentences but also corresponding timestamps, the video content of the corresponding video recording 114 can be searched based on the timestamp of the first text record 116 to generate a second text record 117 related to this video content. That is, if a key spoken record is displayed under a timestamp in the first text record 116, in order to present the situation at the timestamp, a second text record 117 related to the video content of the video recording 114 can be generated based on the video recording 114 under the timestamp. The second text record 117 is used to describe the events in the video content of the video recording 114, and the situation at the timestamp is presented through the text record. In some embodiments, an attention mechanism deep learning model (Transformer Base) can be used to analyze the video content of the video recording 114, infer the context, and generate the second text record 117. However, it is worth noting that the above is only one implementation method. Other types of models can also be used to analyze the image content of image record 114 to generate second text record 117.

[0060] Accordingly, in step 240, a third text record 118 is generated regarding the image content of image record 114. In some embodiments, the third text record 118 is generated by identifying the image content of image record 114 based on a key object, thereby describing whether the image content of image record 114 contains at least one key object. In some embodiments, the at least one key object is an illegal object, such as a gun or drugs. In some embodiments, the image content of image record 114 can be searched and identified based on the timestamp of first text record 116 to generate a third text record 118 regarding whether the image content of image record 114 contains at least one key object. That is, if a key speech record appears to indicate that the person in the first text record 116 is in possession of an illegal object under a timestamp, a third text record 118 regarding whether the image content of image record 114 contains this illegal object can be generated by identifying the image record 114 under the timestamp. In some implementations, a zero-shot object detection model, such as YOLO World or NVIDIA-AI Nanoowl, can be used to detect the image content of image record 114 to generate third text record 118. However, it is worth noting that the above is only one implementation; other types of models can also be used to recognize the image content of image record 114 to generate third text record 118.

[0061] After the first text record 116, the second text record 117, and the third text record 118 are generated, in step 250, the second text record 117 and the third text record 118 can be written into the corresponding timestamp in the first text record 116 to merge the first text record 116, the second text record 117, and the third text record 118 to form a text report 119. In the above embodiment, the second text record 117 and the third text record 118 are records used to record the image content under the timestamp in the image record 114, and are therefore written into the corresponding timestamp in the first text record 116. However, this application is not limited to this. In other embodiments, if the second text record 117 is used to record the image content under the first timestamp in the image record 114, and the third text record 118 is used to record the image content under the second timestamp in the image record 114, then the second text record 117 will be written to the first timestamp in the first text record 116, and the third text record 118 will be written to the second timestamp in the first text record 116 to form a text report 119.

[0062] In summary, this case provides a text report generation device, system, and method. It involves converting the audio recordings in an audio-visual archive word-for-word into a transcript, then using a large-scale language model to analyze the transcript, identifying the corresponding speaker, and marking the speaker to form a first text record. Next, the video recordings in the audio-visual archive are processed, using video event analysis technology to infer the content of the video recording at a given time stamp, generating a second text record describing the events in the video recording. Simultaneously, an object detection technology is integrated to identify the video recording at a given time stamp, generating a third text record indicating whether the video recording contains an illegal object. Finally, the first, second, and third text records are merged using timestamps to form an accurate and complete text report. Besides significantly improving the efficiency and accuracy of data integration, this method also avoids omissions or errors that may occur during manual recording of audio-visual content due to human error, ensuring the consistency and completeness of the record.

[0063] Although the present invention has been disclosed above with reference to embodiments, it is not intended to limit the present invention. Any person skilled in the art may make some changes and modifications without departing from the concept and scope of the present invention. Therefore, the scope of protection of the present invention shall be determined by the claims.

Claims

1. A method for generating a text report, used to generate a text report from an audio-visual recording of an event, characterized in that, Include: Based on a time sequence of the audio-visual recording, the audio recordings in the audio-visual recording are sequentially converted into a first text record; Analyze an image record in the audio-visual record at a timestamp in the time series, and generate a second text record about the image content of the image record; Based on at least one key object, identify the content of the image record to generate a third text record; and The first text record, the second text record, and the third text record are merged based on the timestamp to form the text report.

2. The text report generation method as described in claim 1, characterized in that, Also includes: Translate the audio recording in the video recording into a verbatim transcript; as well as The content of the verbatim transcript is analyzed, the corresponding speaker is identified in the transcript, and the first written record is formed.

3. The text report generation method as described in claim 2, characterized in that, The transcript was analyzed using a large language model, and the corresponding speakers were identified within the transcript.

4. The text report generation method as described in claim 1, characterized in that, The second textual record describes an event presented by the content of the image recording.

5. The text report generation method as described in claim 4, characterized in that, A deep learning model with an attention mechanism is used to analyze the image and record the event presented in the image content to generate the second text record.

6. The text report generation method as described in claim 1, characterized in that, The third textual record describes whether the image content of the image record contains the at least one key object.

7. The text report generation method as described in claim 6, characterized in that, At least one key object is an illegal object.

8. The text report generation method as described in claim 6, characterized in that, The third text record is generated by using a zero-shot object detection model to detect the content of the image record.

9. The text report generation method as described in claim 1, characterized in that, The second and third text records are written into the first text record at the corresponding timestamp, so that the first, second, and third text records are merged to form the text report.

10. A text report generation system, characterized in that, Include: An audio-visual capture device records audio-visual data of an event to generate an audio-visual record; and A text report generation device is coupled to the audio / video capture device for receiving the audio / video recording and generating a text report. The text report generation device further includes: A storage device that stores at least one instruction; and A processor, coupled to the storage device, is used to access and execute the at least one instruction: Based on a time sequence of the audio-visual recording, the audio recordings in the audio-visual recording are sequentially converted into a first text record; Analyze an image record in the audio-visual record at a timestamp in the time series, and generate a second text record about the image content of the image record; Based on at least one key object, identify the content of the image record to generate a third text record; and The first text record, the second text record, and the third text record are merged based on the timestamp to form the text report.

11. The text report generation system as described in claim 10, characterized in that, The processor is also used to access the at least one instruction for execution: Translate the audio recording in the video recording into a verbatim transcript; and The content of the verbatim transcript is analyzed, the corresponding speaker is identified in the transcript, and the first written record is formed.

12. The text report generation system as described in claim 11, characterized in that, The processor is also used to access the at least one instruction to perform a large language model analysis of the verbatim text, identifying the corresponding speaker in the verbatim text.

13. The text report generation system as described in claim 10, characterized in that, The second textual record describes an event that occurred within the content of the image recorded in the image.

14. The text report generation system as described in claim 13, characterized in that, The processor is also used to access the at least one instruction for execution: A deep learning model with an attention mechanism analyzes the event recorded in the video recording to generate the second text record.

15. The text report generation system as described in claim 10, characterized in that, At least one key object is an illegal object.

16. The text report generation system as described in claim 15, characterized in that, The processor is also used to access the at least one instruction for execution: A zero-shot object detection model detects the image content of the image record to generate a third text record of whether the image content contains the at least one key object.

17. The text report generation system as described in claim 10, characterized in that, The processor is also used to access the at least one instruction for execution: The second and third text records are written into the first text record at the corresponding timestamp, so that the first, second, and third text records are merged to form the text report.

18. A text report generation apparatus for receiving an audio-visual recording from an audio-visual capture device to generate a text report, characterized in that, The text report generation device includes: A storage device that stores at least one instruction; and A processor, coupled to the storage device, is used to access and execute the at least one instruction: Based on a time sequence of the audio-visual recording, the audio recordings in the audio-visual recording are sequentially converted into a first text record; Analyze an image record in the audio-visual record at a timestamp in the time series, and generate a second text record about the image content of the image record; Based on at least one key object, identify the content of the image record to generate a third text record; and The first text record, the second text record, and the third text record are merged based on the timestamp to form the text report.